Skip to main content
Glama

BlackVeil DNS & Email Security Scanner

Server Details

DNS and email security scanner with 81 MCP tools for SPF, DMARC, DNSSEC, SSL, and brand audits.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
100.0% over 49 days
Last Tested
Transport
Streamable HTTP · MCP 2025-06-18
URL
Repository
MadaBurns/bv-mcp
GitHub Stars
9
Server Listing
Blackveil DNS

TDQS

A3.6/5.0

Scored across 81 tools

Disambiguation3/5

The set includes many explicitly differentiated tools (e.g., check_dane vs check_dane_https, check_shadow_domains vs check_lookalikes), and async start/status/report triplets are named predictably. However, with 81 tools there are numerous overlapping clusters such as scan_domain vs batch_scan vs compare_domains, check_spf vs resolve_spf_chain vs map_supply_chain, and brand_audit_single vs discover_brand_domains, so an agent can still misselect despite lengthy 'distinct from' notes.

Naming Consistency4/5

Names are overwhelmingly snake_case and follow domain_action patterns (check_*, scan_*, batch_scan_*, brand_audit_*, osint_*). Minor deviations exist—generic 'generate', acronym-led 'sge_quickscan', reversed 'rdap_lookup', and standalone 'cymru_asn'—but the overall convention is readable and consistent.

Tool Count1/5

81 tools is far beyond the well-scoped 3–15 range and exceeds even the heavy 25+ threshold, fragmenting the surface into many atomic checks and async triplets. Although the domain is broad, this count is an extreme mismatch for an MCP tool set.

Completeness5/5

The surface covers a wide lifecycle: single/batch scans, async start/status/report/findings, brand audits and watches, OSINT investigations, compliance mapping, remediation generation, and validate_fix. No obvious core DNS/email-security gap is apparent, though some checks are explicitly unverified by design, such as DANE pin comparison.

Available Tools

81 tools
analyze_driftA
Read-onlyIdempotent
Inspect

Measure whether a domain's DNS security posture improved or regressed by comparing the current state against a prior scan snapshot. Returns a drift classification (improving/stable/regressing/mixed), score delta, and lists of improvements and regressions. Use to answer "did our security score improve or regress since last time?" — distinct from compare_baseline which checks compliance against a fixed policy (not improvement over time).

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to analyze drift for
formatNoOutput verbosity. Auto-detected if omitted.
baselineYesPrior scan reference for drift-over-time analysis: a previous ScanScore JSON STRING, or the literal "cached" to reuse the last cached scan (the default when omitted). NOT a policy/requirements object — for compliance enforcement against required controls, use compare_baseline instead.cached
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds value by specifying the return payload (drift classification, score delta, lists of improvements/regressions) and clarifies the baseline 'cached' reuse, without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with clear information architecture: purpose, output, usage/alternatives. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description covers the return values. It also explains when to use the tool and the key baseline parameter behavior (via schema), making it sufficiently complete for a moderate-complexity tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed per-param explanations (e.g., baseline's 'cached' literal and exclusion of policy objects). The description itself doesn't add parameter-level semantics beyond schema, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool measures DNS security posture change against a prior snapshot, using specific verbs and resources. It explicitly distinguishes from compare_baseline by contrasting improvement-over-time vs compliance-to-policy, differentiating it from a key sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit use case ('did our security score improve or regress since last time?') and an explicit exclusion (not for compliance checks — use compare_baseline instead). The baseline parameter schema reinforces this with 'NOT a policy/requirements object — for compliance enforcement against required controls, use compare_baseline instead.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_spoofabilityA
Read-onlyIdempotent
Inspect

Compute a composite email spoofability risk score (0–100, higher = more spoofable) by combining SPF trust surface, DMARC enforcement, and DKIM coverage. Returns a risk level (minimal→critical), per-control sub-scores, and plain-language summary of how easy it would be to spoof email from the domain. Use when asked how easy it is to spoof email from a domain, or for a composite email spoofing risk score.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=true, idempotentHint=true, and destructiveHint=false, the description adds value by explaining how the score is computed (combining SPF trust surface, DMARC enforcement, DKIM coverage) and what outputs to expect (risk level, sub-scores, plain-language summary). This goes beyond annotation coverage and helps set expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no repetition of schema content, and the most important information (what the tool does and its output) is front-loaded. Every phrase serves a purpose, with the usage note placed last as a natural call to action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description's explanation of returned data (risk level, sub-scores, summary) is valuable. It covers the core behavior and usage context. Minor gaps like handling of invalid domains or network-only data sources exist, but they are not critical for a read-only, idempotent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description doesn't add parameter-level details beyond what the schema already provides (domain, format, force_refresh). It doesn't harm, but it also doesn't compensate further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb ('Compute') and clearly names the resource ('composite email spoofability risk score'), including the 0–100 scale and direction. It explicitly distinguishes itself from sibling tools like check_spf, check_dmarc, and check_dkim by framing this as a combined/composite assessment rather than a single-control check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage triggers: 'Use when asked how easy it is to spoof email from a domain, or for a composite email spoofing risk score.' It implies alternatives (individual control checks) exist among siblings but does not name them or state when NOT to use this tool. This is clear context but lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_scanA
Read-onlyIdempotent
Inspect

Bulk-scan up to 10 domains in parallel. Runs a full security audit on each domain in the list and returns score, NIST-aligned letter grade (6-band A+/A/B/C/D/F), and finding counts per domain. Use when you want to audit multiple domains at once or do a bulk scan of several domains simultaneously — distinct from compare_domains which does a side-by-side analysis of 2–5 domains. Version stamps (hoisted once per batch): 'scoringModelVersion' is the scoring POLICY semver and is INDEPENDENT of 'dnsChecksPackageVersion', the @blackveil/dns-checks npm engine-package version — the model version legitimately lags and the two must not be compared. Record 'scoringConfigHash' when citing scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
domainsYesDomains to scan (max 10 per request)
force_refreshNoBypass cache and run fresh scans.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive traits. The description adds valuable context about version stamps (scoringModelVersion vs dnsChecksPackageVersion independence) and instructs to record scoringConfigHash. No contradictions. It could have mentioned caching behavior (related to force_refresh) but not required given annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized and front-loaded with the core purpose. The version stamp paragraph is essential to prevent misuse, so no fluff. Slightly dense but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description adequately explains the key return values (score, grade, finding counts). It also covers usage distinctiveness and the version gotcha. Lacks explicit response structure or error handling, but sufficient for a batch tool with rich annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and descriptions for each parameter are present. The tool description does not add extra meaning beyond the schema—it repeats the max-10 domain limit and mentions output fields, but that's about return values, not parameters. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Bulk-scan up to 10 domains in parallel. Runs a full security audit on each domain...' and names the return fields (score, NIST-aligned letter grade, finding counts). It also distinguishes from compare_domains, making the tool's scope clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when you want to audit multiple domains at once' and notes distinctness from compare_domains, a side-by-side analysis. It does not explicitly exclude single-domain scans (e.g., scan_domain), but the guidance is sufficient for typical use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_scan_findingsA
Read-onlyIdempotent
Inspect

Fetch owner-scoped findings for a completed asynchronous batch scan.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
job_idYesJob ID returned by batch_scan_start.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond annotations: findings are 'owner-scoped' and the scan must be 'completed'. It does not discuss error behavior for incomplete scans, but this is a minor gap given the strong annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, dense sentence that front-loads the action, resource, scope, and precondition. Every word earns its place and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, scope, and the completion precondition, which is sufficient for an agent to decide to call this tool. Since there is no output schema, a bit more detail about the return shape or format semantics would improve completeness, but the low parameter complexity and strong annotations keep this from being a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes job_id, including its pattern and source, but provides no description for the optional 'format' parameter. The tool description adds no parameter-level meaning, failing to clarify the difference between 'full' and 'compact' or any default behavior. With 50% schema coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Fetch') and resource ('owner-scoped findings'), and specifies the condition 'for a completed asynchronous batch scan'. This clearly differentiates it from siblings like batch_scan_start (starting) and batch_scan_status (status).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'completed asynchronous batch scan' clearly implies this tool is used after a batch_scan_start job has finished, giving solid usage context. It does not explicitly name alternatives or state when not to use it, but the timing condition is unambiguous enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_scan_startAInspect

Start a durable asynchronous scan of 1–10 domains. Returns a stable job ID; replaying the same idempotency key with the same principal, normalized inputs, and scoring versions returns the same job.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
domainsYesDomains to scan (max 10 per request)
force_refreshNoBypass cache and run fresh scans.
idempotency_keyYesCaller-stable replay key for this exact batch request.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral context beyond the annotations: the operation is durable, asynchronous, returns a stable job ID, and has explicit idempotency-replay semantics conditional on principal, normalized inputs, and scoring versions. This complements the annotations (idempotentHint=false) by clarifying that replay is safe only with the same idempotency key, not a general guarantee.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary action and scope are in the first clause, and the idempotency/replay behavior is front-loaded in the second sentence. Every sentence adds necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description adequately covers what the tool does and the key idempotency contract, but it omits guidance on how to retrieve the scan results or monitor progress (e.g., use batch_scan_status or batch_scan_findings). Given there is no output schema, a brief mention of the job ID's role in follow-up steps would make it more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters and their constraints. The description adds no new parameter-level detail; it only references the idempotency key's behavior, which is a behavioral note rather than a semantic explanation of the parameter values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Start') with a clear resource ('durable asynchronous scan of 1–10 domains') and states a concrete outcome (stable job ID). It is semantically clear but does not explicitly differentiate itself from similar batch-start siblings like brand_audit_batch_start or discover_brand_domains_start, nor from the synchronous batch_scan tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for asynchronous batch scans via 'durable asynchronous scan', but it gives no explicit guidance on when to choose this over alternatives, when not to use it, or how it relates to sibling tools like batch_scan_status, batch_scan_findings, or scan_domain. No exclusions or routing hints are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_scan_statusA
Read-onlyIdempotent
Inspect

Read the owner-scoped status of an asynchronous batch scan.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
job_idYesJob ID returned by batch_scan_start.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, covering the safety profile. The description adds 'owner-scoped' and 'asynchronous', which are genuinely useful behavioral context, but it does not disclose how unknown job IDs are handled or what the status response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the core action, resource, and scope. There is no filler or repetition of schema information, making it immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple status-read operation and the annotations cover safety, but with no output schema the description does not explain return values or polling/completion semantics. An agent may not know what 'status' looks like or how to interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 50% of parameters with descriptions; job_id is well-documented with a pattern and a note that it comes from batch_scan_start, and format has a self-explanatory enum. The description itself adds no parameter-level meaning, but the critical job_id parameter is already adequately explained in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a clear resource ('status of an asynchronous batch scan'), and an important scoping qualifier ('owner-scoped'). This distinguishes it from sibling status tools like brand_audit_status or discover_brand_domains_status without requiring schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus the many sibling status tools, nor does it state prerequisites such as 'use after batch_scan_start'. Though 'batch scan' implies context, no alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_audit_batch_startAInspect

Enqueue an async brand audit across up to 50 target domains with optional standard/deep discovery depth, brand aliases, and caller-supplied candidate domains. Returns { auditId, queuedAt, targetCount, etaSeconds } immediately; poll with brand_audit_status and fetch results with brand_audit_get_report once complete. Each target consumes 1 unit of the monthly BRAND_AUDIT_QUOTAS budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoOutput view mode. 'registrar_complement' produces a registrar-complement payload; requires enterprise tier. Default 'standard'.
depthNoDiscovery depth. standard is default; deep expands candidate seeding and enrichment fanout.
formatNoInline output mode. Defaults to "both".
domainsYesDomains to audit (max 50 per batch). Duplicates are merged.
planner_modeNoPlanner mode for staged discovery fanout. observe emits metrics; enforce applies candidate-backed signal caps.
brand_aliasesNoOptional public brand aliases to seed, such as product or legal-entity labels.
discovery_modeNoBrand-discovery pipeline mode. classic = legacy sweep; tiered = tenant/graph/evidence wrappers first (BlackVeil-internal).
min_confidenceNoDrop candidates whose combined confidence falls below this threshold (0-1, default 0.5).
candidate_domainsNoOptional candidate domains supplied by the caller for corroboration.
ownership_verifiedNoCaller attests that the target domains are owned or authorized for scanning. Required when discovery_mode is "tiered" and the caller is not an enterprise/owner/partner principal.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnlyHint=false, destructiveHint=false, openWorldHint=true, idempotentHint=false), and the description adds genuinely useful behavior beyond them: the call is async and returns immediately, results must be polled, and each target consumes 1 unit of the monthly BRAND_AUDIT_QUOTAS budget. That quota cost is exactly the kind of context an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: what it does and its limits, what it returns, how to follow up, and what it costs. Front-loaded with the action and scope, with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter async job-starter with a full output schema, the description covers everything an agent needs: scope limits, async semantics, the follow-up tool chain, and quota impact. Return-value details are correctly left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 10 parameters are already documented in the schema, including enums for depth, discovery_mode, and view. The description names a few of them (discovery depth, brand aliases, candidate domains) and repeats the 50-domain cap, adding marginal value over the structured data. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Enqueue an async brand audit') plus scope (up to 50 target domains), which cleanly separates it from brand_audit_single and the status/report siblings. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent through the workflow: poll with brand_audit_status, fetch results with brand_audit_get_report once complete. It does not explicitly contrast with brand_audit_single, but the 'batch/up to 50 domains' framing makes the choice obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_audit_get_reportA
Read-onlyIdempotent
Inspect

Fetch the result JSON for a completed brand audit. With target set, returns the per-target CheckResult; without, returns the audit-level aggregate. Returns notReady when polling an in-flight audit. When a rendered PDF sidecar exists, metadata includes pdfUrl — an authenticated /reports/ download link (same bearer credential as this call). Completed targets whose PDF is still rendering include pdfPending so callers can poll again.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoSpecific target domain. Omit for audit-level aggregate.
auditIdYesAudit ID returned by brand_audit_batch_start.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations by disclosing the notReady in-flight return, the authenticated /reports/ download link requiring the same bearer credential as this call, and the pdfPending flag for still-rendering PDFs. This is exactly the operational behavior an agent needs to poll and authenticate correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then layers the mode distinction, polling state, and PDF metadata in a tight sequence with no filler. Slightly dense but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema already present, the description only needs to cover the behavioral edges, and it does: the notReady polling state, the credential requirement for pdfUrl, and the pdfPending retry condition. Nothing an agent needs to call and interpret this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented, and the schema itself says 'Omit for audit-level aggregate'. The description restates this distinction rather than adding new syntax or format detail, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the result JSON for a completed brand audit') and immediately scopes the two modes (per-target CheckResult vs audit-level aggregate). An agent can distinguish this from brand_audit_status and brand_audit_batch_start without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional context: set target for per-target results, omit for the aggregate, and expect notReady when polling an in-flight audit. It does not explicitly name which sibling to use instead (e.g. brand_audit_status) when the audit is not yet complete, leaving that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_audit_singleA
Read-onlyIdempotent
Inspect

Run a full brand audit on a single target with optional standard/deep discovery depth, brand aliases, and caller-supplied candidate domains. Discovers brand-related domains, looks up registrar + registrant for each candidate, and classifies each into consolidated, real registrar-sprawl shadowIt, authorized vendor dependency, indeterminate, or impersonation relationships. Gated tier-wide by monthly BRAND_AUDIT_QUOTAS (free/agent=0, developer=50, partner=200, enterprise=500, owner=unlimited).

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoOutput view mode. 'registrar_complement' produces a registrar-complement payload; requires enterprise tier. Default 'standard'.
depthNoDiscovery depth. standard is default; deep expands candidate seeding and enrichment fanout.
domainYesTarget domain to audit (e.g., apple.com).
formatNoInline output mode. Defaults to "both".
planner_modeNoPlanner mode for staged discovery fanout. observe emits metrics; enforce applies candidate-backed signal caps.
brand_aliasesNoOptional public brand aliases to seed, such as product or legal-entity labels.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.
discovery_modeNoBrand-discovery pipeline mode. classic = legacy sweep; tiered = tenant/graph/evidence wrappers first (BlackVeil-internal).
min_confidenceNoDrop candidates whose combined confidence falls below this threshold (0-1, default 0.5).
candidate_domainsNoOptional candidate domains supplied by the caller for corroboration.
ownership_verifiedNoCaller attests that the target domain is owned or authorized for scanning. Required when discovery_mode is "tiered" and the caller is not an enterprise/owner/partner principal.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/openWorld/idempotent/non-destructive, so the bar is lower; the description adds real value beyond them by disclosing monthly quota gating with exact per-tier limits and the five classification outcomes. It does not detail caching or latency behavior, but the quota and output-shape context is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose before process and quota details. Efficient, though the middle sentence's classification list is dense and slightly overlaps what the output schema likely conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich schema (100% coverage), an output schema, and full annotations, the description only needs to add the process and gating context it does. It is essentially complete for calling the tool correctly, missing only explicit alternative-routing guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly. The description references depth, brand_aliases, and candidate_domains but adds no syntax or format meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a full brand audit on a single target') and the 'single' qualifier implicitly contrasts with the brand_audit_batch_start sibling. It also enumerates the concrete work performed (domain discovery, registrar/registrant lookup, classification), so an agent knows exactly what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through 'single target' (vs. batch) and mentions depth/alias/candidate options, but never explicitly says when to choose this over brand_audit_batch_start, discover_brand_domains, or brand_audit_get_report. No when-not guidance or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brand_audit_statusA
Read-onlyIdempotent
Inspect

Poll the status of an enqueued brand audit. Returns audit-level status (queued | running | completed | failed), progress 'N/M', and per-target statuses. Owner-scoped — auditIds owned by other principals surface as notFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
auditIdYesAudit ID returned by brand_audit_batch_start.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description still adds real value by disclosing the return shape (audit-level status enum, 'N/M' progress, per-target statuses) and the owner-scoping rule that foreign auditIds surface as notFound.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with purpose, then return shape, then the ownership edge case. No filler and nothing repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, yet the description still summarizes the payload and, more importantly, flags the non-obvious notFound behavior for audits owned by other principals. Only the relationship to brand_audit_get_report and any polling expectations are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single auditId parameter with 100% schema description coverage that already states it comes from brand_audit_batch_start. The description adds no format, length, or lookup semantics beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Poll') plus resource ('an enqueued brand audit') that makes the tool's job unambiguous, and the status/report split versus brand_audit_get_report is implied. It never explicitly names a sibling to route against, so it falls short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Enqueued' implies this is used after brand_audit_batch_start returns an auditId, so the usage context is inferable. There is no explicit statement of when to poll versus when to fetch the report or start a new audit, and no guidance on polling cadence or termination.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_agent_discoveryA
Read-onlyIdempotent
Inspect

Assess the security posture of IETF BANDAID agent-discovery records (draft-mozleywilliams-dnsop-dnsaid). Detects SVCB agent records under _agents/index.{protocol}._agents, reports whether the discovery zone is DNSSEC-anchored (unsigned = spoofable agent endpoints), evaluates DANE/TLSA binding trust (RFC 6698 §10.1), and checks capability-document integrity (cap / cap-sha256). Read-only; uses Private-Use SVCB param code points pending IANA assignment.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoResolve a single named agent ({name}.{domain}) instead of enumerating the zone.
domainYesDomain to check for published agent-discovery records (e.g., example.com).
formatNoOutput verbosity. Auto-detected if omitted.
protocolNoScope discovery to a single agent protocol index (_index._{protocol}._agents). Omit to sweep the zone.
verify_capNoFetch each declared capability document (cap=) over HTTPS via safeFetch and verify it against the cap-sha256 integrity pin. Default false (declaration/existence check only).
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine extra context beyond the annotations: the unsigned-zone spoofability implication, the DNSSEC-anchoring check, and the caveat that Private-Use SVCB param code points are used pending IANA assignment.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then lists the specific checks in a single dense but well-structured block. It is information-rich rather than padded, though the prose is heavily packed and slightly long for a definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary, and annotations cover the safety profile. Combined with 100% schema coverage, the description is complete enough for correct invocation, though it could more explicitly address caching/pagination behavior of the zone sweep.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters thoroughly. The description ties cap/cap-sha256 to the capability-integrity check and protocol to the index pattern, but adds little beyond what the schema already states. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Assess/Detects/reports/evaluates/checks) combined with a precise, unusual resource (IETF BANDAID agent-discovery SVCB records). The subject matter is so specialized that it is inherently distinguishable from DNS/TLS siblings like check_svcb_https or check_srv.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case (security posture of agent-discovery records) but never states when to choose this over alternatives or when not to use it. There is no explicit routing to or from sibling tools, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_authoritative_dns_infraA
Read-onlyIdempotent
Inspect

Measure authoritative DNS infrastructure posture for a hostname over direct DNS-over-TCP/53 from a single vantage: TCP/53 reachability, the authoritative AA flag, recursion exposure, SOA serial consistency across nameservers, DNSKEY/RRSIG presence (not validation), IPv4/IPv6 answer parity, and unsupported-query handling — plus, for authenticated callers, a zone-transfer refusal test (first response only, no zone data retrieved) and CHAOS version.bind/id.server disclosure. Reports UDP/53 reachability, amplification, EDNS large-response/truncation, DNS cookies/RRL, BGP origin, RPKI, route-leak signals, anycast diversity, vantage latency, and RIR/RDAP as inconclusive — none of those are measured. Uses BV_INFRA_PROBE when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnly, idempotent, openWorld, non-destructive), and the description adds substantial context beyond them: authenticated callers get zone-transfer refusal and CHAOS version.bind tests, the ZT test is 'first response only, no zone data retrieved,' and it uses BV_INFRA_PROBE when available. This discloses auth gating and side-effect boundaries the annotations do not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the description is a single very long, em-dash-dense sentence followed by a lengthy enumeration of non-measured signals. The scope-negation list is useful for routing yet verbose; some items could be trimmed without loss. Appropriately sized in intent but structurally heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and annotations cover safety. The description still supplies full scope, auth-dependent behavior, probe fallback, and a clear boundary of what is not measured — nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with all three parameters documented, so the schema carries the parameter burden. The description adds no syntax or format detail beyond the schema (e.g., it does not explain the full/compact format semantics or when force_refresh is worthwhile), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (measure), resource (authoritative DNS infrastructure posture), scope (for a hostname, over direct DNS-over-TCP/53 from a single vantage), and enumerates the exact signals it covers. The explicit list of what it measures versus what it does not (UDP/53, amplification, BGP, RPKI, etc.) lets an agent separate it from siblings like check_ns and check_dnssec without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The scope delineation implies when the tool is appropriate (authoritative-side posture, TCP/53-based checks), and the 'inconclusive' list helps rule it out for other needs. However, it never names an alternative tool or states an explicit condition for choosing this over check_ns, check_dnssec, or check_zone_hygiene. Usage is inferred rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_bimiA
Read-onlyIdempotent
Inspect

Check the BIMI brand-logo record at default._bimi.. Validates the logo URL (l=) and the presence of mark-certificate authority evidence (a=) — the a= tag is a bare URL, so the certificate type (VMC or CMC) is not determined — and verifies the DMARC enforcement prerequisite (p=quarantine/reject) that mail clients require before displaying a BIMI logo. Returns findings for a missing/malformed record or unmet prerequisites. Use to assess brand-indicator readiness in inboxes. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive, openWorld), so credit goes to added context: the disclosure that a= is a bare URL so VMC/CMC cannot be determined, the DMARC enforcement prerequisite, and the scope of returned findings. It does not address the caching behavior implied by force_refresh, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and validation scope, then caveats, then usage. Dense but each clause earns its place; the trailing 'Part of the scan_domain audit' is the one mildly expendable element.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be re-explained, and the description already summarizes the finding scope. Annotations carry the safety profile and the description carries the BIMI-specific semantics, so nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents domain, format, and force_refresh. The l=/a=/p= references are BIMI record tags, not input parameters, so the description adds no parameter-level meaning beyond the schema. Baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check the BIMI brand-logo record at default._bimi.<domain>') and enumerates exactly what it validates (logo URL l=, authority evidence a=, DMARC p= prerequisite). An agent can distinguish this from check_dmarc or brand_audit_single without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use to assess brand-indicator readiness in inboxes. Part of the scan_domain audit.' gives clear context for when this applies and ties it to the scan_domain flow. It stops short of naming alternatives or exclusions (e.g., when to prefer brand_audit_single), so it's clear context but no explicit when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_caaA
Read-onlyIdempotent
Inspect

Look up CAA records for a domain. Shows which Certificate Authorities are authorized to issue certificates. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered. The description adds the audit-framing context but says nothing about caching behavior, rate limits, or auth that the annotations don't already imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with zero waste; the core action is front-loaded and the CAA explanation and audit tie-in follow. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema present, return values need no explanation, and annotations cover the safety profile, so the description is largely sufficient. It could be stronger only by clarifying its place relative to sibling DNS checks.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters (domain, format, force_refresh) are documented inline, including the enum options. The description adds no parameter detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Look up CAA records for a domain") and even explains what the records mean, making the tool's intent unambiguous. It does not explicitly contrast with neighbors like check_dane or check_bimi, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Part of the scan_domain audit" implies when this tool fits in the workflow, giving some usage context. However, it offers no explicit when-to-use/when-not guidance or named alternatives among the many sibling DNS checkers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_daneA
Read-onlyIdempotent
Inspect

Check DANE/TLSA certificate pinning for SMTP at port 25. Resolves the domain's MX hosts and looks up TLSA records at _25._tcp., validating their syntax, usage/selector/matching-type fields and DNSSEC backing on the MX host's zone. The record is reported as present but UNVERIFIED: there is no certificate probe for port 25/SMTP, so the pinned data is never compared against the certificate the mail server actually serves (the comparable capture-and-compare pipeline exists only for check_dane_https at port 443, and is itself currently kill-switched there — see that tool's description). Use when asked if SMTP mail servers publish DANE/TLSA pinning; this does not confirm the pin matches the live certificate. For HTTPS DANE at port 443, use check_dane_https instead. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent/openWorld annotations by disclosing a critical behavioral caveat: the record is reported as present but UNVERIFIED because there is no port 25 certificate probe, so pinned data is never compared to the served certificate. It also notes the comparable HTTPS pipeline is kill-switched. This is exactly the kind of non-obvious limitation an agent cannot infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and mechanism before the caveat, and every sentence carries weight. The parenthetical about the kill-switched HTTPS pipeline is slightly tangential for this tool but supports the UNVERIFIED claim, so it is defensible rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich annotation set and an output schema present, the description fills the remaining gap — the verification limitation — completely. An agent knows what it does, what it returns conceptually, and what it cannot confirm.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters (domain, format enum, force_refresh) are documented in the schema itself. The description adds no parameter-level syntax or defaults, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Check) and resource (DANE/TLSA certificate pinning for SMTP at port 25), and precisely scopes the mechanism: MX resolution plus TLSA lookup at _25._tcp.<mx-host>. It explicitly distinguishes itself from check_dane_https for HTTPS at port 443, so an agent can separate it from its nearest sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('Use when asked if SMTP mail servers publish DANE/TLSA pinning') and when not to ('this does not confirm the pin matches the live certificate'), and names the correct alternative for the HTTPS case. Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dane_httpsA
Read-onlyIdempotent
Inspect

Verify DANE certificate pinning for HTTPS connections. Looks up TLSA records at _443._tcp.{domain} (port 443) and validates their syntax, usage/selector/matching-type fields and DNSSEC backing. The record is reported as present but UNVERIFIED in every deployment: comparison against the certificate the host actually serves is currently withdrawn because the operator probe vantage cannot observe the origin certificate (finding metadata notAssessedReason probe_vantage_intercepted when the probe is bound), so a stale pin is NOT detected here. Distinct from check_dane which covers SMTP at port 25. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent/openWorld annotations by disclosing a decisive behavioral caveat: the record is reported UNVERIFIED in every deployment, certificate comparison is withdrawn, and a stale pin is NOT detected, with the reason (probe_vantage_intercepted). This is exactly the kind of hidden limitation an agent must know before trusting a result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the method, then the critical caveat, then the sibling distinction. Dense and mostly waste-free, though the middle clause explaining the withdrawn comparison is convoluted enough to slow a reader.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description supplies everything else an agent needs: scope, port/record pattern, the UNVERIFIED limitation, and the sibling boundary. Nothing material is missing for correct invocation or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already documented in the schema. The description only echoes the _443._tcp.{domain} lookup pattern and adds no extra semantics (e.g., caching behavior tied to force_refresh), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify DANE certificate pinning) and resource (HTTPS TLSA records at _443._tcp.{domain}), and explicitly separates itself from the sibling check_dane which handles SMTP port 25. An agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (check_dane for SMTP port 25) and situates itself in the scan_domain audit, which is strong routing context. It stops short of stating when an agent should prefer this over, say, check_ssl or check_svcb_https, so no full when/when-not matrix.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dblB
Read-onlyIdempotent
Inspect

Check domain reputation against DNS-based Domain Block Lists (Spamhaus DBL, URIBL, SURBL). Returns listing status with decoded return codes.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds only that it returns listing status with decoded return codes; it says nothing about query latency, per-provider rate limits, or caching, though the force_refresh param hints at cache behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loading the action and scope before the return behavior. Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description is nearly complete for a read-only check. The only shortfall is the lack of routing against the sibling check_rbl, which matters given how similar the two tools are.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema. The description adds no parameter-level detail beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Check domain reputation against DNS-based Domain Block Lists' and names concrete providers (Spamhaus DBL, URIBL, SURBL). However, it does not distinguish itself from the near-identical sibling check_rbl, so an agent cannot tell them apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the closely related check_rbl or check_mx_reputation alternatives. Usage is only implied by the verb 'Check', leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dkimA
Read-onlyIdempotent
Inspect

Look up DKIM records for a domain. Probes common selectors, validates the signing algorithm used for outgoing email (RSA-1024/2048, Ed25519), and reports key strength. Use to verify that outbound email signatures are cryptographically sound. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
selectorNoDKIM selector. Omit to probe common ones.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/openWorld/idempotent, so the safety profile is covered. The description adds genuine behavioral context beyond that: it probes common selectors automatically, validates the signing algorithm (RSA-1024/2048, Ed25519) and reports key strength. It doesn't mention cache behavior despite a force_refresh param, but the added context is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, purpose-first, each carrying a distinct payload (what it does, what it validates, when to use it, where it fits). No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations carry the safety profile; the description supplies purpose, behavior, and audit context. It is nearly complete for a read-only probe, only lacking failure/error expectations or DNS-resolution caveats an agent might want.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents domain, format, selector, and force_refresh. The description's mention of probing "common selectors" lightly contextualizes the selector param but adds no syntax or format meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (look up) and resource (DKIM records) scoped to a domain, and immediately differentiates itself from the email-auth siblings (check_spf, check_dmarc, check_bimi) by naming DKIM explicitly. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use to verify that outbound email signatures are cryptographically sound" gives a clear use case, and "Part of the scan_domain audit" situates it relative to a sibling. However, it names no exclusions or explicit alternative (e.g., when to prefer check_dmarc or check_bimi instead), so it stops short of full when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dmarcA
Read-onlyIdempotent
Inspect

Look up and validate the DMARC record for a domain. Shows the enforcement level (none/quarantine/reject), alignment mode (strict/relaxed), and aggregate/forensic reporting destinations. Use to determine a domain's DMARC enforcement level, whether it sends aggregate reports, or if it is protected against email impersonation — distinct from check_shadow_domains (which checks TLD variants) and assess_spoofability (composite score). Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/open-world, so the safety profile is covered. The description adds that this is part of the scan_domain audit and what data is surfaced, but does not mention caching behavior implied by force_refresh or refresh latency - a modest gap for a DNS-backed lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and outputs before the disambiguation. The closing 'Part of the scan_domain audit' is slightly incidental but still useful context, not padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be restated, and annotations cover the safety profile; what remains missing is any note on caching/refresh semantics tying force_refresh to real-world behavior. Otherwise the definition is sufficient to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema with enum values and semantics. The description adds no parameter-level detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('look up and validate the DMARC record for a domain') and enumerates the concrete outputs (enforcement level, alignment mode, reporting destinations). It explicitly names the sibling tools it is distinct from, so an agent can route correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions ('determine enforcement level, whether it sends aggregate reports, or if it is protected against email impersonation') and names the two nearest alternatives (check_shadow_domains, assess_spoofability) with the reason each differs. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dnskey_strengthA
Read-onlyIdempotent
Inspect

Audit the cryptographic strength of DNSKEY signing algorithms used for DNSSEC. Reports which algorithm is used for DNSSEC signing keys (RSA/SHA-1, RSA/SHA-256, ECDSA P-256, Ed25519, etc.), flags deprecated algorithms (RSA/SHA-1, DSA), independent of whether the DNSSEC chain validates. Use when asked what algorithm is used for DNSSEC signing keys, or if deprecated DNSKEY algorithms are in use. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: it reports algorithm details and deprecated-flag results independently of whether the DNSSEC chain validates, which tells the agent the result is not gated on chain success.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, purpose front-loaded, then scope and usage. Slight redundancy between the opening audit statement and the later 'Use when asked what algorithm is used' clause keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is unnecessary, and annotations carry the safety profile. The description supplies purpose, when-to-use, algorithm coverage, and its place in the scan_domain audit, leaving nothing an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all three parameters carry descriptions, so the schema does the heavy lifting. The description says nothing extra about domain, format, or force_refresh behavior, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: audits the cryptographic strength of DNSKEY signing algorithms for DNSSEC, and names the specific algorithms it reports and flags. This is clearly differentiated from siblings such as check_dnssec or check_dnssec_chain, which cover chain validation rather than algorithm strength.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions ('what algorithm is used for DNSSEC signing keys', 'deprecated DNSKEY algorithms in use') and notes it is part of the scan_domain audit. It lacks explicit when-not or a direct sibling comparison (e.g., vs check_dnssec), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dnssecA
Read-onlyIdempotent
Inspect

Check DNSSEC status for a domain. Verifies whether DNS is tamper-proof and protected against cache poisoning and DNS spoofing attacks by validating DNSKEY and DS records. Reports whether DNSSEC is enabled and validating. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds useful semantics about what the check validates (DNSKEY/DS) and what it reports, but says nothing about caching behavior, latency, or how results relate to the parent audit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core purpose and with no filler. The trailing audit-context clause is slightly tacked on but still earns its place as routing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers what is checked and reported. The one gap is sibling disambiguation, which matters given the many adjacent DNSSEC-related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema. The description adds no format examples, default behavior, or refresh semantics beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Check DNSSEC status for a domain') and elaborates on what is validated (DNSKEY and DS records) and what is reported (enabled/validating). However, it never distinguishes itself from the close sibling check_dnssec_chain, leaving overlap for the agent to resolve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trailing 'Part of the scan_domain audit' hints at where this fits in a larger workflow, but there is no explicit when-to-use guidance and no exclusion versus check_dnssec_chain, check_dnskey_strength, or check_nsec_walkability. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_dnssec_chainA
Read-onlyIdempotent
Inspect

Walk the full DNSSEC chain of trust from the DNS root down to the target domain, tracing DS/DNSKEY records and algorithm usage at each zone level. Use when asked to trace the chain of trust from the DNS root, or to see the full DNSSEC delegation path step by step.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: the tool performs a full recursive walk from the root, examining DS/DNSKEY records and algorithms per zone, which explains the traversal cost and structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no filler; the purpose is front-loaded and the usage condition follows immediately. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description need not explain return values, and it covers purpose and usage well. It could add one detail about behavior on domains without DNSSEC (error vs. broken chain), but the definition is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema. The description adds no syntax, format, or default details beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (walk/trace) and resource (DNSSEC chain of trust from root to target domain), and specifies the mechanism: tracing DS/DNSKEY records and algorithm usage at each zone level. This is clearly distinguishable from siblings like check_dnssec (validation status) and check_dnskey_strength (key strength).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives an explicit when-to-use: tracing the chain of trust from the root, or seeing the full DNSSEC delegation path step by step. It does not name an alternative tool (e.g., check_dnssec) or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_fast_fluxA
Read-onlyIdempotent
Inspect

Detect fast-flux DNS behavior: performs multiple rounds of A/AAAA queries and checks whether IP addresses are rotating rapidly on each DNS query (a sign of botnet or malicious infrastructure). Compares IP answer sets and TTLs across rounds to identify rapidly rotating infrastructure used to hide malicious activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
roundsNoNumber of query rounds (3-5, default 3).
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safe read-only, idempotent, non-destructive profile, and the description adds real behavioral context: it discloses the multi-round query method and that IP answer sets and TTLs are compared across rounds. It does not mention latency or cost implied by multiple query rounds, but it clearly exceeds what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the operation and its purpose, with no filler. Slightly redundant in restating 'rapidly rotating infrastructure' at the end, but efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage, an output schema handling return values, and annotations covering the safety profile, the description covers everything needed to invoke the tool correctly. Only the absence of when-to-use routing keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, rounds, and force_refresh are all documented in the schema itself. The description adds no syntax or default detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Detect) and resource (fast-flux DNS behavior) and explains the mechanism (multiple rounds of A/AAAA queries, IP rotation across rounds). This clearly distinguishes it from siblings like check_ns, check_dnssec, or check_lookalikes, which target entirely different DNS signals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case (identifying botnet/malicious infrastructure) but never states when to prefer this over alternatives such as check_shadow_domains or check_realtime_threat_feed, nor any prerequisites or exclusions. Usage is only inferable from the threat-detection framing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_http_securityA
Read-onlyIdempotent
Inspect

Audit a domain's browser-facing HTTP security headers over HTTPS. Inspects Content-Security-Policy (flagging unsafe-inline/unsafe-eval/wildcards), X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, and the cross-origin isolation headers (COOP/COEP/CORP), and detects CDN/WAF interception. Returns per-header findings for missing or weak protections against XSS, clickjacking, and cross-origin attacks. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world. The description adds meaningful behavioral context beyond them: it detects CDN/WAF interception (which affects interpretation of results) and returns per-header findings for missing or weak protections. It does not, however, note auth needs, rate limits, or caching interplay beyond force_refresh's schema note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and scope before the header list and the scan_domain tie-in. Dense but each clause carries information; no filler. Slightly list-heavy but well organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety and idempotency, the description need not explain return values, and it correctly focuses on scope, header coverage, and CDN/WAF caveats. Only the absence of sibling routing guidance and caching nuance keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents domain, format, and force_refresh. The description adds no parameter-level detail (e.g., what compact vs full omit), so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: auditing a domain's browser-facing HTTP security headers over HTTPS. Enumerates the exact headers inspected (CSP, X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP/COEP/CORP), which cleanly separates it from adjacent siblings like check_ssl, check_bimi, and check_dane_https.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing phrase 'Part of the scan_domain audit' situates it within a workflow, implying when it's relevant, but there is no explicit when-to-use/when-not or named alternative among the many sibling check_* tools. An agent must infer routing from header scope alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_llms_txtA
Read-onlyIdempotent
Inspect

Inspect a domain's published /llms.txt and /llms-full.txt for links and install instructions an AI agent could inherit from someone else. Parses and dedupes the links (same-origin vs external), sweeps external link hosts for dangling CNAMEs and deprovisioned-service fingerprints (evidence of a dangling service, not proof it can be claimed), and checks package names in npm/npx/pnpm/yarn/pip/uv/pipx install commands against the npm or PyPI registry (an unregistered name is claimable) and OSV malicious-package (MAL-) advisories. Detection only; not scored. Anything not measured is listed under notAssessed.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/openWorld annotations, the description discloses meaningful behavior: caching behavior implied by force_refresh, evidence-strength caveats ('evidence of a dangling service, not proof it can be claimed', 'an unregistered name is claimable'), and the explicit 'not scored' scope plus a notAssessed section for unmeasured items.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph that front-loads the core purpose before listing the sub-checks. Nearly every clause carries information, though the list of checks is long and could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is unnecessary, and the description still covers scope, evidence interpretation, cache/refresh behavior, and a notAssessed out-of-scope section. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all three parameters (domain, format, force_refresh) documented including the enum and the force_refresh purpose. The description adds no parameter-level meaning beyond that, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (a domain's /llms.txt and /llms-full.txt) and precisely enumerates the checks performed: link parsing/dedup, external-host CNAME sweep, and install-command registry/OSV verification. It clearly delineates a distinct capability, though it never explicitly differentiates itself from the adjacent sibling check_agent_discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated - an agent can infer this inspects agent-discovery artifacts, and 'Detection only; not scored' hints at its boundary. But there is no explicit when-to-use condition, no mention of prerequisites, and no routing to siblings such as check_agent_discovery or check_subdomain_takeover.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_lookalikesA
Read-onlyIdempotent
Inspect

Detect active typosquat and lookalike/homoglyph domains that impersonate your brand and could be used in phishing. Identifies character-substitution and visual-confusion domains registered by attackers. Distinct from check_shadow_domains (TLD variants with auth gaps) and discover_brand_domains (legitimate brand portfolio).

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered structurally. The description adds that it targets 'active' attacker-registered domains, which is useful framing, but says nothing about caching behavior, latency, or rate limits for a network-dependent scan.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose and ending on the sibling differentiation. Slight redundancy between 'lookalike/homoglyph' and 'character-substitution and visual-confusion', but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and annotations carry the safety profile. Purpose, detection target, and sibling boundaries are all present, leaving nothing an agent needs in order to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents domain, format, and force_refresh with their own descriptions. The description adds no syntax, format, or refresh semantics beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Detect) and resource (typosquat/lookalike/homoglyph domains) with the impersonation/phishing motive spelled out. It explicitly distinguishes itself from two named siblings, so an agent can route correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternatives and the distinguishing condition: check_shadow_domains for 'TLD variants with auth gaps' and discover_brand_domains for the 'legitimate brand portfolio'. This is a clear when-to-use-this-vs-that signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_mta_stsA
Read-onlyIdempotent
Inspect

Check whether a domain enforces SMTP TLS for inbound mail via MTA-STS, protecting against downgrade attacks. Queries _mta-sts. and fetches the policy file, reports mode (enforce/testing/none) and MX coverage. Use to verify whether inbound SMTP is protected against TLS downgrade or MITM — distinct from check_dane which uses TLSA pinning. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and openWorld, so the safety profile is covered. The description adds genuinely new behavioral context: it performs DNS lookups on _mta-sts.<domain>, fetches a remote policy file, and reports mode plus MX coverage — plus an implicit cache (force_refresh bypasses it).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler: the mechanism leads, the routing guidance follows, and the parent-tool relationship closes. Nothing is repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, and annotations cover the safety profile. The description still names the reported fields (mode, MX coverage) and the sibling boundary, leaving no gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three params (domain, format, force_refresh) are documented there. The description only touches caching obliquely via 'fresh check' and adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (check MTA-STS enforcement) plus the mechanism: queries _mta-sts.<domain> and fetches the policy file, reporting mode and MX coverage. It explicitly distinguishes itself from check_dane, so an agent can route correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the when ('Use to verify whether inbound SMTP is protected against TLS downgrade or MITM') and names the sibling it is not ('distinct from check_dane which uses TLSA pinning'), supplying the discriminating condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_mxA
Read-onlyIdempotent
Inspect

Look up MX records for a domain. Identifies which mail servers receive inbound email for the domain and which email hosting provider is used (Google Workspace, Microsoft 365, Proofpoint, etc.). Use when asked which email provider hosts inbound mail for a domain, or to see MX record configuration. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety and repeatability are covered. The description adds the derived provider-classification behavior and its role inside the scan_domain audit, but says nothing about caching, latency, or rate limits (the force_refresh parameter hints at caching only via the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then the meaning of the output, then the usage triggers. No filler and nothing that repeats the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return formatting need not be explained, and annotations carry the safety profile. Purpose, trigger conditions, and the audit-workflow relationship are all present, leaving nothing an agent needs before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: domain, format enum, and force_refresh are all documented in the schema itself. The description adds no parameter-level detail (no format values, no note that force_refresh matters after DNS changes), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Look up MX records for a domain') and then spells out what the result actually means: inbound mail servers plus the inferred email hosting provider. That inference (Google Workspace, Microsoft 365, Proofpoint) also implicitly separates it from the sibling check_mx_reputation, which is about reputation rather than configuration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions ('Use when asked which email provider hosts inbound mail for a domain, or to see MX record configuration') and places itself in the scan_domain audit workflow. It never states when NOT to use it or names a sibling to prefer instead, so it falls short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_mx_reputationA
Read-onlyIdempotent
Inspect

Check whether the mail server (MX) IP addresses are listed on spam blocklists (Spamhaus, Barracuda, SORBS, and other RBLs). Also verifies reverse DNS for MX hosts. Use when you want to know if your mail server IP is blacklisted, or if your MX is on any blocklist — distinct from check_rbl which checks a specific IP directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by naming specific blocklist providers and mentioning reverse DNS verification. It does not describe rate limits or result interpretation, but those are secondary given the rich annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then adds reverse DNS verification, then usage and sibling differentiation. All three sentences earn their place with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema and comprehensive annotations, so the description need not explain return values. It covers purpose, behavior, usage, and sibling routing sufficiently for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents domain, format, and force_refresh fully. The description does not add parameter syntax or constraints beyond what the schema provides. Baseline 3 is appropriate when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: checks MX IP addresses against spam blocklists and verifies reverse DNS. It explicitly distinguishes itself from check_rbl by explaining that check_rbl checks a specific IP directly. An agent can identify the tool's scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use guidance: 'Use when you want to know if your mail server IP is blacklisted, or if your MX is on any blocklist.' It also names the closest alternative, check_rbl, and states the selection condition. No further routing is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_nsA
Read-onlyIdempotent
Inspect

Audit a domain’s nameserver delegation and redundancy. Identifies the DNS hosting provider and, when the infrastructure probe is available, directly compares parent and child NS sets, verifies authoritative AA responses, and checks required glue addresses. Use to detect stale registrar delegations, lame nameservers, and intermittent resolution risk. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so the safety profile is covered. The description adds real value beyond that by disclosing conditional behavior ('when the infrastructure probe is available') and the scope of what is compared (parent vs child NS sets, AA responses, glue), which an agent could not infer from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences with the core action and scope front-loaded and no filler. The dense enumeration of checks is informative rather than redundant, though the sentence could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover safety and idempotency. The description covers what, why, and workflow placement; the only gap is a slightly vague dependency on 'the infrastructure probe' without stating the fallback behavior when unavailable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema, and the description adds no syntax or format detail. Baseline 3 applies when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Audit a domain's nameserver delegation and redundancy') and then enumerates the concrete checks performed: provider identification, parent/child NS comparison, AA response verification, and glue address checks. This is clearly distinguishable from siblings like check_authoritative_dns_infra or check_root_server_set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit use cases ('detect stale registrar delegations, lame nameservers, and intermittent resolution risk') and situates the tool within the scan_domain audit workflow. It stops short of naming when to prefer an alternative sibling or stating exclusions, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_nsec_walkabilityA
Read-onlyIdempotent
Inspect

Assess zone walkability risk by analyzing NSEC3PARAM configuration. Detects plain NSEC zones, weak NSEC3 parameters, and opt-out flags.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, open-world, non-destructive, so the safety profile is covered. The description adds substantive behavioral context by naming the concrete conditions it flags (plain NSEC zones, weak NSEC3 parameters, opt-out flags), which is more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence naming the purpose first and the detection specifics second, with zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description fully covers the tool's analytical intent. Only the absence of when-to-use routing against the many sibling checks keeps it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already documented in the schema. The description adds no parameter-level detail (e.g., cache implications of force_refresh), so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Assess zone walkability risk') plus the mechanism ('analyzing NSEC3PARAM configuration') and enumerates what it detects. This distinguishes it from generic siblings like check_dnssec or check_zone_hygiene, though it doesn't explicitly name a sibling it replaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the described detection scope (zone enumeration risk), but there is no explicit when-to-use or when-not-to-use guidance, nor any routing to alternatives such as check_dnssec or check_zone_hygiene. An agent can infer intent but gets no explicit selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_ptrA
Read-onlyIdempotent
Inspect

Verify forward-confirmed reverse DNS (PTR/FCrDNS) for mail servers. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds the FCrDNS framing but says nothing about caching, DNS resolution behavior, or failure modes beyond what annotations already provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose and followed by the workflow context. Nothing is wasted and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations carry the safety profile. The description is complete enough for a read-only single-domain check, though a brief note on what a failing FCrDNS result means would have added value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all three parameters documented including the format enum and force_refresh's cache-bypass purpose. The description adds no parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Verify') and resource ('forward-confirmed reverse DNS (PTR/FCrDNS)') with the target scope ('for mail servers'). This distinguishes it from siblings like check_mx, check_spf, and check_dnssec, which operate on different DNS records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Part of the scan_domain audit' gives implied context for when this tool belongs in a workflow, but there are no explicit when-to-use/when-not conditions or named alternatives. An agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_rblA
Read-onlyIdempotent
Inspect

Check MX server IP reputation against 6 DNS-based Real-time Blocklists (SpamCop, UCEProtect, Mailspike, Barracuda, PSBL). Resolves MX hosts to IPs first.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: the MX-to-IP resolution step performed first, and the concrete set of blocklists queried. It omits caching behavior and any rate-limit/query-volume implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with the core operation first and the resolution pre-step second; no filler. The parenthetical blocklist enumeration is defensible but slightly verbose for a description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the safety profile. The description supplies what is checked and the pre-resolution step, leaving only minor gaps such as the cache interaction implied by force_refresh.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are self-documented, so the schema does the heavy lifting. The description adds no parameter-level detail beyond implying that 'domain' is the input whose MX hosts get resolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Check) plus a precise resource (MX server IP reputation) and the mechanism (DNS-based blocklists), with a useful pre-step disclosed. It is not explicitly differentiated from the close sibling check_mx_reputation, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: knowing it evaluates MX-derived IPs against blocklists tells an agent this applies to mail deliverability/blocklist concerns. There is no explicit when-to-use, when-not-to-use, or named alternative (e.g., check_mx_reputation, check_dbl) to route against.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_realtime_threat_feedA
Read-onlyIdempotent
Inspect

Check a domain against BlackVeil real-time threat intelligence (curated intel-gateway feed). Distinct from DNSBL checks. Operator-deploy only; degrades to info when unprovisioned.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/open-world semantics, so the bar is lower. The description adds genuinely useful operational context annotations cannot: it only works when provisioned by the operator, and degrades gracefully to an informational result otherwise. It doesn't mention rate limits or cache TTL, but force_refresh in the schema implies caching.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler; the core purpose is front-loaded and the caveats follow in priority order (sibling distinction, then deployment caveat).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the key operational caveats are covered. Minor gap: what 'degrades to info' concretely returns is left implicit, though the output schema likely resolves it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already documented in the schema. The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Check) and resource (domain) against a named threat feed, and explicitly positions itself against a sibling category: 'Distinct from DNSBL checks' separates it from check_dbl/check_rbl. An agent can identify this as a threat-intel lookup rather than a blocklist lookup without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the discriminating condition versus DNSBL tools and states a deployment precondition ('Operator-deploy only'). It doesn't spell out a positive use case (e.g., 'use when you suspect active targeting'), but the exclusion plus precondition is clear guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_resolver_consistencyA
Read-onlyIdempotent
Inspect

Check DNS consistency across 4 public resolvers.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
record_typeNoRecord type. Omit for A/AAAA/MX/TXT/NS.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, and the description adds that exactly 4 public resolvers are used. However, it does not explain what 'consistency' means in terms of output or behavior (e.g., whether it reports mismatches only, or includes details).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that is front-loaded and contains no wasted words. It conveys the core function quickly and clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description should hint at what the result looks like, but it does not. The tool is simple and annotations cover safety, but the lack of return-value context leaves a gap for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%; each parameter already has a description. The tool description adds no additional parameter semantics beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check') and clearly identifies the resource ('DNS consistency across 4 public resolvers'). It distinguishes this tool from sibling DNS-checking tools like check_dnssec or check_ns by focusing on cross-resolver consistency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when DNS consistency across public resolvers is needed, but it does not explicitly state when to choose this tool over alternatives or when not to use it. No exclusions or alternative-recommendations are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_root_server_setA
Read-onlyIdempotent
Inspect

Query a rotating sample of 3 root servers per call and compare the priming NS set, glue, SOA serials, and cross-root consistency against the embedded official root hints. Uses BV_INFRA_PROBE when available; without it, returns the embedded hints as reference data only.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so safety is covered. The description adds real behavioral context beyond that: it samples exactly 3 rotating root servers per call and depends on the BV_INFRA_PROBE environment, degrading to embedded hints as reference-only data when absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler, and the core behavior is front-loaded before the environment caveat. Slightly jargon-heavy but every clause contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no exposition. The description covers what is compared, the sampling behavior, and the environment dependency, leaving little an agent needs in order to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single, fully-described enum parameter (format), the schema already carries the semantics. The description adds nothing about the format options, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Query/compare) and resource (root servers, priming NS set, glue, SOA serials, cross-root consistency) with the scope of the operation. An agent can tell this inspects the root server set rather than a general domain, though it does not explicitly distinguish itself from neighbors like check_ns or check_resolver_consistency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use vs alternatives guidance. The BV_INFRA_PROBE sentence implies a prerequisite and a degraded mode, which is useful but not framed as usage direction, so the guidance remains implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_shadow_domainsA
Read-onlyIdempotent
Inspect

Find alternate TLD variants of a domain (e.g. example.net, example.co) that have weak or missing email authentication and could be used to spoof email. Use when asked about TLD variants with email auth gaps — distinct from check_lookalikes which detects typosquat/homoglyph impersonation domains.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the detection focus (email-auth gaps on alternate TLDs), which is useful framing, but says nothing about caching, result shape, or rate behavior even though a force_refresh parameter exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero padding, with the core purpose front-loaded and the sibling disambiguation in the second sentence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations plus the 100%-covered input schema carry the rest. The description supplies exactly the remaining need: what this check looks for and how it differs from the nearest sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with domain, format (enum), and force_refresh all documented in the schema itself. The description adds no parameter-level detail (e.g., how format auto-detection works), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource with a concrete example (example.net, example.co) and the outcome it targets (TLD variants with weak/missing email auth). It explicitly names the sibling check_lookalikes and how it differs, so the agent can disambiguate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States exactly when to use it ('when asked about TLD variants with email auth gaps') and gives an explicit exclusion/alternative ('distinct from check_lookalikes which detects typosquat/homoglyph impersonation domains'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_spfA
Read-onlyIdempotent
Inspect

Look up and validate the SPF record for a domain. Lists all IP addresses and third-party senders authorised to send email on behalf of the domain, flags syntax errors, and shows the trust surface (which mail servers are whitelisted). Use when you need to know who is permitted to send email as a domain. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description restates what the tool returns (IP list, third-party senders, syntax errors, whitelisted mail servers), which largely duplicates the output schema rather than adding non-obvious behavior such as cache semantics or fallback handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, with no filler. The parenthetical '(which mail servers are whitelisted)' is mildly redundant with the preceding 'trust surface' phrasing but still clarifies terminology.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety, a 100%-covered schema, an output schema handling return values, and only one required parameter, the description is nearly complete. The one omission is sibling disambiguation against resolve_spf_chain, which matters in a tool set this dense.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description says nothing about the optional 'format' or 'force_refresh' parameters, so it adds no meaning beyond the schema, which already documents all three.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Look up and validate the SPF record for a domain') and enumerates what the result contains. It is clearly distinct from the many unrelated check_* siblings, but it does not distinguish itself from resolve_spf_chain, the closest SPF-related sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use when you need to know who is permitted to send email as a domain' gives a genuine usage trigger, and 'Part of the scan_domain audit' situates it in a workflow. However, there are no when-not conditions and no routing to the obvious alternative resolve_spf_chain, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_srvA
Read-onlyIdempotent
Inspect

Map a domain's DNS-visible service footprint by probing 19 common SRV record prefixes (email, calendar, messaging, directory, web) in parallel. Returns discovered services and flags insecure service advertisements — e.g. plaintext IMAP/POP3/LDAP without an encrypted variant. A domain with no matches among the probed prefixes is not proof the domain has no services at all. Use when asked to map DNS-visible services or flag insecure service advertisements.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, openWorld, idempotent and non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: parallel probing of 19 prefixes, the insecure-advertisement flagging logic, and a false-negative limitation. It stops short of describing rate limits or caching behavior, which is left to the schema's force_refresh parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and scope, with each sentence earning its place including the limitation caveat. The first sentence is dense with category detail but remains readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers action, scope, security-relevant behavior and a key interpretive limitation. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters carry their own descriptions, so the schema does the heavy lifting and a 3 baseline is appropriate. The description adds no parameter-specific meaning (e.g., what 'compact' omits or when the auto-detected format applies).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (map) and resource (DNS-visible SRV service footprint), specifies the exact scope (19 common SRV prefixes across five categories), and describes the flagging behavior. An agent can distinguish this from sibling checkers like check_mx or check_svcb_https without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('map DNS-visible services or flag insecure service advertisements') and adds a meaningful negative caveat that no matches is not proof of no services. It does not, however, name alternatives or state when a sibling tool would be better.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_sslA
Read-onlyIdempotent
Inspect

Check the HTTPS/TLS posture of a domain: HTTPS reachability, HSTS policy, and HTTP-to-HTTPS redirect. Also returns certificate metadata (issuer, expiry date, days remaining, SAN count) read from public Certificate Transparency logs — this describes the most recently LOGGED certificate, which may differ from the one currently served. Origin TLS protocol support and cipher suites are not assessed; legacy-TLS detection is withdrawn because the probe cannot observe the origin handshake. Use to verify HTTPS/HSTS configuration and certificate issuer/expiry. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, idempotent, non-destructive), the description discloses a critical data-provenance caveat: certificate metadata comes from public CT logs and reflects the most recently LOGGED certificate, which may differ from the one currently served. It also explains why legacy-TLS detection is absent. This is exactly the kind of behavioral context an agent needs and could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then the caveats. Every sentence carries information (CT-log limitation, excluded checks, usage). Slightly dense, with the withdrawn-legacy-TLS clause being wordy, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only domain-inspection tool with an output schema already describing returns, the description supplies what remains: scope, provenance caveats, explicit non-coverage, and a usage pointer. An agent has enough to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already fully documented in the schema, including the cache-bypass rationale. The description adds no parameter-level syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Check) plus a precisely scoped resource (HTTPS/TLS posture of a domain) and enumerates the sub-checks: HTTPS reachability, HSTS policy, HTTP-to-HTTPS redirect, and CT-log certificate metadata. This cleanly separates it from siblings like check_http_security, check_dane_https, and check_svcb_https.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use to verify HTTPS/HSTS configuration and certificate issuer/expiry' gives a clear use case, and the description draws explicit boundaries by naming what is NOT assessed (origin TLS protocol support, cipher suites, legacy TLS). It stops short of naming a sibling tool to use for those excluded areas, so it is clear context without full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_subdomailingA
Read-onlyIdempotent
Inspect

Detect SubdoMailing risk: analyzes the SPF include chain for dangling or hijackable subdomains that could let an attacker send email as the domain. Use when you want to know if an SPF include chain can be hijacked through a dangling domain, or to detect subdomain mailing risk hidden in SPF includes. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds that the check inspects the SPF include chain, but says nothing about caching or result semantics beyond what the schema's force_refresh/format params imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the analysis target, but the second sentence largely paraphrases the first ('dangling domain' / 'subdomain mailing risk hidden in SPF includes'), so a full sentence is spent restating rather than adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description plus annotations give enough to invoke the tool correctly. It would be stronger if it distinguished this check from the adjacent resolve_spf_chain tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are all documented in the schema itself. The description adds no syntax, default, or format nuance beyond that, which matches the baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (detect) and a specific analysis target (SPF include chain for dangling/hijackable subdomains). It clearly differentiates from neighboring tools like check_spf and resolve_spf_chain by naming the hijack/dangling-domain angle rather than generic SPF validation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('if an SPF include chain can be hijacked through a dangling domain') and places it in the broader scan_domain audit. It stops short of naming the sibling alternative to choose instead when the concern is plain SPF correctness, so the routing guidance is clear but not complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_subdomain_takeoverA
Read-onlyIdempotent
Inspect

Sweep subdomains for dangling CNAMEs pointing to deprovisioned cloud services that could be claimed by an attacker (subdomain takeover vulnerabilities). Detects 16 provider families (AWS S3/CloudFront, Azure Front Door/CDN/Blob/App Service, GCP Cloud Storage, Heroku, GitHub Pages, Vercel, Firebase, Shopify, etc.). Use when asked if subdomains are pointing to deprovisioned cloud services. Pair with discover_subdomains to widen the candidate set — note that returns a CT sample, not a full inventory.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com).
formatNoOutput verbosity. Auto-detected if omitted.
subdomainsNoOptional explicit subdomain list (full FQDNs or short labels). When provided (deduped, capped at 1000), this list is swept instead of the 15-name built-in. Source from Certificate-Transparency enumeration or brand-audit discovery.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds real behavioral context: 16 provider families are probed, a 15-name built-in list is used when no subdomain list is supplied, and caching exists. It stops short of describing pagination or output shape, but output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and vulnerability class, then usage and pairing guidance. The parenthetical provider list is somewhat long but earns its place by conveying coverage breadth. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only scanning tool with an output schema, the description covers purpose, trigger conditions, companion tooling, detected-provider scope, and the built-in-list fallback. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are documented in the schema itself; baseline 3 applies. The description reinforces the subdomains parameter by suggesting CT enumeration or brand-audit discovery as sources, but adds no syntax beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+outcome: 'Sweep subdomains for dangling CNAMEs pointing to deprovisioned cloud services that could be claimed by an attacker.' It names the vulnerability class explicitly and lists the provider families detected, so an agent can distinguish it from siblings like discover_subdomains or check_shadow_domains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use when asked if subdomains are pointing to deprovisioned cloud services') and names the companion tool, discover_subdomains, plus the reason to pair ('widen the candidate set'), while clarifying that sibling returns only a CT sample. Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_svcb_httpsB
Read-onlyIdempotent
Inspect

Validate HTTPS/SVCB records (RFC 9460) for modern transport capability advertisement. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds only the RFC reference and domain context, but no additional behavioral details like caching behavior (beyond the force_refresh param), authentication requirements, or rate limits. With annotations covering the safety profile, this is minimal added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose followed by contextual placement. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the presence of an output schema, and full annotation coverage, the description is nearly complete. It lacks explicit routing to alternatives, but otherwise an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (domain, format, force_refresh). The description adds no parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Validate') and resource ('HTTPS/SVCB records (RFC 9460)'), and the scope is distinct from sibling checks like check_dane_https or check_srv. The purpose is immediately clear without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says it is 'Part of the scan_domain audit', but gives no explicit guidance on when to use this tool versus alternatives such as check_dane_https or check_srv. It provides no when-to-use or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_tlsrptA
Read-onlyIdempotent
Inspect

Check whether a domain has SMTP TLS Reporting (TLS-RPT) configured. Queries _smtp._tls. for the v=TLSRPTv1 record and validates its reporting destination (rua= mailto:/https:), flagging a missing record, duplicate records, or an invalid/absent reporting URI. Complements MTA-STS by giving visibility into TLS delivery failures. Part of the scan_domain audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered; the description adds real value by naming the exact record queried and the specific failure modes it flags (missing record, duplicates, invalid/absent reporting URI). It does not mention caching, which matters given the force_refresh parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose in the first sentence, then the mechanics, then the sibling relationship. Four sentences, each carrying information, though slightly dense with DNS-specific detail an agent may not need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return formatting need not be described. The description covers what is checked, the record path, the validation rule, and the failure cases, which is everything an agent needs to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Check), resource (SMTP TLS Reporting / TLS-RPT), the exact DNS query (_smtp._tls.<domain> / v=TLSRPTv1), and the validation performed (rua= mailto:/https:). An agent can distinguish this from check_mta_sts and the other check_* siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: it complements MTA-STS and is part of the scan_domain audit. This tells the agent how the tool fits its siblings, but it never states an explicit when-not or a direct alternative choice (e.g. 'use check_mta_sts for policy, this for reporting'), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_txt_hygieneB
Read-onlyIdempotent
Inspect

Audit TXT records for stale entries and SaaS exposure.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations confirm the tool is read-only, idempotent, and non-destructive, so the description does not need to cover safety. It does not add extra behavioral context such as caching behavior (partially covered by the force_refresh parameter schema) or what the audit reports. With annotations covering the core traits and an output schema presumably describing results, a baseline 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that efficiently conveys the purpose and scope. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is concise but adequately conveys the tool's purpose for an agent to select it among TXT-related checks, especially given the rich sibling set. It could benefit from a brief note on what constitutes 'stale' or 'SaaS exposure' but is otherwise complete for a read-only audit with an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters are fully documented in the input schema. The description adds no parameter details beyond what is already present, which matches the baseline for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Audit') and resource ('TXT records') with a clear scope ('stale entries and SaaS exposure'). This distinguishes it from the many sibling check_* tools that look at other DNS record types, though it does not explicitly name an alternative for overlapping goals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to run this versus adjacent tools like check_spf or check_dmarc, which also involve TXT records. No prerequisites or contexts are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_zone_hygieneA
Read-onlyIdempotent
Inspect

Audit DNS zone hygiene: identifies sensitive or forgotten subdomains exposed in DNS, stale SOA records, and zone propagation issues. Use to find any sensitive subdomains that should not be publicly visible, or to audit overall DNS zone cleanliness.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description adds genuine behavioral content by disclosing the audit dimensions it evaluates — sensitive subdomains, stale SOA records, propagation issues — which is not derivable from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose and audit dimensions followed by the usage statement. No filler or redundancy, though the two sentences overlap somewhat in framing the same scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations carry the safety profile. The description supplies the scope and audit dimensions, leaving little an agent would need in order to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters (domain, format, force_refresh) are documented in the schema itself, including the caching behavior of force_refresh. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Audit') and resource ('DNS zone hygiene') and enumerates the concrete checks: sensitive/forgotten subdomains, stale SOA records, propagation issues. It is distinguishable from single-record siblings like check_txt_hygiene, though it never explicitly names the sibling it replaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Two clear use cases are given ('find sensitive subdomains that should not be publicly visible' and 'audit overall DNS zone cleanliness'), which gives an agent solid context for selecting this tool. It stops short of naming alternatives such as discover_subdomains or check_subdomain_takeover, and states no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_baselineA
Read-onlyIdempotent
Inspect

Compare a domain's current security configuration against a fixed policy baseline to determine compliance. Use to check whether a domain meets a policy requirement — not for tracking improvement/regression over time (use analyze_drift) and not for comparing multiple domains (use compare_domains).

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to scan and compare.
formatNoOutput verbosity. Auto-detected if omitted.
baselineYesPolicy/requirements baseline OBJECT for compliance enforcement — "does this domain meet these required controls?" (grade/score floors, require_* flags, max_*_findings). NOT a prior scan. For drift-over-time vs a previous ScanScore (or the literal "cached"), use analyze_drift instead.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds meaningful semantic context by defining 'fixed policy baseline' (compliance requirements) and excluding drift/regression analysis, which prevents misuse and clarifies the tool's scope. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action ('Compare...'), followed immediately by usage guidance with alternative tool names. No redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a nested baseline object and no output schema, but the description plus rich schema provide sufficient context for selection and invocation: it explains the compliance purpose, explicitly excludes drift and multi-domain comparison, and the schema fully documents parameters. A minor gap is not describing the output format, but this is partially mitigated by the 'format' parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, baseline (including all nested require_* flags, max_* fields, grade/score floors), format, and force_refresh are fully documented. The description text does not add parameter-level detail, but the schema carries the burden, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource+outcome: 'Compare a domain's current security configuration against a fixed policy baseline to determine compliance.' It also distinguishes from sibling tools by explicitly naming analyze_drift and compare_domains for other use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use and when-not-to-use guidance is provided: 'Use to check whether a domain meets a policy requirement — not for tracking improvement/regression over time (use analyze_drift) and not for comparing multiple domains (use compare_domains).' The schema's baseline description reinforces this by clarifying the baseline is a policy object, not a prior scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_domainsA
Read-onlyIdempotent
Inspect

Side-by-side security comparison of 2–5 domains. Shows relative scores, category gaps, and unique weaknesses for each domain. Use when comparing your security posture against a competitor, or doing a head-to-head comparison between multiple domains.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
domainsYesDomains to compare (2–5 domains)
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the description need not repeat safety traits. It adds output details (relative scores, gaps, weaknesses) but does not disclose behavioral traits like caching, rate limits, or side effects. This matches the baseline for a description that adds some value without rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with purpose, then output details, then usage guidance. Every sentence earns its place with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must partially explain return values. It does mention 'relative scores, category gaps, and unique weaknesses,' which is a useful summary, though not exhaustive. Annotations cover safety, and schema covers parameters. It is adequately complete for a comparison tool, but lacks a precise output format description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add meaning beyond the schema for domains, format, or force_refresh. The only implied semantic is the 2–5 domain limit, which is already in the schema. No compensation needed, so baseline holds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('comparison') and resource ('domains'), and specifies output: 'relative scores, category gaps, and unique weaknesses.' This distinguishes it from sibling tools like compare_baseline, which likely compares against a baseline rather than other domains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage context is provided: 'Use when comparing your security posture against a competitor, or doing a head-to-head comparison between multiple domains.' However, it does not mention alternatives or when not to use it, so it falls short of the highest score requiring explicit when-not/alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cymru_asnA
Read-onlyIdempotent
Inspect

Map domain IPs to Autonomous System Numbers via Team Cymru DNS. Returns ASN, prefix, country, registry, and organization for each IP. Flags high-risk hosting ASNs.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and open-world behavior, so the safety profile is covered. The description adds genuine context beyond that: it enumerates the returned fields and discloses the high-risk hosting ASN flagging, which is a behavioral trait not visible in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short front-loaded sentences: purpose and source, return contents, then the flagging behavior. Zero filler, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the return value needn't be explained, yet the description still summarizes outputs. Combined with rich annotations, the definition is complete for calling the tool; only the lack of routing guidance against siblings is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents domain, format, and force_refresh. The description adds no parameter-level detail (e.g., what 'compact' vs 'full' means, or when caching matters), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Map domain IPs to Autonomous System Numbers') and even names the data source (Team Cymru DNS). An agent can immediately distinguish this ASN-mapping tool from siblings like rdap_lookup or check_ptr.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the usage context (ASN/prefix/organization enrichment) but never states when to prefer this over alternatives such as rdap_lookup or check_mx_reputation, nor any exclusions. Usage is inferable but not guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_brand_audit_watchA
DestructiveIdempotent
Inspect

Permanently removes a recurring brand-audit watch by watchId. Owner-scoped — a watchId owned by another principal surfaces as notFound. Returns confirmation of deletion.

ParametersJSON Schema
NameRequiredDescriptionDefault
watchIdYesWatch ID returned by register_brand_audit_watch.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds real value beyond that: it warns the deletion is permanent (irreversible), and—more importantly—discloses the owner-scoping rule that a watchId belonging to another principal surfaces as notFound rather than an error. It stops short of covering rate limits or side effects on dependent watches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the destructive action and its keying parameter, followed by the ownership caveat. No filler and nothing repeated from the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with an output schema present, the description covers the essentials: permanence, scoping/auth behavior, and error mapping for foreign IDs. Nothing an agent needs to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is only one parameter, so the baseline would be 3. The description goes beyond that by adding ownership semantics for watchId—cross-principal IDs return notFound—which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Permanently removes'), a specific resource ('recurring brand-audit watch'), and the keying parameter (watchId). Easily distinguished from sibling register_brand_audit_watch and list_brand_audit_watches without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'removes' implies the use case, and the owner-scoping note hints at when this will fail, but there is no explicit statement of when to reach for this tool versus list_brand_audit_watches or register_brand_audit_watch. Usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_brand_domainsA
Read-onlyIdempotent
Inspect

Discover all domains that belong to a brand's portfolio by aggregating certificate, DNS, redirect, and mail-policy signals. Use when asked what domains are part of a brand portfolio, or to find all domains related to a brand. Pass the EXACT seed domain verbatim — do NOT normalize or substitute a canonical domain.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoDiscovery depth. standard is default; deep expands candidate seeding and enrichment fanout.
domainYesThe exact seed domain to expand, scanned verbatim (e.g., example.com). Do NOT normalize, resolve, or substitute a brand's canonical/main domain — pass the literal domain the user named (e.g. pass `clau.de`, not `anthropic.com`). Use `brand_aliases` for related brand labels.
formatNoOutput verbosity. Auto-detected if omitted.
signalsNoSignal modules to invoke. Defaults to all 12 discovery/enrichment signals.
planner_modeNoPlanner mode for staged discovery fanout. observe emits metrics; enforce applies candidate-backed signal caps.
brand_aliasesNoOptional public brand aliases to seed, such as product or legal-entity labels.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.
discovery_modeYesDiscovery mode. "classic" (default, BSL-licensed) runs the public signal-sweep pipeline. "tiered" layers Tier 0 (tenant-declared portfolio), Tier 1 (infrastructure-graph), and Tier 2 (declared-evidence) lookups in front of the legacy sweep, falling back to Tier 3 (the existing sweep) only on cache miss / very_stale fingerprint / uncovered caller candidates. Tiered mode requires private BlackVeil service bindings — BSL self-hosts should leave this on "classic".classic
dkim_selectorsNoOptional DKIM selectors to probe. Defaults to a built-in common-selector list.
min_confidenceNoDrop candidates whose combined confidence falls below this threshold (0-1, default 0.5).
candidate_domainsNoOptional candidate domains supplied by the caller for corroboration.
ownership_verifiedNoCaller attests that the seed domain is owned or authorized for scanning. Required when discovery_mode is "tiered" and the caller is not an enterprise/owner/partner principal. Prevents unauthorized mass reconnaissance via deep tier lookups.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered structurally. The description adds real value beyond that by disclosing the seed-handling contract ("scanned verbatim, do NOT normalize or substitute a canonical domain"), which is a behavioral trap that would otherwise silently produce wrong results. It omits cost/latency and the tiered-mode authorization constraint, which live only in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all front-loaded and free of filler, ending on the highest-risk instruction (verbatim seed). The middle sentence's two clauses ("what domains are part of a brand portfolio" / "all domains related to a brand") restate the same trigger twice and could be collapsed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover the safety profile. The description adequately conveys purpose and triggering context for a 12-parameter tool. It stops short of flagging that `discovery_mode: tiered` requires private service bindings and `ownership_verified`, which are the main ways a caller can misuse this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 12 parameters, so the schema already carries the semantics and a baseline of 3 is warranted. The description reinforces the critical `domain` constraint but merely echoes the schema's own wording, and adds nothing for the other eleven parameters (depth, signals, discovery_mode, ownership_verified, confidence thresholds).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (Discover) and resource (all domains in a brand's portfolio) and even names the aggregation mechanism — certificate, DNS, redirect, and mail-policy signals. That is well beyond a restatement of the name. It does not, however, distinguish this synchronous tool from its sibling family (discover_brand_domains_start/_status/_findings), so an agent cannot tell from the text alone which entry point to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states clear triggering conditions: "Use when asked what domains are part of a brand portfolio, or to find all domains related to a brand." There is no explicit when-not guidance and no named alternative (e.g., discover_subdomains, check_shadow_domains, or the async start/status siblings), so routing among them is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_brand_domains_findingsA
Read-onlyIdempotent
Inspect

Fetch the ranked candidate domains (the discovery CheckResult) for an async run started with discover_brand_domains_start. Returns notReady while the discovery is still in-flight; the discovery result once complete. Owner-scoped.

ParametersJSON Schema
NameRequiredDescriptionDefault
operationIdYesOperation ID returned by discover_brand_domains_start.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and non-destructive behavior. The description adds valuable behavioral context beyond annotations: the notReady interim state, eventual result delivery, and 'Owner-scoped' access restriction. This transparently sets expectations for an async polling operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no superfluous text. Every phrase earns its place: what is fetched, the source run, the polling states, and the scoping constraint. Information is front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter async polling tool, the description covers the core lifecycle (in-flight vs. complete), the data returned, and the owner scope. It lacks an explicit output structure, but given no output schema exists, the reference to 'discovery CheckResult' offers enough context for a competent agent. Sibling tools further clarify its role.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single parameter (operationId) that is already described as 'Operation ID returned by discover_brand_domains_start.' The description reinforces this by referencing the start tool, but adds no new semantic detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Fetch') and resource ('ranked candidate domains (the discovery CheckResult)') tied to an async run. It explicitly names the starting tool (discover_brand_domains_start), distinguishing this findings-retrieval tool from sibling tools like the start/status variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool: after starting an async run, and it describes the polling behavior ('Returns notReady while the discovery is still in-flight; the discovery result once complete'). It does not explicitly name alternatives or exclusions, but the lifecycle context and sibling names make the usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_brand_domains_startAInspect

Start an async brand-domain discovery for the EXACT seed domain provided (the async sibling of discover_brand_domains, which can run ~24s and time out interactive clients). Same args as discover_brand_domains. Returns { auditId, queuedAt, etaSeconds } immediately; poll with discover_brand_domains_status and fetch ranked candidates with discover_brand_domains_findings once complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoDiscovery depth. standard is default; deep expands candidate seeding and enrichment fanout.
domainYesThe exact seed domain to expand, scanned verbatim (e.g., example.com). Do NOT normalize, resolve, or substitute a brand's canonical/main domain — pass the literal domain the user named (e.g. pass `clau.de`, not `anthropic.com`). Use `brand_aliases` for related brand labels.
formatNoOutput verbosity. Auto-detected if omitted.
signalsNoSignal modules to invoke. Defaults to all 12 discovery/enrichment signals.
planner_modeNoPlanner mode for staged discovery fanout. observe emits metrics; enforce applies candidate-backed signal caps.
brand_aliasesNoOptional public brand aliases to seed, such as product or legal-entity labels.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.
discovery_modeYesDiscovery mode. "classic" (default, BSL-licensed) runs the public signal-sweep pipeline. "tiered" layers Tier 0 (tenant-declared portfolio), Tier 1 (infrastructure-graph), and Tier 2 (declared-evidence) lookups in front of the legacy sweep, falling back to Tier 3 (the existing sweep) only on cache miss / very_stale fingerprint / uncovered caller candidates. Tiered mode requires private BlackVeil service bindings — BSL self-hosts should leave this on "classic".classic
dkim_selectorsNoOptional DKIM selectors to probe. Defaults to a built-in common-selector list.
min_confidenceNoDrop candidates whose combined confidence falls below this threshold (0-1, default 0.5).
candidate_domainsNoOptional candidate domains supplied by the caller for corroboration.
ownership_verifiedNoCaller attests that the seed domain is owned or authorized for scanning. Required when discovery_mode is "tiered" and the caller is not an enterprise/owner/partner principal. Prevents unauthorized mass reconnaissance via deep tier lookups.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate non-read-only, non-destructive, open-world, and non-idempotent. The description adds valuable beyond-annotation detail: it returns { auditId, queuedAt, etaSeconds } immediately, indicates queued execution, and mandates polling. This gives a concrete async lifecycle picture, though it does not mention side effects like the actual background DNS scanning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences: the first states the core action and the sibling rationale, the second gives the exact return contract and follow-up tools. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async start tool with 12 parameters and no output schema, this description provides the entire agent-facing workflow: immediate response shape, polling via status, and retrieval via findings. It is complete enough for an agent to invoke and track the operation correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover 100% of parameters, so the baseline is 3. The description only says 'Same args as discover_brand_domains' and emphasizes EXACT seed domain, which partially echoes the schema's existing 'Do NOT normalize' note. It adds no meaningful parameter-level information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Start' and the resource 'async brand-domain discovery,' clearly distinguishing it from the synchronous discover_brand_domains. It also emphasizes the EXACT seed domain requirement, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names the async sibling discover_brand_domains and explains why to use this version (sync can run ~24s and time out interactive clients). It also prescribes the post-start workflow: poll with discover_brand_domains_status and fetch with discover_brand_domains_findings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_brand_domains_statusA
Read-onlyIdempotent
Inspect

Poll the status of an async brand-domain discovery started with discover_brand_domains_start. Returns status (queued | running | completed | failed) and progress. Owner-scoped — operationIds owned by other principals surface as notFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
operationIdYesOperation ID returned by discover_brand_domains_start.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the tool is read-only, idempotent, and non-destructive. The description adds valuable behavioral context: the returned status enum, progress reporting, and the owner-scoped behavior (operationIds of other principals surface as notFound). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the primary purpose, followed by return values and a key behavioral caveat. Every sentence carries necessary information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description appropriately discloses return content (status and progress) and the notFound behavior. It lacks a bit of detail on the progress format or next-step guidance after completion, but for a simple status polling tool, it is largely complete and well-scoped.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description enhances the operationId parameter by explaining that operationIds owned by other principals surface as notFound, which is not in the schema. This adds semantic value beyond the schema's own description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Poll') and clearly identifies the resource ('status of an async brand-domain discovery'). It distinguishes itself from sibling tools by explicitly referencing discover_brand_domains_start as the initiation point, making its role as a status poller unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly indicates that this tool is for polling status after starting an async discovery with a specific companion tool. The owner-scoping note provides useful context about access, but it does not explicitly explain when to use this tool versus other status or findings tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_subdomainsA
Read-onlyIdempotent
Inspect

Find subdomains of a domain using Certificate Transparency logs. Reveals shadow IT, forgotten services, and unauthorized certificate issuance. Returns a CT SAMPLE, not an asset inventory: the count is a lower bound, a host with no publicly-logged certificate never appears, and the result carries a per-source coverage record stating what was actually consulted. countBasis says whether totalSubdomains is the tool’s normal reach (sample) or a floor from a run whose recall was cut (then minSubdomainsObserved is present); concreteSubdomains excludes wildcard patterns.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds substantial behavior beyond those: the result is a CT sample (lower bound), unseen hosts are silently absent, a per-source `coverage` record reports what was consulted, and `countBasis` distinguishes normal reach from a recall-cut `floor`. This is exactly the kind of open-world and sampling nuance an agent needs and could not infer from annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded: purpose in the first sentence, use cases in the second, and the critical sampling caveat in the third. Every sentence earns its place given the tool has no output schema and needs to communicate subtle recall semantics. The final sentence packs a lot of field-level detail into one long clause, so it is not as crisp as a 5, but it is far from bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description is unusually complete: it covers what the result represents (a sample, not an inventory), how to interpret counts (lower bound, floor vs. normal), which fields to expect (coverage, countBasis, minSubdomainsObserved, concreteSubdomains), and the open-world limitation. The schema covers all parameters, and the annotations cover safety. Nothing an agent needs to call and interpret this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema. The description adds value by explaining output semantics tied to tool behavior (coverage, countBasis, concreteSubdomains) but does not add parameter-level meaning beyond the schema. This matches the baseline 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+method: 'Find subdomains of a domain using Certificate Transparency logs.' It goes further by stating the investigative use cases (shadow IT, forgotten services, unauthorized certificate issuance), which distinguishes it from siblings like check_subdomain_takeover (takeover risk) and discover_brand_domains (brand-related discovery). An agent can tell this tool apart from its siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (shadow IT/reconnaissance via CT) and an explicit exclusion: 'Returns a CT SAMPLE, not an asset inventory.' This tells the agent this tool is the wrong choice when a complete asset inventory is required. However, it does not name any alternative sibling tools directly, so the routing is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_findingB
Read-onlyIdempotent
Inspect

Explain a finding with impact and remediation.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
statusYesFinding severity or status.
detailsNoAdditional detail from check result.
checkTypeYesCheck type (e.g., 'SPF', 'DMARC').

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds minimal behavioral context (that output includes impact and remediation) but does not describe format, pagination, or external data access. With strong annotations, this is acceptable but not enriching.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to understanding the tool's purpose. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, strong annotations, and 100% schema coverage, the description adequately covers the essentials. It mentions the output content (impact and remediation) but does not detail return format or edge cases. No output schema exists, so a bit more detail on the response structure could be helpful, but it is not a critical gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, as every parameter (checkType, status, format, details) has a description. The tool description does not add meaning beyond the schema, but the schema fully carries the parameter semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Explain a finding with impact and remediation' clearly states the verb (explain), resource (finding), and the value delivered (impact and remediation). It is concise and distinguishes from sibling check_* tools that likely run checks rather than explain them, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description does not mention typical use cases (e.g., after a check returns a finding) or exclude cases where other tools are more appropriate. Without this, an agent may struggle to choose between explain_finding and similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateA
Read-onlyIdempotent
Inspect

Generate a DNS/email security remediation artifact. Artifact types: spf_record (build a new SPF record), dmarc_record (create a DMARC policy), dkim_config (DKIM key setup), mta_sts_policy (generate an MTA-STS policy file), fix_plan (prioritized remediation plan for all findings), or rollout_plan (phased DMARC enforcement timeline). Use when asked to generate or create a record or policy.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
policyNodmarc_record: policy (default "reject").
artifactYesWhich artifact to generate (e.g., "dmarc_record", "fix_plan").
mx_hostsNomta_sts_policy: MX hosts. Omit to detect from DNS.
providerNodkim_config: provider (e.g., "google"). Omit for generic.
timelineNorollout_plan: rollout speed (default: standard).
rua_emailNodmarc_record: report email. Default: dmarc-reports@{domain}.
force_refreshNofix_plan: bypass cache and run a fresh scan.
target_policyNorollout_plan: target DMARC policy (default: reject).
include_providersNospf_record: providers to include (e.g., ["google"]).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, which signal a safe, non-mutating operation. The description adds artifact-type context but doesn't disclose additional behavioral traits (e.g., output format, side effects). It is consistent with annotations, so no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that introduces the tool's purpose, lists artifact types, and gives a usage note. The list adds length but is necessary for clarity; no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema, the description covers the main purpose, artifact types, and usage context. The schema handles parameter details, and annotations cover safety. It doesn't explain return format, but the tool's function is straightforward enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already describes every parameter, including enums and defaults. The description repeats the artifact enum in prose but adds no meaningful extra semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates DNS/email security remediation artifacts and enumerates all six artifact types (spf_record, dmarc_record, etc.). It distinguishes itself from sibling check/analyze tools by being the generation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use when asked to generate or create a record or policy.' This provides a clear usage trigger. It does not mention alternatives or exclusions, but given the sibling tool names (all check_*/analyze_*), the context is still clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_benchmarkA
Read-onlyIdempotent
Inspect

Get industry benchmark data: shows what percentile a domain's security score ranks at within its sector or country cohort, the mean score, and the most common DNS security failures across the industry. Use when asked how a score compares to the industry average, what percentile a score is in, or what the most common security failures are in an industry or sector.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
profileNoProfile to benchmark (default "mail_enabled").

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that the tool shows benchmark comparisons by sector/country cohort, which is useful context, but it does not add much beyond the annotations and the basic output description. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the core function, and every sentence serves a purpose: the first states what the tool does, the second lists clear use cases. No unnecessary words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only 2 optional parameters, a fully documented schema, and no output schema, the description provides sufficient context for an AI to select and invoke the tool. It explains the output contents (percentile, mean score, failures) and the use cases. It lacks only explicit alternative guidance, but overall it is complete for the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of the parameters, with enums and descriptions for both 'format' and 'profile'. The description does not add any additional semantic detail beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Get industry benchmark data' and enumerates the specific data points returned (percentile, mean score, common failures). It is specific about the resource and scope, but it does not explicitly distinguish itself from sibling tools like get_domain_rank or get_provider_insights, which limits it to a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes an explicit 'Use when' section listing three concrete scenarios: comparing to industry average, asking about percentiles, or asking about common security failures. This provides clear context for when to use the tool, though it does not mention when not to use it or name specific alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_domain_rankA
Read-onlyIdempotent
Inspect

Rank a domain against its country or global cohort using the GSI benchmark corpus. Accepts a domain score (from scan_domain) and optional country/sector; returns a percentile: "scores better than X% of peers". Owner-gate exempt — public cohort data only.

ParametersJSON Schema
NameRequiredDescriptionDefault
scoreYesDomain score (0–100) from scan_domain. Used to compute the cohort percentile.
domainYesDomain to rank against its cohort (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
sectorNoSector label (e.g., "finance"). Forwarded to the cohort endpoint; sector filtering is planned for a future release.
countryNoISO 3166-1 alpha-2 country code to use the country cohort (e.g., "NZ"). Omit for global cohort.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint, idempotentHint, and destructiveHint false. The description adds useful behavioral context beyond that: 'Owner-gate exempt — public cohort data only' clarifies permission requirements and data scope. It also explicitly mentions the return format, which helps set expectations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary verb and resource, and every clause earns its place. It avoids redundancy and fluff, making it highly scannable and concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema, the description takes responsibility for explaining return values ('returns a percentile: "scores better than X% of peers"'). It covers the core inputs (score, country/sector), the dependency on scan_domain, and the permission exemption. The description is fully sufficient for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baselines are 3. The description does add marginal meaning by linking the score to scan_domain and summarizing country/sector as optional cohort selectors, but it largely restates what the schema already documents. No significant additional parameter semantics provided beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Rank a domain against its country or global cohort using the GSI benchmark corpus.' It clearly distinguishes the tool's purpose from siblings by emphasizing cohort ranking and the benchmark corpus, and it states the exact output (a percentile). This fully clarifies what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear contextual guidance: it references scan_domain as the source of the required score, and explains how country/sector alter the cohort. However, it does not explicitly name alternatives or state when not to use this tool (e.g., vs. get_benchmark). This is a minor omission, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_provider_insightsA
Read-onlyIdempotent
Inspect

Get security benchmarks and common configuration issues for a specific email or DNS service-provider cohort (e.g. Google Workspace customers, Microsoft 365 customers). Use when asked how an email service provider compares to competitors on security posture, or to see typical misconfigurations for a named vendor's customers.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput verbosity. Auto-detected if omitted.
profileNoProfile (default "mail_enabled").
providerYesProvider (e.g., "google workspace").

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds useful context by explaining the cohort-based nature (aggregated over provider customers) and that it surfaces typical misconfigurations, going beyond the generic 'read-only' annotation. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, with the first sentence concisely stating the core purpose and the second giving precise usage guidance. Every word earns its place, and the structure is front-loaded with the action and outcome.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage, and parameter context adequately. However, with no output schema, it does not describe the return format (e.g., how benchmarks are presented or what 'full' vs 'compact' affects), which would be helpful for an agent setting expectations. Still, the description is largely complete for a read-only lookup tool with strong annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all three parameters. The description adds a little contextual meaning (e.g., provider refers to an email or DNS service vendor) but does not materially improve understanding beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns security benchmarks and common configuration issues for a specific provider cohort, with concrete examples (Google Workspace, Microsoft 365). It uses a specific verb ('Get') and resource, and effectively distinguishes from sibling check_* tools by focusing on provider-level insights rather than domain-level checks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'Use when asked how an email service provider compares to competitors on security posture, or to see typical misconfigurations for a named vendor's customers.' It does not mention when-not-to-use or alternatives, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_brand_audit_watchesA
Read-onlyIdempotent
Inspect

Returns the caller's recurring brand-audit watches: watchId, domain, interval, webhook presence, last-run time, and active state. Owner-scoped. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, idempotentHint=true, openWorldHint=true), so 'Read-only' is largely a restatement. The genuinely additive detail is 'Owner-scoped', which clarifies the authorization/visibility boundary the annotations do not express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, with the core purpose and the returned field list front-loaded ahead of the scope and safety qualifiers. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only listing tool with rich annotations and an output schema, the description covers purpose, scope, and safety adequately. Only edge behavior (what an empty watch list looks like, any cap on the number returned) is unaddressed, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there are no parameter semantics to document. The description's field enumeration is aimed at return values rather than inputs, and an output schema already exists to carry that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the caller's recurring brand-audit watches') and even previews the returned fields, so an agent knows exactly what this produces. It never names the sibling pair register_brand_audit_watch / delete_brand_audit_watch to draw the list-vs-manage boundary explicitly, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Owner-scoped' gives an implied usage constraint (only the caller's own watches are returned), which is useful routing context. However, there is no explicit when-to-use statement, no exclusion for the register/delete siblings, and nothing about when this is preferable to brand_audit_status or brand_audit_get_report.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_complianceA
Read-onlyIdempotent
Inspect

Map scan findings to compliance frameworks: NIST 800-177, PCI DSS 4.0, SOC 2, CIS Controls. Shows pass/fail/partial status per control.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds useful context about the compliance frameworks and pass/fail/partial statuses, but does not disclose deeper behaviors such as caching or external calls, which is acceptable given the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the main purpose and listing the frameworks and output format. No redundant words or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey return values. It does mention pass/fail/partial status per control, but lacks clarity on prerequisites (e.g., whether findings must already exist), result grouping, or relationship to other scan tools. This is adequate but leaves gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all three parameters (domain, format, force_refresh), including a description for each. The tool description adds no parameter-specific information beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool maps scan findings to specific compliance frameworks (NIST 800-177, PCI DSS 4.0, SOC 2, CIS Controls) and shows status per control. This distinguishes it from sibling tools like map_supply_chain and assess_coverage, using a specific verb and resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you have scan findings and need compliance mapping, giving clear context. However, it does not explicitly mention when not to use it or name alternative tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_supply_chainA
Read-onlyIdempotent
Inspect

Map DNS-visible third-party service dependencies for a domain. Correlates SPF, NS, TXT verifications, SRV services, and CAA records to reveal which third-party vendors can send email as the domain, control DNS, or access integrated services. Use when asked to map third-party or supply-chain dependencies — not for listing who can send email (use check_spf for that).

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal read-only, idempotent, and non-destructive behavior. The description adds valuable context by detailing the correlation methodology (SPF, NS, TXT, SRV, CAA) and the insights it produces (vendor email sending, DNS control, integrated services). It goes beyond annotations without contradicting them, earning a solid score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences: main purpose, methodology, and usage guidance. It is front-loaded with the core function, and every sentence contributes meaning without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mapping tool with no output schema, the description explains what it does and how, and gives usage boundaries. It does not describe return value structure, but the input schema is rich and sibling differentiation is clear. This is sufficiently complete for an AI to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage for all three parameters (domain, format, force_refresh) with descriptions. The tool description does not add parameter-specific semantics, but the schema already carries the load. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Map DNS-visible third-party service dependencies for a domain.' It clearly differentiates from siblings by explicitly stating it is 'not for listing who can send email (use check_spf for that).' This makes the tool's purpose unambiguous and distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance: 'Use when asked to map third-party or supply-chain dependencies — not for listing who can send email (use check_spf for that).' This gives both positive and negative usage context and names an alternative tool, satisfying the highest bar for usage guidelines.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigate_domain_startAInspect

Start an async OSINT investigation for a domain. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns an investigationId immediately — poll with osint_investigation_status and retrieve results with osint_investigation_report.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, open-world, non-idempotent, non-destructive, but the description adds meaningful behavior beyond them: it is async, returns an investigationId immediately, degrades to info when unprovisioned, and requires operator authorization. The auth and degradation details are exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the action front-loaded, then preconditions, then the workflow handoff. No filler, and the most important routing information (start/poll/report) is stated compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async start tool with no output schema, the description covers the essentials: async semantics, immediate return value (investigationId), the poll/report workflow, and authorization prerequisites. The only real gap is parameter format guidance for query.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter (query) at 0% schema coverage, so the description must compensate. Saying the investigation is 'for a domain' implies the query is a domain string, but the description never states the expected format, validation, or how the query relates to the maxLength 253 constraint. Baseline for the un-compensated gap is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Start) plus the resource (async OSINT investigation for a domain), and the 'for a domain' qualifier distinguishes it from the sibling osint_investigate_username_start/email_start/infrastructure_start tools without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real selection context (operator-deploy only, targets must be operator-authorized on the recon watchlist) and routes the agent to the follow-up siblings osint_investigation_status and osint_investigation_report. It does not explicitly contrast against the other osint_investigate_*_start variants, so it stops just short of a full when-to-use map.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigate_email_startAInspect

Start an async OSINT investigation for an email address (breach exposure, account correlation). Owner/enterprise tier only — people-centric OSINT is restricted to prevent misuse. Returns an investigationId immediately — poll with osint_investigation_status and retrieve results with osint_investigation_report.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=false (indicating mutating/start) and idempotentHint=false, but description clarifies it starts an async process and returns an investigationId immediately, which is beyond annotations. It also discloses the restricted access (owner/enterprise tier) and the restriction on people-centric OSINT to prevent misuse. This adds valuable behavioral context not present in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no redundancy. Front-loads the action and purpose, then adds restrictions and next steps. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Contextual complexity is moderate: async operation, restriction, and follow-up tools are mentioned. No output schema, but description explicitly tells how to get results via other tools, so return value is covered. Minor gap: doesn't specify what happens if query is invalid or requires additional auth beyond tier, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has only one parameter 'query' with min/max length but no description; coverage is 0%, so description must compensate. Description implies 'query' is the email address, but doesn't explicitly state format or validation. However, given only one param, it's fairly inferable. Baseline for 0 params would be 4, but single param with implied type is okay; still, explicit mention would improve.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states specific verb and resource: 'Start an async OSINT investigation for an email address' with explicit scope ('breach exposure, account correlation'). It distinguishes from siblings by noting the email-specific variant and the async nature, setting it apart from domain, infrastructure, supply chain, and username variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context: mentions it is restricted to owner/enterprise tier, which is a usage condition. It also directs to poll with osint_investigation_status and retrieve with osint_investigation_report, providing next steps. However, it doesn't explicitly state when NOT to use this tool (e.g., for non-email queries) or alternatives beyond the sibling tools, though the sibling context implies alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigate_infrastructure_startAInspect

Start an async deep-infrastructure OSINT investigation for a query (domain, IP, or org). Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns an investigationId immediately — poll with osint_investigation_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (non-read-only, open-world, non-idempotent, non-destructive), and the description adds real behavioral context beyond them: it is async, returns an investigationId immediately, requires operator deploy plus watchlist authorization, and degrades to 'info' when unprovisioned. It stops short of saying what happens on a duplicate start or what the async job does internally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, each earning its place: what it does, what gates it, and what it returns. The 'degrades to info when unprovisioned' clause is terse but load-bearing. Slight density risk, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description tells the agent the immediate return (investigationId) and the polling tool, plus deploy and authorization prerequisites. Complete enough to invoke correctly, though it omits failure modes for unauthorized or unprovisioned requests.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'query' parameter has no description in the schema, so the description must carry the burden — and it does by naming the accepted input types (domain, IP, org). It does not add format constraints beyond the schema's length bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Start an async deep-infrastructure OSINT investigation'), the resource type, and the accepted query forms (domain, IP, org). This cleanly separates it from siblings like osint_investigate_username_start and osint_investigate_domain_start, which target narrower input types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the follow-up path ('poll with osint_investigation_status') and states the deployment prerequisite and watchlist authorization requirement. It does not, however, explain when to prefer this broad infrastructure investigation over narrower siblings such as osint_investigate_domain_start or map_supply_chain.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigate_supply_chain_startAInspect

Start an async supply-chain OSINT investigation for a query. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns an investigationId immediately — poll with osint_investigation_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnly=false, openWorld=true, idempotent=false, destructive=false. The description adds real context beyond these: it is async and returns an investigationId immediately, it requires operator-deploy and an authorized watchlist target, and it 'degrades to info when unprovisioned'. It doesn't disclose rate limits or what the investigation actually consumes, but the precondition/degradation disclosure is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with purpose, then constraints, then return/poll behavior. No filler and every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the immediate return (investigationId) and the next step. Combined with the operator/watchlist preconditions it is nearly complete, only lacking guidance on the query payload itself.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'query' parameter, so the description carries the burden. It only says 'for a query', never clarifying what the query should contain (domain, company name, format) despite the schema imposing a 1-253 char limit. Minimal compensation for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Start an async supply-chain OSINT investigation') and scopes it to 'a query'. It clearly positions itself in the async-start family and distinguishes the follow-up via the named sibling osint_investigation_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states prerequisites ('Operator-deploy only', targets 'must be operator-authorized on the recon watchlist') and the follow-up action ('poll with osint_investigation_status'). It stops short of explicitly contrasting with map_supply_chain or the other osint_investigate_*_start variants, so no when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigate_username_startAInspect

Start an async OSINT investigation for a username (cross-platform presence, breach correlation). Owner/enterprise tier only — people-centric OSINT is restricted to prevent misuse. Returns an investigationId immediately — poll with osint_investigation_status and retrieve results with osint_investigation_report.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate openWorldHint=true and idempotentHint=false, but the description adds crucial behavioral details: the tool is asynchronous and returns an investigationId immediately while the investigation runs in the background. It also discloses access restrictions (owner/enterprise tier). This goes beyond the annotations, though it does not describe edge cases like authorization failures or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each with a clear role: purpose, restriction, and workflow. It is front-loaded with the main action, avoids filler, and every sentence contributes essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple starter tool with one parameter and no output schema, the description adequately covers the full workflow: what it does, the access restriction, and the follow-up steps (poll and retrieve). It does not detail report contents, but that is handled by the report tool. Minor missing details like rate limits or concurrent investigations, but not critical for this async starter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter (query) and 0% schema description coverage, the description must compensate. It clarifies that the query is a username for investigation, which adds meaning beyond the raw schema. However, it does not specify format details (e.g., case sensitivity, whether '@' is needed) or confirm it accepts only one username. It suffices but is not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Start' and the resource 'async OSINT investigation for a username', and it specifies the scope (cross-platform presence, breach correlation). This distinguishes it from sibling starters like osint_investigate_email_start or osint_investigate_domain_start by explicitly focusing on username investigations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage context by stating that this is owner/enterprise tier only and that people-centric OSINT is restricted to prevent misuse. It also explains the async workflow: returns an investigationId immediately, poll with status and retrieve with report. However, it does not explicitly compare to alternative investigation starters (e.g., email, domain), though this is implied by the username focus.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigation_reportB
Read-onlyIdempotent
Inspect

Retrieve the final report of a completed OSINT investigation by investigationId. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned or not yet complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
investigationIdYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: a deployment restriction, an authorization precondition on the target, and a documented fallback ('degrades to info') when the investigation is not ready.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler; the retrieval action and its key are front-loaded, and the constraints follow. Nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a report-fetching tool with no output schema, the description covers invocation prerequisites and the degraded case but says nothing about what the report contains or how large/structured it is. It is adequate, not complete, for a tool whose return value is undefined elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden for the single parameter, and it does no more than restate the field name ('by investigationId'). It does not say where the ID originates (e.g. from an osint_investigate_*_start call) or what form it takes, leaving the parameter under-explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: retrieve the final report of a completed OSINT investigation, keyed by investigationId. It is clearly distinguishable from the sibling osint_investigation_status by the word 'final report', though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real preconditions (operator-deploy only, target must be operator-authorized on the recon watchlist) and notes the degraded outcome when unprovisioned or incomplete. However, it never tells the agent to check osint_investigation_status first or otherwise routes between the two investigation tools, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

osint_investigation_statusA
Read-onlyIdempotent
Inspect

Poll the status of an OSINT investigation by investigationId. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns current status (running | completed | failed) and progress metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
investigationIdYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld). The description adds genuinely non-derivable behavior: operator-deploy only, targets must be operator-authorized on the recon watchlist, and graceful degradation to 'info' when unprovisioned. A hidden constraint worth surfacing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, zero filler, front-loaded action first and constraints/return values after. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the return shape (status enum running|completed|failed plus progress metadata) and states provisioning/auth constraints. The main remaining gap is that it doesn't explain how to obtain investigationId, which is the one thing an agent needs before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single investigationId parameter, so the description had room to compensate. It only restates the parameter name ('by investigationId') without saying where the id originates (e.g., returned by an osint_investigate_*_start call) or its accepted format, leaving semantics underexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Poll the status of an OSINT investigation by investigationId.' That clearly separates it from osint_investigate_*_start (which launch) and osint_investigation_report (which retrieves results), even though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Poll the status' implies iterative checking after a start call, and the operator-deploy/ watchlist-authorization prerequisites set context. However, it never states when to prefer this over osint_investigation_report or what to do when status is failed, so usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prioritize_portfolio_leadsA
Read-onlyIdempotent
Inspect

Rank a brand’s portfolio (or an explicit domain set) into prioritized registrar-partner sales leads by product-gap value × severity. Multi-domain, paid. Reuses map_registrar_products per domain, then ranks. Distinct from map_registrar_products (per-domain product mapping) and batch_scan (raw scores).

ParametersJSON Schema
NameRequiredDescriptionDefault
brandNoBrand seed apex; discovers the portfolio, derives ownership buckets, then ranks the top candidates.
formatNoOutput verbosity. Auto-detected if omitted.
domainsNoExplicit domain set to rank (max 10). Ownership bucket = "unknown".
force_refreshNoBypass cache and run fresh scans.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the agent knows this is a safe read operation. The description adds context beyond annotations: it is paid, multi-domain, reuses another tool, computes value × severity, and ranks. It also implies caching behavior via force_refresh parameter, which is useful. It could mention output shape or failure modes, but with annotation coverage this is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The core purpose is front-loaded, the key differentiators are explicit, and the sibling references are compact. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only ranking tool with 100% schema coverage and no output schema, the description covers purpose, alternatives, and parameter semantics sufficiently. The main gap is not describing the output shape or the 'prioritized leads' format, but given the complexity and existing structured fields, the definition is largely complete. It could also clarify what 'value × severity' means in output terms.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already thoroughly documents all four parameters. The description adds light context: 'Multi-domain' aligns with domains array, and 'brand ... discovers the portfolio, derives ownership buckets, then ranks' adds meaning to the brand parameter. But it doesn't add much beyond the schema; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Rank'), a resource ('a brand's portfolio or an explicit domain set'), and the output ('prioritized registrar-partner sales leads by product-gap value × severity'). It also explicitly distinguishes itself from map_registrar_products and batch_scan, making it easy for an agent to select among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description openly states its context: multi-domain, paid, reuses map_registrar_products per domain, then ranks. It names two distinct alternatives and the differentiator ('Distinct from map_registrar_products... and batch_scan'). This gives an agent clear guidance on when to pick this tool over neighbors.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rdap_lookupA
Read-onlyIdempotent
Inspect

Fetch domain registration data via RDAP (modern WHOIS replacement). Returns the domain registrar (the company the domain was registered with), registrant contact, creation/expiration dates, EPP status codes, and domain age. Use when asked who registered the domain, who the registrar is, or when the registration expires — distinct from check_ns which identifies the DNS nameserver provider.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so the safety profile is fully covered. The description adds only 'modern WHOIS replacement' framing and a list of return fields, which the output schema already conveys; it says nothing about rate limits, latency, or partial-data cases. With annotations and output schema doing the heavy lifting, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, followed by return contents and then usage routing, in three tight sentences. The enumeration of return fields is slightly redundant with the existing output schema, costing a little efficiency but not readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema, rich annotations, and 100% parameter coverage, the description only needs to establish purpose and routing, which it does fully. It is complete for an agent's needs, though it could note caching behavior rather than leaving that entirely to the force_refresh param.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so domain, format, and force_refresh are already documented with examples and enum values. The description adds no parameter-level meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch domain registration data via RDAP') and enumerates exactly what is returned: registrar, registrant contact, dates, EPP status codes, and domain age. It explicitly distinguishes itself from the sibling check_ns, so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use triggers ('who registered the domain, who the registrar is, or when the registration expires') and names the alternative check_ns along with its distinct purpose (DNS nameserver provider). Nothing about routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_brand_audit_watchAInspect

Creates a recurring brand-audit watch for a domain on a daily/weekly/monthly cadence. Each run enqueues a fresh brand_audit_batch_start and (when a webhook is configured) POSTs a diff webhook on classification drift. Returns the new watchId. Owner-scoped; per-principal cap of 20 active watches.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to watch.
intervalYesRecurrence interval.
webhook_urlNoOptional webhook URL — POSTed on classification drift. Re-validated for SSRF at both register and delivery time.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
passedYes
partialNo
categoryYes
findingsYes
checkStatusNo
verdictWithheldNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the recurring run behavior, the enqueued brand_audit_batch_start per run, conditional webhook delivery on classification drift, the returned watchId, owner scoping, and a per-principal cap of 20 active watches. These are exactly the operational facts an agent needs before invoking a non-idempotent write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with what is created and the cadence, followed by run-time side effects and limits. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present (so return values need not be explained), full schema coverage, and annotations covering the safety profile, the description still contributes scope, side effects, and quota limits. Nothing needed to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with per-parameter descriptions, so the schema already defines domain, interval, and webhook_url semantics. The prose echoes cadence and webhook behavior but adds no new syntax or format detail beyond what the schema supplies, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Creates), resource (recurring brand-audit watch), and scope (per domain, on daily/weekly/monthly cadence). An agent can distinguish it from siblings like brand_audit_single or list_brand_audit_watches without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the effect of use (enqueues a batch start each run, POSTs a drift webhook) and the constraints (owner-scoped, 20 active watch cap), which implies when this is appropriate. However, it never explicitly contrasts with alternatives such as brand_audit_single or brand_audit_batch_start, so selection guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_spf_chainA
Read-onlyIdempotent
Inspect

Trace the full SPF include chain for a domain. Recursively resolves all includes, shows lookup count, tree depth, and flags circular includes or exceeding the 10-lookup limit.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds valuable behavioral details beyond annotations: recursive resolution, output metrics (lookup count, tree depth), and detection of circular includes and the 10-lookup limit. This is more context than typical and aligns with the read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long and front-loads the primary action. Every sentence adds substantive information (what it does, what it shows, and what it flags), with no waste or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description explains the key output aspects (lookup count, tree depth, flags) and edge cases (circular includes, limit). It covers the core behavior adequately. It could mention output format or caching behavior (given force_refresh exists), but overall it gives enough context for an agent to understand the tool's capabilities.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (all three parameters have descriptions). The tool description provides no additional parameter-specific meaning beyond the schema. It mentions output characteristics but not parameter details, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Trace the full SPF include chain for a domain.' It specifies the exact resource (SPF include chain) and the specific actions (recursively resolves includes, shows lookup count, tree depth, flags circular includes/limit). This distinguishes it from siblings like check_spf or check_dnssec_chain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: when you need to trace SPF includes and check for circular includes or lookup limits. However, it does not explicitly mention when not to use it or contrast with alternative tools like check_spf. The context is clear but lacks explicit exclusions or alternative references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_buckets_findingsA
Read-onlyIdempotent
Inspect

Retrieve findings from a completed cloud-bucket discovery scan by scanId. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. The scanId is required so reads can be owner-scoped; target and provider filters are optional.

ParametersJSON Schema
NameRequiredDescriptionDefault
scanIdYes
targetNo
providersNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the description's added value is the operator-deploy restriction, watchlist authorization, owner-scoped reads, and the degrade-to-info-when-unprovisioned behavior. These are real behavioral disclosures beyond the annotations, though the degraded output format itself is undocumented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action and the scanId rationale front-loaded, and prerequisites appended. Density is good; the compressed clause stack ('Operator-deploy only; targets must be operator-authorized...; degrades to info when unprovisioned') borders on jargony but earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should give some sense of what findings are returned and in what shape; it only says 'Retrieve findings'. Combined with zero param coverage and no output schema, there is a real gap, though annotations carry the safety profile and the scan-family context is intact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description carries the load. It explains why scanId is required (owner-scoping) and marks target/providers as optional filters, which is useful, but adds no format, matching, or value-semantics guidance for any of the three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (findings from a completed cloud-bucket discovery scan), keyed by scanId. This clearly distinguishes it from scan_buckets_start and scan_buckets_status within the same family, so an agent can place it in the scan lifecycle without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Conveys the key precondition that the scan must be 'completed' and that this is operator-deploy only with operator-authorized targets, which is strong context. It stops short of naming sibling alternatives (e.g. scan_buckets_status) or stating when-not to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_buckets_startAInspect

Start an async cloud-bucket discovery scan for a target domain. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns a scanId immediately — poll progress with scan_buckets_status and retrieve results with scan_buckets_findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
providersNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=false), and the description adds real behavioral context: async execution, immediate scanId return, authorization requirement, and info-level degradation when unprovisioned. It does not describe rerun/duplicate-submission semantics, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and followed by constraints and the polling/fetch routing. No filler, though the semicolon clause is dense enough to be slightly harder to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The lifecycle is well covered (start, poll, fetch) and the scanId return is stated despite no output schema, but the undocumented 'providers' parameter and the absence of any return-shape detail leave a gap for a tool with 0% schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with 2 parameters and no enum hints, so the schema does nothing. The description only implies the 'target' parameter via 'target domain' and never mentions the 'providers' parameter at all, leaving it undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope ('Start an async cloud-bucket discovery scan for a target domain'), and the async-start framing distinguishes it cleanly from the scan_buckets_status and scan_buckets_findings siblings it names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States preconditions explicitly (operator-deploy only, targets must be operator-authorized on the recon watchlist), the fallback behavior when unprovisioned, and routes the agent to status/findings for the follow-up steps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_buckets_statusA
Read-onlyIdempotent
Inspect

Poll the status of a cloud-bucket discovery scan by scanId. Operator-deploy only; targets must be operator-authorized on the recon watchlist; degrades to info when unprovisioned. Returns scan status (running | completed | failed) and progress metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
scanIdYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds value the annotations cannot: the operator-deploy restriction, the watchlist authorization requirement, and the unprovisioned degradation path. Return states are also disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, zero filler: the action, the authorization constraints, and the return shape are each front-loaded in order of importance. Nothing is repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by enumerating the return states (running | completed | failed) plus progress metadata. Combined with the deployment and authorization caveats, an agent has enough to call and interpret the tool, though polling cadence and error behavior are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single scanId parameter, so the description carries the burden, and it only says the scan is identified 'by scanId' with no format, provenance, or length guidance. With one param and no enum, this is minimally adequate rather than rich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Poll') and resource ('status of a cloud-bucket discovery scan') keyed by scanId, which cleanly separates it from scan_buckets_start (launch) and scan_buckets_findings (results). However, it never names those siblings explicitly, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real prerequisites ('Operator-deploy only', 'targets must be operator-authorized on the recon watchlist') and a fallback behavior ('degrades to info when unprovisioned'), which is better than nothing. But it never says when to poll vs. call scan_buckets_findings, or how often to poll, so the operational guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_domainA
Read-onlyIdempotent
Inspect

Run a full DNS and email security audit for a single domain. Aggregates every scan-included check in parallel (SPF, DKIM, DMARC, DNSSEC, TLS/SSL, MTA-STS, CAA, BIMI, subdomain takeover, and more) and returns an overall security score, NIST-aligned letter grade (6-band A+/A/B/C/D/F), maturity stage, and prioritized findings. Use for a comprehensive single-domain audit, to get a domain's overall security grade, or to assess email security maturity. Version stamps: 'scoringModelVersion' is the scoring POLICY semver (changes only when weights/thresholds/severities change, so it advances slowly) and is INDEPENDENT of — never comparable to — 'dnsChecksPackageVersion', the @blackveil/dns-checks npm engine-package version, which moves every release; a lower model version is expected, not a version gap. When citing a score, record 'scoringConfigHash' — it identifies the exact scoring configuration that produced the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
profileNoScoring profile. Default "auto" detects.
force_refreshNoBypass cache and run a fresh scan. Useful after DNS changes.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description adds value by detailing the aggregation of checks and the output (score, grade, maturity, findings). It also discloses the critical nuance about version stamps (scoringModelVersion vs dnsChecksPackageVersion being independent and not comparable) and advises recording scoringConfigHash. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured, starting with a clear purpose sentence and then providing usage context and version stamp details. Every sentence adds value, though the version stamp explanation could be condensed. It is efficient for a complex tool with many checks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (aggregates many checks) and lack of an output schema, the description is quite complete: it lists the main checks, explains the output (overall score, NIST grade, maturity, findings), and gives explicit usage guidance plus version stamp handling. It covers the essential context for using the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema describes all parameters (domain, format, profile, force_refresh). The description does not add significant parameter-specific semantics beyond what the schema provides—it focuses on overall tool behavior. The version stamp explanation is not about parameters. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Run a full DNS and email security audit for a single domain' with a specific verb and resource. It distinguishes itself from sibling check_* tools (e.g., check_spf) by emphasizing 'full audit' and aggregation, and from batch_scan by explicitly saying 'single domain'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit use cases: 'Use for a comprehensive single-domain audit, to get a domain's overall security grade, or to assess email security maturity.' This gives clear context on when to invoke it, though it does not explicitly state when not to use it or name alternative tools (like individual checks).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sge_quickscanA
Read-onlyIdempotent
Inspect

Answer, for ONE domain, whether it meets the New Zealand Secure Government Email (SGE) requirements agencies must satisfy by October 2026. Reports all seven SGE controls — DMARC p=reject, SPF -all, DKIM, SMTP transport TLS, MTA-STS enforce, TLS-RPT, full sub-domain coverage — each as satisfied, not satisfied, or NOT MEASURED, with the structured evidence behind every verdict. Neither SMTP transport TLS nor sub-domain coverage can be observed from a single domain scan, so a DNS-only result tops out at INDETERMINATE, which is not a pass. Distinct from map_compliance, which maps findings to NIST/PCI/SOC 2/CIS.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral value beyond annotations: the tool cannot observe SMTP transport TLS or sub-domain coverage from a single scan, DNS-only results cap at INDETERMINATE (not a pass), and verdicts include a NOT MEASURED state with structured evidence. It stops short of disclosing cache behavior or rate limits, but the added constraints are materially useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, output structure, key limitation, and sibling distinction. It front-loads the core purpose and scoping before the limitation and differentiation. Slightly long, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description explains the return semantics well: the seven controls, the three verdict states (satisfied, not satisfied, NOT MEASURED), structured evidence, and the INDETERMINATE ceiling. Combined with fully documented parameters and safety annotations, little an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (domain, format, force_refresh) are already documented in the schema. The description reinforces the domain parameter's scope with 'ONE domain' but adds little beyond that; the baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Answer, for ONE domain, whether it meets the New Zealand Secure Government Email (SGE) requirements' with a deadline (October 2026) and lists the seven controls evaluated. It explicitly names map_compliance as the sibling it is not, so an agent can distinguish it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: it is scoped to ONE domain, and it explicitly names the alternative map_compliance for NIST/PCI/SOC 2/CIS mapping. It also tells the agent when the tool is insufficient — a single-domain DNS-only scan cannot observe SMTP transport TLS or sub-domain coverage and tops out at INDETERMINATE. It does not spell out when to prefer batch_scan or scan_domain, leaving that partially implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_attack_pathsA
Read-onlyIdempotent
Inspect

Analyze current DNS posture and enumerate specific attack paths an adversary could exploit, with severity, feasibility, steps, and mitigations.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to check (e.g., example.com)
formatNoOutput verbosity. Auto-detected if omitted.
force_refreshNoBypass cache and run a fresh check. Useful after DNS changes.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds value by explaining that the tool evaluates attack paths with severity, feasibility, steps, and mitigations, which provides insight into the nature of the analysis. However, it does not disclose additional behavioral traits such as caching, performance characteristics, or whether it relies on external data sources beyond what annotations and schema imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that is front-loaded with the primary action ('Analyze current DNS posture') and enumerates key output aspects in a compact list. Every word earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (simulating attack paths), the description provides a solid overview of purpose and expected output, and the annotations cover safety and idempotency. The schema documents all parameters. It lacks explicit mention of caching or relationship to other tools, but for the scope of this tool, the description is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% description coverage for all 3 parameters, so the baseline is 3. The description does not add parameter-specific information beyond what the schema already provides, but it does mention 'DNS posture' which loosely relates to the 'domain' parameter. No additional clarity is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Analyze' and 'enumerate') and a specific resource ('current DNS posture'), and it distinguishes itself from sibling check_* tools by focusing on attack path enumeration with concrete outputs (severity, feasibility, steps, mitigations). It unambiguously describes what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (use when you want to analyze attack paths from DNS posture), but it does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or alternative tools. Given the large sibling list, some explicit guidance would help, but the purpose is understandable enough to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_fixA
Read-onlyIdempotent
Inspect

Re-check a specific security control after applying a fix, to confirm the finding is now resolved. Use only when a fix has already been applied and you want to verify or confirm the remediation was successful — not for initial inspection of a record.

ParametersJSON Schema
NameRequiredDescriptionDefault
checkYesCheck name to re-run (e.g., "dmarc", "spf")
domainYesDomain to validate the fix for
formatNoOutput verbosity. Auto-detected if omitted.
expectedNoExpected DNS record value to verify against

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the tool as read-only and non-destructive. The description adds the context that it is a follow-up verification step after remediation, which is useful behavioral context. However, it does not go deeper into output behavior or any special side effects, so a mid-range score is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences front-loaded with the core purpose, followed by explicit usage conditions. No wasted words; every sentence adds meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a clear post-fix verification purpose, and the description covers when to use it and what it does. With annotations providing safety attributes and the schema covering all parameters, the description is complete for this tool's complexity. No output schema exists, so return values are not expected to be documented here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides descriptions for all four parameters, including enums and examples, so the description does not need to compensate. The schema coverage is 100%, and the description adds no parameter-specific details, so a baseline score of 3 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('re-check') and resource ('security control') and clearly states the purpose is to confirm remediation after applying a fix, distinguishing from initial inspection tools. It effectively communicates the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use ('only when a fix has already been applied') and when not to use ('not for initial inspection'), providing clear exclusion criteria. This guides the agent away from using it for initial scans.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 45 tool updates
    • Changedbrand_audit_batch_start1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedbrand_audit_get_report1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedbrand_audit_single1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedbrand_audit_status1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_agent_discovery1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_authoritative_dns_infra1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_bimi1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_caa1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dane1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dane_https1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dbl1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dkim1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dmarc1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dnskey_strength1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dnssec1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_dnssec_chain1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_fast_flux1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_http_security1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_llms_txt1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_lookalikes1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_mta_sts1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_mx1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_mx_reputation1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_ns1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_nsec_walkability1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_ptr1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_rbl1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_realtime_threat_feed1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_root_server_set1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_shadow_domains1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_spf1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_srv1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_ssl1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_subdomailing1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_subdomain_takeover1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_svcb_https1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_tlsrpt1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_txt_hygiene1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcheck_zone_hygiene1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedcymru_asn1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changeddelete_brand_audit_watch1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changeddiscover_brand_domains1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedlist_brand_audit_watches1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedrdap_lookup1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedregister_brand_audit_watch1 field changed
      • addedOutput schema / properties / verdictWithheld
        Added value: +{
        +  "type": "boolean"
        +}
  2. 1 tool update
    • Addedcheck_llms_txt

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables DNS and email security analysis through passive and active scanning capabilities. Provides comprehensive domain security checks including SPF, DMARC, DNSSEC validation, MX record analysis, and SMTP connectivity testing.
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that checks SPF, DKIM, DMARC, and MX records for a domain, returning a health verdict and specific DNS fixes to improve email deliverability.
    1
    MIT
  • A
    license
    A
    quality
    F
    maintenance
    Provides comprehensive tools for real-time DNS queries across 53 record types, global propagation checks, and SSL certificate analysis. It also enables domain security scans for SPF/DKIM/DMARC configurations and HTTP uptime monitoring.
    8
    72 npm
    22
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.