Skip to main content
Glama

Server Quality Checklist

92%
Profile completionA complete profile improves this server's visibility in search results.
  • Latest release: v1.36.0

  • Disambiguation4/5

    Most tools have clearly distinct purposes, with detailed descriptions guiding usage. Some overlap exists among domain analysis tools (domain_report, audit_domain, dns_lookup) and CVE tools (cve_lookup, cve_search, cve_leading), but they are differentiated by scope and use case. The detailed 'Use for' hints in descriptions help disambiguate.

    Naming Consistency4/5

    Tool names predominantly use snake_case with a verb_noun or noun_verb pattern (e.g., asn_lookup, check_headers, bulk_cve_lookup). A few names like 'email_security_posture' deviate from the common pattern, but overall consistency is high.

    Tool Count3/5

    54 tools is heavily weighted for a single server. While each tool serves a niche cybersecurity domain (domain, IP, CVE, MITRE, code scanning, email, etc.), the breadth could overwhelm agents and suggests possible consolidation. The count is at the high end of appropriate for a comprehensive threat intelligence platform.

    Completeness4/5

    The tool surface covers major cybersecurity workflows: domain reconnaissance, IP intelligence, threat feed correlation, vulnerability management, MITRE framework navigation, email/phone validation, and basic code scanning. Minor gaps exist (e.g., no port scanning beyond ip_lookup, no malware sandboxing), but the set is well-scoped for a read-only enrichment service.

  • Average 4.7/5 across 54 of 54 tools scored.

    See the Tool Scores section below for per-tool breakdowns.

    • 3 of 3 community issues answered or closed in the last 6 months
    • 70 commits in the last 12 weeks
    • Last stable release on
    • No critical vulnerability alerts
    • No high-severity vulnerability alerts
    • No code scanning findings
    • CI is failing
  • This repository is licensed under MIT License.

  • This repository includes a README.md file.

  • No tool usage detected in the last 30 days. Usage tracking helps demonstrate server value.

    Tip: use the "Try in Browser" feature on the server page to seed initial usage.

  • This repository includes a glama.json configuration file.

  • This server has been verified by its author.

How to sync the server with GitHub?

Servers are automatically synced at least once per day, but you can also sync manually at any time to instantly update the server profile.

To manually sync the server, click the "Sync Server" button in the MCP server admin interface.

How is the quality score calculated?

The overall quality score combines two components: Tool Definition Quality (70%) and Server Coherence (30%).

Tool Definition Quality measures how well each tool describes itself to AI agents. Every tool is scored 1–5 across six dimensions: Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), and Contextual Completeness (10%). The server-level definition quality score is calculated as 60% mean TDQS + 40% minimum TDQS, so a single poorly described tool pulls the score down.

Server Coherence evaluates how well the tools work together as a set, scoring four dimensions equally: Disambiguation (can agents tell tools apart?), Naming Consistency, Tool Count Appropriateness, and Completeness (are there gaps in the tool surface?).

Tiers are derived from the overall score: A (≥3.5), B (≥3.0), C (≥2.0), D (≥1.0), F (<1.0). B and above is considered passing.

Tool Scores

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate read-only, idempotent, non-destructive. The description adds important behavioral details: batch processing, capped at 500 entries with extra dropped, rate limits (30/hr Free, 500/hr Pro), and the exact fields in the return object. This provides sufficient transparency beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is a single paragraph of five sentences, efficient and front-loaded with the main action. It includes only necessary information: purpose, return fields, usage hint, rate limits. No fluff, but could be more structured (e.g., bullet points for return fields).

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (batch, multiple return fields, rate limits) and the presence of an output schema, the description covers the main aspects: purpose, return object keys, usage, and constraints. It does not detail the format of coverage_by_tactic, but the output schema likely provides that. Overall, it is sufficiently complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The input schema already covers the parameter 'attack_technique_ids' with full description (100% coverage), including maxItems and example. The description only reiterates the cap and example, adding no new meaning. So baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool takes a list of ATT&CK T-codes and returns a coverage breakdown including defense counts per D3FEND tactic and identifies undefended techniques. It explicitly says this is for assessing defensive posture of an entire campaign or threat model, distinguishing it from siblings like d3fend_defense_for_attack.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides clear usage context: 'Use to assess the defensive posture of an entire attack campaign or threat model in one call' and suggests pairing with cve_search for gaps. However, it does not explicitly state when not to use it or compare to alternatives, so it falls short of full guidance.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnly and idempotent hints. The description adds useful behavioral details: returns 404 for missing slug, describes response fields including attack_techniques, and states rate limits. No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is well-structured with main action first, then supportive details, usage guidance, error handling, rate limits, and return schema. Every sentence adds value without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a simple lookup tool with good annotations, this description covers purpose, usage context, error behavior, rate limits, and output structure. It is comprehensive relative to complexity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Input schema has one parameter with a detailed description and 100% coverage. The tool description does not add extra parameter information, so baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool looks up a MITRE D3FEND defense technique, explains what D3FEND is, and distinguishes from sibling tools by advising to use after d3fend_defense_search.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Description explicitly says to use after d3fend_defense_search and mentions 404 for missing slugs, providing clear context. It does not explicitly list alternatives or when not to use, but the guidance is sufficient.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false. Description adds valuable behavioral context: TXT records are returned raw with honest count, rate limits (30/hr free, 500/hr pro), and return structure. Does not contradict annotations; scores 4 for adding beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is a single packed sentence but effectively front-loads purpose and key details. Every phrase adds value, no fluff. Could be slightly more structured, but overall concise.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given simple input schema (1 param, fully described), presence of output schema, and annotations, the description provides sufficient context: explains return format, rate limits, and distinguishes from sibling. No gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with parameter 'domain' described clearly ('Root domain to query, without protocol or path'). The description does not add additional parameter meaning beyond what the schema provides, so baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description is specific: 'Query all DNS record types (A, AAAA, MX, NS, TXT, CNAME, SOA) for a domain.' Lists clear use cases like mail routing, nameserver verification, SPF/DMARC checks, and explicitly distinguishes from sibling 'domain_report' for filtered TXT views.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use this tool vs alternatives: 'Use for mail routing inspection, nameserver verification, or SPF/DMARC checks; for full overview use domain_report.' Also clarifies that TXT records are returned raw and 'total_txt_records' is honest, directing to domain_report for filtered view.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Goes beyond annotations by explaining DKIM status logic ('verified' vs 'unverifiable'), grading formula under both DKIM conditions, and that DKIM absence is not penalized. Also mentions rate limits. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Starts with core purpose, followed by alternative, rate limits, and detailed behavioral notes. Each sentence adds value. Slightly long but well-organized and front-loaded.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Covers what the tool returns (list of fields) and explains grading and DKIM logic. The 'issues' field is mentioned but not detailed, which is acceptable with output schema present. Adequate for the tool's complexity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema already provides clear description for the only parameter 'domain'. The tool description adds no further parameter meaning; baseline 3 applies since schema coverage is 100%.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Clearly states the tool analyzes email security covering MX records, SPF, DMARC, DKIM, mail provider, and grade. Distinguishes from sibling 'domain_report' by specifying this is for verification and phishing risk while domain_report is for full audit.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says to use for verifying email-auth setup and phishing risk, and for full audit use domain_report. Provides rate limits but does not mention other siblings like email_security_posture. Clear context with one alternative.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations declare readOnlyHint, idempotentHint, etc. Description goes far beyond: explains response fields, wildcard_status and crtsh_status interpretation, rate limits, and how to handle partial results. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness3/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is lengthy but well-organized with clear structure. Front-loads purpose. However, some detail about response fields could be shortened or moved to output schema documentation.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given complexity (multiple sources, status fields, rate limits, output schema present but not covering all interpretation), description is very complete. Covers use case, method, limitations, result interpretation, and actions for partial results.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema has one parameter (domain) with description, coverage 100%. Description does not add significant new info for the parameter beyond what schema provides, but schema already covers it well.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states the tool discovers subdomains using passive methods (Certificate Transparency logs + DNS brute-force) and maps attack surface. Distinguishes from sibling tools by specifying the passive, non-intrusive nature.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly describes when to use (map organization's attack surface) and mentions rate limits. Does not explicitly say when not to use or list alternatives, but context is clear.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds value with 404 error handling, rate limits (30/hr Free, 500/hr Pro), and response size variants (slim vs full). No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise at 5 sentences, front-loading the purpose and then logically covering variants, usage tips, error case, and limits. Every sentence is informative, though it could be slightly tighter.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the presence of an output schema, the description adequately covers return fields (case_study_id, name, description, techniques_used, next_calls), error handling, rate limits, and usage chaining. No gaps identified.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    With 100% schema description coverage, the description enhances parameter meaning by clarifying the default slim response, the effect of passing 'full', and the exact ID format (AML.CS####). This adds context beyond the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it looks up a MITRE ATLAS case study by ID, contrasting with sibling tools like atlas_case_study_search for searching and atlas_technique_lookup for techniques. It specifies the resource and action with details on response format and usage context.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly advises use after atlas_technique_search to find incidents exercising a technique, and suggests bulk_atlas_technique_lookup for full techniques_used. It provides clear context but lacks explicit 'when not to use' statements.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds significant value by specifying rate limits (30/hr free, 500/hr pro) and a critical behavioral trait: the carrier field is omitted for certain regions and carrier_status must be interpreted accordingly. This prevents incorrect inferences by the agent.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is well-structured with purpose first, then format requirements, companion tools, rate limits, and return behavior. It is concise but slightly redundant with the schema examples. Every sentence contributes, but it could be more streamlined.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the output schema exists, the description sufficiently explains return values and adds crucial behavioral context about carrier omission and carrier_status interpretation. It covers edge cases (MNP-restricted regions) and rate limits, making it complete for an agent to use correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a detailed parameter description including format examples. The description reinforces the E.164 format but does not add new semantic meaning beyond the schema. Baseline is 3 for high coverage.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description uses a specific verb 'Validate and analyze' and clearly lists expected outputs (country, region, carrier, line_type, timezone, formats). It distinguishes itself from siblings by mentioning companion tools username_lookup and email_disposable for identity investigation, making its role in the toolkit clear.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly states when to use the tool: to verify phone legitimacy and detect fraud risks. It mentions companion OSINT tools for identity investigation, providing context. However, it lacks explicit 'when not to use' guidance or alternatives besides the companions, which would prevent misuse.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, idempotentHint, and no destructiveness. The description adds significant behavioral context: identical for all tiers, uses local DB mirrors, no tier gating, cost 10 tokens, rate limits (30/hr for free, 500/hr for Pro), and that IPs/internal hostnames are rejected. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is front-loaded with the core composite purpose in the first sentence. It is comprehensive but somewhat dense, covering multiple aspects concisely. Every sentence adds value, though it could be slightly more structured for easier scanning.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (composite, multiple steps), the description is thorough: it explains all steps, tier behavior, cost, rate limits, and explicitly lists return fields (domain, technologies, cves_by_tech, etc.). An output schema exists, so return values are well-covered. No gaps noted.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with one parameter 'domain'. The schema description adds meaning beyond type/format by specifying 'Target domain to fingerprint and CVE-audit (e.g. 'example.com'). IPs and internal hostnames are rejected.' This provides clear constraints and examples.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it is a 'Composite tech-stack + CVE audit' tool that detects technologies, queries CVEs, enriches with KEV deadlines, and checks exploit availability. This specific verb+resource combination distinguishes it from sibling tools like 'tech_fingerprint' (just fingerprinting) and 'cve_lookup' (just CVE details).

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines3/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description implies usage for auditing a domain's technology stack and vulnerabilities, but does not explicitly state when to use this tool versus alternatives (e.g., separate tech_fingerprint or cve_lookup). It mentions 'MCP-only, no REST endpoint' but lacks clear when-not or alternative guidance.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations indicate readOnly, idempotent, non-destructive. Description adds that affected_products truncates to first 20 by default, references to first 10, and boolean flags to override. Also mentions next_calls for chaining and return shape. No contradictions. Provides valuable behavioral defaults and optional behaviors.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is dense with information but each sentence adds value. Could be slightly better organized (e.g., grouping defaults, rate limits, chaining). However, it is front-loaded with core purpose and constraints, making it effective.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given full schema coverage, output schema, and annotations, description covers all needed context: rate limits, return shape (results, total, etc.), chaining guidance, and default behaviors. Complete for a batch query tool.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage 100% but description adds significant meaning: cve_ids format and max, boolean parameters explained with default values, truncation behavior, and use-case guidance (e.g., set include_affected_products for dependency audits). Each parameter's effect is clear.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Clearly states batch query for multiple CVEs up to 50 per call, retrieving full CVE details. Distinguishes from sibling cve_lookup for single CVE. Verb 'bulk query' and resource 'CVE details' are specific.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says to use for dependency audits or bulk vulnerability enrichment, and to use cve_lookup for single CVE. Also mentions chaining with kev_detail, cwe_lookup, exploit_lookup based on result fields. Rate limits provided (30/hr Free, 500/hr Pro). No explicit when-not-to-use, but context is clear.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    The description adds behavioral context beyond annotations: it states 'No external requests' (confirming readOnlyHint and destructiveHint), explains default truncation to 500 chars with include='full' option, and describes the return structure {total, by_severity, findings}. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise at 4 sentences, front-loading the core purpose, then guidelines, then details. Every sentence provides unique value with no redundancy or wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the presence of an output schema (not shown but mentioned in context) and the tool's simplicity, the description covers purpose, usage, alternatives, limits, truncation behavior, return structure, and safety. It feels complete for an agent to decide and use correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Both parameters have schema descriptions covering their purpose, types, and constraints (100% coverage). The description adds a note to include only security-relevant headers and explains the include parameter slightly further, but the schema already does a good job. Baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool validates HTTP security headers (CSP, HSTS, etc.) against best practices. It distinguishes from sibling tool scan_headers by specifying when to use each: use check_headers for offline testing, use scan_headers to fetch live headers.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly says when to use this tool (test config before deployment, validate non-public servers) and mentions the alternative scan_headers. It also provides rate limits (30/hr free, 500/hr pro). It could be more explicit about when not to use, but the information is clear.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false) are complemented by extensive behavioral details: null-explicit response fields, two-tier cloud_provider detection, tor_exit false on fetch failure vs genuine absence, next_calls logic, version-breaking changes for vulns field format, and severity handling. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness3/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is long and dense with information. While every sentence adds value, the sheer length and run-on style reduce readability. It could benefit from structured formatting (e.g., bullet points) or segmentation by topic. Front-loading is good but overall conciseness suffers.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the complexity of the tool (many response fields, conditional behaviors, version history), the description is thorough. It explains every field's behavior, null handling, and next_calls, making it complete for agent usage. The output schema exists but is not needed given the detailed prose.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The single parameter 'ip' is described in schema (IPv4 or IPv6). The description adds examples ('8.8.8.8', '2606:4700::1111') and context about its purpose, providing value beyond the schema's brief description. Schema coverage is 100%, so baseline is 3; the extra examples increase it to 4.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description uses a specific verb ('Query') and resource ('comprehensive IP intelligence'), listing numerous data points (reverse DNS, ASN, ports, vulnerabilities, etc.). It clearly distinguishes from siblings like asn_lookup, ioc_lookup, and threat_report by stating 'Use for IP investigation; for orchestrated IP+reputation use threat_report.'

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly says 'Use for IP investigation; for orchestrated IP+reputation use threat_report.' It includes rate limits (Free 30/hr, Pro 500/hr) and mentions triggers for next_calls (asn_lookup, ioc_lookup, threat_report). However, it does not explicitly state when not to use this tool beyond the alternative mention.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true and destructiveHint=false. Description adds context about 404 return behavior and rate limits, but does not contradict annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is front-loaded with action, followed by return details, usage guidance, and rate limits. Every sentence adds value without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given tool complexity (one parameter, output schema exists), description covers purpose, return values, error handling, rate limits, and follow-up steps. Fully adequate for agent selection.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with clear parameter description including format examples. Description does not add significant meaning beyond listing return fields, which is not parameter-specific.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description explicitly states 'Look up CISA KEV full record for a CVE,' using a specific verb and resource. It distinguishes from siblings like cve_lookup by noting it returns 404 for non-KEV CVEs.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit guidance: use for in_kev=true CVEs after cve_lookup or cve_search, and alternatives like cve_lookup for non-KEV CVEs. Also mentions chaining with cwe_lookup and rate limits.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate read-only (readOnlyHint=true), idempotent (idempotentHint=true), and non-destructive behavior. The description adds value by explaining the analysis method (HTTP headers + HTML) and the return structure, which goes beyond the annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Two concise sentences: first explains what the tool does, second provides usage context and return format. No unnecessary words, front-loaded with key information.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a simple tool with one parameter, the description covers purpose, methodology, usage guidance, rate limits, and return format (with a hint at the structure). There is no missing critical information for an AI agent to decide when and how to use this tool.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% for the single parameter 'domain', with a clear description in the schema. The tool description does not add additional semantic information about the parameter beyond what is in the schema, so a baseline score of 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool's purpose: 'Detect website technology stack' with specific examples (CMS, frameworks, etc.) and methodology (HTTP headers + HTML analysis). It distinguishes from sibling tool 'audit_domain' by noting this is for passive reconnaissance, while a full audit should use that tool.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly says when to use ('passive reconnaissance') and when not to ('for full audit use audit_domain'), providing an alternative tool. It also includes rate limits (Free: 30/hr, Pro: 500/hr), guiding usage expectations.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already signal readOnly, openWorld, idempotent, non-destructive. The description adds the return format {username, total_found, platforms: [...]} and rate limits, fully disclosing behavior beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Two sentences, front-loaded with platform list and key details. No wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the single parameter, existing output schema, and richness of annotations, the description covers purpose, usage, output structure, and limitations (rate) fully.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The input schema already describes the username parameter with examples (without '@'). The tool description does not add new parameter meaning beyond what's in the schema, so baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states that the tool searches for a username across 15+ platforms, listing specific examples (GitHub, Reddit, etc.) and mentions OSINT and identity verification as use cases. This distinguishes it from sibling tools like asn_lookup or dns_lookup.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly recommends use for OSINT investigations and identity verification, and notes rate limits (30/hr free, 500/hr Pro). While it doesn't specify when not to use or name alternatives, the context is clear enough for an agent.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true, destructiveHint=false, etc. Description adds return structure details and rate limits, providing context beyond annotations. No contradiction. Moderate added value.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Three concise sentences, each serving a purpose: purpose/data, usage/alternative, return/limits. Front-loaded and no wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a single-parameter, read-only tool with annotations covering safety and output schema hinted in description, all necessary context is provided: purpose, usage, rate limits, return fields. Complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only one parameter 'domain' with schema description covering 100%. Description does not add additional semantic details beyond what's in the input schema. Baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description starts with 'Retrieve WHOIS registration data: registrar, creation/expiry dates, nameservers, status.' This clearly states the action (retrieve) and resource (WHOIS data) with specific fields. It distinguishes from sibling 'domain_report' by noting 'for full audit use domain_report'.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use: 'Use to verify domain ownership, age, expiration' and when not: 'for full audit use domain_report'. Also includes rate limits: 'Free: 30/hr, Pro: 500/hr', aiding in resource management.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Beyond annotations (readOnly, idempotent, non-destructive), the description discloses default truncation (first 50 prefixes), the effect of setting include_full_prefixes=True, and that ipv4_count and ipv6_count are honest pre-truncation totals. It also includes rate limits. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is succinct (three sentences) with the main purpose front-loaded. Every sentence adds useful information, and there is no redundant or extraneous content.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    The description covers inputs, outputs (with a clear return format hint), and default behavior. Given the presence of an output schema, the description is sufficiently complete. It does not cover error cases, but the tool is simple enough.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Both parameters have schema descriptions (100% coverage). The description adds value by explaining the default behavior for include_full_prefixes and providing a concrete example (Cloudflare AS13335) to illustrate when to use it.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool looks up ASN for a domain or IP, specifying what it returns (AS number, organization, IPv4/IPv6 prefixes). It distinguishes from sibling lookup tools by focusing on network operator identification.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides context on when to use the tool (e.g., for network mapping or BGP route audits) and mentions default vs. full prefix behavior. It does not explicitly mention when not to use it or list alternatives, but the context is clear enough.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already provide readOnlyHint, idempotentHint, destructiveHint. Description adds value by noting 404 on missing ID, rate limits (30/hr free, 500/hr pro), sub-technique inheritance of tactics, and return structure fields. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is somewhat long but well-structured and front-loaded with purpose. Every sentence adds value, though a slight tightening could improve conciseness without losing information.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the single parameter and presence of an output schema (implied by return description), the description is complete: covers use cases, error handling, rate limits, return fields, and relationships to other tools and ATT&CK/D3FEND.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only one parameter (technique_id) with 100% schema description coverage. The schema already describes the format with examples. The tool description does not add additional parameter meaning beyond what's in the schema, so baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it looks up MITRE ATLAS techniques, which are AI/ML adversarial attack TTPs. It distinguishes from siblings like atlas_technique_search (search vs lookup) and bulk_atlas_technique_lookup (single vs bulk).

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says when to use: when user asks about AI/ML threats, LLM red-teaming, or adversarial ML. Also advises using bulk version for multiple techniques and mentions pivoting to D3FEND and CVE search.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, openWorldHint, idempotentHint, destructiveHint. Description adds rate limits (30/hr Free, 500/hr Pro) and token cost (6 tokens), plus next_calls behavior, providing extra transparency beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is front-loaded with purpose and efficient, but slightly dense as a single paragraph. Could benefit from more structure (e.g., bullet lists) for complex details like rate limits and filtering, but overall concise and informative.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given tool complexity (combines multiple scans), description covers all essential aspects: purpose, filtering logic, usage guidance, chaining, rate limits, and return structure (explicitly lists output fields). No gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 100%. Description adds context: explains default TXT filtering (strips vendor verification strings) and when to set include_all_txt=True, going beyond schema descriptions.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states 'Perform comprehensive domain audit' and explicitly lists combined components: domain_report, live HTTP security headers, technology fingerprinting. Distinguishes from sibling domain_report by specifying active vs passive assessment.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit when-to-use ('Use when you need the full picture') and alternative ('use domain_report for passive-only assessment'). Also suggests chaining with subdomain_enum and ssl_check for further recon.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, idempotentHint, destructiveHint. The description adds details about response shape (status per item, error handling), case-insensitivity, normalization, deduplication, and next_calls. No contradiction.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is single paragraph but packed with essential information. All sentences are meaningful, though slightly dense. Could be improved with bullet points.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a bulk tool with one parameter and output schema, the description covers motivation, follow-up use case, response structure, rate limits, and per-item status. No gaps given the context.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a clear description of technique_ids format. The description adds value by noting each id counts as one request toward rate limit, which is not in schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool retrieves full records for up to 50 techniques in a single request, distinguishing it from the sibling atlas_technique_lookup. The verb 'retrieve' and resource 'ATLAS techniques' are specific, and the alternative use case is explicitly mentioned.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states it is the natural follow-up to atlas_case_study_lookup and implies alternative (atlas_technique_lookup for single lookups). Rate limits and per-item request counting provide clear usage context.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable behavioral context: auto-detection of IOC type, per-type source coverage (hash → ThreatFox, etc.), and mention that each result carries verdict.sources_queried / sources_unavailable for partial failure visibility. This goes beyond what annotations provide, though it doesn't detail all side effects (none exist).

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is a single paragraph that efficiently packs key information—purpose, constraints, source coverage, rate limits, and return format—without unnecessary words. It is front-loaded with the primary use case. No fluff, but slight improvement could be made with bullet points for clarity.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (multiple IOC types, different source mappings, batch behavior, partial failures), the description covers all essential aspects: input limits, per-type source coverage, return structure with {results, total, successful, failed, timed_out, partial, summary}, and rate limits. An output schema exists (not shown), so return values are documented. The description is comprehensive for an AI agent to use correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The input schema has 100% coverage for the single required parameter 'indicators'. The description adds meaning beyond the schema: maximum 50 per request, auto-detection of indicator type, and the per-indicator behavior. These details help the agent construct valid inputs and understand constraints.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it is for batch querying multiple IOCs (IP/domain/URL/hash) up to 50 per call, with auto-detection of type. It distinguishes itself from the sibling tool ioc_lookup by specifying batch vs. single indicator use.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit guidance: use for SOC alert triage or batch enrichment, and use ioc_lookup for single indicator. Also includes rate limits (Free 30/hr, Pro 500/hr) and max batch size, helping the agent decide when to invoke this tool.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    The annotations (readOnlyHint: true, idempotentHint: true, destructiveHint: false) indicate a safe, read-only operation. The description adds extensive behavioral context: the formula with weights, multiplicative boosters, clamping to [0,100], label bands, urgency encoding, the source of PoC data, methodology attribution, and rate limits (30/hr free, 500/hr pro). No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is detailed and front-loaded with the core purpose and formula. While it includes necessary details like boosters and rate limits, it is somewhat lengthy. However, every sentence adds value, and the structure is logical. It earns a 4 for being informative without being excessively verbose.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given that the tool has an output schema (implied by the return fields listed), the description covers the full context: purpose, usage guidelines, behavioral details, parameter format, output structure, and limitations. It leaves no critical gaps for an agent to misuse the tool.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The input schema has one required parameter (cve_id) with a description of its format. The description does not add additional semantics for this parameter beyond what the schema already provides. With schema description coverage at 100%, the baseline is 3, and the description does not exceed it.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states that this tool calculates a composite CVE risk score using CVSS, EPSS, KEV, and PoC, with a detailed formula. It explicitly distinguishes itself from siblings like cve_lookup and exploit_lookup by stating: 'Use to triage a single CVE without orchestrating cve_lookup + exploit_lookup separately' and that for full exploit detail, one should call exploit_lookup separately.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides explicit guidance on when to use this tool (for triaging a single CVE) and when not to use it (for multi-source exploit detail, call exploit_lookup). It also notes the limitation that the PoC signal is from the local ExploitDB mirror only, which helps agents decide if they need additional tools.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations indicate read-only and non-destructive. The description adds behavioral context: 'No data stored' and details rate limits (Free: 30/hr, Pro: 500/hr), going beyond annotations to disclose important operational traits.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is a single, well-structured paragraph that front-loads the purpose and efficiently packs essential information without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (injection scanning), the description covers purpose, usage, behavioral traits, parameters, and output format. The output schema exists, so the return description is supplemental. Complete enough for effective use.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100%. The description adds value by listing the supported languages and explaining the 'generic' option, which supplements the enum and descriptions in the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool scans source code for injection vulnerabilities (SQL injection, command injection, path traversal) and lists supported languages. It distinguishes from sibling tools by mentioning companion tools like check_secrets.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly says 'Use to detect input-handling bugs; for secrets use check_secrets' and lists other companion tools. It also provides context about rate limits and output format, but does not explicitly state when not to use it.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnly, idempotent, openWorld, and non-destructive traits. The description adds significant behavioral context: it performs passive DNS-only probing, tests common DKIM selectors plus custom ones, and outputs a score 0-100 with grades A+-F. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is three sentences, front-loaded with the core purpose and outputs. Every sentence adds essential information: first sentence covers function and output, second sentence covers use cases, third covers technical details and rate limits. No wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness4/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    The description covers the key aspects: what the tool does (SPF, DMARC, DKIM), how it works (passive DNS, custom selectors), output format (score and grades), and rate limits. Since there is an output schema, return values are covered. Minor details like the exact grading scale or finding structure are omitted, but the tool is sufficiently described for correct invocation.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Input schema has 100% description coverage for both parameters. The description reinforces the schema by explaining the purpose of custom selectors and the probing behavior, adding value beyond the schema's basic descriptions. This justifies a score above the baseline of 3.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool analyzes domain email authentication posture (SPF, DMARC, DKIM) and produces a numeric score with findings. It uses a specific verb ('Analyze') and resource ('domain email authentication posture'), effectively distinguishing it from sibling tools like dns_lookup or domain_report that do not focus on email security.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly mentions dual-use for red-team and blue-team, indicates it is passive DNS-only with no SMTP probe, and provides rate limit information. It explains when custom DKIM selectors are needed. While it does not explicitly state when not to use or name alternatives, the context is clear enough for correct tool selection.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Discloses response structure with field details, Shodan refs limit (default 200, truncated flag), and explains next_calls logic. No contradictory annotations; readOnlyHint=true aligns with the search nature.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is dense with useful information but slightly verbose. It is well-structured, covering sources, usage, and returns, though a few sentences could be trimmed without loss.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the rich annotations and output schema, the description fully covers the tool's behavior, integration points, and constraints. No missing information for an effective agent invocation.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only one parameter (cve_id) with full schema coverage (100%). The description reinforces the format and gives examples, but adds no substantial meaning beyond the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it searches public exploits/PoC for a specific CVE across three named sources. It distinguishes itself from sibling tools like cve_lookup, kev_detail, and cwe_lookup by specifying its unique function.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicit guidance: 'Use to assess if a vulnerability has weaponized exploits in the wild; run after cve_lookup'. Provides context on when to pair with kev_detail and cwe_lookup, and mentions rate limits (30/hr free, 500/hr Pro).

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already mark the tool as read-only and idempotent. The description adds rate limits (30/hr free, 500/hr pro) and details the return structure (fields like base_score, metrics object). No contradictions. A score of 4 reflects the added value beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is well-structured and front-loaded with the core action. It is relatively long but each sentence adds value, covering purpose, return fields, usage guidelines, and rate limits. Could be slightly more concise but maintains good density.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the output schema exists, the description comprehensively lists all return fields (canonicalized vector, version, base_score, base_severity, metrics, temporal/environmental scores, summary, verdict). It also covers error handling for v2 vectors. Fully complete for the tool's complexity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% and the schema already describes the vector parameter. The description adds context by providing an example vector string and explicitly stating v2 vectors are rejected. This adds meaningful guidance beyond the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it parses a CVSS v3.x vector string into a per-metric breakdown and recomputed base score. It distinguishes itself from siblings like cve_lookup and exploit_lookup by focusing on CVSS vector parsing, which no other sibling tool does.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says when to use: to translate raw CVSS strings and verify upstream NVD scoring. Also states v2 vectors are rejected and directs users to cvss_v2_field on cve_lookup for v2 details. Clear when-not-to-use guidance.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Discloses detailed behavioral logic (active evidence criteria, historical retention, stale state) beyond annotations, and includes rate limits. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Well-structured with front-loaded purpose, then detailed logic, companion tools, and output format. Slightly lengthy but every sentence adds value; no wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Covers purpose, usage context, behavioral details, rate limits, output structure, and sibling differentiation. Output schema is described inline, making it complete despite 1 param complexity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a clear description of the 'url' parameter including examples. Description adds behavioral context but not new parameter meaning beyond schema, per baseline.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Clearly states it queries URLhaus for a URL and host, defines is_malicious logic, and explicitly distinguishes from siblings like threat_intel for domain-level checks.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use (URL-level threat assessment), when not (use threat_intel for domains), lists companion tools, and includes rate limits (30/hr free, 500/hr Pro).

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Beyond annotations (readOnlyHint, idempotentHint, destructiveHint), the description adds significant behavioral details: SSRF-guarded with IP re-validation, loop detection with loop_detected flag, truncation at 10 hops, per-target eTLD+1 throttle with specific rate limits (60 req/min), and hard fetch failure handling. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is a single paragraph but well-structured: it opens with the core action and output, then lists use cases, followed by technical details (SSRF, limits, errors). Every sentence adds value without redundancy. It could be slightly more scannable, but it is compact and informative.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a tool with one parameter and a detailed output schema (described in text), the description covers all essential aspects: purpose, use cases, technical behavior (SSRF, loops, truncation), rate limits, error handling (502, 429 with Retry-After). It provides a complete picture without requiring the user to infer anything.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The single parameter 'url' is already well-documented in the input schema (100% coverage). The description adds practical usage guidance: 'Pass the URL exactly as you'd `curl -L` it; the server handles encoding.' This enriches the schema description with an intuitive analogy and reassurance about encoding, providing added value.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states the tool walks an HTTP redirect chain hop-by-hop, specifying the exact per-hop data returned {url, status_code, location, latency_ms}. It lists concrete use cases (deobfuscating URL shorteners, audit suspicious links, trace marketing redirects), which distinguishes it from sibling tools that are primarily lookups and scans.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Description provides explicit use cases for when to use the tool (deobfuscate shorteners, audit links, trace redirects). It also details limits (10 hops, rate limits) and error responses (502, 429). However, it does not explicitly state when not to use it or mention alternatives, though sibling tools don't offer similar functionality.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already provide readOnlyHint, openWorldHint, idempotentHint, destructiveHint. Description adds user-agent identification, rate limiting, free/pro limits, return fields, and error handling details, going beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is thorough but slightly lengthy. However, it is well-structured with front-loaded main purpose and subsequent details. Every sentence adds value.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the output schema defines the return structure, the description covers all necessary context: prerequisites, error handling, rate limits, and sibling tool relationships. Highly complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only one parameter (domain) with 100% schema coverage. Description adds constraints (no scheme/path/port, subdomain handling, HTTP fallback) that clarify usage beyond schema descriptions.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool fetches and parses robots.txt, listing extracted elements (sitemaps, rules, etc.). It uses specific verbs and resources, and distinguishes itself from sibling tools by being the only one dealing with robots.txt.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly instructs to use BEFORE crawling/scraping for seo_audit, brand_assets, redirect_chain. Explains when 404 means implicit allow-all and when 502 is an upstream failure, not 'no robots'.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate read-only and idempotent. Description adds useful behavioral context: truncation to 500 chars, CSP can exceed 4 KB, and return format. No contradiction.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is dense but efficient, front-loading purpose and including all key details in a compact form. Could be slightly better structured, but every sentence adds value.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given two parameters and an output schema, the description fully covers tool behavior, usage scenarios, rate limits, and parameter options. No gaps for effective usage.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100%, so baseline 3. Description adds meaning by explaining the effect of the include parameter and the return fields, going beyond schema descriptions.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states tool performs live HTTP GET and analyzes security headers, listing specific headers. Distinguishes from sibling check_headers by specifying when to use each.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says use for auditing live website headers and check_headers for validating existing headers. Provides rate limits and details on truncation behavior with the include parameter.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations indicate read-only, open-world, idempotent, non-destructive. Description adds valuable behavioral details: how invalid certs are handled (returns valid=false+validation_errors, not endpoint failures), grade definitions, and rate limits (30/hr free, 500/hr Pro). No contradiction.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is fairly long but each sentence adds value, front-loaded with purpose. Could be slightly more concise but still effective. No waste.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given that an output schema exists (indicated), description covers key aspects: purpose, behavior on invalid certs, grade definitions, rate limits, and return fields. Complete for this tool's complexity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only one parameter (domain) with 100% schema coverage. Description does not add meaning beyond schema, which is sufficient. Baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states the tool analyzes SSL/TLS certificates and lists specific aspects (grade, protocol, etc.). It explicitly distinguishes from sibling audit_domain with 'for full domain audit use audit_domain', showing specific verb+resource+scope.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use: 'Use to audit certificate validity and detect expiring certs' and provides an alternative: 'for full domain audit use audit_domain'. Clear usage context and exclusion.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds critical behavioral details: the meaning of 'status' values, that 'total_snapshots' is omitted when unavailable (not zero), error code mapping to warnings, and heavy domain fallback. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is moderately long but highly informative, front-loading the purpose and then providing usage guidelines and behavioral caveats. Every sentence adds value, though it could be slightly more concise.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (one parameter, rich output, error handling, rate limits, alternative tool), the description covers return value structure, status meanings, error codes, and usage context. The presence of an output schema (though not detailed here) is noted, and the description provides sufficient context for an agent to select and use the tool correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The single parameter 'domain' has 100% schema coverage with a clear description. The tool description does not add new semantic details beyond the schema, so a baseline score of 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool retrieves Wayback Machine snapshots for a domain, specifying what is returned (first capture, latest, total count, snapshot list). It distinguishes from the sibling tool 'domain_report' by mentioning 'for full audit use domain_report'.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use this tool ('investigate domain history and age') and provides an alternative ('for full audit use domain_report'). Also includes rate limits and warnings about heavy domains, guiding proper invocation.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds 'No data stored' and explains the suppression rule for generic password-assignment, providing context beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise (5 sentences), front-loaded with the main action, and includes necessary details without fluff.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    With an output schema provided, the description doesn't need to explain return values. It covers purpose, usage, behavioral traits, and parameters sufficiently.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with detailed descriptions for both parameters. The description lists supported languages but does not add significant new meaning beyond the schema. Baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description explicitly states 'Scan source code (or snippet) for hardcoded secrets' and lists types of secrets (cloud provider keys, API tokens, etc.), distinguishing it from the sibling 'check_injection' for injection detection.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    It provides clear when-to-use guidance: 'Use to detect leaked credentials before commit' and contrasts with an alternative: 'for injection detection use check_injection'. Rate limits (30/hr free, 500/hr Pro) are also included.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, openWorldHint, idempotentHint true, and destructiveHint false. The description adds valuable behavioral context beyond annotations: rate limits (30/hr free, 500/hr Pro) and a summary of returned fields, which helps the agent understand side effects and limits.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise (6-7 sentences) and front-loaded: it starts with purpose, then usage guidelines, then rate limits, then return fields. No redundant or irrelevant information; every sentence adds value.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool has an output schema, the description does not need to explain return values in detail, but it still lists the key fields. The description covers all essential aspects: purpose, usage context, behavioral traits, and parameter constraints, making it fully actionable for an AI agent.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 100%, so the baseline is 3. The schema already fully documents the file_hash parameter with accepted formats and examples. The description does not add any new meaning beyond what is in the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool queries MalwareBazaar for file hashes (MD5/SHA1/SHA256) and returns malware family, file type, size, tags, dates, and download count. It explicitly distinguishes itself from sibling tools like ioc_lookup (multi-source) and others, satisfying the 'specific verb+resource, distinguishes from siblings' criterion.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides explicit guidance: 'Use to check if file hash is known malware; use ioc_lookup for auto-detection of all IOC types.' It also lists companion tools with their purposes, giving clear when-to-use and when-not-to-use context.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior4/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint, openWorldHint, idempotentHint, destructiveHint. The description adds behavioral details: auto-detection of type, per-type source coverage, verdict fields (sources_queried, sources_unavailable), and failure handling. No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Every sentence serves a purpose: purpose, source coverage, usage guidance, rate limits, output format. No filler, well-organized, and front-loaded with key information.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (multiple input types, multiple feeds, output schema exists), the description covers input possibilities, feed mapping, failure behavior, and output structure. It is fully informative for an agent to invoke correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a detailed description of the indicator parameter. The description adds context about auto-detection and feed behavior, enriching the parameter semantics beyond the schema alone.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description specifies the verb 'Enrich' and the resource 'Indicator of Compromise' with explicit types (IP/domain/URL/hash) and auto-detection. It distinguishes from siblings like threat_intel (domain-only) and hash_lookup (richer MalwareBazaar data), making the purpose clear and unique.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicit guidance: 'Use as primary IOC triage tool when type unknown' and alternatives for specific cases. Rate limits (30/hr free, 500/hr Pro) are provided, helping the agent decide when to use this tool vs. others.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already provide readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds behavioral context: it is a single-source check, returns specific fields, and has rate limits. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise (3 sentences) and well-structured: main purpose first, then alternatives, rate limits, and return format. Every sentence adds value.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the presence of an output schema, the description is complete: it specifies the source, alternatives, rate limits, and the fields returned. No gaps remain for an AI agent to understand usage.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 100% for the single parameter 'domain'. The description does not add additional meaning beyond what the schema provides, so baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool checks a domain against abuse.ch URLhaus for known malware-distribution URLs, using a specific verb ('Check') and resource ('domain'). It distinguishes from siblings like ioc_lookup and phishing_check.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly states when to use this tool (fast domain-level threat assessment) and when to use alternatives: ioc_lookup for multi-feed correlation and phishing_check for specific URLs. It also provides rate limit details.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already declare readOnlyHint=true, etc. The description adds significant behavioral context beyond annotations: rate limits (30/hr Free, 500/hr Pro), breaking change in enrichment.vulns (list of strings pre-1.16 vs list of VulnInfo objects), and the return shape. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is concise and well-structured: first sentence states purpose and scope, following sentences add usage guidance, limitations, breaking changes, and rate limits. Every sentence adds value, and the information is front-loaded. No wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Despite having an output schema (not shown but mentioned), the description still outlines the return shape ('Returns {ip, enrichment, abuseipdb, shodan, asn, threat_level}'). It covers all important aspects: what it does, when to use it, behavioral details, limitations, and rate limits. For a one-parameter tool with high schema coverage and annotations, this is complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 100% and the schema already explains the ip parameter adequately (must be public IPv4/IPv6, private/reserved rejected). The tool description adds minimal extra meaning ('Use for IP investigation'), but does not provide new syntax or format details beyond what the schema already states. Baseline 3 is appropriate.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it queries a comprehensive threat profile for an IP, listing specific data sources (Shodan, AbuseIPDB, ASN/geolocation, open ports). It also explicitly distinguishes from the sibling 'domain_report' by saying 'for domain data use domain_report'. This provides specific verb+resource and differentiates among siblings.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description gives explicit usage context: 'Use for IP investigation and SOC alert triage' and directs to a sibling tool for domain data. It also notes limitations (ASN prefixes at most 50, breaking change in vulns format) and rate limits, helping the agent decide when to use this tool versus alternatives.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations indicate readOnly, idempotent, non-destructive. Description adds details on return structure {findings, total, by_severity, summary} and explains when fixed_in is omitted, providing behavioral context beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is somewhat long but well-structured: first sentence summarizes purpose, then provides usage, rate limits, return format, and special cases. Every sentence adds value, though slight trimming possible.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given complexity of multiple package managers, bulk query limit, and return format, description covers all essential aspects including edge cases (open-ended version ranges). Output schema exists but description complements it.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with detailed description of packages parameter. Description adds context on supported package managers (npm/PyPI etc.) and CVE database source, augmenting schema meaning.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states 'Audit project dependencies against CVE database: find known vulnerabilities in your package list.' It uses specific verbs and resources, and distinguishes from siblings like cve_lookup by specifying bulk query of packages.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states 'Use for dependency security scanning; use cve_lookup for single CVE.' Also provides rate limits for Free and Pro tiers, guiding appropriate usage.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations declare readOnlyHint, openWorldHint, idempotentHint, and not destructive. The description adds key behavioral details: active outbound fetch, per-target throttle, rejection of bare IPs/private domains, and token costs. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is a single dense sentence that packs many details. It is front-loaded with the main purpose. While slightly verbose, every sentence earns its place, but could be more readable with breaks.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (11 modules, many return fields), the description covers purpose, usage, behavioral constraints, and output structure. The presence of an output schema (mentioned) and full parameter schema coverage makes the description complete enough for correct invocation.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The schema covers the single parameter 'domain' with description. The description goes beyond by clarifying that bare IPs and private-resolving domains are rejected. With 100% schema coverage, baseline is 3; the extra context raises it to 4.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly identifies the tool as an active website security scan using the ContrastScan C engine with 11 modules. It specifies the resource (live site) and the output (severity-ranked findings and letter grade). It also distinguishes itself from sibling tools like audit_domain and scan_headers, fulfilling the specificity requirement.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicit guidance: 'Use for a hands-on misconfiguration scan; use audit_domain for passive recon... and scan_headers for headers only.' Also mentions rate limits (60 req/min) and token costs, providing clear when-to-use and when-not-to-use context.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations only hint at read-only and idempotent behavior. The description adds concrete behavioral details: default slim response (saves ~60 chars/row, ~30%), include='full' option, rate limits (30/hr free, 500/hr Pro), and return format specification.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The single-paragraph description is well-structured: purpose, default behavior, parameters, usage example, chaining, rate limits, return format. Every sentence adds necessary information without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (multiple filters, chaining, rate limits, differing response formats) and the presence of output schema, the description covers all relevant aspects thoroughly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    All 6 parameters are already described in the schema (100% coverage). The description adds value by explaining the default include behavior with char savings, the chaining use of exclude_id, and providing examples for keyword and artifact.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool searches the MITRE D3FEND catalog by keyword, tactic, or artifact. It differentiates itself from siblings like d3fend_defense_lookup by specifying chaining behavior.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides explicit use cases (e.g., hardening access tokens) and chaining instructions with exclude_id. It lacks explicit when-not-to-use, but context with sibling tools implies appropriate usage.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate safe, read-only, idempotent behavior. The description adds detailed behavioral context: checks performed, role_address detection details (admin@, info@, etc.), free_provider detection, and mentions that Gmail-style +tag is stripped. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is informative but slightly long; however, every sentence adds value. Well-structured: purpose first, then usage, limitations, details. Front-loaded.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given complexity with multiple checks, the description covers input, output fields (listed in description), rate limits, exclusions (no SMTP probing), making it fully complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Only parameter 'email' has 100% schema coverage. Description adds examples and requirement ('Must contain @'), providing additional meaning beyond the schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description explicitly states it is a one-call email validation combining syntax, MX, disposable, role, and free-provider checks. It distinguishes itself from sibling tools like email_mx and email_disposable by noting it replaces 2-3 tool calls.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides clear when-to-use guidance ('BEFORE adding an email to a contact list...') and explicit what-it-does-not-do (SMTP RCPT TO deliverability probing) with alternatives (Hunter.io/ NeverBounce). Also mentions rate limits (30/hr Free, 500/hr Pro).

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already provide readOnlyHint, openWorldHint, idempotentHint. The description adds substantial behavioral context: fully deterministic (no LLM queried), honors robots.txt with specific error codes (403, 502), cache behavior, throttling per eTLD+1, and detailed output fields. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is lengthy but well-structured, front-loading purpose and then detailing rules, usage, ethics, and output. Every sentence provides necessary information for a complex tool, though it could be slightly more concise without losing clarity.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (7 weighted rules, multiple output fields, error states, rate limits, ethical considerations), the description is highly complete. It covers input constraints, behavioral details, output format, and edge cases. The output schema is mentioned, so return values are sufficiently documented.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The single parameter 'domain' has 100% schema description coverage. The description adds meaning by specifying it must be registrable, no scheme/path/port, homepage-only, with HTTP fallback. This adds value beyond the schema's own description.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states this tool performs a 'Deterministic GEO / AI-visibility readiness audit' of a domain's homepage, producing a 0-100 score and fix list. It distinguishes itself from siblings like seo_audit by focusing specifically on AI assistant discoverability, and from audit_domain by being homepage-only.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explicitly states when to use the tool: 'to triage why a brand is absent from AI recommendations, as a pre-flight before GEO/AEO content work, or to score a prospect's AI-readiness.' It also states what not to do: 'Strictly homepage-only — we do NOT crawl.' It provides context on ethical handling of robots.txt and rate limits.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds behavioral details: default response is SLIM (description truncated to 240 chars), include='full' for verbose summary, and return structure. No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Four sentences that are front-loaded: first sentence states purpose and parameters, second gives use case, third mentions alternative, fourth notes rate limits and return structure. No wasted words.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given full schema, annotations, and output schema existence, the description is complete. It explains default behavior, alternatives, and return fields. No gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100%, so baseline is 3. Description adds context about default slim behavior for 'include' and provides example values for 'keyword' and 'technique_id' (e.g., 'evasion', 'AML.T0051'), enhancing understanding beyond the schema's own descriptions.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool searches ATLAS case studies by keyword or technique referenced. It distinguishes from siblings by mentioning 'Drill via atlas_case_study_lookup for the full procedure list', and the sibling list includes atlas_technique_search and atlas_case_study_lookup, making differentiation clear.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says 'Useful when the user has a technique in hand and wants to see incidents that exercised it' and 'Drill via atlas_case_study_lookup for the full procedure list'. Also mentions rate limits (Free: 30/hr, Pro: 500/hr), providing when-to-use and alternatives.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Discloses that default response is SLIM with truncated descriptions, and that passing include='full' returns verbose records which can be large. Provides return structure and notes idempotent, read-only behavior beyond annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is dense and informative but slightly lengthy. It front-loads the main purpose and then provides details. Every sentence is valuable, but could be tightened slightly without losing clarity.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Covers all necessary aspects: purpose, parameter details, usage flow (chaining with lookup), response structure, rate limits, and cross-referencing with other tools. No gaps given the tool's complexity and presence of output schema and annotations.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Despite 100% schema coverage, the description adds significant value: explains the effect of include (default vs. full, size warning), gives examples for keyword and tactic, clarifies maturity enum meanings, and describes exclude_id usage for chaining.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states it searches the MITRE ATLAS catalog of AI/ML attack techniques by keyword, tactic, or maturity. It distinguishes itself from sibling tools like atlas_technique_lookup by being a discovery/search tool.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly advises when to use this tool (discover techniques for threat-model questions), how to drill into atlas_technique_lookup, and how to cross-reference with D3FEND. Also mentions exclude_id for chaining and rate limits (free vs. pro).

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnlyHint, openWorldHint, idempotentHint, and non-destructive. The description adds critical behavioral details: robots.txt handling with 403, cache-control respect, rate limiting (60 req/min), security warnings about untrusted URL fields, and error responses (502, 403). No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is dense but well-structured, front-loading the main action and use cases then adding details. Every sentence provides useful information, but length could be slightly reduced without losing clarity. Minor improvement possible.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the complexity and presence of output schema, the description covers all relevant aspects: ethical compliance, caching behavior, rate limits, security considerations, error responses. It is fully adequate for an agent to understand the tool's behavior and constraints.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a single required 'domain' parameter. The description adds meaning: specifies no scheme/path/port, explains HTTP fallback, and that it fetches https://<domain>/>, providing valuable context beyond the schema's type and description.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool scrapes a domain's homepage <head> for public brand assets like favicon, og:image, etc. It lists specific assets and use cases (enrich CRM, build UI, correlate lead site), distinguishing it from sibling tools by emphasizing 'homepage-only' and no crawling.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description explains when to use (enrich CRM, build UI) and when not (strictly homepage-only, respects robots.txt, ethical floor). It provides clear context including rate limits and error handling, guiding appropriate usage without the need for manual screenshots.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnly and idempotent. Description adds extensive behavioral details: default truncation of affected_products and references, optional flags to expand, reference tags behavior, severity breakdown, patch detection, rate limits. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is very detailed and well-structured, but somewhat verbose. However, every sentence adds necessary context for a complex tool. Slightly longer than ideal but not wasteful.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity, schema coverage, output schema presence, and annotations, the description covers all aspects: default behavior, optional parameters, response structure, chaining, rate limits. No gaps identified.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100%, but description goes beyond by explaining default values, use cases for each boolean parameter (e.g., 'Set True for bulk audits'), and the impact on response. Adds significant value beyond schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it retrieves detailed CVE data by ID, listing all key fields. It explicitly distinguishes from sibling cve_search, stating 'Use for single-CVE details; use cve_search for queries by product/severity.'

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit context on when to use this tool vs cve_search, and describes chaining with kev_detail, cwe_lookup, exploit_lookup via next_calls. Also mentions rate limits for free vs Pro tiers.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Discloses behavioral traits beyond annotations: default SLIM response vs include='full', exact match token requirements (e.g., nginx vs nginx_open_source), explanation of low/zero counts meaning token mismatch, and response structure (verdict at root, not per-row). No contradictions with annotations (readOnlyHint, idempotentHint, destructiveHint).

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is comprehensive but slightly lengthy; however, every sentence adds value for a complex tool with 14 parameters. Front-loaded with purpose and filters, then details alternatives and special behavior. Could be slightly more concise but remains effective.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Complete for the tool's complexity: covers all filter types, default response, alternative tools, token matching caveats, pagination hints, and response structure (verdict, hint). Despite no full output schema provided, the description adequately explains return fields and usage patterns.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Adds significant meaning beyond schema: explains product/vendor EXACT match with examples (nginx_open_source, log4j), details include parameter's slim vs full options and token cost implications, describes vendor filter's cross-row matching behavior. All 14 parameters are covered and enriched.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states 'Search CVE database with filters' and lists specific filter criteria, distinguishing it from siblings like cve_lookup (single CVE) and check_dependencies (auto-normalized searches). It precisely defines the tool's scope and output format.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly guides when to use alternatives: for dependency/package lists use check_dependencies, for domain's tech stack use tech_stack_cve_audit, for single CVE use cve_lookup, for kev deadlines use kev_detail. Also advises using cwe_id for enumerating weaknesses paired with cwe_lookup. Provides clear scenarios and exclusions.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations indicate read-only, idempotent, non-destructive. Description elaborates on behavior: default slim response (first 3 mitigations/examples, extended_description null), total counts always honest, cve_count as lower bound (only primary CWE matches), 404 when not in research view 1000. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Every sentence serves a purpose. Information is front-loaded: main purpose, optional mode, return fields, usage chain, rate limits, error condition. No redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Covers all aspects: input format, output fields, behavior nuances (slim vs full, cve_count lower bound), error handling (404), rate limits, and chaining suggestions. Output schema exists, but description still details return fields.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100%, but description adds value: lists common CWE IDs for cwe_id, explains default vs full for include, and describes total_mitigations/total_examples as honest pre-truncation counts. Adds clarity beyond schema.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Clearly states 'Look up MITRE CWE catalog record from research view 1000', specifying the resource and scope. Differentiates from siblings like cve_lookup and cve_search by its focus on weaknesses. The description of default slim vs. full mode adds precision.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly recommends using after cve_lookup or kev_detail and chaining with cve_search. Provides rate limits (30/hr free, 500/hr pro). No ambiguity about when to use.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Discloses key behaviors beyond annotations: truncation logic (`defenses` capped at `limit`, `total` and `coverage_by_tactic` always full), default slim response, rate limits (30/hr free, 500/hr Pro), and 200 response with empty list for unmapped T-codes. No contradiction with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is front-loaded with the core purpose, but is somewhat long. However, every sentence contributes meaning; a minor trim would improve conciseness.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's complexity (4 params, API behaviors, truncation, and rate limits), the description covers all necessary aspects including return structure (`defenses` fields, `coverage_by_tactic`, `next_calls`). Output schema handles the rest.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Each parameter description adds value beyond the schema: `limit` explains why cap matters for popular T-codes, `include` quantifies character savings, `exclude_id` details chaining use case, and `attack_technique_id` gives format examples.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it's a reverse lookup from ATT&CK T-code to D3FEND defenses, distinguishing it as the bridge from offensive intelligence to defensive playbook. It explicitly mentions pairing with cve_lookup or atlas_technique_lookup, differentiating it from siblings like d3fend_defense_lookup.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit when-to-use guidance: call this tool when you have an ATT&CK id from CVE/ATLAS lookups. Also explains when to pass `exclude_id` when chaining from d3fend_defense_lookup, and when to use `include='full'` for verbose records.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations (readOnlyHint, idempotentHint, etc.) are complemented by descriptions of default TXT filtering, rate limits (30/hr free, 500/hr Pro), next_calls behavior, and the rationale for stripping vendor verification strings. No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is dense but front-loaded with purpose. While slightly long, every sentence provides essential information for correct tool use, balancing completeness with clarity.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    With an output schema (not shown), returns are presumed documented. The description covers parameters, usage context, behavioral details, rate limits, and chaining instructions, leaving no gaps for effective agent invocation.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% and descriptions add value: domain parameter clarifies 'without protocol or path'; include_all_txt explains default filter and when to use True, including examples of stripped strings.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states it queries DNS, WHOIS, SSL, subdomains, and threat intel for a domain in one call. It distinguishes from sibling audit_domain by specifying it as a starting point and names other tools for deeper analysis.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly says 'Use as a starting point for domain investigations; use audit_domain for live headers + tech stack.' Provides guidance on include_all_txt parameter and suggests chaining with specific tools via next_calls.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false. Description adds return fields {disposable, domain, provider}, consistent with read-only behavior. No contradiction.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is compact: purpose, examples, use-case, companion tools, rate limits, and return structure in three sentences. No fluff, front-loaded with essential info.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's simplicity (1 param, output schema exists, annotations cover safety), the description fully covers what the agent needs: purpose, when to use, return fields, and rate limits. Complete.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Single parameter email is fully described in schema (100% coverage). Description adds example formats ('user@tempmail.com', 'test@guerrillamail.com'), which is helpful but not critical. Baseline 3, plus 1 for extra context.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description uses specific verb 'Check if email address uses a known disposable/temporary provider' and lists examples (Guerrilla Mail, Temp Mail, Mailinator). It clearly distinguishes from sibling tools like threat_intel, email_mx, and domain_report by stating their different use cases.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use ('input validation to detect throwaway signups') and when not ('for domain reputation use threat_intel'), and mentions companion tools. Also provides rate limits (Free: 30/hr, Pro: 500/hr), aiding selection.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations declare readOnlyHint, idempotentHint, destructiveHint=false; description adds that no data is stored and the full hash never leaves the tool, reinforcing the privacy-preserving k-anonymity approach. Rate limits also disclosed.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is concise, front-loaded with purpose and method, and includes usage guidance, companion tools, rate limits, and return format—all in a few sentences with no superfluous content.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a single-parameter tool with output schema, the description covers what it does, how it works, when to use it, and what it returns, leaving no significant gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters4/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema covers the sole parameter (sha1_hash) with 100% documentation. Description adds behavioral nuance (only 5-char prefix used) that helps the agent understand how the parameter is processed, adding extra context.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool checks if a SHA-1 hash appears in the HIBP breach dataset using k-anonymity (5-char prefix). It distinguishes from sibling tools like hash_lookup by specifying different namespaces.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly mentions use for password breach audits, read-only nature, and lists companion tools (hash_lookup, email_disposable, username_lookup) as alternatives for different investigation areas. Rate limits are provided.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Beyond annotations (readOnlyHint, etc.), the description discloses ethical handling (robots.txt honoring, cache-control respect), rate limits (60 req/min per eTLD+1), error codes (502, 403), and the 'cache_respected' flag. No contradictions with annotations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is detailed and well-structured, starting with a high-level summary then diving into specifics. While slightly verbose, every sentence adds value and no redundancy is present.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the existence of an output schema covering return values, the description fully covers behavior, constraints, error handling, ethical considerations, and rate limiting. It addresses all likely agent questions for correct invocation.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The only parameter 'domain' is thoroughly described: registrable domain, no scheme/path/port, strictly homepage-only. The description adds constraints beyond the schema (e.g., HTTP fallback, no trailing slash). Schema coverage is 100%.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description explicitly states it performs a one-shot SEO audit of a domain's homepage with a 0-100 composite score and a list of concrete fixes. It distinguishes from siblings like 'audit_domain' by emphasizing the homepage-only scope and the specific audit rules.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Clearly specifies when to use: before pitching SEO work, triaging lead marketing maturity, or as a pre-flight before deeper tools. Also clarifies what not to do: strictly homepage-only and not a site crawl.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already confirm readOnly, idempotent, not destructive. Description adds rate limits (30/hr Free, 500/hr Pro), daily refresh at 02:00 UTC, and return structure including next_calls with follow-up tool suggestions. No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is comprehensive yet well-structured: purpose, return fields, usage, follow-ups, rate limits. Each sentence adds value. Slightly verbose in enumerating status levels, but overall efficient.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    With output schema mentioned and only 1 parameter, the description fully covers behavior and output. Includes next_calls for contextual follow-ups. No gaps given the tool's simplicity.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema coverage is 100% with a well-described parameter. Description adds value by explaining where to obtain the UUID (sigma_rule_search or external SIEM) and providing an example. Exceeds basic schema information.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Clearly states 'Look up a single Sigma detection rule by UUID' with specific source (SigmaHQ corpus), size, and refresh. Distinguishes from sibling tools like bulk_sigma_rule_lookup and sigma_rule_search.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Provides explicit use cases: 'fetch a known rule for context (e.g., a SIEM detection that fired)' or 'inspect a rule discovered via REST sigma_rule_search'. Implies alternatives (search, bulk) and includes rate limits as usage constraints.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Annotations already indicate readOnly and idempotent, but the description adds critical behavioral details: per-item error handling (status per ID), quota consumption (1 unit per rule_id, hourly limits), batch cap of 50, and rate limit handling (skipped_due_to_rate_limit). No contradictions.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is efficient and well-structured: first sentence gives core purpose, then use case, then output shape, then quota details. Every sentence adds unique value without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's single parameter, rich annotations, and output schema, the description covers all necessary aspects: purpose, usage context, quota semantics, error handling, and batch behavior. No gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The input schema covers 100% of parameters (only rule_ids), but the description significantly enhances understanding: adds max 50 items, RFC 4122 format, quota counting per ID, and per-item validation outcomes (invalid_format, not_found).

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly states the tool does a bulk lookup of Sigma rules by UUIDs, retrieving full records for up to 50 IDs in a single call. It explicitly differentiates from the sibling sigma_rule_lookup by noting the bulk capability.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description provides explicit guidance: 'Designed for triage workflows where multiple rule ids are known (e.g., from a SIEM alert batch or a tagged detection bundle)' and contrasts with N separate sigma_rule_lookup calls.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior5/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    Beyond annotations (readOnlyHint, idempotentHint, destructiveHint), description adds behavioral context: early-warning nature, slim default results, include='full' for full payload, verdict at root, response hint pointing to cve_lookup, and rate limits. No contradiction.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    Description is comprehensive but well-structured, front-loading core purpose, then usage guidelines, parameter details, response shape, and rate limits. Every sentence adds value without redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness5/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    Given the tool's low complexity (3 parameters, no required, output schema exists), the description is extremely complete. It covers response shape, verdict, hints, follow-up tools, and rate limits. No gaps.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters5/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    All three parameters have detailed schema descriptions (100% coverage). Description adds extra context for the 'include' parameter, explaining slim vs full payload trade-offs and why slim is default.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    Description clearly states the tool lists CVEs indexed from MITRE/GHSA before NVD publication, serving as an early-warning for emerging threats. It explicitly distinguishes from siblings cve_lookup and cve_search.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines5/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    Explicitly states when to use (threat intelligence on emerging CVEs) and when not (for drill-down on single CVE prefer cve_lookup, for published NVD data use cve_search). Also mentions rate limits and payload options.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

GitHub Badge

Glama performs regular codebase and documentation scans to:

  • Confirm that the MCP server is working as expected.
  • Confirm that there are no obvious security issues.
  • Evaluate tool definition quality.

Our badge communicates server capabilities, safety, and installation instructions.

Card Badge

contrastapi MCP server

Copy to your README.md:

Score Badge

contrastapi MCP server

Copy to your README.md:

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/UPinar/contrastapi'

If you have feedback or need assistance with the MCP directory API, please join our Discord server