recon-mcp
Detects subdomain takeover risks for services hosted on Fastly, such as unclaimed resources.
Detects subdomain takeover risks for services hosted on GitHub Pages, such as unclaimed repositories.
Detects subdomain takeover risks for services hosted on Heroku, such as unclaimed apps.
Detects subdomain takeover risks for services hosted on Shopify, such as unclaimed storefronts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@recon-mcprun a security recon report on example.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
recon-mcp
English | 繁體中文
An MCP server that gives AI coding agents — Claude Code, Codex, Cline, and any MCP client — safe, structured network and security reconnaissance tools.
Most MCP servers wrap CRUD APIs. recon-mcp instead exposes the kind of
read-only recon an engineer reaches for when investigating an asset, and returns
clean JSON — with a graded verdict — so the agent can reason over results
instead of parsing console output.
⚠️ Authorized use only. These tools are for security testing of assets you own or have explicit written permission to assess, for CTF practice, and for education. Do not point them at third-party infrastructure without authorization. You are responsible for how you use this software.
Tools
Tool | What it does |
| Start here. One call → DNS, TLS, and HTTP headers checked together, with an overall grade |
| DNS + WHOIS + email security (SPF/DMARC/DKIM), graded |
| Discover subdomains via DNS brute-force and/or Certificate Transparency logs |
| Check subdomains for a dangling-CNAME takeover risk against known services |
| Certificate, protocols, ciphers, and known TLS vulnerabilities, graded |
| HTTP security headers (CSP, HSTS, X-Frame-Options, …), graded |
| Redirect chain + cookie flags (Secure / HttpOnly / SameSite), graded |
| CORS policy probe — flags arbitrary-Origin reflection and wildcard misuse |
| Fingerprint the web stack (server, CDN/WAF, language, framework, CMS, JS) from one GET |
| Report which HTTP methods a server allows and grade the risk (TRACE/PUT/DELETE) |
| Fetches & parses |
| Resolves the host and enriches its IP via RDAP (owner, country, CIDR, abuse) |
| TCP port scan of one host (≤1024 ports/call), open ports + services |
Related MCP server: Bug Bounty Assistant MCP
Example
Just ask your agent: "run a security recon report on example.com." It calls
recon_report once and gets a graded overview it can act on:
{
"domain": "example.com",
"overall_grade": "F",
"summary": "Overall posture F: email A, TLS B, headers F; 13 actionable issue(s).",
"components": {
"email": { "grade": "A", "issues": [] },
"tls": { "grade": "B", "issues": [] },
"headers": { "grade": "F", "issues": [
{ "severity": "high", "label": "Missing Content-Security-Policy", "detail": "CSP not set; cannot restrict resource load sources" }
] }
}
}Need more detail on one area? The agent can call dns_recon, subdomain_enum,
subdomain_takeover, tls_check, http_headers_audit, cookie_audit,
cors_check, tech_detect, http_methods_audit, well_known_audit,
ip_info, or port_scan directly.
Install
Requires Python ≥ 3.10. Runs on Linux, macOS, and Windows (tested in CI).
Recommended — no clone, via uv:
uvx recon-kit-mcpOr from source (for development):
git clone https://github.com/nan786521/recon-mcp
cd recon-mcp
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS / Linux
source .venv/bin/activate
pip install -e .Use with Claude Code
Add the server (stdio transport). With uvx you don't need an absolute path:
claude mcp add recon -- uvx recon-kit-mcpOr add it manually to any MCP client config:
{
"mcpServers": {
"recon": {
"command": "uvx",
"args": ["recon-kit-mcp"]
}
}
}(From a source checkout, point the command at /absolute/path/to/.venv/bin/recon-kit-mcp instead.)
Then just ask: "run a security recon report on example.com" — or target one area, e.g. "check the email security of example.com."
The server also ships a security_recon prompt: pick it from your client's
prompt menu and pass a domain for a guided, severity-sorted audit.
Tool reference
recon_report(domain, timeout?) -> dict
Runs DNS/email, TLS, HTTP-header, web-stack (tech_detect), and apex
subdomain-takeover checks together and returns overall_grade (as weak as the
weakest component, capped at F if a live takeover is found), a one-line
summary, components (email / tls / headers, each with its grade and
actionable issues), a tech section (detected technologies + any version
disclosure), and a takeover section when the apex is at risk. Uses a fast
single-handshake TLS check for speed — call tls_check for the full
cipher/vulnerability analysis. The best starting point; use the tools below for
raw detail.
dns_recon(domain, checks?, timeout?) -> dict
records — A, AAAA, MX, NS, TXT, SOA, CNAME, CAA records
whois — parsed registration fields + raw WHOIS text
email — SPF, DMARC, and DKIM posture, plus advisory MTA-STS, TLS-RPT, BIMI, and DNSSEC signals, and a graded
assessment(letter grade A–F, a summary, and per-check findings with severity and a recommended fix). The advisory signals surface as findings but don't move the core SPF/DKIM/DMARC grade.
checks is any subset of ["records", "whois", "email"]; omit it to run all.
subdomain_enum(domain, wordlist?, source="dns", timeout?) -> dict
Discovers subdomains from two complementary sources:
source="dns"(default) — resolves candidate labels via DNS.wordlistis comma-separated labels ("www,api,dev"); omit it for a built-in common list. Capped at 512 candidates per call. Returns resolvedips.source="ct"— queries public Certificate Transparency logs (crt.sh) for every name ever certified for the domain. Fully passive; finds real hosts no wordlist would guess.source="both"— runs both and merges, recording which source(s) saw each host.
Returns sources, found_count, and found (each with subdomain, the
sources that saw it, and ips when resolved).
subdomain_takeover(hosts, timeout?) -> dict
Checks subdomains for a dangling-CNAME takeover — a subdomain that CNAMEs to
a third-party service (GitHub Pages, S3, Heroku, Azure, Fastly, Shopify, …)
whose resource was deleted or never claimed, letting anyone who registers that
resource serve content on the victim's subdomain. For each host it resolves the
CNAME, recognizes known takeover-prone services, fetches the page, and flags
the provider's "unclaimed resource" fingerprint and/or a CNAME target that no
longer resolves. hosts is one hostname or a comma-separated list (capped at
100). Read-only — DNS lookups plus one HTTP GET per host. Pair it with
subdomain_enum: enumerate first, then check the interesting hosts.
Returns checked, vulnerable_count, and results (each with host, cname,
service, status, vulnerable, severity, and detail). status is one of
not_applicable, not_vulnerable, potential, dangling_cname, or
vulnerable.
tls_check(host, port=443, timeout?) -> dict
Returns grade, certificate (validity / expiry / key algorithm),
protocols (flags legacy SSLv3 / TLS 1.0 / 1.1), cipher info,
forward_secrecy, hsts, vulnerabilities (each with a vulnerable flag),
and a findings list.
http_headers_audit(host, port?, use_ssl=True, timeout?) -> dict
Returns grade, score, the observed security headers, and a findings
list with a recommendation per header. Defaults to HTTPS (port 443).
cookie_audit(host, port?, use_ssl=True, timeout?) -> dict
Follows the redirect chain from the host (capped at 10 hops, flagging any
HTTPS→HTTP downgrade) and audits every Set-Cookie seen for the Secure,
HttpOnly, and SameSite flags. Returns redirect_chain, final_url,
cookies (flags only — values are never returned), cookie_grade,
cookie_score, and a findings list.
cors_check(host, port?, use_ssl=True, timeout?) -> dict
Sends one GET with an untrusted Origin and inspects the
Access-Control-Allow-Origin / -Allow-Credentials response. Reflecting an
arbitrary Origin with credentials is high severity (any site can read
authenticated responses); a wildcard or trusted null origin are lesser issues.
Returns acao, allows_credentials, reflects_origin, wildcard, severity,
and findings.
tech_detect(host, port?, use_ssl=True, timeout?) -> dict
Fingerprints the technology stack behind a website from one HTTP GET. It
matches response headers, set cookies, the HTML body, and the
<meta name="generator"> tag against a signature table to identify the web
server, reverse proxy / CDN, WAF, programming language, web framework, CMS,
JavaScript framework, and analytics. Where a version is exposed it is captured
and flagged (info) — a precise version eases known-CVE lookup. Read-only.
Returns status, technology_count, technologies (each with name,
category, version when known, and evidence), and a findings list noting
any version disclosure.
http_methods_audit(host, port?, use_ssl=True, path="/", timeout?) -> dict
Reports which HTTP request methods a server allows and grades the risk. Enabled
write/diagnostic methods widen the attack surface: TRACE enables Cross-Site
Tracing (XST), and PUT / DELETE can allow file upload or deletion under
weak access control. Safe by design — it never sends a mutating request: it
actively probes only OPTIONS, HEAD, and TRACE (TRACE merely echoes), and reads
PUT / DELETE / PATCH / CONNECT from the OPTIONS Allow header as advertised,
never invoking them.
Returns grade, score, allow_header, advertised_methods, trace_enabled,
dangerous_methods, and a findings list (each with the method, severity, and a
recommendation).
well_known_audit(host, timeout?) -> dict
Fetches and parses security.txt (RFC 9116, tried at /.well-known/ then the
legacy path) and robots.txt. Returns security_txt (parsed fields, structural
issues, location) and robots_txt (sitemaps, disallow/allow paths,
user_agents), each with a present flag.
ip_info(host, timeout?) -> dict
Resolves the host's IP and looks it up in the public RDAP registry (via
rdap.org's bootstrap to the right RIR). Returns ip and rdap (handle,
name, country, cidr, org, abuse_email).
port_scan(host, ports?, timeout?) -> dict
TCP connect scan of a single host. ports is a string — "22,80,443", a
range "1-1024", or a mix — and omitting it scans a built-in common-port set.
Hard-capped at 1024 ports per call (single-host recon, not mass scanning).
Returns host, ip, scanned, open_count, and open_ports (port +
service). Scan only hosts you are authorized to assess.
License
Available Tools
13 toolscookie_auditA
Follow a host's redirect chain and audit the cookies it sets.
Walks each redirect hop (capped at 10) from the host, recording status and Location and flagging any HTTPS->HTTP downgrade. Every Set-Cookie seen along the way is checked for the Secure, HttpOnly, and SameSite flags and graded. Cookie values are never returned (they may be secrets).
Args: host: Hostname to inspect, e.g. "example.com". port: TCP port. Defaults to 443 when use_ssl is True, else 80. use_ssl: Start the chain over HTTPS (default True). timeout: Per-hop network timeout in seconds.
Returns: A dict with host, redirect_chain, final_url, cookies (flags only), cookie_grade, cookie_score, and a findings list.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| port | No | ||
| timeout | No | ||
| use_ssl | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool follows redirects (capped at 10), flags downgrades, checks cookie flags, and explicitly states that cookie values are never returned because they may be secrets. This is strong behavioral disclosure, though it does not mention non-destructive nature or error behavior explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with a one-line summary, followed by clear sections for behavior, arguments, and return value. Every sentence provides essential information without fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no output schema, and no annotations, the description is highly complete. It explains the tool's operation, parameter semantics, and return structure, including the fields in the returned dict. An agent has all necessary information to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does so thoroughly with an 'Args:' section explaining each parameter, including defaults and interdependencies (e.g., port default depends on use_ssl). This adds significant meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Follow a host's redirect chain and audit the cookies it sets.' It further details the audit of each redirect hop, recording status, Location, and HTTPS->HTTP downgrades, and checking cookie flags. This clearly distinguishes it from sibling tools like http_headers_audit or tls_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (when you want to audit cookies and redirects) but does not explicitly state when to use this tool over alternatives, nor does it mention when not to use it. There is no reference to sibling tools or exclusions, so the guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cors_checkA
Probe a host's CORS policy with a crafted Origin and flag misconfigurations.
Sends one GET with an untrusted Origin header and inspects the Access-Control-Allow-Origin / -Allow-Credentials response. Reflecting an arbitrary Origin while allowing credentials is high severity (any site can read authenticated responses); a wildcard or a trusted 'null' origin are lesser issues. One request, read-only.
Args: host: Hostname to test, e.g. "example.com". port: TCP port. Defaults to 443 when use_ssl is True, else 80. use_ssl: Connect over HTTPS (default True). timeout: Network timeout in seconds.
Returns: A dict with host, port, test_origin, acao, allows_credentials, reflects_origin, wildcard, severity, and a findings list.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| port | No | ||
| timeout | No | ||
| use_ssl | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It explicitly states it sends one GET with an untrusted Origin header, inspects Access-Control-Allow-Origin and Allow-Credentials, classifies severity, and is read-only. This gives the agent a reliable behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with purpose and behavior, followed by concise parameter and return documentation. Every sentence adds useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides everything needed to call the tool correctly: behavior, parameter semantics, severity interpretation, and the exact return keys. Without an output schema, the explicit return dictionary fills the gap completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate. It explains every parameter: host with an example, port default behavior based on use_ssl, use_ssl semantics, and timeout units. This goes well beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Probe a host's CORS policy with a crafted Origin and flag misconfigurations.' It clearly states what the tool does and the behavior that distinguishes it from probing or audit siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: to test a host for CORS misconfigurations. It also notes 'One request, read-only,' implying a lightweight check. It does not explicitly name alternatives or state when not to use it, but the context is still clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dns_reconA
Passive DNS/WHOIS reconnaissance for a domain using only public data.
Args: domain: The domain to inspect, e.g. "example.com". checks: Which checks to run. Any of "records", "whois", "email". Defaults to all three when omitted. timeout: Per-query network timeout in seconds.
Returns: A structured dict keyed by the requested checks: - records: DNS records grouped by type (A, AAAA, MX, NS, TXT, SOA, CNAME) - whois: parsed registration fields plus the raw WHOIS text - email: SPF / DMARC / DKIM posture, plus advisory MTA-STS, TLS-RPT, BIMI, and DNSSEC signals, with a graded assessment
| Name | Required | Description | Default |
|---|---|---|---|
| checks | No | ||
| domain | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states the tool is passive and uses only public data, which is a meaningful non-intrusive guarantee, and it documents a per-query timeout. It could additionally cover failure modes or rate-limit risks, but the core behavioral profile is clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, then organized into Args and Returns sections without redundancy. The Returns detail is somewhat lengthy but justified because there is no output schema, and every line adds operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description supplies a structured return contract keyed by checks, including the nested DNS record types and email security signals. All three parameters are documented, the default behavior is explained, and the passive/public-data constraint is stated, making the tool fully callable by an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully explain the parameters, and it does: domain gets an example, checks are enumerated with a default behavior, and timeout is defined as seconds. This fully compensates for the empty schema descriptions and adds meaning beyond the field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Passive DNS/WHOIS reconnaissance for a domain using only public data.' It clearly distinguishes this tool from active siblings like port_scan and tls_check by emphasizing passivity and public data, and the Args/Returns sections further define its exact scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Passive... using only public data' phrasing provides clear context for when to choose this tool over active reconnaissance tools, and the checks list tells the agent what subsets of work it can request. It does not explicitly name alternatives or state when not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
http_headers_auditA
Audit a web server's HTTP security response headers and grade them.
Inspects headers such as Content-Security-Policy, Strict-Transport-Security (HSTS), X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, and the COEP/COOP/CORP isolation headers. Returns a letter grade plus per-header findings with recommendations.
Args: host: Hostname or IP to audit, e.g. "example.com". port: TCP port. Defaults to 443 when use_ssl is True, else 80. use_ssl: Connect over HTTPS (default True). timeout: Per-connection network timeout in seconds.
Returns: A structured dict with: grade, score, the observed headers, and a findings list.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| port | No | ||
| timeout | No | ||
| use_ssl | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It explains the connection context (port defaults based on use_ssl, timeout), the scope of inspection, and the return structure. It stops short of describing failure modes or redirect handling, but for a read-only audit tool the behavioral picture is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a clear opening, a list of inspected headers, and separate Args/Returns sections. Every sentence adds useful information; the header enumeration is purposeful, and there is no filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is remarkably complete: it covers all four parameters, explains defaults and dependencies, and specifies the returned structure. An agent has enough to invoke the tool correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description is the only source of parameter meaning. It fully compensates by explaining each argument, including the port default behavior tied to use_ssl and the timeout as a per-connection value in seconds. This goes well beyond what the raw schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Audit a web server's HTTP security response headers and grade them.' It lists the exact header families inspected and names a concrete outcome (letter grade plus findings), which clearly distinguishes it from sibling tools like cookie_audit, http_methods_audit, and tls_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's purpose clear and implies its use case, but it never explicitly states when to choose this tool over sibling alternatives such as cookie_audit or well_known_audit, nor does it give any exclusions. Context is present, but alternative routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
http_methods_auditA
Audit which HTTP request methods a server allows and grade the risk.
Enabled write/diagnostic methods widen the attack surface: TRACE enables Cross-Site Tracing (XST), and PUT / DELETE can allow file upload or deletion under weak access control. This is read-only and never sends a mutating request: it actively probes only OPTIONS, HEAD, and TRACE (TRACE merely echoes the request); PUT, DELETE, PATCH, and CONNECT are read from the OPTIONS Allow header and reported as advertised, never invoked.
Args: host: Hostname to audit, e.g. "example.com". port: TCP port. Defaults to 443 when use_ssl is True, else 80. use_ssl: Connect over HTTPS (default True). path: Request path to test (default "/"). timeout: Network timeout in seconds.
Returns: A dict with host, url, grade, score, allow_header, advertised_methods, trace_enabled, dangerous_methods, and a findings list. An error field on fetch failure.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| path | No | / | |
| port | No | ||
| timeout | No | ||
| use_ssl | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly discloses that the tool is read-only, does not send mutating requests, and details the exact HTTP methods probed (OPTIONS, HEAD, TRACE) versus those merely read from the Allow header (PUT, DELETE, etc.). It explains the XST risk and error handling, providing significant behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but well-structured. It front-loads the purpose and safety, then explains methods, provides Args, and Returns sections. Each sentence adds value, and the formatting makes it scannable, though it is longer than necessary for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security audit tool with no annotations and no output schema, the description is thorough. It explains the method usage, return fields, defaults, and error handling, giving an agent everything needed to invoke correctly and interpret results. The output is specified with a list of fields, which compensates for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, meaning property descriptions are missing. The description fully compensates by explaining each parameter: host (Hostname to audit), port (defaults based on use_ssl), use_ssl (HTTPS connection), path (request path), and timeout (in seconds). It also explains defaults and the relationship between port and use_ssl, which the schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Audit') with a specific resource: HTTP request methods on a server. It explains the risk of enabling methods and mentions the active probes (OPTIONS, HEAD, TRACE) and the passive reporting of others, distinguishing it from siblings like 'http_headers_audit' which likely focus on headers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says the tool is 'read-only and never sends a mutating request' and lists which methods are probed vs reported, helping agents decide when to use it. It doesn't name alternatives, but the context of being a security audit tool among siblings is clear, and the description prevents misuse by explaining the safe probing behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ip_infoA
Resolve a host and enrich its IP with RDAP registry ownership data.
Looks up the IP in the public RDAP registry (via rdap.org's bootstrap to the right RIR) and reports who owns the address block, the country, the CIDR range, and the abuse-reporting contact. Read-only registry query; nothing is sent to the target.
Args: host: Hostname or IP, e.g. "example.com". timeout: Per-request network timeout in seconds. Defaults high because rdap.org's bootstrap redirect can take 10-15s on its own.
Returns: A dict with host, ip, and rdap (handle, name, country, cidr, org, abuse_email). An unresolved host or RDAP failure is reported via an error field.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It explicitly states 'Read-only registry query; nothing is sent to the target,' which is a key non-interference trait. It also discloses timeout behavior and default rationale, plus error handling via an error field. This is comprehensive and adds significant value beyond the bare schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a summary paragraph, behavior notes, and clearly labeled Args and Returns sections. Every sentence serves a purpose—purpose, safety, timeout rationale, return format, and error behavior. It is front-loaded with the core function and stays focused without unnecessary verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema, the description is fully complete. It explains parameters, return structure (dict with host, ip, and rdap fields), error handling, and behavioral nuance (read-only, no target contact). An agent has all necessary information to call this tool correctly without needing to infer anything.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does so effectively in the Args section: 'host' is clarified as 'Hostname or IP' with an example, and 'timeout' is explained as per-request seconds with a note on why the default is high due to rdap.org's bootstrap redirect. This adds meaning beyond the schema's bare titles and defaults. A small gap is that it doesn't specify timeout bounds or accepted formats, but it's well above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Resolve a host and enrich its IP with RDAP registry ownership data.' It clearly states the tool's function and distinguishes it from sibling recon tools (port_scan, dns_recon, etc.) by focusing on RDAP registry ownership lookup. The additional detail about being a read-only registry query reinforces its unique purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context about the tool's behavior (read-only, nothing sent to target, timeout rationale) but does not explicitly state when to use this tool versus alternatives. It implies usage for ownership/abuse contact lookup but offers no direct comparison or exclusion guidance relative to the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
port_scanA
TCP connect scan of a single host, reporting open ports and services.
Scoped to a single host with a hard cap of 1024 ports per call — it is recon for one authorized target, not a mass scanner. Only scan hosts you own or have explicit permission to assess.
Args: host: Hostname or IP to scan, e.g. "example.com". ports: Ports to scan as a string: "22,80,443", a range "1-1024", or a mix "1-100,443,8080". Omit to scan a built-in set of common ports. timeout: Per-port connection timeout in seconds.
Returns: A dict with: host, ip, scanned (count), open_count, and open_ports (each with port and service). Returns an error field on bad input or DNS failure.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| ports | No | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral transparency. It discloses the scan type (TCP connect), the port cap, the per-port timeout, and error conditions on bad input or DNS failure. It appropriately warns about authorization, but does not detail rate-limiting, timing behavior, or side effects beyond the 1024-port cap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized: a one-sentence core definition, a scoping/authorization paragraph, an Args section, and a Returns section. Every sentence adds necessary information, and the most important scoping constraint is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no annotations and no output schema, the description gives enough context for an agent to invoke the tool correctly: it defines inputs, output structure, error handling, and usage boundaries. Nothing essential for selecting and calling this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate. It explains all three parameters in detail: host with an example, ports with explicit formats ('22,80,443', '1-1024', '1-100,443,8080') and its default behavior, and timeout as a per-port connection timeout. This is far more informative than the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('TCP connect scan of a single host') and states the precise output ('open ports and services'). It distinguishes itself from a 'mass scanner', making its scope and purpose unmistakable relative to the sibling recon tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly scopes the tool to a single host, notes the 1024-port hard cap, and strongly states the proper use case: 'recon for one authorized target, not a mass scanner'. It also adds the authorization requirement ('only scan hosts you own or have explicit permission to assess'). It does not explicitly name a sibling alternative, but the domain is specific enough that this is a minor omission.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recon_reportA
One-shot security posture report for a domain.
Runs DNS/email, TLS, HTTP-header, web-stack (tech_detect), and apex subdomain-takeover recon concurrently and returns a single graded overview: an overall grade (as weak as the weakest component, and capped at F if a live takeover is found), each component's grade, the actionable issues found, the detected technology stack, and any takeover risk. Use this for a quick full picture; call the individual tools for raw detail.
Args: domain: The domain to assess, e.g. "example.com". timeout: Per-connection network timeout in seconds.
Returns: A dict with domain, ip, overall_grade, summary, components (email / tls / headers, each with a grade and issues), a tech section (detected technologies + any version disclosure), and a takeover section when the apex is at risk. A check that errors is reported without breaking the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses several behavioral traits beyond what annotations provide (no annotations exist): it runs checks concurrently, returns a single graded overview, caps the overall grade at F if a live takeover is found, and reports errors without breaking the rest. It also explains the grading model ('as weak as the weakest component'). This is rich behavioral context for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear summary, an Args section, and a Returns section. It front-loads the core purpose and scope. It is slightly long but every sentence earns its place by explaining the grading model, the concurrency, and the error-handling behavior. The only minor inefficiency is the repetition of 'recon' and 'report' concepts, but overall it is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (aggregating multiple recon checks) and the absence of an output schema, the description is quite complete. It explains what the return dict contains, how grades are computed, and how errors are handled. It could add a note about rate limits or permission requirements, but for a read-only recon tool, the description covers the essential context an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'domain' as 'The domain to assess, e.g. "example.com"' and 'timeout' as 'Per-connection network timeout in seconds.' This adds meaning beyond the bare schema types, though it doesn't elaborate on timeout defaults or edge cases. The description also clarifies the return structure, which helps the agent understand what the parameters produce.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: a one-shot security posture report for a domain. It names the specific verb ('runs', 'returns'), the resource ('domain'), and the scope (DNS/email, TLS, HTTP-header, web-stack, apex subdomain-takeover). It also distinguishes itself from sibling tools by saying 'Use this for a quick full picture; call the individual tools for raw detail.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool ('for a quick full picture') and when not to ('call the individual tools for raw detail'). It also names the sibling tools implicitly by category (DNS/email, TLS, HTTP-header, tech_detect, subdomain_takeover), which helps the agent choose between this aggregate tool and the individual ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subdomain_enumA
Discover subdomains of a domain via DNS brute-force and/or CT logs.
Two complementary sources:
"dns": probe candidate labels with DNS A lookups (active but light, capped at 512 candidates). Returns resolved IPs.
"ct": query public Certificate Transparency logs (crt.sh) for every name ever certified for the domain — fully passive, and finds real hosts no wordlist would guess.
"both": run both and merge, marking which source saw each host.
Enumerate only domains you are authorized to assess.
Args: domain: The base domain, e.g. "example.com". wordlist: Comma-separated labels for the DNS source (e.g. "www,api,dev"). Omit to use a built-in list of common labels. Ignored for "ct". source: "dns" (default), "ct", or "both". timeout: Per-query DNS timeout in seconds (the CT query uses its own longer timeout since crt.sh can be slow).
Returns: A dict with domain, sources, found_count, and found (each with subdomain, the source(s) that saw it, and resolved ips when known).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | ||
| source | No | dns | |
| timeout | No | ||
| wordlist | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to rely on, the description fully carries behavioral disclosure. It explains active vs passive enumeration, the 512-candidate cap, the use of crt.sh, the timeout distinction for DNS vs CT queries, and how 'both' marks source provenance. This gives an agent realistic expectations about side effects and scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a concise summary, bullet-pointed source options, an authorization note, an Args section, and a Returns section. It is detailed but every sentence earns its place, and important operational constraints are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description thoughtfully includes a Returns section outlining the dictionary structure. Combined with thorough parameter semantics and source behavior, this is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate, and it does. Every parameter is explained: domain with an example, wordlist with comma-separated labels, source with enumerated valid values, and timeout with per-query semantics. It even clarifies that wordlist is ignored for ct, which is not inferable from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, action-oriented statement: 'Discover subdomains of a domain via DNS brute-force and/or CT logs.' This clearly distinguishes it from siblings like ip_info or port_scan, and even from dns_recon by focusing on subdomain enumeration rather than general DNS reconnaissance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use each source: dns is active but light, ct is fully passive and finds hosts wordlists miss, and both merges results. It also includes an authorization caveat. However, it does not explicitly name sibling tools or state when this tool should not be used in favor of another.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subdomain_takeoverA
Check subdomains for a dangling-CNAME takeover risk.
A subdomain is takeover-prone when it CNAMEs to a third-party service (GitHub Pages, S3, Heroku, Azure, ...) whose resource was deleted or never claimed — anyone who registers that resource then controls the subdomain. For each host this resolves the CNAME, recognizes known takeover-prone services, fetches the page, and flags the provider's "unclaimed resource" fingerprint and/or a CNAME target that no longer resolves.
Read-only recon (DNS lookups + one HTTP GET per host). Pair it with subdomain_enum: enumerate first, then pass the interesting hosts here. Only check domains you are authorized to assess.
Args: hosts: One hostname or a comma-separated list, e.g. "blog.example.com,shop.example.com". Capped at 100 per call. timeout: Per-probe network timeout in seconds.
Returns: A dict with checked, vulnerable_count, and results (one entry per host with host, cname, service, status, vulnerable, severity, and detail). status is one of not_applicable, not_vulnerable, potential, dangling_cname, or vulnerable.
| Name | Required | Description | Default |
|---|---|---|---|
| hosts | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It thoroughly discloses behavior: resolves CNAME, recognizes takeover-prone services, fetches the page, flags fingerprints, and notes it is read-only recon (DNS + one HTTP GET per host). It also mentions the host cap and the output structure, leaving little to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a lead sentence, a detailed explanation, usage pairing, and an Args section. It is slightly verbose but every sentence carries weight—explaining the method, safety, and output. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description is remarkably complete. It explains the output dict fields (checked, vulnerable_count, results) and the possible statuses. It also covers limitations (cap) and usage context. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates. It explains 'hosts' with format, example, and cap ('comma-separated list, e.g. ... Capped at 100 per call'), and 'timeout' as 'Per-probe network timeout in seconds.' This adds meaning far beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check subdomains for a dangling-CNAME takeover risk.' It clearly distinguishes itself from siblings like subdomain_enum (enumerates) and dns_recon (DNS lookups) by focusing on takeover risk. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs when to use this tool: 'Pair it with subdomain_enum: enumerate first, then pass the interesting hosts here.' It also states a condition for use: 'Only check domains you are authorized to assess.' This provides clear routing to the tool and its complementary sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tech_detectA
Fingerprint the technology stack behind a website from one HTTP GET.
Passively identifies the web server, reverse proxy / CDN, WAF, programming language, web framework, CMS, JavaScript framework, and analytics by matching response headers, set cookies, the HTML body, and the meta-generator tag against a signature table. Disclosed versions are captured and flagged (a precise version eases known-CVE lookup). One read-only HTTP GET.
Args: host: Hostname to fingerprint, e.g. "example.com". port: TCP port. Defaults to 443 when use_ssl is True, else 80. use_ssl: Connect over HTTPS (default True). timeout: Network timeout in seconds.
Returns: A dict with host, url, status, technology_count, technologies (each with name, category, version when known, and evidence), and a findings list noting any version disclosure. An error field on fetch failure.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| port | No | ||
| timeout | No | ||
| use_ssl | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and handles it well: it declares a single read-only HTTP GET, passive identification via response headers/cookies/body/meta tags, and version disclosure with CVE relevance. It does not discuss redirects, rate limits, or authorization, but the core safety profile and side effects of the tool are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: a one-sentence purpose, a behavior paragraph, then neatly separated Args and Returns sections. Each sentence provides useful information, with only the minor repetition of 'from one HTTP GET' and 'One read-only HTTP GET' as slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there are no annotations and no output schema, the description serves as the sole documentation and is thorough enough: it covers what the tool does, how it works, all parameters, return fields, and the error field on fetch failure. An agent has everything needed to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by explaining every parameter: host with an example, port with default behavior conditional on use_ssl, use_ssl with its default, and timeout in seconds. It also documents the return dict structure, which is entirely absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific action and resource: 'Fingerprint the technology stack behind a website from one HTTP GET.' The detailed list of detected categories (web server, CDN/WAF, language, framework, CMS, analytics) makes the tool's role clear, but it never explicitly names a sibling, so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys the appropriate context by emphasizing 'Passively identifies' and 'One read-only HTTP GET,' which suggests when this tool is safe and suitable. However, it gives no explicit when-to-use or when-not-to-use guidance and does not mention alternatives such as port_scan or http_headers_audit, leaving the agent to infer placement among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tls_checkA
Inspect a host's SSL/TLS configuration and grade it.
Checks the certificate (validity, expiry, key algorithm), supported protocol versions (flagging legacy SSLv3/TLS 1.0/1.1), cipher suites and forward secrecy, TLS compression, HSTS, OCSP stapling, and known protocol vulnerabilities. Returns a letter grade plus structured findings.
Args: host: Hostname or IP to inspect, e.g. "example.com". port: TLS port (default 443). timeout: Per-connection network timeout in seconds.
Returns: A structured dict with: grade, certificate, protocols, cipher info, forward_secrecy, hsts, vulnerabilities, and a findings list.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| port | No | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It clearly states what is checked (including flagging legacy protocols), what is returned (a grade and structured findings), and the fact that it is a read-only inspection. It does not mention any side effects or rate limits, but none are expected for a network inspection tool. The disclosure is thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured into clear sections (summary, checks, args, returns) and front-loads the primary purpose. It is somewhat lengthy due to the enumeration of checks, but every sentence adds value. It avoids fluff and is appropriately detailed for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that inspects many aspects of TLS, the description lists every check performed, describes the return dict structure, and explains all parameters. It is complete enough for an agent to call the tool correctly without needing additional context. The absence of an output schema is mitigated by the description's return format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% since the schema only provides types and defaults, but the description's 'Args' section fully explains each parameter: host with an example, port with default, and timeout with its purpose. This completely compensates for the schema gap, giving the agent all needed semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Inspect a host's SSL/TLS configuration and grade it.' It enumerates the exact checks performed (certificate, protocols, ciphers, HSTS, etc.), which clearly distinguishes it from sibling tools like port_scan or http_headers_audit. An agent can immediately understand the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for TLS configuration assessment but never explicitly states when to use it over alternatives, nor does it provide exclusion criteria. For example, it doesn't say 'use this when you need TLS details, not for general port reachability (use port_scan).' The guidance is implicit through the detailed checks, but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
well_known_auditA
Fetch and parse a host's security.txt and robots.txt.
Both are standard public files. security.txt (RFC 9116) gives the vulnerability-disclosure contact, policy, and encryption key; its absence is itself a finding for a security-conscious site. robots.txt lists the paths the operator asks crawlers to skip — frequently admin/internal areas worth noting during recon.
Args: host: Hostname to inspect, e.g. "example.com". timeout: Per-request network timeout in seconds.
Returns: A dict with host, security_txt (present flag, parsed fields, structural issues, location), and robots_txt (present flag, sitemaps, disallow/allow paths, user_agents).
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly succeeds: it discloses that both files are fetched and parsed, that absence of security.txt is itself a finding, and that return values include present flags and parsed fields. It does not cover edge behaviors like redirects or network errors, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, and each paragraph earns its place by adding context, argument details, or return structure. There is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and no annotations, the description is self-sufficient: it explains inputs, the full return dict shape, and domain-specific interpretation. An agent has enough information to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema description coverage is 0%, the Args section fully defines host with an example and timeout with units ('Per-request network timeout in seconds.'). This adds real meaning beyond the schema's bare property names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names exact verbs and resources: 'Fetch and parse a host's security.txt and robots.txt.' This is unambiguous and distinguishes the tool from sibling recon tools focused on headers, DNS, TLS, HTTP methods, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear use context by explaining that security.txt is evaluated as a security finding and robots.txt reveals 'admin/internal areas worth noting during recon.' It does not explicitly name alternatives or exclusions, but the intended recon use case is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.13.0- First observed
cookie_audit - First observed
cors_check - First observed
dns_recon - First observed
http_headers_audit - First observed
http_methods_audit - First observed
ip_info - First observed
port_scan - First observed
recon_report - First observed
subdomain_enum - First observed
subdomain_takeover - First observed
tech_detect - First observed
tls_check - First observed
well_known_audit
TDQS
Scored across 13 tools
Each tool maps to a distinct reconnaissance technique — IP ownership, port scanning, DNS records/WHOIS, TLS, HTTP security headers, cookies, CORS, HTTP methods, subdomain enumeration, takeover checks, tech detection, well-known files, and the aggregate report — so misselection risk is low. The several *_audit tools are clearly differentiated by target (headers, cookies, methods, well-known files).
Names are uniformly lower_snake_case and mostly follow a predictable `<target>_<operation>` pattern (port_scan, dns_recon, tls_check, cookie_audit). A few entries use noun-like suffixes (ip_info, recon_report, subdomain_takeover) rather than a strict verb, so the convention is consistent but not perfectly uniform.
Thirteen tools is well within the ideal size for a security-recon server; every tool covers a distinct phase or technique and none feels redundant. This scope comfortably supports both targeted checks and the aggregate one-shot report.
The server covers the main recon lifecycle: discovery, DNS/WHOIS/email checks, TLS and HTTP posture, subdomain enumeration/takeover, tech fingerprinting, and a summary report. It omits deeper optional techniques such as banner grabbing, directory brute-forcing, and web vulnerability scanning, but those are reasonable workarounds rather than blocking gaps.
Maintenance
Related MCP Connectors
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.
MCP server for building and testing AI agents with multi-model experimentation and insights.
AI-security knowledge as MCP: standards-mapped tools (OWASP, NIST, MITRE) for AI agents.
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server that lets AI agents use the Online Cyber Tools catalogue as a set of native MCP tools.1009 npm1MIT
- AlicenseAqualityBmaintenanceAn MCP server that provides passive and low-impact active reconnaissance tools for authorized bug bounty and security assessments, enabling LLMs to perform structured recon and generate reports.11Apache 2.0
- AlicenseCqualityCmaintenanceAI-powered security scanning MCP server that exposes 25+ professional tools, enabling penetration testing and security assessments through natural language interaction with AI agents like Claude Desktop.31MIT
- FlicenseNot gradedqualityCmaintenanceA production-style MCP server providing AI models with cybersecurity tools including port scanning, WHOIS, DNS, threat intelligence, CVE lookup, and more.-