WebDataTools Domain & website intelligence MCP server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@WebDataTools Domain & website intelligence MCP serverrun a security audit on stripe.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebDataTools Domain & website intelligence MCP server
webdatatools-domain-mcp
An MCP server with 11 domain & website intelligence tools for AI agents — Claude Desktop, Cursor, Cline or any MCP client. WHOIS/RDAP, DNS and e-mail security (SPF, DKIM, DMARC), TLS and security headers, subdomains, tech stack, contacts, SEO, Core Web Vitals, sitemaps and Wayback page diffs for any domain.
This server uses your own Apify API token. Every tool call runs a WebDataTools Actor under your Apify account and is billed to your Apify credit — pay per result, the price is in each tool description. Your token is only sent to Apify's API.
Quick start
Requires Node.js 18+.
APIFY_TOKEN=apify_api_... npx -y github:paulet4a-commits/webdatatools-domain-mcpGet a free token (the free plan includes monthly credit): https://console.apify.com/settings/integrations
Related MCP server: CyberMCP
Claude Desktop / Cursor
Add this to claude_desktop_config.json (Claude Desktop) or .cursor/mcp.json (Cursor):
{
"mcpServers": {
"webdatatools-domain": {
"command": "npx",
"args": [
"-y",
"github:paulet4a-commits/webdatatools-domain-mcp"
],
"env": {
"APIFY_TOKEN": "apify_api_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
}
}
}
}Tools (11)
Tool | What it does | Price (free plan) | Backing Actor |
| Email Extractor — Website Contact & Social Finder | $0.01 / result | |
| Tech Stack Detector — Wappalyzer & BuiltWith Alternative | $0.02 / result | |
| Domain DNS & Email Security Checker | $0.005 / result | |
| Domain Security Audit (TLS, HTTP headers, redirects, robots) | $0.005 / result | |
| Subdomain Finder (Certificate Transparency) | $0.0005 / result | |
| Bulk Core Web Vitals & PageSpeed Audit | $0.005 / result | |
| On-Page SEO Audit | $0.002 / result | |
| Sitemap URL Extractor & Change Monitor | $0.0002 / result | |
| Wayback Machine Snapshot & Page Change Tracker | $0.003 / result | |
| Bulk Domain WHOIS & RDAP Lookup | $0.003 / Domain | |
| Web Scraper — CSS Selector & Data Extractor | $0.002 / Page |
More WebDataTools MCP servers
webdatatools-mcp-server — the 10 most popular tools in one server
webdatatools-rag-mcp — Web content for AI & RAG
webdatatools-social-mcp — Search, video & social data
webdatatools-leads-mcp — Leads, jobs & company data
webdatatools-dev-mcp — Developer, app & research data
License
MIT
Available Tools
11 toolscontact_extractorA
Email Extractor finds every e-mail address, phone number and social profile on a website — one clean row per domain, ready for outreach. Billed to your own Apify account: ~$0.01 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| maxDepth | No | Max link depth — Enter how many links deep to follow from the start URL, e.g. 2. Set 0 to crawl only the start page. | |
| startUrls | Yes | Websites — Enter the websites to crawl, one row is returned per website. Use bare domains or full URLs, e.g. example.com or https://example.com. Example: [{"url":"https://apify.com"}]. | |
| extractEmails | No | Extract e-mail addresses — Turn this on to collect e-mail addresses from mailto: links and page source, filtered through a false-positive blocklist, e.g. info@example.com. | |
| extractPhones | No | Extract phone numbers — Turn this on to collect phone numbers from tel: links only, e.g. +902120000000, which avoids the false positives you get from scraping raw text. | |
| extractSocials | No | Extract social profiles — Turn this on to collect social profile links, e.g. https://linkedin.com/company/example, covering LinkedIn, X/Twitter, Instagram, Facebook, YouTube, TikTok, GitHub, Telegram and WhatsApp. | |
| followSubdomains | No | Follow subdomains — Turn this on to also crawl subdomains such as blog.example.com and shop.example.com, not just the main hostname. | |
| maxPagesPerDomain | No | Max pages per website — Enter how many pages to crawl per website before stopping, e.g. 30. Contact details usually live on the home, about and contact pages, so 20-50 is plenty. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses the billing model and per-result cost (~$0.01, charged to the caller's own Apify account), which is real value beyond structured fields, but it says nothing about runtime, crawl politeness/robots handling, rate limits, or failure behavior for a multi-page scraper.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler: capability and output shape are front-loaded, and the cost caveat follows. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the essential return shape (one row per domain) and the cost model, which is enough to call the tool correctly given 100% schema coverage on all 7 parameters. Only operational details (runtime, crawl limits in practice) are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter already has a detailed description with defaults, ranges and examples, so the schema does the heavy lifting. The description adds no parameter-level meaning beyond the crawl/output summary, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb (finds) and specific resources (e-mail addresses, phone numbers, social profiles) scoped to a website, plus the output shape ('one clean row per domain'). It does not explicitly differentiate itself from the nearest sibling, dns_email_security_checker, so an agent must infer that this one scrapes pages while the other inspects DNS records.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via 'ready for outreach', which signals a lead-generation context but gives no explicit when-to-use, when-not-to-use, or alternative-selection rule against the other ten sibling tools. Nothing tells the agent when a DNS/security checker would be the better choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
core_web_vitals_auditA
Bulk Core Web Vitals and PageSpeed audit: Lighthouse performance, SEO and accessibility scores, LCP, CLS, INP, TBT and the top fixes — one row per URL and device. Billed to your own Apify account: ~$0.005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs to audit — Enter the page URLs to audit, one per row, e.g. https://apify.com. Each URL returns one row per device strategy. Audit the exact URL you care about (a landing page, not only the homepage) - PageSpeed Insights follows redirects but scores the page it lands on. Example: ["https://apify.com"]. | |
| locale | No | Report locale — Enter the language for Lighthouse titles and display values, e.g. en, de or tr. Only the wording of titles changes, never the metrics. | en |
| strategy | No | Device strategy — Select the device Lighthouse should emulate, e.g. mobile. Google ranks on mobile, so audit mobile unless you specifically need desktop numbers. Choose 'Both' to get one row per device (two billable rows per URL). Options: mobile = Mobile (recommended); desktop = Desktop; both = Both (2 rows per URL). | mobile |
| categories | No | Lighthouse categories — Select which Lighthouse categories to score, e.g. performance, seo. Every extra category makes the Google request slower; performance alone already returns all Core Web Vitals and the optimisation opportunities. Options: performance = Performance; accessibility = Accessibility; best-practices = Best practices; seo = SEO. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so partially: it discloses the billing model (~$0.005 per result, charged to the user's own Apify account, free-plan price) and the row multiplicity per URL/device. It omits execution traits such as expected runtime, rate limits, or failure behavior on unreachable URLs, which matters for a bulk external fetch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and dense — the audit scope and its outputs come first, with the cost caveat as a short second sentence. The first sentence is a long comma-run list of metric acronyms, which is slightly heavy but each element earns its place by telling the agent what results to expect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema tool, the description supplies the crucial return shape (one row per URL and device) and the billing implication of 'both'. Parameters are fully covered by the schema, so the only meaningful gap is the absence of runtime/rate-limit expectations for bulk audits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already richly documented with examples, defaults, enums and rationale. The tool description adds no parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — a bulk Core Web Vitals / PageSpeed audit via Lighthouse — and enumerates exactly the metrics returned (LCP, CLS, INP, TBT, top fixes). The scope 'one row per URL and device' distinguishes it from the sibling seo_page_audit by making clear this is a per-URL performance audit rather than a general SEO crawl.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool's purpose but gives no explicit when-to-use vs when-not guidance and never names or contrasts with siblings such as seo_page_audit. The substantive usage advice (prefer mobile, performance category returns everything, 'both' doubles rows/billing) lives in the schema property descriptions rather than the tool description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
css_selector_extractorA
Web Scraper pulls any CSS selector off any page, returning text, HTML, attributes or match counts — one row per URL, no browser required. Billed to your own Apify account: ~$0.002 per Page (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs — Enter one URL per row to fetch and run the selectors against, e.g. https://example.com. Each URL gets exactly one result row, even if the fetch fails. Example: ["https://example.com"]. | |
| selectors | Yes | Selectors — Enter the fields to extract as a JSON array of {"name", "selector", "type"} objects. "type" is one of text, html, attr or count; an "attr" entry also needs "attribute" (e.g. {"name":"canonicalUrl","selector":"link[rel='canonical']","type":"attr","attribute":"href"}). Capped at 25 entries per run. Example: [{"name":"title","selector":"h1","type":"text"}]. | |
| firstMatchOnly | No | First match only — Keep this on to return only the first element each selector matches. Turn it off to return every match (up to Max matches per selector) as an array. | |
| trimWhitespace | No | Trim whitespace — Keep this on to collapse runs of whitespace and trim the ends of extracted text and attribute values. Turning it off returns text and attr values exactly as found in the page. | |
| maxMatchesPerSelector | No | Max matches per selector — Enter the most matches to return for one selector when First match only is off, e.g. 20. Does not limit the reported match count, only how many values are returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure. It adds useful billing and execution context (charged to the user's Apify account, no browser required) and notes one row per URL, but omits failure behavior, rate limits, and explicit read-only safety details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences front-load the core purpose and then add pricing context. Every sentence earns its place by clarifying capability, output shape, or cost, with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does well to describe return types and row structure, and it covers billing and browserless execution. It could be more complete by explaining error/failure output format or authentication specifics, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all five parameters. The description adds no parameter-specific syntax or constraints beyond what is already in the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: pulling CSS selectors off pages, with return types (text, HTML, attributes, match counts) and deployment context (no browser required). It is clear but does not explicitly differentiate this tool from sibling scrapers like contact_extractor or tech_stack_detector.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'pulls any CSS selector off any page', but there is no explicit guidance on when to prefer this tool over alternatives, nor any exclusions. The description gives enough context to infer it is for custom selector-based extraction, but lacks when/when-not instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dns_email_security_checkerA
Domain DNS & Email Security Checker returns SPF, DKIM, DMARC, MX, MTA-STS, BIMI, CAA, nameservers, registrar and domain age for every domain you give it — one scored row per domain. Billed to your own Apify account: ~$0.005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| domains | Yes | Domains — Enter the domains to check, one row is returned per domain, e.g. apify.com. Full URLs and www. prefixes are accepted and stripped automatically (https://www.apify.com/store becomes apify.com), so you can paste a list straight out of your CRM. Example: ["apify.com"]. | |
| checkDkim | No | Check DKIM selectors — Keep this on to probe common DKIM selectors such as google and selector1. DKIM has no discoverable record name, so it can only be found by guessing selectors; turning this off makes each domain 16 DNS queries cheaper but drops 15 points from the score. | |
| dkimSelectors | No | DKIM selectors to probe — Enter the DKIM selectors to try, e.g. google, selector1, k1. Each one is looked up as <selector>._domainkey.<domain>. Leave the defaults unless you know the provider your targets use; a shorter list means fewer DNS queries per domain. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose genuinely useful behavioral facts: results are scored, and cost is billed to the caller's own Apify account at ~$0.005 per result with lower paid-plan pricing. It says nothing about how non-resolving domains, timeouts, or partial failures are reported, which leaves real gaps for a network-probing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The output enumeration is front-loaded in sentence one and the pricing caveat is deferred to sentence two, so an agent reads the capability before the commercial detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by enumerating the returned record types and stating the one-row-per-domain shape. Parameters are fully covered by the schema, so nothing needed for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters in depth (DKIM selector mechanics, cost trade-off of disabling checkDkim). The description adds no parameter detail beyond the general notion of supplying domains, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource — it 'returns SPF, DKIM, DMARC, MX, MTA-STS, BIMI, CAA, nameservers, registrar and domain age' for each supplied domain, plus a scored row. That is far more concrete than the tool name alone, though it never distinguishes itself from the sibling domain_security_audit, which appears to overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'for every domain you give it — one scored row per domain', which makes the batch-report pattern obvious. There is no explicit when-to-use statement and no routing away from the overlapping domain_security_audit sibling, so an agent must infer the choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_security_auditA
Domain Security Audit checks TLS certificate validity, HSTS/CSP and other HTTP security headers, the redirect chain and robots.txt/llms.txt AI-bot rules for any list of domains — one scored, graded row per domain. Billed to your own Apify account: ~$0.005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| domains | Yes | Domains — Enter the domains to audit, one row is returned per domain, e.g. apify.com. Full URLs and www. prefixes are accepted and stripped automatically (https://www.apify.com/store becomes apify.com), so you can paste a list straight out of your CRM. Example: ["apify.com"]. | |
| checkTls | No | Check TLS certificate — Connect on port 443 and read the peer certificate's validity dates, issuer, subject, SAN count, days until expiry and whether the chain is trusted. Turn off to skip this and save a few seconds per domain. | |
| checkRobots | No | Check robots.txt and llms.txt — Fetch /robots.txt (disallow-all detection, sitemap count, and which AI crawlers such as GPTBot, ClaudeBot and Google-Extended are blocked) and /llms.txt (presence, title, size). | |
| checkHeaders | No | Check HTTP security headers — Read HSTS, Content-Security-Policy, X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, Server, X-Powered-By, Cache-Control and cookie flags from the final response after redirects. | |
| followRedirects | No | Follow redirect chain — Start at http://<domain>/ and follow Location headers manually (up to 10 hops) to find the final URL, whether it upgrades to HTTPS and whether it adds or drops the www. prefix. Turn off to check the headers of http://<domain>/ directly with no redirect following. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose meaningful operational traits: billing to the user's own Apify account at ~$0.005 per result, acceptance/normalization of full URLs and www prefixes, and that disabling checks saves seconds per domain. It stops short of describing auth requirements or how the score/grade is computed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense capability sentence plus one short billing sentence; the purpose and scope are front-loaded and nothing is wasted. The first sentence is packed but remains readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description gives a rough return shape ('one scored, graded row per domain') plus cost, which is enough for an agent to decide and invoke. It does not explain what fields or scale the score/grade uses, a minor gap but not disqualifying.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters with defaults and behavior. The description adds only marginal value here (cost framing, per-domain row), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (checks) and enumerates the exact resources examined: TLS certificate validity, HSTS/CSP and other security headers, redirect chain, and robots.txt/llms.txt AI-bot rules. It also states the output shape ('one scored, graded row per domain'), which cleanly separates it from siblings like dns_email_security_checker and seo_page_audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumeration of checks, so an agent can infer the tool applies to bulk domain security posture assessment. However, it never states when to prefer this over siblings (e.g. dns_email_security_checker for DNS/email, seo_page_audit for on-page SEO) nor any explicit when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_whois_rdapA
Bulk domain WHOIS lookup via the official RDAP protocol — registrar, registration/expiry dates, domain age, nameservers, DNSSEC and status codes for any domain, no API key, no rate limits. Billed to your own Apify account: ~$0.003 per Domain (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| domains | Yes | Domains — Enter one domain per line, e.g. github.com. A full URL such as https://github.com/apify also works — the path, protocol and www. prefix are stripped automatically. Example: ["github.com"]. | |
| includeRawEvents | No | Include raw RDAP events — Keep this on to also store the raw, unparsed RDAP events array (registration, last-changed, expiration, transfer, etc.) alongside the parsed date fields, useful when a registry publishes an event type this Actor does not map to a named field. | |
| includeNameservers | No | Include nameservers — Keep this on to list every nameserver RDAP returns for the domain in the nameservers field. Turn it off to skip that field and leave it null, which slightly shrinks the output. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does unusually well: it discloses the protocol (official RDAP), the absence of API keys/rate limits, and the billing model (~$0.003 per domain, charged to the caller's Apify account). It does not describe failure behavior for malformed or unregistered domains, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence pairs the action and the payload fields before appending cost/limit facts. No filler, no restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, and the description compensates by enumerating the returned fields and the billing model. It stops short of explaining error handling or output mass/volume expectations for large domain lists, which would fully close the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters in depth (URL/path stripping, raw events, nameserver toggling). The description adds no parameter-level meaning beyond what the schema provides, making baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (bulk lookup), resource (domain WHOIS/RDAP registration data), and enumerates the exact fields returned (registrar, registration/expiry dates, domain age, nameservers, DNSSEC, status codes). This is clearly distinguishable from siblings like subdomain_finder or domain_security_audit, which cover different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (fetch registration data for domains in bulk) and the operational constraints ('no API key, no rate limits') are stated, but there is no explicit when-to-use vs. sibling guidance — e.g. no pointer that domain_security_audit or dns_email_security_checker are alternatives when the goal is security rather than registration metadata.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seo_page_auditA
On-Page SEO Audit checks title, meta description, headings, images, links, Open Graph, Twitter Card and Schema.org data for any list of URLs — one scored row per page. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs — Enter the page URLs to audit, one row is returned per URL, e.g. https://apify.com. A bare domain like apify.com is accepted and gets https:// added automatically. Example: ["https://apify.com"]. | |
| checkLinks | No | Check internal links for broken ones — Keep this off for a fast audit. Turn it on to HEAD-check up to 50 internal links per page and list the broken ones in the `brokenLinks` field — this is slower and adds extra requests per page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden, and it does add real value: per-result billing (~$0.002, Apify account) and the cost/speed tradeoff of checkLinks (extra HEAD requests, up to 50 links per page). It still omits auth requirements, failure behavior, and rate limits, so the safety/operational picture is incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero padding: capability and output shape first, billing second. Every clause earns its place and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, no-output-schema tool, the description covers capability, output granularity, cost, and the checkLinks tradeoff — enough for an agent to call it correctly. Only minor gaps remain, such as what a failure or empty result looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are fully documented in the schema itself (URL normalization rules, checkLinks semantics). The description adds no syntax or format detail beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource (On-Page SEO Audit) and enumerates exactly what it inspects — title, meta, headings, images, links, OG, Twitter Card, Schema.org — with the output shape stated as 'one scored row per page'. It is clearly distinguishable from siblings like core_web_vitals_audit and tech_stack_detector.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool fits (on-page SEO element auditing for a list of URLs) and notes that checkLinks should stay off for a fast audit, but it never states when to choose it over siblings or any exclusion conditions. Usage is inferred rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sitemap_extractorA
Sitemap URL extractor that reads robots.txt, sitemap indexes, .xml.gz and plain-text sitemaps and returns one row per URL with lastmod, changefreq, priority — plus a new/removed diff between runs. Billed to your own Apify account: ~$0.0002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode — Choose list to dump every URL found right now, or diff to compare this run with the previous one and label each URL new, unchanged or removed. Diff mode needs at least two runs of the same source to be useful. Options: list = list — every URL in the sitemap; diff = diff — new / removed / unchanged since the last run. | list |
| sources | Yes | Sitemaps or domains — Enter the sitemaps to read, or just the domains, e.g. https://apify.com or https://www.allbirds.com/sitemap.xml. For a bare domain the Actor reads the Sitemap: lines of /robots.txt and falls back to /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /sitemap.xml.gz and /sitemap.txt. Sitemap indexes are followed into their child sitemaps automatically. Example: ["https://apify.com"]. | |
| urlFilter | No | URL filter (regex) — Optional JavaScript regular expression; only URLs matching it are kept, e.g. /products/ for a Shopify catalogue or \.pdf$ for documents. Leave empty to keep every URL. The pattern is matched against the full URL and is case-sensitive, so write [Pp]roducts when you need both cases. | |
| maxUrlsPerSource | No | Max URLs per source — Enter how many URLs to keep per source, e.g. 5000. Counted across all child sitemaps of an index, so a 200,000-URL e-commerce site stops as soon as the cap is reached. Each URL is one billed dataset row. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well on several fronts: it discloses the billing model and approximate per-result cost on the user's own Apify account, the output shape (one row per URL with lastmod/changefreq/priority), and the cross-run diff behavior. It omits auth prerequisites beyond the Apify account mention and any rate-limit details, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that front-load the core capability, then add scope and pricing. The pricing sentence is extra but genuinely useful for a billed Apify tool; little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by naming the returned fields (lastmod, changefreq, priority) and the diff labels. Combined with the fully documented input schema, an agent has enough to call it correctly, though auth and pagination behavior remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the property descriptions are already rich (mode enum, source fallbacks, regex case-sensitivity, per-source cap with billing note), so the baseline of 3 applies. The top-level description adds nothing about parameters beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Sitemap URL extractor that reads robots.txt, sitemap indexes, .xml.gz and plain-text sitemaps') and specifies the output granularity ('one row per URL'). The sitemap-URL scope is clearly distinct from the SEO/diff/whois siblings without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are named relative to the sibling extractors. The mention that 'diff mode needs at least two runs of the same source to be useful' is a usage condition, but it lives in the schema description rather than the tool description and no exclusions or sibling routing are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subdomain_finderA
Subdomain Finder enumerates every subdomain of a domain from Certificate Transparency logs (crt.sh, Cert Spotter) and resolves each one — one row per subdomain, no proxies needed. Billed to your own Apify account: ~$0.0005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| domains | Yes | Domains — Enter the root domains to enumerate subdomains for, e.g. apify.com. Bare domains and full URLs both work — https://www.apify.com/pricing is reduced to apify.com. One row is returned per subdomain found. Example: ["apify.com"]. | |
| resolveDns | No | Resolve DNS — Turn this on to look up an A record for every subdomain found, so you can tell live hosts from dead certificate records. Costs one extra DNS query per subdomain and fills in resolves, ipv4 and cname. Turn it off for a pure certificate list. | |
| includeWildcards | No | Include wildcard names — Turn this on to also return wildcard certificate names such as *.example.com. They cannot be resolved, so they are excluded by default. | |
| maxSubdomainsPerDomain | No | Max subdomains per domain — Enter the maximum number of subdomains to return per domain, e.g. 500. Big brands have thousands of certificate names; when the cap is hit the subdomains with the most recently issued certificates are kept. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does disclose meaningful traits: the two upstream CT sources, the one-row-per-subdomain contract, that no proxies are needed, and the cost model (~$0.0005 per result billed to the caller's own Apify account). It omits rate-limit/pagination behavior and any auth prerequisite detail beyond the Apify billing note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the core action and data sources, with pricing relegated to the end. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description supplies the important context an agent needs: sources, row granularity, resolution semantics, and cost. Minor gaps remain around pagination/limits at the API level, but an agent can call this correctly from what is given.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters with defaults, ranges, and examples. The description adds no syntax or format information beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('enumerates every subdomain of a domain') plus the data sources (Certificate Transparency logs from crt.sh and Cert Spotter) and the output shape (one row per subdomain, DNS-resolved). This clearly separates it from siblings like dns_email_security_checker or domain_security_audit, which audit rather than enumerate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description (CT-log enumeration, optional DNS resolution) and the resolveDns flag is framed as 'turn it off for a pure certificate list', which is a mild when-to-use hint. However, no sibling alternative is named and no explicit when-not condition is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tech_stack_detectorA
Tech Stack Detector finds the CMS, e-commerce platform, analytics, ad pixels, chat widget and payment provider any website runs — one row per domain. Billed to your own Apify account: ~$0.02 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| startUrls | Yes | Websites — Enter the websites to fingerprint, one row is returned per website. Use bare domains or full URLs, e.g. example.com or https://example.com. Example: [{"url":"https://apify.com"}]. | |
| categories | No | Technology categories — Select which technology categories to detect, e.g. cms, ecommerce. Leave empty to detect everything; narrowing the list makes runs cheaper and keeps the output to the columns you care about. Options: cms = CMS; ecommerce = E-commerce platform; analytics = Analytics; advertising = Advertising & retargeting pixels; emailMarketing = E-mail & marketing automation; chat = Chat & support widgets; payments = Payment providers; frameworks = Front-end frameworks; cookieConsent = Cookie consent; cdn = CDN & hosting; server = Web server. | |
| followSubdomains | No | Follow subdomains — Turn this on to also inspect subdomains such as shop.example.com, which often run a different platform than the main site. | |
| maxPagesPerDomain | No | Max pages per website — Enter how many pages to fetch per website, e.g. 3. The homepage exposes most of the stack; 2-4 pages also catches checkout and marketing scripts that only load on shop or contact pages. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that runs are billed to the caller's own Apify account at roughly $0.02 per result, which implies an Apify account/token prerequisite and gives cost predictability. It stops short of explaining execution behavior (e.g. asynchronous actor runs, result delivery, rate limits) or whether the scrape is non-destructive, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the capability statement is front-loaded, and the billing note follows as secondary but relevant operational context. No filler or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four simple parameters, full schema coverage and no output schema, the description covers purpose, output shape ('one row per domain') and cost. It would be fully complete if it hinted at the returned columns or the execution model, but nothing essential for calling the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter fully documented in the schema including examples, defaults, ranges and enum meanings, so the description is not required to compensate. The description adds no parameter detail beyond the schema — it only echoes the result-row semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it fingerprints a website's CMS, e-commerce, analytics, ad pixels, chat widget and payment providers, with 'one row per domain' defining the output granularity. This is clearly distinct from all siblings (contact_extractor, dns_email_security_checker, seo_page_audit), none of which detect installed technologies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the tool's narrow purpose, and the description notes that narrowing `categories` makes runs cheaper, which is useful operational guidance. However, it never says when to prefer this over a sibling such as core_web_vitals_audit or seo_page_audit, and gives no exclusions or prerequisites beyond the billing note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wayback_page_diffA
Wayback Machine Snapshot & Page Change Tracker lists every Internet Archive capture of a URL and diffs the visible text of the oldest vs. newest snapshot in range — one scored change row per URL, or a full snapshot list. Billed to your own Apify account: ~$0.003 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | To date — Only consider captures on or before this date, e.g. 2024-06-30 or 20240630. Leave empty for no upper bound. | |
| from | No | From date — Only consider captures on or after this date, e.g. 2023-01-01 or 20230101. Leave empty for no lower bound. | |
| mode | No | Mode — "Diff" compares the oldest and newest archived snapshot in range and returns what changed. "Snapshots" lists every capture Wayback has for the URL, newest first, with no comparison. Options: diff = Diff (compare oldest vs. newest capture); snapshots = Snapshots (list all captures). | diff |
| urls | Yes | URLs — Enter the page URLs to look up in the Wayback Machine, e.g. https://apify.com/pricing. Bare domains are accepted too. One row (diff mode) or up to Max snapshots per URL rows (snapshots mode) comes back per URL. Example: ["https://apify.com/pricing"]. | |
| maxSnapshots | No | Max snapshots per URL — Snapshots mode only: the most captures returned per URL (newest first). Ignored in diff mode, which always emits exactly one row per URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningfully discharge it: it discloses that the run is billed to the caller's own Apify account at roughly $0.003 per result, and that diff mode always emits exactly one row per URL while snapshots mode emits many. It does not cover rate limits, failure handling for URLs with no captures, or what the change score represents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, and the second sentence carries only the cost disclosure, which is genuinely decision-relevant. The first sentence is long and packs three ideas (listing, diffing, output cardinality), slightly hurting scanability, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, no-output-schema tool, the description covers what comes back per mode and the cost model, which is the main missing structured information. Remaining gaps are minor: sorting of diff rows and the meaning/range of the change score are unstated, and there is no note on behavior when Wayback has no captures.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the schema already documents from/to, mode, urls, and maxSnapshots in detail including format examples and enum meanings. The description adds no parameter-level detail beyond the schema, which is acceptable but earns no credit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource set: it lists Internet Archive captures of a URL and diffs the visible text of the oldest vs. newest snapshot in range. The dual output shape ('one scored change row per URL, or a full snapshot list') makes the tool's two behaviors unambiguous, and no sibling tool overlaps with Wayback/archival work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated: the mention of a 'scored change row' versus a 'full snapshot list' hints at tracking changes vs. auditing history, but the description never says when to pick this tool, nor does it name any alternative or exclusion criteria. With no close siblings, the omission is low-risk but still a gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
contact_extractor - First observed
core_web_vitals_audit - First observed
css_selector_extractor - First observed
dns_email_security_checker - First observed
domain_security_audit - First observed
domain_whois_rdap - First observed
seo_page_audit - First observed
sitemap_extractor - First observed
subdomain_finder - First observed
tech_stack_detector - First observed
wayback_page_diff
TDQS
Scored across 11 tools
Each tool targets a distinct website or domain intelligence aspect: contacts, tech stack, DNS/email security, HTTP/TLS security, subdomains, Core Web Vitals, on-page SEO, sitemaps, archive diffs, WHOIS, and custom CSS scraping. However dns_email_security_checker, domain_whois_rdap, and domain_security_audit all return domain-level registration/security metadata, so generic domain lookup requests could cause minor misselection.
All tool names use lower snake_case and form descriptive noun phrases with stable suffixes such as -extractor, -detector, -checker, -audit, and -finder. There are no mixed conventions or vague one-word names.
The server has 11 tools, which is well within a focused range for domain and website intelligence. Each tool represents a distinct capability and none appears redundant or out of scope.
The surface covers contacts, technology detection, DNS/email security, TLS/header security, subdomains, performance, SEO, sitemaps, Wayback diffs, WHOIS, and custom CSS scraping. Minor gaps remain for full-site crawling/link graph analysis, backlink or traffic intelligence, and integrated site-wide content extraction.
Maintenance
Related MCP Connectors
Domain intel for AI agents: RDAP registration, DNS, email deliverability, tech stack.
Domain intel for AI agents: RDAP registration, DNS, email deliverability, tech stack.
Live web checks for AI agents: sitemaps, robots.txt, URL status, broken links, feeds, citations.
20 domain recon tools for AI agents: DNS, SSL, headers, email, subdomains, lookalikes, changes.
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server for domain intelligence — WHOIS, DNS records, SSL certificate inspection, SPF/DMARC validation, security-header audits, and blacklist/reputation checks, callable by AI agents. Powered by domainintel.app; runs server-side, no local setup.746 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to perform cybersecurity analysis including RDAP lookup, DNS analysis, SSL inspection, security header detection, and more, returning structured security reports.-
- AlicenseAqualityBmaintenanceProvides tools for AI agents to audit websites, including stack detection, DNS snapshots, and security checks.817 npmMIT
- AlicenseAqualityBmaintenanceEnables AI agents to perform passive security scans on domains, checking email spoofing (DMARC/SPF/DKIM), TLS weaknesses, security headers, exposed files, and subdomain-takeover risk without needing an API key.777 npmMIT