Skip to main content
Glama
paulet4a-commits

WebDataTools Domain & website intelligence MCP server

contact_extractor

Crawl websites to extract emails, phone numbers, and social profiles into one clean row per domain for outreach.

Instructions

Email Extractor finds every e-mail address, phone number and social profile on a website — one clean row per domain, ready for outreach. Billed to your own Apify account: ~$0.01 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
maxDepthNoMax link depth — Enter how many links deep to follow from the start URL, e.g. 2. Set 0 to crawl only the start page.
startUrlsYesWebsites — Enter the websites to crawl, one row is returned per website. Use bare domains or full URLs, e.g. example.com or https://example.com. Example: [{"url":"https://apify.com"}].
extractEmailsNoExtract e-mail addresses — Turn this on to collect e-mail addresses from mailto: links and page source, filtered through a false-positive blocklist, e.g. info@example.com.
extractPhonesNoExtract phone numbers — Turn this on to collect phone numbers from tel: links only, e.g. +902120000000, which avoids the false positives you get from scraping raw text.
extractSocialsNoExtract social profiles — Turn this on to collect social profile links, e.g. https://linkedin.com/company/example, covering LinkedIn, X/Twitter, Instagram, Facebook, YouTube, TikTok, GitHub, Telegram and WhatsApp.
followSubdomainsNoFollow subdomains — Turn this on to also crawl subdomains such as blog.example.com and shop.example.com, not just the main hostname.
maxPagesPerDomainNoMax pages per website — Enter how many pages to crawl per website before stopping, e.g. 30. Contact details usually live on the home, about and contact pages, so 20-50 is plenty.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses the billing model and per-result cost (~$0.01, charged to the caller's own Apify account), which is real value beyond structured fields, but it says nothing about runtime, crawl politeness/robots handling, rate limits, or failure behavior for a multi-page scraper.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler: capability and output shape are front-loaded, and the cost caveat follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the essential return shape (one row per domain) and the cost model, which is enough to call the tool correctly given 100% schema coverage on all 7 parameters. Only operational details (runtime, crawl limits in practice) are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter already has a detailed description with defaults, ranges and examples, so the schema does the heavy lifting. The description adds no parameter-level meaning beyond the crawl/output summary, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete verb (finds) and specific resources (e-mail addresses, phone numbers, social profiles) scoped to a website, plus the output shape ('one clean row per domain'). It does not explicitly differentiate itself from the nearest sibling, dns_email_security_checker, so an agent must infer that this one scrapes pages while the other inspects DNS records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied via 'ready for outreach', which signals a lead-generation context but gives no explicit when-to-use, when-not-to-use, or alternative-selection rule against the other ten sibling tools. Nothing tells the agent when a DNS/security checker would be the better choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.