webdatatools-leads-mcp
The MCP server is distributed via GitHub and its tools are backed by Apify Actors, but it does not provide integration with GitHub's own APIs or services; only installation/distribution.
Provides hiring signals data by scraping Greenhouse job boards, enabling AI agents to detect hiring activity from companies using Greenhouse ATS.
Extracts OpenStreetMap points of interest (shops, amenities, etc.) via the Overpass API, allowing AI agents to retrieve location-based POI data.
Enriches entities and companies with Wikidata facts, IDs, and links, enabling AI agents to retrieve structured knowledge from Wikidata.
Scrapes Y Combinator company and founder data, allowing AI agents to retrieve information about YC-backed startups and their founders.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@webdatatools-leads-mcpget the full company profile for stripe.com and check their hiring signals"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebDataTools Leads, jobs & company data MCP server
webdatatools-leads-mcp
An MCP server with 8 leads, jobs & company data tools for AI agents — Claude Desktop, Cursor, Cline or any MCP client. Company profiles from a domain, hiring signals from ATS boards, Y Combinator companies, Wikidata enrichment, e-mail validation, OpenStreetMap places, market quotes and remote jobs.
This server uses your own Apify API token. Every tool call runs a WebDataTools Actor under your Apify account and is billed to your Apify credit — pay per result, the price is in each tool description. Your token is only sent to Apify's API.
Quick start
Requires Node.js 18+.
APIFY_TOKEN=apify_api_... npx -y github:paulet4a-commits/webdatatools-leads-mcpGet a free token (the free plan includes monthly credit): https://console.apify.com/settings/integrations
Related MCP server: apollo-io-mcp-server
Claude Desktop / Cursor
Add this to claude_desktop_config.json (Claude Desktop) or .cursor/mcp.json (Cursor):
{
"mcpServers": {
"webdatatools-leads": {
"command": "npx",
"args": [
"-y",
"github:paulet4a-commits/webdatatools-leads-mcp"
],
"env": {
"APIFY_TOKEN": "apify_api_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
}
}
}
}Tools (8)
Tool | What it does | Price (free plan) | Backing Actor |
| Company 360: full company profile from a domain | $0.05 / company | |
| Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) | $0.002 / result | |
| Y Combinator Companies & Founders Scraper | $0.002 / result | |
| Wikidata Entity & Company Enrichment (facts, IDs, links) | $0.001 / result | |
| Email Validator & Verifier — Bulk Email Check | $0.0005 / email | |
| OpenStreetMap POI Extractor (Overpass API: shops, amenities) | $0.0005 / place | |
| Stock, Crypto & FX Quotes | $0.001 / Quote | |
| Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) | $0.001 / Job |
More WebDataTools MCP servers
webdatatools-mcp-server — the 10 most popular tools in one server
webdatatools-domain-mcp — Domain & website intelligence
webdatatools-rag-mcp — Web content for AI & RAG
webdatatools-social-mcp — Search, video & social data
webdatatools-dev-mcp — Developer, app & research data
License
MIT
Available Tools
8 toolscompany_360A
Company 360 turns a domain into one rich profile row: contacts, tech stack, DNS/e-mail security, TLS/security headers, on-page SEO, hiring signals and company facts (HQ, employees, revenue, founders) merged from our own scraper suite. Billed to your own Apify account: ~$0.05 per company (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| domains | Yes | Domains — Enter the company domains to profile, one row is returned per domain, e.g. apify.com. Full URLs and www. prefixes are accepted and stripped automatically, so you can paste a list straight out of your CRM. Example: ["apify.com"]. | |
| includeSeo | No | Include SEO audit — Call our On-Page SEO Audit Actor for the home page's 0-100 SEO score and a list of issues found. | |
| includeHiring | No | Include hiring signals — Call our Hiring Signals Actor for open job counts and hiring velocity. Only returns data when the domain's ATS job board could be found under its own name (e.g. a Greenhouse/Lever/Ashby board named after the company) - most domains will not match and get a warning instead, not an error. | |
| includeContacts | No | Include contacts — Fetch the homepage (and a common contact-page guess) for e-mail addresses, phone numbers and social-profile links. | |
| includeTechStack | No | Include tech stack — Call our Tech Stack Detector Actor for CMS, e-commerce platform, analytics, ad pixels, chat widget and payment providers. | |
| includeCompanyFacts | No | Include company facts — Call our Wikidata Entity Enrichment Actor for headquarters, employee count, revenue, founders, founding date and Crunchbase/LinkedIn IDs. Works best for companies with a Wikidata entry. | |
| includeEmailSecurity | No | Include DNS & e-mail security — Look up SPF/DMARC policy, mail provider, DNS provider, registrar and domain age directly (no sub-Actor call). | |
| includeSecurityAudit | No | Include security audit — Call our Domain Security Audit Actor for TLS certificate validity/expiry and a 0-100 security grade based on HTTP security headers. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and adds meaningful context: billing is charged to the user's own Apify account at roughly $0.05 per company, most sub-data comes from named scraper Actors, and the hiring-signals component returns warnings rather than errors when no ATS board is found. It still does not state authentication requirements or explicitly confirm read-only behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and followed by pricing. The long first sentence is an efficient enumeration of included data categories, and no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 8-parameter aggregate with no annotations and no output schema, the description explains the row-level output, the main data categories, the one-row-per-domain behavior, and the cost model. It is nearly complete, though it omits explicit auth requirements and alternative-tool routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the input schema, including the domains format and each include flag. The description adds no parameter-level syntax or constraints beyond what the schema provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it turns a domain into a rich company profile row. It enumerates the merged data categories (contacts, tech stack, DNS/e-mail security, TLS, SEO, hiring, company facts) and makes clear this is an aggregate rather than a single-purpose sibling like hiring_signals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but gives no explicit when-to-use, when-not-to-use, or alternative guidance. It does not say whether to prefer this aggregate over individual sibling tools such as hiring_signals or wikidata_entity_enrichment when only one data category is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_validatorA
Email Validator checks syntax, MX/A DNS records, disposable/free-provider/role classification and typo suggestions for a list of email addresses — one scored row per address, no SMTP mailbox probing. Billed to your own Apify account: ~$0.0005 per email (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| emails | Yes | Emails — Enter the email addresses to validate, one row is returned per address, e.g. jane.doe@example.com. Paste straight out of a CRM export or lead list; blank lines and duplicates are ignored. Example: ["support@apify.com"]. | |
| checkMx | No | Check MX / A records — Keep this on to look up the domain's MX records (falling back to an A record) over DNS-over-HTTPS. Turning it off skips DNS entirely and scores purely on syntax and list matches — faster, but cannot tell whether the domain can receive mail at all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses cost and billing model (~$0.0005/email billed to your own Apify account), output cardinality (one scored row per address), the DNS-over-HTTPS mechanism, and the key limitation of not probing SMTP mailboxes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the capability list and followed by a compact pricing note; both earn their place. Slightly dense but free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema or annotations, the description supplies the return shape (one scored row per address) and the classification dimensions a consumer can expect, plus cost and method. It stops short of enumerating the exact fields in each scored row, which is a minor residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both the emails array and checkMx flag are already thoroughly documented in the schema (including defaults, examples, and behavior). The description adds no parameter-level detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validates) and resource (list of email addresses) and enumerates the exact checks performed: syntax, MX/A DNS records, disposable/free-provider/role classification, and typo suggestions. It also scopes out what it does not do (no SMTP mailbox probing), so an agent knows precisely what it gets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context of use — validating a batch/lead list, one row per address — and a meaningful boundary (no SMTP mailbox probing). It lacks explicit when-to-use-vs-alternative guidance, but the sibling tools are unrelated Apify actors, so there is no real alternative to route against.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hiring_signalsA
Scrape company job postings and hiring signals from Greenhouse, Lever, Ashby and Workable job boards — one row per job, or one hiring summary per company. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| companies | Yes | Companies — Enter the companies whose job board you want to scrape. Use the ATS slug on its own, e.g. stripe, or paste a full job-board URL, e.g. https://boards.greenhouse.io/stripe, https://jobs.lever.co/spotify, https://jobs.ashbyhq.com/ashby, https://apply.workable.com/blueground/, https://jobs.smartrecruiters.com/smartrecruiters, https://bunq.recruitee.com, https://personio.jobs.personio.de, https://storytel.teamtailor.com, https://euna.bamboohr.com/careers/list or https://<tenant>.wd<N>.myworkdayjobs.com/<site>. A bare slug is probed against every board except Workday (which needs a full URL — a tenant and a site, not just a slug) and every board that actually has postings is returned. Example: ["stripe"]. | |
| outputMode | No | Output mode — Choose what one dataset row means. Pick "One row per job" to get every open posting (best for job aggregators and recruiters), or "One row per company" to get a single hiring summary per company with job counts, department and location breakdowns and a hiring-velocity score (best for sales intelligence and competitor watching). Options: jobs = One row per job; companies = One row per company (hiring summary). | jobs |
| maxJobsPerCompany | No | Max jobs per company — Enter how many of the newest postings to keep per company, e.g. 200. This caps both cost and row count in "One row per job" mode. It is deliberately ignored in "One row per company" mode, where the summary always counts the whole board so totals and hiring velocity stay accurate. | |
| includeDescriptions | No | Include job descriptions — Turn this on to add the full job description as plain text (HTML stripped, truncated to 5,000 characters). Leave it off for faster, smaller runs — on Workable it also saves one extra request per job, because that board only serves descriptions from a per-job endpoint. SmartRecruiters and Workday never return a description here (their list endpoints carry no job body), regardless of this setting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full behavioral burden, and it delivers unusually well: it discloses billing to the user's own Apify account with a per-result price, notes that maxJobsPerCompany is deliberately ignored in company mode, and warns that SmartRecruiters/Workday never return descriptions. It omits failure modes and any rate/quota behavior beyond cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core purpose and pricing front-loaded; every clause carries information. It is slightly run-on, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter scraper with no output schema and no annotations, the description covers output shape, cost, and per-source limitations well enough to call it correctly. Return-field details are left implicit, which is a minor gap given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters in detail; the description adds only the cost framing around result count. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Scrape company job postings and hiring signals') and names the exact ATS sources (Greenhouse, Lever, Ashby, Workable). The output-shape clause ('one row per job, or one hiring summary per company') further distinguishes it from the sibling remote_jobs_aggregator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Each outputMode option carries an explicit 'best for' context (job aggregators/recruiters vs. sales intelligence/competitor watching), which is genuine when-to-use guidance. It stops short of naming or excluding sibling tools, so it is clear context without full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
market_quotesA
Live-delayed stock, ETF, index, cryptocurrency and foreign-exchange quotes in one call — price, change, day range, volume and optional history, from free public endpoints with no API key. Billed to your own Apify account: ~$0.001 per Quote (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| symbols | Yes | Symbols — Enter one symbol per line: a stock/ETF/index ticker (AAPL, ^GSPC), a crypto ticker or pair (BTC, BTC-USD), or an FX pair (EUR/USD, EURUSD=X, USD/EUR). Example: ["AAPL","BTC-USD","EUR/USD"]. | |
| assetType | No | Asset type — Choose auto to detect each symbol's type from its shape (recommended), or force every symbol in this run to be read as stock, crypto or fx. Options: auto = Auto-detect; stock = Stock / ETF / index; crypto = Cryptocurrency; fx = Foreign exchange. | auto |
| historyDays | No | History days — Enter how many days of daily closes to include in history when includeHistory is on, e.g. 30. | |
| includeHistory | No | Include price history — Keep this off for a lean row per symbol. Turn it on to add a history array of daily closing prices going back historyDays days. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and discloses important traits: delayed rather than real-time data, no API key requirement, billing to the user's Apify account, and approximate per-quote cost. It omits rate limits, error behavior, and data coverage caveats, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two well-structured sentences, front-loading what the tool does and then adding billing context. Every clause contributes useful information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a rich input schema and no output schema, the description supplies the key return fields, data delay, source, and cost. It could elaborate slightly on output shape or symbol limits, but it is complete enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters thoroughly. The description only alludes to optional history and does not add syntax, constraints, or enum meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource—live-delayed quotes for stocks, ETFs, indices, crypto, and FX—and specifies the scope of one call returning price, change, day range, volume, and optional history. This clearly distinguishes it from the unrelated sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful context such as live-delayed data, free public endpoints, no API key, and Apify billing, but it does not explicitly say when to use this tool versus alternatives or when not to use it. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
overpass_poi_extractorA
OpenStreetMap POI Extractor returns shops, cafes, restaurants and any tagged place within a radius, a bounding box, or a named area — one row per point of interest, powered by the free Overpass API. Billed to your own Apify account: ~$0.0005 per place (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| bbox | No | Bounding box (bbox mode) — Bounding box as "south,west,north,east" decimal degrees, e.g. "50.05,14.35,50.13,14.55" for central Prague. Required when Search mode is "bbox". Example: "50.05,14.35,50.13,14.55". | |
| mode | Yes | Search mode — How to define the search area: "radius" around a point, "bbox" (a bounding box), or "area" (a place name resolved via Nominatim, e.g. a city or district). Options: radius = Radius around a point; bbox = Bounding box; area = Named area (city/district). Example: "radius". | |
| areaName | No | Area name (area mode) — A place name to search inside, e.g. "Prague, Czechia" or "Brooklyn, New York". Resolved to an OpenStreetMap area via Nominatim. Required when Search mode is "area". Example: "Prague, Czechia". | |
| latitude | No | Latitude (radius mode) — Center point latitude for radius mode, e.g. 50.0755 for Prague's city centre. Ignored in bbox and area mode. Example: 50.0755. | |
| longitude | No | Longitude (radius mode) — Center point longitude for radius mode, e.g. 14.4378 for Prague's city centre. Ignored in bbox and area mode. Example: 14.4378. | |
| categories | Yes | OSM categories (key=value tags) — OpenStreetMap tags to search for, each as "key=value", e.g. "amenity=cafe" or "shop=supermarket". Every category is searched and merged into one result list. Example: ["amenity=cafe","amenity=restaurant"]. | |
| maxResults | No | Max results — Maximum number of POIs to return, e.g. 500. Overpass results beyond this count are dropped before they are pushed to the dataset. | |
| radiusMeters | No | Radius in meters (radius mode) — How far from the center point to search, in meters, e.g. 1000 for a 1 km radius. Larger radii return more results and take longer to query. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a good job: it discloses the data source (free Overpass API), the cost model (~$0.0005 per place, billed to the caller's Apify account), and the result cardinality (one row per POI). It omits auth prerequisites and rate/quota limits, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the tool returns and how the search area is defined, followed by the billing caveat. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description states the return shape ('one row per point of interest') and the cost implication, covering what an agent needs to decide and call. Missing only ancillary details such as how categories combine with modes and any hard caps beyond maxResults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (mode, bbox, areaName, latitude/longitude, categories, maxResults, radiusMeters) is already fully documented with examples. The description only restates the three area modes and the merge behavior for categories, adding little beyond the schema — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('returns shops, cafes, restaurants and any tagged place') with explicit scope (radius, bbox, named area) and output granularity (one row per POI). No sibling tool (hiring_signals, market_quotes, etc.) overlaps with POI extraction, so differentiation is implicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the three search-area modes and the backing service, which tells the agent how to frame a request, but offers no explicit when-to-use/when-not or alternative tools to prefer for other geodata needs. Context is clear; exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_jobs_aggregatorA
Merge and de-duplicate remote job listings from RemoteOK, We Work Remotely and Hacker News' Who is Hiring thread into one clean row per job — title, company, salary, tags and apply link. Billed to your own Apify account: ~$0.001 per Job (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| maxJobs | No | Max jobs — Enter the maximum number of job rows to return after merging and de-duplicating every feed, e.g. 100. | |
| sources | No | Feeds to aggregate — Choose which public feeds to merge. remoteok reads RemoteOK's JSON API, weworkremotely reads its main RSS feed, and hackernews reads the current month's "Who is hiring?" thread on Hacker News. Options: remoteok = RemoteOK; weworkremotely = We Work Remotely; hackernews = Hacker News (Who is hiring?). | |
| keywords | No | Keywords (optional) — Enter words or phrases to keep only matching jobs, e.g. python, remote react. Matching is case-insensitive against the job title, company name and tags. Leave empty to keep every job. | |
| ashbyBoards | No | Ashby boards (optional) — Enter Ashby job-board slugs to fold into the same feed, e.g. ashby (from https://jobs.ashbyhq.com/ashby). A wrong slug returns one row with an error message instead of failing the run. | |
| greenhouseBoards | No | Greenhouse boards (optional) — Enter Greenhouse board slugs to fold into the same feed, e.g. stripe (from https://boards.greenhouse.io/stripe). A wrong slug returns one row with an error message instead of failing the run. | |
| postedWithinDays | No | Posted within (days) — Enter how many days back to keep a job by its posted date, e.g. 30. A job with no known posted date is always kept. Set to 0 to disable this filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: results are merged and de-duplicated across feeds, billing hits the caller's own Apify account at ~$0.001/job, and the schema adds that bad board slugs return an error row instead of failing. It stops short of covering rate limits, runtime, or partial-failure semantics, but the billing and de-dup disclosures go well beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and resource before the pricing note. No filler, no repetition of schema content, and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, all-optional aggregation tool with no annotations and no output schema, the description covers the source set, the de-duplication/row model, the returned fields, and the cost model — enough for an agent to call it correctly. Minor gaps remain around failure modes and result volume, but those are largely covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (maxJobs, sources, keywords, ashbyBoards, greenhouseBoards, postedWithinDays) is already fully documented in the schema, including examples and edge-case behavior. The description adds no additional parameter semantics beyond restating the aggregated fields, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (merge and de-duplicate) applied to a specific resource (remote job listings) from three explicitly named sources, and describes the output shape (one clean row per job with title, company, salary, tags, apply link). No sibling tool overlaps this function, and an agent can tell exactly what it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (aggregating remote job feeds) and gives useful cost/billing context, but never states when to use it versus alternatives or when not to use it (e.g. vs. a single-source scraper). Usage is inferable rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wikidata_entity_enrichmentB
Wikidata Entity & Company Enrichment resolves a QID, company name, or Wikipedia URL to its Wikidata item and returns facts (country, HQ, founders, CEO, employees, revenue), identifiers (ISIN, Crunchbase, LinkedIn) and social/media links — one row per entity. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| entities | Yes | Entities — Enter a Wikidata QID (e.g. Q95), an entity name to search for (e.g. Apify), or a Wikipedia article URL (e.g. https://en.wikipedia.org/wiki/Stripe,_Inc.) — one row is returned per entity. Names are resolved via Wikidata search, so the closest match wins. Example: ["Q95"]. | |
| language | No | Language — Enter the Wikidata/Wikipedia language code to use for labels, descriptions and search, e.g. en, de, fr. Falls back to English internally when a translation is missing. | en |
| entityType | No | Entity type filter — When resolving a name search, prefer a candidate whose Wikidata instance-of (P31) matches this type over the plain top search result. Leave as 'any' to just take the best text match. Options: any = Any; company = Company; person = Person; place = Place; product = Product. | any |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does add genuinely useful behavioral context — per-result billing to the caller's own Apify account (~$0.001/result) and the 'one row per entity' output shape — but it omits auth/credential requirements, rate limits, and failure behavior on unresolved names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resolution capability and output contents before the billing note. Efficient and readable, with no padding or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately lists the return contents (facts, identifiers, social/media links, one row per entity) and notes the billing model. For a read-only enrichment tool this is nearly complete, though credential expectations for the Apify account are unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents entities, language, and entityType with examples and enum meanings. The description adds no parameter-level detail beyond what the schema supplies, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('resolves a QID, company name, or Wikipedia URL to its Wikidata item') and enumerates the returned fact categories, so the agent knows exactly what the tool produces. It does not, however, distinguish itself from overlapping siblings like company_360, which also appears to produce company enrichment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what inputs are accepted but never states when to choose this tool over the alternatives (company_360, yc_companies_scraper, etc.), nor any exclusions or prerequisites beyond the billing note. Usage is only weakly implied by the input-type enumeration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
yc_companies_scraperA
Y Combinator company directory scraper — batch, industry, hiring status, one-liner, description and founder names/titles/LinkedIn/Twitter for every YC-funded startup, one row per company. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode — Pick "Browse companies" to list/filter the directory by batch, industry, status or hiring status, or "Search by keyword" to additionally match the query text against each company's name and description. Options: companies = Browse companies; search = Search by keyword. | companies |
| query | No | Search text — Enter keywords to match against the company name, one-liner and description, e.g. "AI agents". Only used when mode is "Search by keyword", or as an extra filter on top of the other fields in "Browse companies" mode. Leave empty to skip text search. | |
| batches | No | Batches — Enter YC batch codes to keep, e.g. W25, S25, F26. Leave empty to include every batch. Codes are case-insensitive and matched exactly against the company's batch field. | |
| isHiring | No | Is hiring — Turn on to keep only companies currently marked as hiring on YC's directory, or leave this switch off (unset) to include both hiring and non-hiring companies. There is no third state for "only non-hiring" — untoggled means "do not filter". | |
| statuses | No | Company statuses — Select which company statuses to keep. Leave empty to include every status (most companies are Active). Options: Active = Active; Acquired = Acquired; Public = Public; Inactive = Inactive. | |
| industries | No | Industries — Enter industry names to keep, e.g. B2B, Fintech, Healthcare. A company matches if it has at least one of the listed industries. Leave empty to include every industry. | |
| maxCompanies | No | Max companies — Enter the maximum number of matching companies to return, e.g. 200. This is also the billed unit — the run stops as soon as it collects this many matches or runs out of companies to check. | |
| includeFounders | No | Include founders — Keep this on to fetch each matched company's public YC profile page and extract founder name, title, LinkedIn and Twitter/X links. Turning it off skips one extra request per company and returns founders as null, which is much faster for large runs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses billing to the user's own Apify account (~$0.002 per result), explains that maxCompanies is the billed unit, and notes that includeFounders triggers an extra profile request and can return null for speed. It does not cover authentication setup or rate limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and output fields, followed by the cost model. Every phrase earns its place, and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description usefully summarizes the return shape as 'one row per company' and lists the included fields; it also covers non-obvious cost behavior. It lacks details on auth/rate limits, but it is complete enough for selection given the rich input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all 8 parameters with defaults, enums, and constraints. The description adds no parameter-level syntax or behavior beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Y Combinator company directory scraper,' and enumerates the returned fields (batch, industry, hiring status, one-liner, description, founder details), so the agent knows exactly what it does and returns. It does not explicitly differentiate itself from siblings such as company_360 or hiring_signals, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers no when-to-use or when-not-to-use guidance relative to alternatives, and does not name or exclude any sibling tool. The mode and query options are parameter-level, not routing guidance, so an agent must infer its place in the toolset entirely from the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
company_360 - First observed
email_validator - First observed
hiring_signals - First observed
market_quotes - First observed
overpass_poi_extractor - First observed
remote_jobs_aggregator - First observed
wikidata_entity_enrichment - First observed
yc_companies_scraper
TDQS
Scored across 8 tools
Most tools target distinct data sources or use cases, but there is some overlap in company enrichment (company_360, wikidata_entity_enrichment, yc_companies_scraper) and job data (hiring_signals, remote_jobs_aggregator, plus hiring signals inside company_360). Descriptions clarify boundaries by specifying input types and sources, so an agent can usually pick correctly, but a few selections require careful reading.
All tool names use snake_case and are descriptive noun phrases, which is consistent and readable. Minor deviations exist in suffix patterns (scraper, extractor, aggregator, validator, enrichment, signals, quotes) and one name includes a number (company_360), but overall the convention is predictable.
Eight tools is well within the ideal 3–15 range for a data enrichment and leads toolkit. Each tool covers a distinct data source or processing need, so the count feels scoped and purposeful.
The surface covers company profiles, hiring signals, email validation, local POIs, market quotes, and YC data, which is broad for a leads-oriented server. Minor gaps remain, such as generic company discovery by criteria or person-level contact enrichment beyond what company_360 offers, but agents can work around them.
Maintenance
Related MCP Connectors
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
100+ MCP tools for AI agents: content metadata, trade intelligence, business-expertise analysis.
- geoOAuthco.thinair
Geocoding, routing, isochrones, traffic, weather, and place search for AI agents. 19 MCP tools.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceExposes over 19,000 Apify Actors as MCP tools for web scraping, data extraction, and OSINT automation. It enables AI agents to dynamically discover and execute scrapers to collect structured data and crawl web content.MIT
- FlicenseBqualityDmaintenanceExposes Apollo.io API functionalities as MCP tools for people and organization enrichment, search, and job postings. Enables natural language interaction with Apollo.io data.516-
- AlicenseNot gradedqualityCmaintenanceSeven remote MCP servers exposing 51 published Apify scrapers as agent tools: company diligence, social listening, recruiting, real estate, lead generation, e-commerce and academic research. Billed per result, and a call that returns nothing is never charged.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables B2B lead intelligence through seven tools for company lookup, domain enrichment, person search, email finding, and startup discovery, accessible via MCP for AI agents.-