WebDataTools Developer, app & research data MCP server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@WebDataTools Developer, app & research data MCP servercheck the health of the npm package express"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebDataTools Developer, app & research data MCP server
webdatatools-dev-mcp
An MCP server with 13 developer, app & research data tools for AI agents — Claude Desktop, Cursor, Cline or any MCP client. npm/PyPI/Crates package health, GitHub repo health and trending, VS Code and Chrome Web Store extensions, Google Play and App Store apps, CrossRef DOIs, openFDA recalls, iCal feeds, Shopify products, Hacker News and Stack Exchange.
This server uses your own Apify API token. Every tool call runs a WebDataTools Actor under your Apify account and is billed to your Apify credit — pay per result, the price is in each tool description. Your token is only sent to Apify's API.
Quick start
Requires Node.js 18+.
APIFY_TOKEN=apify_api_... npx -y github:paulet4a-commits/webdatatools-dev-mcpGet a free token (the free plan includes monthly credit): https://console.apify.com/settings/integrations
Related MCP server: AXE Fleet MCP Server
Claude Desktop / Cursor
Add this to claude_desktop_config.json (Claude Desktop) or .cursor/mcp.json (Cursor):
{
"mcpServers": {
"webdatatools-dev": {
"command": "npx",
"args": [
"-y",
"github:paulet4a-commits/webdatatools-dev-mcp"
],
"env": {
"APIFY_TOKEN": "apify_api_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
}
}
}
}Tools (13)
Tool | What it does | Price (free plan) | Backing Actor |
| npm, PyPI & Crates.io Package Health Checker | $0.002 / result | |
| GitHub Repository Health & Activity Report | $0.002 / result | |
| VS Code Marketplace Extension Scraper (installs, ratings) | $0.0015 / result | |
| Chrome Web Store Extension Scraper (installs, ratings) | $0.0015 / result | |
| Google Play Scraper | $0.002 / app | |
| App Store (iOS) App Metadata, Ratings & Top Charts Lookup | $0.001 / result | |
| CrossRef DOI & Citation Metadata Lookup | $0.001 / result | |
| FDA Recalls & Adverse Events Monitor (openFDA) | $0.002 / result | |
| iCal / ICS Calendar Feed to Events Extractor | $0.0005 / result | |
| Shopify Store Products Scraper | $0.001 / result | |
| Hacker News Search & Front Page Scraper | $0.0005 / Story | |
| GitHub Trending Repositories Scraper | $0.002 / Repository | |
| Stack Overflow & Stack Exchange Q&A Scraper | $0.0005 / Question |
More WebDataTools MCP servers
webdatatools-mcp-server — the 10 most popular tools in one server
webdatatools-domain-mcp — Domain & website intelligence
webdatatools-rag-mcp — Web content for AI & RAG
webdatatools-social-mcp — Search, video & social data
webdatatools-leads-mcp — Leads, jobs & company data
License
MIT
Available Tools
13 toolsapp_store_lookupA
App Store lookup for iOS apps: ratings, price, version, update date, size, genres and description from any app id, bundle id, store URL or search term — plus Top Free and Top Paid charts. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| apps | Yes | Apps to look up — Enter one app per row in any of four forms: a numeric App Store id (324684580), a bundle id (com.spotify.client), an App Store URL (https://apps.apple.com/us/app/spotify-music/id324684580), or a search term prefixed with search: (search:meditation, which returns up to searchLimit apps). Used in lookup mode only. Example: ["324684580"]. | |
| mode | No | Mode — Choose lookup to resolve the apps you list below, or topCharts to download a storefront's Top Free or Top Paid chart. In topCharts mode the apps field is ignored and chart, chartLimit and country decide the output. Options: lookup = Lookup apps I list; topCharts = Top charts for a country. | lookup |
| chart | No | Chart — Pick which chart to download in topCharts mode. Apple's public RSS feed serves only these two for apps — there is no top-grossing feed (it returns HTTP 404). Options: top-free = Top Free Apps; top-paid = Top Paid Apps. | top-free |
| country | No | Storefront country — Enter the two-letter country code of the App Store storefront to read, e.g. us, gb, de, jp. Prices, ratings, descriptions and chart positions are all storefront-specific, so us and de can return different numbers for the same app. | us |
| chartLimit | No | Chart size — Enter how many ranked apps to return in topCharts mode, e.g. 50. Apple's feed only serves 10, 25, 50 or 100 positions; any other number (including this field's own min/max) is rounded up to the nearest one of those four. | |
| searchLimit | No | Results per search term — Enter how many apps each search: term should return, e.g. 25. Apple's Search API caps this at 200. Only affects rows starting with search: — an id or bundle id always returns exactly one row. | |
| enrichChartApps | No | Enrich chart rows with full metadata — Keep this on so every charted app also gets its ratings, price, version, size and description from the iTunes lookup API. Apple's chart feed alone only carries name, developer, icon and release date. Turning it off makes a 100-app chart two requests cheaper but leaves those columns null. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers an important non-obvious behavior: billing to the caller's own Apify account with a per-result price (~$0.001) that varies by plan. That is genuine cost/behavioral context an agent cannot get from the schema. It does not, however, describe auth requirements, rate limits, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the resource and returned fields, then appends the pricing clause. No filler, no redundancy, and the most decision-relevant facts come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, two-mode scraper with no output schema and no annotations, the description covers purpose, accepted input forms, chart support, and cost. It does not describe the shape of returned records or behavior when a lookup fails, which a mutation-free read tool of this size would ideally mention, keeping it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already defines every parameter with examples, ranges, and rounding rules. The description's mention of the four input forms and the two charts largely restates what the schema documents, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('App Store lookup for iOS apps') and enumerates the returned fields (ratings, price, version, update date, size, genres, description) plus lookup vs chart coverage. It is clearly distinguishable from the Android-oriented sibling google_play_scraper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description frames the two operating modes ('from any app id, bundle id, store URL or search term — plus Top Free and Top Paid charts'), which tells an agent what inputs are acceptable. It stops short of explicit when-to-use/when-not guidance or naming an alternative tool for edge cases, so it lands below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_web_store_extensionsA
Chrome Web Store Extension Scraper returns users, rating, version, size, category and developer contact for any Chrome extension id or store URL — one row per extension. Billed to your own Apify account: ~$0.0015 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| extensions | Yes | Extensions — Enter Chrome extensions to look up, one row is returned per extension. Accepts a 32-character extension id such as kbfnbcaeplbcioakkpcpgfkobkghlhen (uBlock Origin), or a full Chrome Web Store URL such as https://chromewebstore.google.com/detail/grammarly/kbfnbcaeplbcioakkpcpgfkobkghlhen — the slug in the URL is ignored, only the id matters. Example: ["kbfnbcaeplbcioakkpcpgfkobkghlhen"]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavioral context: one row per extension, billing to the user's own Apify account, and the per-result price (~$0.0015). It omits auth/token requirements, rate limits, and error behavior for invalid ids, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads purpose plus returned fields, the second adds the cost disclosure. No filler, and the most decision-relevant information (what it returns) comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup with no output schema, the description compensates by enumerating the returned fields and the cost model. The only gap is operational detail (auth, limits, failure modes) that an agent invoking it might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter already documents accepted id and full-URL formats with an example, so the schema does the heavy lifting. The description adds no syntax or format detail beyond it, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Chrome Web Store Extension Scraper returns users, rating, version, size, category and developer contact'), and the subject matter ('Chrome Web Store') naturally distinguishes it from the sibling vscode_marketplace_extensions / google_play_scraper. It stops short of explicitly naming a sibling to route between, so it lands at clear-but-not-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: pass an extension id or store URL to fetch metadata. There is no explicit when-to-use vs when-not-to-use and no mention of the near-identical sibling store scrapers, leaving the agent to infer selection from the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crossref_doi_lookupA
CrossRef DOI & Citation Metadata Lookup resolves DOIs, DOI URLs or citations via the CrossRef API and returns title, authors, journal, citation count, abstract and license — one row per item. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | DOIs, DOI URLs or citations — Enter one item per line: a bare DOI (10.1038/nature12373), a DOI URL (https://doi.org/10.1145/3292500.3330919), or free text to search — prefix free text with "search:" (search:Attention is all you need) or just paste a title/citation and it is searched automatically when it does not look like a DOI. Example: ["10.1038/nature12373"]. | |
| mailto | No | Contact e-mail (polite pool) — Optional. Enter your e-mail, e.g. you@example.com, to be added to the User-Agent and as a mailto= query param so CrossRef routes your requests to its faster "polite pool". Leave empty to use the public pool. | |
| searchRows | No | Search results per query — Enter how many candidate works CrossRef should return for each free-text search item, e.g. 3. Only the best match (highest score) becomes the row; a higher number only improves matching, it does not create extra rows. Max 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses the billing model (~$0.001 per result, charged to the caller's Apify account) and the output shape ('one row per item'), but says nothing about rate limits, failure behavior, or whether results are deduplicated/ranked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: purpose and return fields first, then cost. No filler, and the most decision-relevant facts (what it does, what it costs) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields, and it adds cost transparency. For a three-parameter lookup tool with 100% schema coverage this is nearly complete, though error/pagination behavior remains unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents items, mailto, and searchRows thoroughly. The description adds no parameter meaning beyond what the schema provides, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('resolves DOIs, DOI URLs or citations via the CrossRef API') and enumerates the returned fields (title, authors, journal, citation count, abstract, license). It is unmistakably distinct from all siblings, which cover unrelated domains (package health, GitHub, app stores).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the input description ('Enter one item per line... prefix free text with search:'), but the top-level description gives no when-to-use/when-not guidance or named alternatives. A reader must infer the context from the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_repo_healthA
GitHub Repository Health & Activity Report returns stars, commit activity, contributors, latest release and community-health files for any repo list or org: — one scored row per repo. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| repos | Yes | Repositories or organisations — Enter GitHub repositories to check, one row is returned per repository, e.g. apify/crawlee. Full GitHub URLs also work (https://github.com/expressjs/express), and org:<name> (e.g. org:apify) expands to that organisation's public repositories, up to Org repo limit. Example: ["apify/crawlee"]. | |
| orgRepoLimit | No | Org repo limit — Enter the maximum number of repositories to pull from each org:<name> entry, e.g. 30. Repositories are sorted by most recently updated first, so a low limit still gets the organisation's most active projects. | |
| includeReleases | No | Include latest release — Keep this on to fetch the 5 most recent releases per repository (one extra API request per repo) and report the latest one. Turn it off for repositories that don't use GitHub Releases to save requests. | |
| includeContributors | No | Include contributor count — Keep this on to fetch an estimated contributor count for each repository (one extra API request per repo). Turn it off to save requests when you only need stars, activity and release data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses billing ('billed to your own Apify account: ~$0.002 per result') and the hidden cost behavior of the optional flags ('one extra API request per repo'). It does not cover auth requirements or failure handling for invalid repos, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource and returned fields before the pricing tail. Dense but nearly every clause earns its place; the pricing detail is slightly tacked on but remains relevant to the agent's invocation decision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema or annotations, so the description must convey return contents, which it does by listing the reported metrics and the one-row-per-repo shape. It is complete enough to call correctly, though it omits error behavior for unreachable or private repos.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters thoroughly, including defaults and limits. The description adds only marginal meaning (org expansion, cost implications of releases/contributors), consistent with the baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (GitHub repository health & activity), enumerates exactly what is returned (stars, commit activity, contributors, latest release, community-health files), and states the output shape ('one scored row per repo'). It also handles the org:<name> input mode inline, which distinguishes it from github_trending_scraper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the org:<name> expansion and gives per-flag cost tradeoffs ('turn it off to save requests'), which implies usage context, but it never states when to prefer this over github_trending_scraper or when the report is the wrong tool. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_trending_scraperA
Scrapes github.com/trending: the daily, weekly or monthly trending repositories or developers, by programming language and spoken language. One row per repo or developer with stars, forks, stars-in-period and description — no API key, no proxies. Billed to your own Apify account: ~$0.002 per Repository (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode — Choose repositories for github.com/trending (one row per trending repo), or developers for github.com/trending/developers (one row per trending developer). Options: repositories = Trending repositories; developers = Trending developers. | repositories |
| since | No | Time range — Choose the trending window: daily, weekly or monthly, matching the tabs on github.com/trending. Options: daily = Daily; weekly = Weekly; monthly = Monthly. | daily |
| languages | Yes | Programming languages — Enter one GitHub language slug per row, e.g. javascript or python. Leave a row empty to fetch trending across all languages. Example: [""]. | |
| maxResults | No | Max results per language — Enter how many rows to return for each entry in languages, e.g. 25. GitHub's trending page lists at most 25 repositories or 25 developers per language. | |
| spokenLanguage | No | Spoken language — Enter a two-letter spoken-language code to filter by, e.g. en or ja, matching GitHub's "Spoken Language" trending filter. Leave empty for all spoken languages. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the output grain ('one row per repo or developer with stars, forks, stars-in-period and description'), that no API key or proxies are required, and the exact cost model (~$0.002 per Repository, billed to the caller's Apify account). It does not cover pagination, error behavior, or rate limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is scraped and how results are shaped, followed by the cost/auth caveat. Dense but each clause (filters, output grain, billing) carries information; the dash-separated list is slightly packed but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does the heavy lifting: it names the return fields, the billing model, and the no-key/no-proxy setup for a 5-parameter tool. The main residual gap is that return format details (types, ordering, empty-result behavior) are left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the enum options, defaults, and bounds for mode/since/languages/maxResults/spokenLanguage are already fully documented in the schema. The description only echoes the language and time-window concepts without adding syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — scraping github.com/trending for repositories or developers — with the supported time windows and filter dimensions (programming and spoken language). It is immediately distinguishable from siblings like github_repo_health, which inspect a specific repo rather than the trending feed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'trending' framing and the cost note, and the description implies when this tool fits (discovery of what's hot). However, it never states when NOT to use it or points to an alternative such as github_repo_health for per-repo metrics, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
google_play_scraperA
Google Play Scraper returns app title, developer, rating, price, installs, category, media links, and optional reviews for any package id, Play Store URL or search term — one row per app. Billed to your own Apify account: ~$0.002 per app (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| apps | Yes | Apps — Enter the apps to scrape, one row is returned per app. Use the package id, e.g. com.spotify.music, a full Play Store URL, e.g. https://play.google.com/store/apps/details?id=com.spotify.music, or search:<term>, e.g. search:meditation, to scrape the top search results for a term. Example: ["com.spotify.music"]. | |
| country | No | Country — Enter the 2-letter country code for the Play Store storefront to read, e.g. us, gb, de. Pricing, availability and installs shown can differ by country. | us |
| language | No | Language — Enter the 2-letter language code for the app title, description and other text, e.g. en, es, de. | en |
| searchLimit | No | Search result limit — Enter how many apps to return per search:<term> entry, e.g. 20. A single Play Store search page load returns roughly 20-30 unauthenticated results with no further pagination, so this is a ceiling, not a guarantee. | |
| includeReviews | No | Include reviews — Turn this on to add the newest reviews for each app (capped by "Max reviews per app"). Uses the same unauthenticated batchexecute endpoint the open-source google-play-scraper npm package uses for reviews — no login required, but it is an undocumented Google endpoint and could change without notice. | |
| maxReviewsPerApp | No | Max reviews per app — Enter how many of the newest reviews to fetch per app when "Include reviews" is on, e.g. 50. Ignored when "Include reviews" is off. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: it discloses that the tool is billed to the caller's own Apify account at ~$0.002 per app, and that output is one row per app. It still omits auth prerequisites and failure/rate-limit behavior, but the cost disclosure is substantive value beyond structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the capability statement is front-loaded ahead of the pricing note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description correctly takes on the return-value summary (field list) and the cost model. Parameters are fully covered by the schema. It is nearly complete, missing only auth/prerequisite details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, including input formats, country/language codes, search limits and the reviews toggle. The description adds no syntax or semantics beyond that, which is the expected baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (returns) and resource (Google Play app data), and enumerates the returned fields (title, developer, rating, price, installs, category, media links, reviews) plus the accepted input forms. The 'Google Play' resource inherently distinguishes it from the iOS-oriented sibling app_store_lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to pick this tool over alternatives, nor any prerequisites or exclusions. Input-format hints (package id / URL / search term) are parameter documentation, not usage guidance, so the agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hacker_news_scraperA
Search Hacker News stories, comments, Show HN and Ask HN posts by keyword, or pull the current front page — via the free Algolia HN API, no API key. One row per story or comment with points, author, comment count and text. Billed to your own Apify account: ~$0.0005 per Story (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| author | No | Author username — Enter a Hacker News username, e.g. pg, to only return items posted by that user. Leave empty to include every author. | |
| sortBy | No | Sort by — relevance uses Algolia's default ranking (Hacker News' own /search), date returns the newest items first (/search_by_date), and points sorts the matched items by score, highest first. Ignored for front_page queries. Options: relevance = Relevance; date = Newest first; points = Highest points first. | relevance |
| queries | Yes | Search queries — Enter one search term per row, e.g. apify or rust. Leave a row empty, or type front_page, to fetch the current Hacker News front page instead of running a search. Example: ["front_page"]. | |
| minPoints | No | Minimum points — Enter the minimum score an item must have to be included, e.g. 50. Set to 0 to disable this filter. | |
| sinceDays | No | Only items from the last N days — Enter how many days back to search, e.g. 7 for the last week. Set to 0 to disable this filter and search all time. | |
| maxResults | No | Max results per query — Enter how many rows to return for each entry in queries, e.g. 50. Algolia serves at most 1000 hits per query. | |
| contentType | No | Content type — Choose which kind of Hacker News item to search: story (links and text posts), comment, show_hn, ask_hn, poll, or all to search every type. Ignored for front_page queries, which always return front-page stories. Options: story = Stories; comment = Comments; show_hn = Show HN; ask_hn = Ask HN; poll = Polls; all = All types. | story |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does solid work: it discloses the data source (Algolia HN API), that no API key is needed, the billing model (~$0.0005 per Story on Apify's free plan), and the returned row shape. Gaps remain around rate limits, pagination behavior, and error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core capability, then output shape, then cost. Dense but every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter scraper with full schema coverage and no output schema, the description supplies the missing return-value context (row fields) plus sourcing and pricing. Only minor operational details (pagination, cost at scale) are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (author, sortBy, queries, minPoints, sinceDays, maxResults, contentType) is documented in the schema itself. The description repeats the high-level search semantics without adding syntax, defaults, or edge cases beyond what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource+source: 'Search Hacker News stories, comments, Show HN and Ask HN posts by keyword, or pull the current front page — via the free Algolia HN API.' It also states the output shape (one row per story/comment with points, author, comment count, text), so an agent can distinguish it from siblings like stackexchange_scraper or github_trending_scraper at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the two operating modes (keyword search vs. front-page pull) and notes the front_page sentinel. It does not name alternatives or exclusions, but sibling tools target entirely different platforms, so the routing need is minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ical_calendar_extractorA
iCal / ICS Calendar Feed to Events Extractor reads any ICS or webcal:// URL and returns one row per event instance, expanding RRULE recurring events within a date window, with times converted to UTC. Billed to your own Apify account: ~$0.0005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | To date — Only return event instances starting on or before this ISO date, e.g. 2026-12-31. Leave blank to default to 365 days after today. Example: "2026-12-31". | |
| from | No | From date — Only return event instances starting on or after this ISO date, e.g. 2026-01-01. Leave blank to default to 30 days before today. Example: "2026-01-01". | |
| feeds | Yes | ICS feed URLs — Enter the ICS/iCal feed URLs to read, e.g. https://www.officeholidays.com/ics/usa. A webcal:// URL is accepted too and automatically rewritten to https://. One row is returned per event instance found in the feed. Example: ["https://www.officeholidays.com/ics/usa"]. | |
| expandRecurring | No | Expand recurring events — Keep this on to expand RRULE recurring events (e.g. a weekly meeting) into one row per occurrence inside the from/to window, capped at 500 instances per event. Turn it off to get a single row per recurring event at its original start time. | |
| maxEventsPerFeed | No | Max event rows per feed — Enter the maximum number of event-instance rows to keep per feed, e.g. 1000. The earliest instances within the window are kept first when a feed has more than this. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose meaningful behaviour: RRULE expansion, UTC normalisation, one row per instance, and a per-result billing model tied to the caller's Apify account (~$0.0005/result). It does not mention rate limits, failure behaviour for malformed feeds, or auth requirements, but the cost and output-shape disclosure is genuinely useful beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the tool does and its output granularity, then cost. Every clause earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description supplies the missing essentials: the row-per-instance return shape, recurrence expansion semantics, UTC conversion, and billing. Only failure modes and feed-validity prerequisites are absent, which is a minor gap for a read-only extraction tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including defaults, bounds and examples is already documented in the schema; baseline 3 applies. The description only restates the window/expansion behaviour at a high level and adds no syntax or format detail the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('ICS Calendar Feed to Events Extractor reads any ICS or webcal:// URL and returns one row per event instance') and states the transformation performed. Nothing in the sibling list overlaps, so the agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the usage context (reading ICS/webcal feeds) but never states when to use it versus alternatives, nor any when-not conditions or prerequisites such as needing a valid public feed URL. For a tool with no near-neighbours this is adequate but leaves the trigger condition implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
openfda_recall_monitorA
FDA Recalls & Adverse Events Monitor returns food, drug and device recalls, drug and device adverse-event reports, and drug labels from openFDA — one flat, scored row per record for a date range you choose. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| dataset | Yes | Dataset — Choose which openFDA dataset to query. Enforcement datasets are product recalls; event datasets are adverse-event reports; drug-label returns FDA drug label (SPL) documents. Options: food-enforcement = Food recalls (enforcement); drug-enforcement = Drug recalls (enforcement); device-enforcement = Device recalls (enforcement); drug-event = Drug adverse events (FAERS); device-event = Device adverse events (MAUDE); drug-label = Drug labels (SPL). Example: "food-enforcement". | |
| maxItems | No | Max items — Enter the maximum number of records to return, e.g. 100. Results are fetched 100 at a time (limit=100&skip=...) until this many are collected or the dataset runs out. | |
| sinceDays | No | Since days — Enter how many days back to search, e.g. 30. Filters the dataset's report_date (enforcement) or receivedate/date_received (events) field; ignored for drug-label, which has no report date. | |
| searchQuery | No | openFDA search query — Optional. A raw openFDA search expression, ANDed with the date range, e.g. recalling_firm:"Nestle" or classification:"Class I". Leave blank to fetch every record in the date range. See https://open.fda.gov/apis/query-syntax/ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two genuinely useful traits beyond structured fields: billing runs against the caller's own Apify account at ~$0.002/result, and output is one flat, scored row per record. It stops short of describing auth setup or rate limits, so it is good but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the capability before the pricing note, with no filler. Slightly dense and the undefined term 'scored' costs a little clarity, but it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, so the description must carry the load; it names the six datasets, the return shape, and cost. The one notable omission is that 'scored row' is never defined, leaving the output field semantics unclear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents dataset, maxItems, sinceDays and searchQuery in detail. The description adds only the framing of 'a date range you choose' (mapping to sinceDays) and does not extend parameter meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and a rich resource set: returns food/drug/device recalls, adverse-event reports, and drug labels from openFDA, with the scoping unit (one flat, scored row per record for a chosen date range). No sibling tool overlaps this FDA-data domain, so an agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the tool for FDA recall/adverse-event data, but there is no explicit when-to-use or when-not-to-use guidance, and no named alternative. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
package_health_checkerA
Package health checker for npm, PyPI and Crates.io — deprecation, last publish date, licence, weekly downloads, maintainers, GitHub stars and a 0-100 health score, one row per package. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| packages | Yes | Packages — Enter the packages to check, one row is returned per package. Prefix the name with the registry, e.g. npm:react, pypi:requests, crates:serde. Bare names such as react or @scope/name use the default registry below, and full registry URLs work too, e.g. https://pypi.org/project/requests/. Example: ["react"]. | |
| enrichGithub | No | Enrich with GitHub data — Keep this on to add GitHub stars, forks, open issues, the archived flag and the last push date whenever the package points at a GitHub repository. Turn it off to run faster and avoid GitHub rate limits. | |
| defaultRegistry | No | Default registry — Select which registry a bare package name belongs to, e.g. npm for react. Entries that already carry a prefix (pypi:requests) or a registry URL ignore this setting. Options: npm = npm (JavaScript); pypi = PyPI (Python); crates = Crates.io (Rust). | npm |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose real behavioral traits: billing to the caller's own Apify account at ~$0.002/result, one row returned per package, and that GitHub enrichment triggers rate limits. It stops short of stating auth requirements, run/async semantics, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the capability and then the cost caveat; every clause earns its place. The long metric list and the trailing 'one row per package' clause are slightly crammed but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter scraper with no output schema and no annotations, the description compensates by enumerating returned fields and the cost model. It leaves out execution semantics (async run, timing) that an Apify actor tool would benefit from disclosing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters in detail (prefix syntax, defaultRegistry enum, enrichGithub trade-off). The description adds no parameter-level syntax or default beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('Package health checker for npm, PyPI and Crates.io') and enumerates the exact signals returned (deprecation, last publish, licence, downloads, maintainers, stars, 0-100 score). The registry-level scope clearly separates it from the repo-level sibling github_repo_health.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use statement, prerequisites, or named alternatives. The use case (assessing dependency health before adoption) is only implied by the list of returned metrics and the per-result pricing note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_products_scraperA
Shopify Store Products Scraper reads any Shopify store's public products.json feed — title, price, compare-at price, discount %, stock, SKU and images, one row per product or variant. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| stores | Yes | Store URLs — Enter the Shopify store URLs to scrape, e.g. https://www.allbirds.com. A bare domain such as allbirds.com also works. Every store is checked against its public products.json feed; non-Shopify or password-protected stores get one error row instead of crashing the run. Example: ["https://www.allbirds.com"]. | |
| onlyAvailable | No | Only available (in-stock) items — Keep this off to include out-of-stock variants and products. Turn it on to drop any variant (or, in product mode, any product) that is not currently available for sale. | |
| includeVariants | No | One row per variant — Keep this on to get one dataset row per product variant (size/color), with its own price, SKU and stock. Turn it off to get one row per product with aggregated minPrice, maxPrice and anyAvailable instead. | |
| collectionHandles | No | Collection handles (optional) — Optionally enter collection handles to scrape instead of the whole catalog, e.g. mens-shoes. Each handle is fetched from /collections/<handle>/products.json. Leave empty to scrape every product in the store via /products.json. | |
| maxProductsPerStore | No | Max products per store — Enter the maximum number of products to fetch per store, e.g. 250. Products are paged 250 at a time from products.json, so 250 is a single page; raise it for a full catalog export, up to 20000. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does add real value: it discloses cost (~$0.001 per result, billed to the caller's own Apify account, lower on paid plans) and that reads come from a public feed. It does not state rate limits or how non-Shopify stores fail, though that error behavior is covered in the parameter schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the capability and followed by the pricing note; there is no filler. The response is dense but readable and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly enumerates what comes back (title, price, compare-at price, discount %, stock, SKU, images). Combined with the fully covered parameter schema and disclosed billing, an agent has enough to invoke it correctly, though failure modes and pagination behavior are only implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already thoroughly documented in the schema (store URLs, onlyAvailable, includeVariants, collectionHandles, maxProductsPerStore). The description adds only the 'one row per product or variant' framing and does not add syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('reads any Shopify store's public products.json feed') and enumerates the returned fields, cleanly distinguishing it from the other scraper siblings (GitHub, Hacker News, app stores). An agent immediately knows what data source this targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes the context of use (any public Shopify store, one row per product or variant) and the billing model, but gives no explicit when-to-use/when-not guidance or named alternatives. The store-mode vs collection-mode choice, if anything, lives in the schema rather than the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stackexchange_scraperA
Search Stack Overflow or any Stack Exchange site's free public API by keyword or tag and get one row per question — score, tags, author, view/answer counts and the accepted answer's text. Billed to your own Apify account: ~$0.0005 per Question (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| site | No | Site — Enter the Stack Exchange site's API parameter, e.g. stackoverflow, superuser, serverfault, askubuntu, math.stackexchange, or any other site slug from https://stackexchange.com/sites. | stackoverflow |
| sortBy | No | Sort by — Choose how results are ordered before Max results per query is applied. relevance ranks by text match to the query, votes by score, creation by newest first, and activity by most recently active. Options: relevance = Relevance; votes = Votes (score); creation = Creation date; activity = Last activity. | votes |
| tagged | No | Tags (optional) — Enter tags to narrow every query to, e.g. python, playwright. Leave empty to search all tags. Combined with each entry in Search queries. | |
| queries | Yes | Search queries — Enter one search phrase per row, e.g. web scraping, playwright timeout. Each query is run separately against the chosen site and returns its own set of question rows. Example: ["web scraping"]. | |
| minScore | No | Minimum score — Enter the minimum question score (upvotes minus downvotes) to keep, e.g. 0. Questions scoring lower are dropped after fetching. | |
| maxResults | No | Max results per query — Enter the maximum number of questions to return per query, e.g. 50. | |
| includeAnswers | No | Include accepted answer — Turn this on to add the accepted answer's full text as acceptedAnswerBody. Leave off for faster, smaller runs — most questions do not have one, and this field stays null when they don't. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the pricing model (~$0.0005 per Question, billed to the caller's Apify account) and that most questions lack an accepted answer, but says nothing about rate limits, pagination, or authentication flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is retrieved and followed by cost. Dense but every clause carries information; no filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly summarizes the return shape (one row per question with score, tags, author, counts, and optional accepted answer text), which is what an agent needs. Cost disclosure compensates for the missing annotations, though operational details like rate limits remain uncovered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (site, sortBy, tagged, queries, minScore, maxResults, includeAnswers) is already documented with examples and constraints. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (Stack Overflow / any Stack Exchange site's public API) with concrete scope: keyword or tag search returning one row per question. Names the exact returned fields (score, tags, author, view/answer counts, accepted answer text), so it is unmistakable against siblings like hacker_news_scraper or github_trending_scraper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (keyword or tag search against a chosen site) but never states when to prefer this tool over an alternative, nor any prerequisites such as needing an Apify account. With sibling tools covering entirely different sources, there is little routing ambiguity, so this is adequate rather than strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_marketplace_extensionsA
VS Code Marketplace Extension Scraper returns installs, ratings, version, categories, tags, repository and license for any VS Code extension id, marketplace URL or search keyword — one row per extension. Billed to your own Apify account: ~$0.0015 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| queries | Yes | Extensions or searches — Enter what to look up, one row is returned per matched extension. Three formats work: an extension id like ms-python.python, a full marketplace URL like https://marketplace.visualstudio.com/items?itemName=esbenp.prettier-vscode, or search:<keyword> (e.g. search:prettier) to return the top matching extensions for that keyword. Example: ["ms-python.python"]. | |
| searchLimit | No | Search result limit — Enter the maximum number of extensions to return per search:<keyword> entry, e.g. 25. Extension id and URL lookups always return exactly one row and ignore this setting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses billing behavior ('Billed to your own Apify account: ~$0.0015 per result') and row cardinality, but omits authentication requirements, rate limits, and failure modes for invalid ids.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence front-loads the returned fields, followed by pricing. Efficient, though the pricing clause is secondary and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly enumerates the returned fields and per-row granularity, which is what an agent needs to interpret results. Missing only auth/billing-setup details for the Apify account.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the three accepted query formats and the searchLimit semantics are already fully documented in the schema. The description's 'one row per extension' adds a small cardinality hint but nothing the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('returns installs, ratings, version, categories, tags, repository and license') scoped to VS Code extensions, which clearly separates it from sibling store scrapers like chrome_web_store_extensions and app_store_lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when the tool applies (lookup by id, URL, or keyword) but never states when to choose it over siblings or what prerequisites exist. No explicit when/when-not guidance or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
app_store_lookup - First observed
chrome_web_store_extensions - First observed
crossref_doi_lookup - First observed
github_repo_health - First observed
github_trending_scraper - First observed
google_play_scraper - First observed
hacker_news_scraper - First observed
ical_calendar_extractor - First observed
openfda_recall_monitor - First observed
package_health_checker - First observed
shopify_products_scraper - First observed
stackexchange_scraper - First observed
vscode_marketplace_extensions
TDQS
Scored across 13 tools
Each tool targets a distinct data source or platform (npm/PyPI, GitHub, VS Code, Chrome, Google Play, App Store, CrossRef, openFDA, iCal, Shopify, Hacker News, GitHub trending, Stack Exchange). The only slight overlap is between github_repo_health and github_trending_scraper, but they serve different purposes (health metrics vs trending repos), so confusion is minimal.
Names follow a general noun/action pattern but mix conventions: some use underscores (package_health_checker), some use camelCase-like concatenation (github_repo_health), and some are verb phrases (app_store_lookup). The inconsistency is noticeable but each name remains readable and descriptive.
13 tools is a reasonable count for a multi-source data aggregation server. Each tool covers a distinct data source, so the set is well-scoped, though it could be considered slightly heavy if more sources are added without grouping.
The server covers a broad range of developer, app, and research data sources with read-only scraping/lookup operations. However, it lacks tools for write operations (e.g., posting to Hacker News, managing repositories) and some potential sources (e.g., Reddit, Product Hunt), but for a data retrieval server, the surface is fairly complete.
Maintenance
Related MCP Connectors
Give your agent live data from Twitter, Reddit, the web and GitHub. No API keys, no scraping stack.
Discover, inspect and run 63,000+ agent tools from one balance. Pay per call, no subscriptions.
Direct access to 60+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
30+ marketing data tools for AI agents: keywords, SERP, backlinks, AI visibility, app store & commerce intelligence, Reddit/LinkedIn/Facebook/YouTube research, web search, page extraction, site audit, image & video generation, one-shot marketing apps. Bring your own API key from supamarketers.com — per-tool pricing in Credits.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with real-time web capabilities including live search results, markdown web scraping, business lead generation, and detailed company information. It enables agents to bypass knowledge cutoffs by accessing current web data through a monetized Apify Actor.-
- AlicenseNot gradedqualityNot gradedmaintenanceExposes over 19,000 Apify Actors as MCP tools for web scraping, data extraction, and OSINT automation. It enables AI agents to dynamically discover and execute scrapers to collect structured data and crawl web content.MIT
- FlicenseNot gradedqualityCmaintenanceProvides AI agents with 6 tools for searching Hugging Face models, GitHub trending repos, analyzing GitHub repositories, fetching dev.to articles, Show HN launches, and Product Hunt daily launches. Built for dev tooling research and AI ecosystem analysis.-
- AlicenseNot gradedqualityCmaintenanceEnables agents to access 30 pay-per-event Apify actors for ad, video, reel, and audio transcripts; Google Trends and keyword demand; job boards; Amazon product data; and Instagram data, charging only per delivered item.MIT