ScrapingBee
Server Details
Fetch any web page as text, markdown, HTML or a screenshot — JavaScript, CAPTCHAs and anti-bot blocking handled — and extract structured data; JSON tools for Google, Amazon, Walmart, YouTube, ChatGPT and Gemini.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 18 tools
The URL-scraping family (get_page_html, get_page_text, extract_page_data, get_screenshot, get_file) targets the same input but is cleanly separated by output format, and fast_search vs get_google_search_results is explicitly prioritized as primary/fallback. Amazon pricing vs product details is the only borderline pair, but descriptions keep boundaries clear.
A strong dominant get_* pattern runs through most tools (get_page_text, get_amazon_search_results, get_youtube_video_metadata, etc.). Minor deviations like ask_chatgpt, fast_search, and extract_page_data are still readable and follow intuitive verb_noun forms.
18 tools is reasonable for a scraping API spanning web scraping, search, Amazon/Walmart e-commerce, YouTube, LLM queries, and account usage. Each tool earns its place by covering a distinct site or output format, though it sits near the heavier end.
Coverage is broad across scraping formats, search, major e-commerce sites, YouTube, LLM, and account usage. The main gaps (batch, crawl, file output, scheduling) are explicitly deferred to the CLI, which is a deliberate scope choice rather than an oversight.
Available Tools
18 toolsask_chatgptAInspect
This tool allows you to send a prompt to OpenAI's ChatGPT via ScrapingBee and get back the generated response.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| prompt | Yes | The prompt to send to ChatGPT. | |
| search | No | Whether to enable web search capability for the prompt. | |
| add_html | No | Whether to include the full HTML of the page in the results. | |
| country_code | No | ISO country code the request should originate from (affects web results when search is enabled). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that results go 'into the conversation' and hints at HTML inclusion via the schema, but doesn't cover auth requirements, rate limits, or error behavior of the scraping layer. Adequate but incomplete for a mutation-by-proxy tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs, front-loaded with the core purpose, then the scope/alternative note. Efficient, though the CLI sentence with a full command is slightly tangential to calling this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return-value explanation isn't required. For a 5-param tool with no annotations, the description covers purpose and routing well; the main gap is behavioral detail (permissions, quotas) that an agent would need before relying on it at scale.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter (tag, prompt, search, add_html, country_code) is already documented in the schema. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (send a prompt) and resource (OpenAI's ChatGPT via ScrapingBee) with clear output (generated response). The one-query scope further distinguishes it from batch alternatives like the CLI.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the single-query scope and names the alternative (ScrapingBee CLI) for batch operations with a concrete command example. This is exactly the when-to-use/when-not guidance that helps an agent route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ask_geminiAInspect
Ask Gemini through ScrapingBee and receive citation objects when available.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| prompt | Yes | The prompt to send to Gemini. | |
| add_html | No | Whether to return the full HTML of searched pages. | |
| country_code | No | ISO country code to set request geolocation. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the one-query-per-call constraint and that citations are only returned 'when available' (a useful non-determinism warning), but says nothing about cost/credit consumption, auth requirements, or rate limits for a paid scraping-backed backend.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and return behavior, then the scope constraint, then the escape hatch to the CLI. Two tight sentences with no filler; the embedded command syntax is earned because it changes routing decisions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not needed, and the description covers scope plus the batch alternative. The main remaining gap is the absence of any hint about cost or resource consumption, which matters given the sibling get_scrapingbee_usage tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents prompt, tag, add_html and country_code, and the description adds no additional meaning beyond what is there. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Ask Gemini through ScrapingBee') and clarifies the return shape ('citation objects when available'). It differentiates from the CLI sibling implied in the text but never distinguishes itself from the similarly-purposed ask_chatgpt tool, so an agent must infer the choice between the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit scope rule ('one query per call') and a clear alternative with concrete selection conditions: many queries or writing to disk should go through the ScrapingBee CLI, with the exact command shown. It stops short of stating when this tool is preferable to ask_chatgpt or fast_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_page_dataAInspect
Scrape a URL using ScrapingBee and extract specific data using CSS or XPath selectors.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the page to scrape. | |
| wait | No | Milliseconds to wait after page load (0–35000). | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| max_cost | No | Optional credit ceiling for auto_mode (integer >= 1). The request will not escalate to a configuration costing more than this many credits. 0 means no ceiling. Ignored unless auto_mode is True. | |
| wait_for | No | CSS or XPath selector to wait for before returning. | |
| auto_mode | No | On by default — ScrapingBee automatically picks the cheapest configuration that successfully fetches the page, escalating proxy strength only as needed (you are charged only for the winning config). To choose a configuration yourself instead, use one of: auto_mode=False for the classic tier (no proxy, cheapest), premium_proxy=True, or stealth_proxy=True — the two proxy flags override auto_mode on their own, so auto_mode=False is only needed for the classic tier. Setting render_js also switches to manual. | |
| render_js | No | Whether to use headless-browser rendering. Setting this explicitly switches the request to manual configuration (disables auto_mode). | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| ai_selector | No | CSS selector to focus AI extraction on part of the page. | |
| js_scenario | No | Stringified JSON object of browser interaction instructions. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True. | |
| wait_browser | No | domcontentloaded (default), load, networkidle0, networkidle2. | domcontentloaded |
| window_width | No | Viewport width (default 1920). | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| extract_rules | Yes | A JSON string defining extraction rules. SIMPLE SYNTAX: Use {"key_name": "css_or_xpath_selector"} format. - CSS selector example: {"title": "h1", "price": "span.price"} - XPath selector example: {"title": "//h1", "items": "//div[@class='item']"} - Extract attribute value of an element using @<attribute> or if you are using extended syntax <element>@<attribute>, for example: {"image": "img@src"}, "link": {"selector": "a","output": "@href"} Note: Selectors starting with "/" are treated as XPath, otherwise CSS. EXTENDED SYNTAX: For more control, use a dict with these options: - "selector": CSS or XPath selector string (required) - "type": "item" (single element, default) or "list" (multiple elements) - "output": It is also possible to add extraction rules inside the output option in order to create powerful extractors. Extended example: { "title" : "h1", "subtitle" : "#subtitle", "articles": { "selector": ".card", "type": "list", "output": { "title": ".post-title", "link": { "selector": ".post-title", "output": "@href" }, "description": ".post-description" } } } | |
| json_response | No | Return the full JSON envelope instead of the raw body. | |
| premium_proxy | No | Manual option: use a premium proxy (middle tier of classic/premium/stealth). Setting this disables auto_mode. | |
| stealth_proxy | No | Manual option: use a stealth proxy (strongest tier of classic/premium/stealth). Setting this disables auto_mode. | |
| window_height | No | Viewport height (default 1080). | |
| block_resources | No | Block images/CSS to speed up rendering. | |
| ai_extract_rules | No | JSON string of AI extraction rules. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It adds the one-URL-per-call scope and the fact that results return into the conversation, which are useful. But it doesn't disclose cost implications, rate limits, auth requirements, or the auto_mode escalation behavior for an operation that can incur credits — gaps left mostly to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, then a clearly delimited Scope section with alternatives. Some redundancy in the CLI example list, but every line serves routing or scope; no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity 23-param tool with an output schema present (so return values needn't be explained) and 100% schema coverage, the description covers the key decision boundary (single vs batch/crawl) well. Minor gap: no note on credit/cost behavior or auth, but coverage is otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 23 parameters (including auto_mode, max_cost, and the proxy flags with their own trade-offs). The description adds no parameter syntax or format detail beyond the structured fields, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scrape) + resource (a URL) and the mechanism (ScrapingBee) plus the differentiated capability (extract specific data via CSS/XPath selectors). This clearly separates it from siblings like get_page_html and get_page_text, which presumably return whole-page content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit scope statement ('one URL per call, result returned into the conversation') and a detailed when-not list: many URLs, whole site, file output, resumable jobs, scheduled re-runs — each routed to a named CLI alternative with concrete command examples. Nothing left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fast_searchAInspect
Search the web using ScrapingBee's Fast Search API and return structured results. This is the PRIMARY tool for web searches — it returns clean organic results and top stories in under one second at a lower credit cost.
USE THIS TOOL FIRST for any general web search. Only fall back to get_google_search_results if you need specialized search types such as news, maps, Google Lens, shopping, image search, or Google AI mode.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| page | No | The page number to return. Default is 1, use page=2 for the second page, etc. | |
| search | Yes | The search query to use. Make sure to URL-encode special characters such as :, & or +. | |
| language | No | Language of the search results (Example: en, fr, de, etc.). | en |
| country_code | No | ISO 3166-1 alpha-2 country code to localize results (Example: us, fr, de, in, gb, br, etc.). | us |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses latency (<1 second), relative cost (lower credit cost), output shape (clean organic results and top stories), and the one-query-per-call scope limit. It does not mention auth/API-key requirements or rate limits, so it falls short of a full behavioral profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the primary purpose and cost/latency benefit, then the routing rule, then the scope constraint with the CLI escape hatch. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with a full schema and an output schema already covering return values, the description supplies what structured fields cannot: primary-vs-alternative routing, cost/latency expectations, and the single-query scope limit. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (search, page, language, country_code, tag) are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search the web using ScrapingBee's Fast Search API') and explicitly positions itself as the PRIMARY general web-search tool. It names the sibling it is not (get_google_search_results) and the conditions that select it, so an agent can pick it apart from siblings without reading either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('USE THIS TOOL FIRST for any general web search') and a clear when-not with the alternative named ('Only fall back to get_google_search_results if you need ... news, maps, Google Lens, shopping, image search, or Google AI mode'). Routing guidance is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_amazon_pricingAInspect
Retrieve current Amazon product pricing and seller offers for a given ASIN.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| asin | Yes | Required 10-character ASIN from an Amazon product URL. | |
| domain | No | Amazon domain (default com). | com |
| country | No | ISO country code for localization. Do not combine with a matching domain (e.g., country=fr&domain=fr). | |
| add_html | No | Include page HTML in the response. | |
| currency | No | ISO 4217 display currency. | |
| language | No | ISO language code. | |
| zip_code | No | ZIP code for delivery localization. | |
| light_request | No | Use a light request (default True) or browser-rendered (False). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It usefully discloses that the tool handles one query per call and that results otherwise enter the conversation, with batch and disk-output alternatives named. It does not mention authentication, rate limits, or response behavior beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded, with the core purpose first and the scope/CLI alternative second. Every sentence earns its place and no unnecessary detail is included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with full schema descriptions and an output schema, the description provides enough context to invoke it correctly: it names the required ASIN and explains the one-query-per-call model. It is slightly incomplete because it does not clarify when to prefer this tool over similar product-detail siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters, including the required ASIN. The description only restates the ASIN input and adds no further parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: retrieve Amazon product pricing and seller offers for a given ASIN. It is clear on its own, but it does not explicitly distinguish itself from the sibling get_amazon_product_details tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear scope constraint (one query per call) and routes bulk or disk-output use cases to the ScrapingBee CLI. It does not, however, compare against sibling MCP tools such as get_amazon_product_details.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_amazon_product_detailsAInspect
Scrape Amazon product details using ScrapingBee.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| query | Yes | Search term (must be a valid 10-character ASIN code) | |
| domain | No | The Amazon domain to use for the search (e.g., com, co.uk, de). | |
| country | No | The country code for localization (e.g., us, uk, de). Do not combine with a matching domain (e.g., country=fr&domain=fr). | |
| add_html | No | Whether to return the HTML along with the product details. | |
| currency | No | The currency code (ISO 4217) to display results (e.g., USD, GBP, EUR). | |
| language | No | The language code to display results (e.g., en, fr, de). | |
| zip_code | No | The zip code to use for delivery localization. | |
| screenshot | No | Force a browser screenshot (returns base64 image). | |
| light_request | No | Whether to use a light request or not. | |
| autoselect_variant | No | Whether to automatically select product variant if applicable. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses scoping (one query per call) and where output lands (conversation vs disk via CLI), but says nothing about authentication/API key needs, credit consumption or rate limits for the ScrapingBee backend, or error behavior—notable gaps for a paid scraping call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, both earning their place: the first states purpose, the second states scope and the batching/disk alternative with a concrete command example. It is front-loaded and free of repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the schema fully documents parameters. But with zero annotation coverage on a network-scraping tool with 11 parameters, the description leaves authentication, cost/credit implications, and failure modes unaddressed, which is the main completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all 11 parameters documented in the schema itself (including the ASIN requirement, domain/country conflict rule, and localization options). The description adds no parameter-level meaning beyond the generic word 'query', so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Scrape Amazon product details using ScrapingBee'), which is clearly distinct from the generic page/search siblings. However, it never names the closest siblings (get_amazon_search_results, get_amazon_pricing, get_walmart_product_details), so the agent must infer the boundary rather than being told it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the usage constraint ('one query per call') and names the alternative (the ScrapingBee CLI) with the conditions that select it: many queries in one pass, or writing to disk instead of the conversation. It does not, however, explain when to prefer this over sibling tools like get_amazon_search_results or get_amazon_pricing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_amazon_search_resultsAInspect
Scrape Amazon search results using ScrapingBee.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| pages | No | The number of pages to scrape. | |
| query | Yes | The search query to perform. | |
| domain | No | The Amazon domain to use for the search (e.g., com, co.uk, de). | com |
| country | No | The country code for localization (e.g., us, uk, de). Do not combine with a matching domain (e.g., country=fr&domain=fr). | |
| sort_by | No | The sort order to use for the search (most_recent, price_low_to_high, price_high_to_low, featured, average_review, bestsellers). | |
| add_html | No | Whether to return the HTML along with the search results. | |
| currency | No | The currency code (ISO 4217) to display results (e.g., USD, GBP, EUR). | |
| language | No | The language code to display results (e.g., en, fr, de). | |
| zip_code | No | The ZIP code to use for delivery localization. | |
| screenshot | No | Force a browser screenshot (returns base64 image). | |
| start_page | No | The page number to start scraping from. | |
| category_id | No | The category ID to use for the search. | |
| merchant_id | No | The merchant ID to use for the search. | |
| light_request | No | Whether to use a light request or not. | |
| autoselect_variant | No | Automatically select the default/most-popular variant. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full behavioral burden, and it does disclose the 'one query per call' constraint and the disk-vs-conversation tradeoff. However, for a scraping tool it omits cost/rate limits, credential requirements, and anti-bot/throttling behavior, leaving meaningful gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads verb+resource, then adds the scope constraint and CLI escape hatch in a compact flow with no filler. The embedded CLI command is the only element that inflates it slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema documents all parameters. But with no annotations and no output/permission context, the description leaves behavioral details (auth, rate limits, cost) uncovered for a 16-parameter scraping tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 16 parameters, including the non-obvious country/domain conflict rule, so the schema does the heavy lifting. The description adds no parameter-level meaning beyond it, which is the expected baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Scrape') and resource ('Amazon search results') plus the backing service (ScrapingBee). This distinguishes it reasonably from get_amazon_product_details and get_amazon_pricing, though it never explicitly names those siblings to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly defines the scope ('one query per call') and routes the agent to an alternative (the ScrapingBee CLI) with the precise conditions that select it (many queries in one pass, or write to disk). It stops short of comparing against the sibling search/product tools, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fileBInspect
Fetches a file from a URL and returns it in a format suitable for FastMcP.
This tool can be used to fetch any file type, including images, PDFs, etc. The framework will handle the encoding and packaging of the file.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the file to fetch. | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| render_js | No | Whether to use headless-browser rendering (default False). | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy is True. | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| premium_proxy | No | Whether to use a premium proxy for the request. Use it only if you are getting blocked by the website. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that output is returned into the conversation, that the framework handles encoding/packaging, and that batch/crawl/file-output/scheduling are unsupported — all useful negative traits. It omits permission/auth requirements, rate limits, failure behavior, and payload/size limits for large files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the usage constraints follow in a logical order. The four CLI examples plus the closing 'This server has no batch, crawl, file-output or scheduling equivalent' are somewhat redundant with each other, costing a little efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description covers scope and explicit non-capabilities but leaves out return-shape expectations, error handling, and auth prerequisites. It is adequate, not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 9 parameters, so the schema already documents defaults, ranges, and semantics thoroughly. The description adds only 'any file type' and 'one URL per call,' which do not extend meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb+resource ('fetches a file from a URL') and states the return contract ('format suitable for FastMcP'). The second sentence clarifies it works for any file type including images and PDFs, which implicitly separates it from page-text/HTML siblings, but it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear 'when-not' guidance for the CLI (batch, crawl, file-output, scheduling) with concrete command examples, which is genuinely useful. However, it offers no routing guidance against the in-server siblings an agent would actually weigh it against (get_page_text, get_page_html, get_screenshot, extract_page_data), leaving the most likely confusion unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_google_search_resultsAInspect
This is a FALLBACK tool — use fast_search first for general web searches. Only use this tool when you need specialized search types that fast_search does not support: news, maps, Google Lens, shopping, image search, or Google AI mode.
Scrape Google search results using ScrapingBee and return the results. This tool can scrape normal results, news results, maps results, search using google lens, shopping related results, image results, and get the result from Google's AI mode. It can even return the HTML along with the search results.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| nfpr | No | Whether to disable Google's autocorrection feature or not. | |
| page | No | The page number to return. | |
| pages | No | Number of pages to aggregate (max 10, default 1). | |
| device | No | The device to use for the search (desktop, mobile). | desktop |
| radius | No | Radius in meters (requires latitude and longitude). | |
| search | Yes | The search query to use. | |
| sort_by | No | Shopping sort order: relevance, reviews, price_asc, price_desc. | |
| add_html | No | Whether to return the HTML along with the search results. | |
| language | No | The language to use for the search (Example: en, fr, de, etc.). | en |
| latitude | No | Latitude in decimal degrees for geographic searches. | |
| longitude | No | Longitude in decimal degrees for geographic searches. | |
| max_price | No | Shopping maximum price filter. | |
| min_price | No | Shopping minimum price filter. | |
| date_range | No | Filter by date range: past_hour, past_day, past_week, past_month, past_year. Only for classic, news, and images. | |
| search_type | No | The search type to use (classic: normal results, news: news results [not available if device is mobile], maps: maps results, lens: search using google lens [requires image url as search parameter], shopping: shopping related results, images: image results, ai_mode: get the result from Google's AI mode, ads: paid-ad results). | classic |
| country_code | No | The country code to use for the search (Example: us, fr, de, etc.). | us |
| extra_params | No | Extra parameters to pass to the search (Example: tbs=qdr:d&udm=7). | |
| light_request | No | Whether to use a light request or not. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the scraping dependency (ScrapingBee), the option to return raw HTML, and that scope is one query per call. It does not mention rate limits, auth/API-key needs, pagination behavior, or failure modes for a network-scraping tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Fallback routing is front-loaded, then capability and scope follow, which is the right order. It is slightly verbose — the capability sentence largely restates the search_type enum — and the inline CLI example is somewhat tangential, but nothing is wasted outright.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the description covers the routing decision, capability envelope, and per-call scope. Combined with a 100%-documented 19-param schema, an agent has enough to invoke it correctly; only operational caveats like rate limits are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 19 well-described params, so the schema already explains search types, date_range, geo, and shopping filters. The description reinforces a few of these (search types, one query per call) but adds no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scrape) and resource (Google search results), plus the range of result types (news, maps, lens, shopping, images, ai_mode) it can produce. It also explicitly positions itself against the sibling fast_search, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Front-loads the routing rule: try fast_search first, use this only for specialized search types fast_search doesn't support, and lists those types. It also names the ScrapingBee CLI as the alternative for multi-query/batch-to-disk work, covering the when-not case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_page_htmlAInspect
Scrape a URL using ScrapingBee and return the page content in HTML.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the page to scrape. | |
| wait | No | Milliseconds to wait after page load (0–35000). | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| ai_query | No | Natural-language question for AI extraction. | |
| max_cost | No | Optional credit ceiling for auto_mode (integer >= 1). The request will not escalate to a configuration costing more than this many credits. 0 means no ceiling. Ignored unless auto_mode is True. | |
| wait_for | No | CSS or XPath selector to wait for before returning. | |
| auto_mode | No | On by default — ScrapingBee automatically picks the cheapest configuration that successfully fetches the page, escalating proxy strength only as needed (you are charged only for the winning config). To choose a configuration yourself instead, use one of: auto_mode=False for the classic tier (no proxy, cheapest), premium_proxy=True, or stealth_proxy=True — the two proxy flags override auto_mode on their own, so auto_mode=False is only needed for the classic tier. Setting render_js also switches to manual. | |
| render_js | No | Whether to use headless-browser rendering. Omit to use the API default (true). Setting this explicitly switches the request to manual configuration (disables auto_mode). | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| ai_selector | No | CSS selector to focus AI extraction on part of the page. | |
| js_scenario | No | Stringified JSON object of browser interaction instructions. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True. | |
| wait_browser | No | domcontentloaded (default), load, networkidle0, networkidle2. | domcontentloaded |
| window_width | No | Viewport width (default 1920). | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| premium_proxy | No | Manual option: use a premium proxy (middle tier of classic/premium/stealth). Setting this disables auto_mode. | |
| stealth_proxy | No | Manual option: use a stealth proxy (strongest tier of classic/premium/stealth). Setting this disables auto_mode. | |
| window_height | No | Viewport height (default 1080). | |
| block_resources | No | Block images/CSS to speed up rendering. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it does provide useful scope and result-delivery context (one URL per call, result into the conversation). However, it omits key behavioral traits like credit consumption, proxy escalation, or whether the operation is purely read-only, leaving gaps that the rich input schema only partially fills.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then provides a structured scope statement and CLI alternatives. It is somewhat longer than strictly necessary due to multiple CLI command examples, but every sentence supports routing or scope decisions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of a full output schema and rich parameter descriptions, the description is largely complete for calling the tool correctly. It covers scope and server-vs-CLI boundaries, though it does not address sibling tools like get_page_text or extract_page_data, which could matter for agent tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for 21 parameters, so the schema already documents parameter meanings in detail. The description adds only marginal semantic value, notably that one URL is processed per call, which is already implied by the single required url parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (scrape), resource (URL), and output format (HTML), making the tool's function immediately clear. It also implicitly distinguishes itself from sibling tools like get_page_text by specifying HTML output rather than extracted text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly limits scope to one URL per call and routes batch/crawl/file-output/scheduling use cases to the ScrapingBee CLI with concrete command examples. It does not, however, mention when to use sibling tools such as get_page_text or extract_page_data instead, leaving some alternative selection ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_page_textAInspect
Scrape a URL using ScrapingBee and return the page content in Markdown or Text format.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the page to scrape. | |
| wait | No | Milliseconds to wait after page load (0–35000). | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| ai_query | No | Natural-language question for AI extraction. | |
| max_cost | No | Optional credit ceiling for auto_mode (integer >= 1). The request will not escalate to a configuration costing more than this many credits. 0 means no ceiling. Ignored unless auto_mode is True. | |
| wait_for | No | CSS or XPath selector to wait for before returning. | |
| auto_mode | No | On by default — ScrapingBee automatically picks the cheapest configuration that successfully fetches the page, escalating proxy strength only as needed (you are charged only for the winning config). To choose a configuration yourself instead, use one of: auto_mode=False for the classic tier (no proxy, cheapest), premium_proxy=True, or stealth_proxy=True — the two proxy flags override auto_mode on their own, so auto_mode=False is only needed for the classic tier. Setting render_js also switches to manual. | |
| render_js | No | Whether to use headless-browser rendering. Omit to use the API default (true). Setting this explicitly switches the request to manual configuration (disables auto_mode). | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| ai_selector | No | CSS selector to focus AI extraction on part of the page. | |
| js_scenario | No | Stringified JSON object of browser interaction instructions. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True. | |
| wait_browser | No | domcontentloaded (default), load, networkidle0, networkidle2. | domcontentloaded |
| window_width | No | Viewport width (default 1920). | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| premium_proxy | No | Manual option: use a premium proxy (middle tier of classic/premium/stealth). Setting this disables auto_mode. | |
| stealth_proxy | No | Manual option: use a stealth proxy (strongest tier of classic/premium/stealth). Setting this disables auto_mode. | |
| window_height | No | Viewport height (default 1080). | |
| block_resources | No | Block images/CSS to speed up rendering. | |
| return_page_text | No | Whether to return the page content in Text format. | |
| return_page_markdown | No | Whether to return the page content in Markdown format. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it declares the one-URL-per-call scope, states the result is returned into the conversation, and explicitly disclaims batch, crawl, file-output and scheduling capability. It still says nothing about auth requirements, credit cost, or error behavior that an agent invoking a paid scraper would want.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and scope are front-loaded in two tight sentences, then the boundary cases. The inline CLI command examples are slightly verbose but each illustrates a distinct unsupported capability, so they largely earn their space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 23-parameter tool with a fully documented schema and an output schema, the remaining burden is scope and capability boundaries, which the description covers well. What is missing is the relationship to sibling scraping tools, the main ambiguity an agent faces here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 23 parameters, so the schema already documents wait, cookies, auto_mode, proxies, and return format in detail. The description only reiterates the Markdown/Text output distinction and adds no syntax or format meaning beyond structured data, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Scrape a URL using ScrapingBee") plus the return format (Markdown or Text), which implicitly distinguishes it from the raw-HTML sibling. However, it never names get_page_html or extract_page_data, so the agent must infer differentiation itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-not-this-tool: many URLs, whole sites, file output, resumable jobs, scheduled re-runs, each with the exact CLI alternative and command. It stops short of routing between sibling MCP tools (get_page_html vs extract_page_data vs this one), which is the more likely agent decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scrapingbee_usageAInspect
Read real-time credit and concurrency usage for the authenticated ScrapingBee account.
Note: This endpoint is rate-limited to 6 calls per minute and does not count against the account's concurrency limit.
Returns: max_api_credit, used_api_credit, max_concurrency, current_concurrency, renewal_subscription_date.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and discloses a key operational trait: the rate limit of 6 calls per minute and that calls do not count against concurrency. It stops short of details like error behavior on limit exceed, but covers the most important runtime constraint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then a useful rate-limit note, then return fields. Every sentence contributes operational or output context with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with an output schema, the description is complete: it states the resource, the account context, a critical rate limit, and the returned fields. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so no parameter semantics are needed. The baseline for zero-parameter tools is 4, and the description does not misrepresent the absence of inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (Read) and resource (real-time credit and concurrency usage) for the authenticated ScrapingBee account. It is immediately clear what the tool does, and no sibling tool covers the same purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning the authenticated account and real-time nature, but offers no explicit when-to-use or when-not-to-use guidance. Since no alternative tool exists for this exact resource, the implied context is adequate but not fully developed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screenshotAInspect
Scrape a URL using ScrapingBee and return it's screenshot.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the page to scrape. | |
| wait | No | Milliseconds to wait after page load (0–35000). | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| wait_for | No | CSS or XPath selector to wait for before returning. | |
| render_js | No | Whether to use headless-browser rendering. | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| js_scenario | No | Stringified JSON object of browser interaction instructions. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True. | |
| wait_browser | No | domcontentloaded (default), load, networkidle0, networkidle2. | domcontentloaded |
| window_width | No | Viewport width (default 1920). | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| json_response | No | Deprecated, removed in Version 3. Return the full JSON envelope (page content plus the base64 image under its "screenshot" key) instead of the raw image. | |
| premium_proxy | No | Whether to use a premium proxy for the request. Use it only if you are getting blocked by the website, that is when you get a 500 status code. | |
| stealth_proxy | No | Whether to use a stealth proxy for the request. Use it only if you are getting blocked by the website while using premium_proxy, that is when you get a 500 status code. | |
| window_height | No | Viewport height (default 1080). | |
| block_resources | No | Block images/CSS to speed up rendering. | |
| screenshot_selector | No | A CSS selector to take a screenshot of a specific element. | |
| screenshot_full_page | No | Whether to take a screenshot of the full page or just the viewport. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the key behavioral traits: the result is returned into the conversation and the server has no batch, crawl, file-output, or scheduling equivalent. However, it omits auth requirements, rate limits, and failure behavior, and the screenshot format details are left entirely to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and scope are front-loaded in the opening sentences, and the CLI examples earn their place by concretely illustrating the out-of-scope cases. It is slightly longer than necessary for a definition whose remaining detail lives in a fully documented schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter tool with no output schema, the description does a reasonable job: it says what comes back (a screenshot into the conversation) and bounds what the tool cannot do. It stops short of describing the return format or image encoding, which an agent might want given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 20 parameters are already documented in the schema with defaults, ranges, and enum-like values. The description adds no parameter-level meaning (e.g., which combination yields a full-page vs viewport shot), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scrape a URL) and the exact resource it returns (a screenshot), which cleanly distinguishes it from siblings like get_page_html and get_page_text that also scrape URLs but return different outputs. An agent can identify this as the screenshot-returning tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly defines scope (one URL per call, result into the conversation) and names the alternative for out-of-scope work (the ScrapingBee CLI) with concrete conditions: many URLs, whole-site crawl, file output, resumable jobs, scheduled re-runs. It does not, however, contrast against in-server siblings such as get_page_html or get_page_text, so the routing guidance is one-sided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_walmart_product_detailsAInspect
Scrape Walmart product details using ScrapingBee.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| device | No | The device to use for the request (desktop, mobile, tablet). | desktop |
| domain | No | The domain to use for the search (Example: com, ca, com.mx, etc.) for localization. | |
| add_html | No | Whether to return the HTML along with the product details. | |
| store_id | No | Specific Walmart store ID for localization. | |
| product_id | Yes | The unique identifier for the Walmart product. | |
| screenshot | No | Force a browser screenshot (returns base64 image). | |
| delivery_zip | No | The zip code to use for delivery localization. | |
| light_request | No | Whether to use a light request or not. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the third-party scraping backend and the one-query-per-call scope, which is useful, but says nothing about credential/auth requirements, credit or rate-limit costs, or failure/timeout behavior for a paid scraping call. Adequate but materially incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the scope constraint and the CLI alternative. The inline CLI example is slightly verbose but earns its place by being immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the parameter surface is fully documented. However, for a 9-parameter paid scraping tool with zero annotations, the description omits cost/credit implications, auth prerequisites, and error behavior, leaving gaps an agent should know before invoking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 9 well-documented parameters, including defaults for device, domain, store_id, delivery_zip, add_html, screenshot, and light_request. The description adds no parameter meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Scrape Walmart product details') and names the backing service (ScrapingBee), which is enough to distinguish it from sibling listing tools like get_walmart_search_results. It stops short of explicitly contrasting with those siblings, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage constraint ('Scope: one query per call') and names a concrete alternative with its condition (use the ScrapingBee CLI for many queries or disk output). It does not address when to prefer this over get_walmart_search_results or get_amazon_product_details, so guidance is clear but not complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_walmart_search_resultsAInspect
Scrape Walmart search results using ScrapingBee.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| query | Yes | The search query to perform. | |
| device | No | The device to use for the search (desktop, mobile, tablet). | desktop |
| domain | No | The domain to use for the search (Example: com, ca, com.mx, etc.) for localization. | |
| sort_by | No | The sort order to use for the search (best_match, price_low, price_high, best_seller). | best_match |
| add_html | No | Whether to return the HTML along with the search results. | |
| store_id | No | Specific Walmart store ID for localization. | |
| max_price | No | The maximum price to use for the search. | |
| min_price | No | The minimum price to use for the search. | |
| screenshot | No | Force a browser screenshot (returns base64 image). | |
| start_page | No | The page number to return (default 1). | |
| delivery_zip | No | The zip code to use for delivery localization. | |
| light_request | No | Whether to use a light request or not. | |
| fulfillment_type | No | The fulfillment type to use for the search (in_store). | |
| fulfillment_speed | No | The fulfillment speed to use for the search (today, tomorrow, 2_days, anytime). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that results land in the conversation rather than on disk, and that the call is a single-query scope, but says nothing about rate limits, latency, cost, paging behavior across start_page, or failure modes for a tool that hits an external scraping service.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then scope, in two tight sentences with no filler. The embedded CLI invocation is slightly extraneous for an agent that cannot shell out, but it is short and frames the batching alternative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the schema fully documents all 15 parameters. The scope and batching guidance cover the main decision an agent faces, leaving only minor behavioral gaps for an unannotated external-scrape tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 15 parameters, so the schema already documents query, device, domain, sort_by, price bounds, localization, and screenshot behavior. The description adds no parameter-level detail, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Scrape") and resource ("Walmart search results") and names the backend (ScrapingBee). This is distinguishable at a glance from siblings like get_walmart_product_details or get_amazon_search_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit scope rule (one query per call) and routes batch/multi-query work to an alternative (the ScrapingBee CLI) with a concrete condition. It does not, however, clarify when to prefer this over the sibling get_walmart_product_details, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_youtube_search_resultsAInspect
Search YouTube and return structured results.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| hd | No | Return only HD videos. | |
| hdr | No | Return only HDR videos. | |
| tag | No | Response-header label only. | |
| live | No | Return only live streams. | |
| is_3d | No | Return only 3D videos. | |
| is_4k | No | Return only 4K videos. | |
| vr180 | No | Return only VR180 videos. | |
| is_360 | No | Return only 360-degree videos. | |
| search | Yes | The search query. | |
| sort_by | No | Sort order: 'relevance' (default), 'rating', 'view_count', 'upload_date'. | relevance |
| duration | No | Filter by duration: '<4' (short), '4-20' (medium), '>20' (long). | |
| location | No | Return only videos with location metadata. | |
| purchased | No | Return only purchased movies. | |
| subtitles | No | Return only videos with subtitles/captions. | |
| result_type | No | Filter by type: 'video', 'channel', 'playlist', 'movie'. | |
| upload_date | No | Filter by date: 'today', 'last_hour', 'this_week', 'this_month', 'this_year'. | |
| creative_commons | No | Return only videos with Creative Commons license. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the one-query-per-call constraint and the batch alternative, but says nothing about return format, pagination, or rate limits; the safety profile is only inferable from it being a search (read) operation. Adds some context but leaves meaningful behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and followed by scope plus the batch alternative. No filler, and the CLI command example earns its place, though the scope note is slightly terse relative to the pool of filters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter search tool with 100% schema coverage and an output schema (so return values need not be described), the description covers purpose, per-call scope, and the batch alternative. The main omission is sibling differentiation among the YouTube tools, but the structured data carries the rest.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — all 17 parameters (filters like hd, live, duration, sort_by, upload_date) are fully documented in the schema. The description adds no parameter meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search YouTube and return structured results'), immediately telling the agent this is a search tool. It does not, however, explicitly distinguish itself from siblings like get_youtube_video_metadata or get_youtube_video_subtitles, so differentiation is only implicit via the 'search' verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the scope ('one query per call') and names the alternative path (the ScrapingBee CLI) with a concrete condition for choosing it — running many queries or writing to disk. It gives no guidance against the YouTube-specific siblings, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_youtube_video_metadataAInspect
Fetch structured metadata for a YouTube video including title, description, views, likes, channel info, and technical details.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | The unique YouTube video ID (e.g., 'rfscVS0vtbw'). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the important 'one query per call' scope constraint, but says nothing about auth/API key requirements, rate limits, or failure behavior for invalid video IDs. Return values are covered by the output schema, so that gap is excused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, with the scope constraint and CLI alternative following. Every sentence earns its place, though the CLI command snippet is slightly bulky for a tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with a full output schema and 100% schema description coverage, the description is nearly complete. It covers what is returned, the call-scope constraint, and the batch alternative; only peripheral details like auth requirements are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and its schema description fully documents the expected video ID format with an example ('rfscVS0vtbw'). Schema coverage is 100%, so the schema does the work and the description needs to add nothing; baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('structured metadata for a YouTube video') and enumerates the returned fields (title, description, views, likes, channel info, technical details). An agent can distinguish it from get_youtube_search_results and get_youtube_video_subtitles without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes the tool to one query per call and routes batch/disk-output use cases to the ScrapingBee CLI with a concrete command example. It stops short of pointing to the sibling MCP tools (e.g., get_youtube_search_results) that cover adjacent intents, so it is clear context rather than full when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_youtube_video_subtitlesAInspect
Retrieve subtitles for one YouTube video.
Scope: one query per call. To run many queries in one pass, or to write
results to disk instead of into the conversation, use the ScrapingBee CLI —
scrapingbee <command> --input-file queries.txt --output-dir results.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Response-header label only. | |
| language | No | Language code (e.g., 'en', 'es', 'fr'). Default is 'en'. | en |
| video_id | Yes | The unique YouTube video ID. | |
| subtitle_origin | No | 'auto_generated' (default) or 'uploader_provided'. | auto_generated |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the one-query-per-call scope limit and where results land (conversation vs. disk), which is genuine added context, but it says nothing about failure modes (missing subtitles, unavailable language), rate limits, or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, followed by the scope constraint and the alternative. The embedded CLI snippet is slightly heavy for the payload but is load-bearing and earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and all four parameters are documented in the schema. The only real gap is behavioral detail around language/subtitle_origin fallback, which is minor given the structured data available.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so tag, language, video_id, and subtitle_origin are all already documented. The description adds no syntax or behavior detail for language fallback or subtitle_origin selection, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Retrieve") plus resource ("subtitles") plus scope ("one YouTube video") makes the purpose instantly clear. It does not explicitly distinguish itself from nearby siblings like get_youtube_video_metadata or get_youtube_search_results, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operating constraint (one query per call) and names an escape hatch — the ScrapingBee CLI — with the two conditions that select it (many queries, or writing to disk). The alternative offered is a CLI rather than a sibling MCP tool, so it gives clear context without exhausting the when/when-not space.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
- First observed
ask_chatgpt - First observed
ask_gemini - First observed
extract_page_data - First observed
fast_search - First observed
get_amazon_pricing - First observed
get_amazon_product_details - First observed
get_amazon_search_results - First observed
get_file - First observed
get_google_search_results - First observed
get_page_html - First observed
get_page_text - First observed
get_scrapingbee_usage - First observed
get_screenshot - First observed
get_walmart_product_details - First observed
get_walmart_search_results - First observed
get_youtube_search_results - First observed
get_youtube_video_metadata - First observed
get_youtube_video_subtitles
Publisher details
- Operator
- ScrapingBee · Publisher source
- Operator website
- https://www.scrapingbee.com · Publisher source
- Vendor relationship
- First-party · Publisher source
- Documentation
- https://mcp.scrapingbee.com · Publisher source
- Trust center
- Not available
- Restrictions
- Not applicable
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1624 npm1MIT
- AlicenseCqualityBmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs1114 npm40 PyPIMIT
- AlicenseAqualityCmaintenanceRevnuvo Company Intelligence tells AI agents what changed at a company, with evidence. It observes company websites, technologies, and DNS over time and returns timestamped, confidence-aware changes, signals, and monitoring.9MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.