Skip to main content
Glama

web_scrape_page

Read-only

Scrape Page — Extract structured data from a public web page at a user-provided http(s) URL: CSS-selector mode returns text per selector; readability mode returns the main article as clean markdown. [category: web]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYeshttp(s) URL to scrape
modeNoReadability gives you the page's main article as clean text. Selectors gives you only the specific bits you name below.readability
selectorsNoOne CSS selector per thing you want, named: {"headline": "h1", "price": ".price"}. Needed only when you pick Selectors above.
timeout_msNoHow long to wait for the page before giving up, in milliseconds (30000 = 30 seconds).
user_agentNoAdvanced: how we introduce ourselves to the site. Left blank we identify as JohnsEssentialsBot.
acknowledge_robotsNoBusiness plan: scrape the page even when the site's robots.txt asks bots to stay away. On any other plan this switch does nothing and the page is still refused.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / timeout_ms / x-ui
      Added value: +{
      +  "unit": "ms"
      +}
  2. Changed8 schema fields changed
    • changedInput schema / properties / acknowledge_robots / description
      Previous value: -"Business+ only: proceed even when robots.txt disallows the page."New value: +"Business plan: scrape the page even when the site's robots.txt asks bots to stay away. On any other plan this switch does nothing and the page is still refused."
    • changedInput schema / properties / mode / description
      Previous value: -"Chooses the response shape: selectors returns per-key text under 'data'; readability returns the article object under 'article'."New value: +"Readability gives you the page's main article as clean text. Selectors gives you only the specific bits you name below."
    • changedInput schema / properties / selectors / description
      Previous value: -"key → CSS selector map — REQUIRED (non-empty) when mode=selectors"New value: +"One CSS selector per thing you want, named: {\"headline\": \"h1\", \"price\": \".price\"}. Needed only when you pick Selectors above."
    • addedInput schema / properties / selectors / x-show-when
      Added value: +{
      +  "mode": [
      +    "selectors"
      +  ]
      +}
    • addedInput schema / properties / timeout_ms / default
      Added value: +30000
    • changedInput schema / properties / timeout_ms / description
      Previous value: -"Optional fetch timeout override in milliseconds."New value: +"How long to wait for the page before giving up, in milliseconds (30000 = 30 seconds)."
    • changedInput schema / properties / timeout_ms / minimum
      Previous value: -1New value: +1000
    • changedInput schema / properties / user_agent / description
      Previous value: -"Optional custom User-Agent header."New value: +"Advanced: how we introduce ourselves to the site. Left blank we identify as JohnsEssentialsBot."
  3. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, so the description doesn't need to restate safety. It adds meaningful behavioral context: robots.txt handling (acknowledge_robots parameter), user-agent identification, and the distinction between modes. It also discloses that the tool only works on public pages and that robots.txt refusal is enforced on non-business plans, which is valuable beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core purpose and mode distinction, then adds a category tag. It's efficient and doesn't waste words. It could be slightly more structured (e.g., separating the mode explanation), but it's appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, 100% schema coverage, and no output schema, the description covers the key decision points: mode selection, robots.txt behavior, and user-agent. It doesn't describe return format or error cases, but the schema covers parameters and the description covers the main behavioral choices. The lack of output schema means the description could mention what the result looks like, but the mode descriptions partially cover that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds value by explaining the mode semantics ('readability gives you the page's main article as clean text') and the selectors parameter's purpose ('One CSS selector per thing you want'). It doesn't add much beyond the schema, but the mode explanation helps disambiguate the two modes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scrape'), a resource ('public web page at a user-provided http(s) URL'), and explicitly distinguishes two modes (CSS-selector mode vs readability mode). It clearly differentiates from siblings like web_fetch and web_extract_table by focusing on structured extraction from a page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two modes and when each is appropriate ('CSS-selector mode returns text per selector; readability mode returns the main article as clean markdown'). It doesn't explicitly name alternatives like web_fetch or web_extract_table, but the mode guidance gives clear context for when to use this tool. It lacks explicit 'when not to use' exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources