Skip to main content
Glama

web_extract_table

Read-only

Extract Table — Extract an HTML table from a public web page at a user-provided http(s) URL, as JSON rows or CSV. [category: web]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYeshttp(s) URL containing the table
outputNoHow the table comes back: JSON gives rows you can feed to the next step, CSV gives one block of comma-separated text. Either way it is data, not a downloadable file.json
timeout_msNoHow long to wait for the page before giving up, in milliseconds (30000 = 30 seconds).
user_agentNoAdvanced: how we introduce ourselves to the site. Left blank we identify as JohnsEssentialsBot.
table_indexNoWhich table on the page, counting from 0 for the first one. If the page has fewer tables than this, the error tells you how many it found.
acknowledge_robotsNoBusiness plan: read the table even when the site's robots.txt asks bots to stay away. On any other plan this switch does nothing and the page is still refused.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / timeout_ms / x-ui
      Added value: +{
      +  "unit": "ms"
      +}
  2. Changed9 schema fields changed
    • changedInput schema / properties / acknowledge_robots / description
      Previous value: -"Business+ only: proceed even when robots.txt disallows the page."New value: +"Business plan: read the table even when the site's robots.txt asks bots to stay away. On any other plan this switch does nothing and the page is still refused."
    • changedInput schema / properties / output / description
      Previous value: -"Both values return a JSON envelope: 'json' puts a rows matrix in 'rows'; 'csv' puts one CSV string in the 'csv' field — never a file."New value: +"How the table comes back: JSON gives rows you can feed to the next step, CSV gives one block of comma-separated text. Either way it is data, not a downloadable file."
    • changedInput schema / properties / table_index / description
      Previous value: -"0-based table index"New value: +"Which table on the page, counting from 0 for the first one. If the page has fewer tables than this, the error tells you how many it found."
    • addedInput schema / properties / table_index / maximum
      Added value: +50
    • addedInput schema / properties / table_index / minimum
      Added value: +0
    • addedInput schema / properties / timeout_ms / default
      Added value: +30000
    • changedInput schema / properties / timeout_ms / description
      Previous value: -"Optional fetch timeout override in milliseconds."New value: +"How long to wait for the page before giving up, in milliseconds (30000 = 30 seconds)."
    • changedInput schema / properties / timeout_ms / minimum
      Previous value: -1New value: +1000
    • changedInput schema / properties / user_agent / description
      Previous value: -"Optional custom User-Agent header."New value: +"Advanced: how we introduce ourselves to the site. Left blank we identify as JohnsEssentialsBot."
  3. First observed

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description and parameter docs add genuine behavioral context beyond that: the public-page restriction, the bot identifying itself as 'JohnsEssentialsBot' when user_agent is blank, robots.txt acknowledgment gated to the Business plan, timeout behavior, and the table_index error message when fewer tables exist. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a tight single sentence with the core action front-loaded. Minor deductions: the opening 'Extract Table —' duplicates the tool title, and the '[category: web]' tag is low-value metadata that adds noise rather than guidance. Otherwise efficient and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The combination of the concise description and the unusually thorough parameter documentation covers the essentials for correct invocation: URL format, output shape, timeout limits, robot policy, and table selection errors. There is no output schema, but the output parameter adequately conveys the return format. The main gap is that the core description alone is thin; completeness relies heavily on the parameter docs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter carries a rich description — output distinguishes JSON (feedable to next step) from CSV (one text block), timeout_ms gives a worked example, table_index explains error behavior, and acknowledge_robots clarifies plan gating. The main description adds little beyond the schema since the schema already does the heavy lifting, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Extract') plus a clear resource ('an HTML table from a public web page at a user-provided http(s) URL') and names the output forms ('JSON rows or CSV'). This cleanly distinguishes it from sibling tools like web_fetch and web_scrape_page, which handle general page retrieval rather than structured table extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'public web page' implicitly scopes usage to publicly accessible pages, and the parameter docs reveal plan-gated behavior for robots.txt. However, the description does not explicitly name alternatives (e.g., web_fetch or web_scrape_page) or state when this tool is preferred over them. Usage context is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources