Skip to main content
Glama

Bob Research Tools

bob_extract

Read-only

Returns the main text of one public web page as clean markdown, unchanged and not summarized, up to 8,000 characters, as {url, markdown, truncated}, usually in 2 to 5 seconds. Use it when you need the page's own words; use bob_summarize for a short version. Pages that need a login cannot be read; long pages are cut at 8,000 characters and flagged truncated. Each call needs an x402 payment; if the page cannot be read or the URL is invalid, it fails and you are not charged.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPublic http or https address of the page to extract; private or local addresses are refused.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe page that was read.
markdownYesMain content of the page as markdown, up to 8,000 characters.
truncatedNoTrue if the page was longer than the cap.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "additionalProperties": true,
      +  "properties": {
      +    "markdown": {
      +      "description": "Main content of the page as markdown, up to 8,000 characters.",
      +      "type": "string"
      +    },
      +    "truncated": {
      +      "description": "True if the page was longer than the cap.",
      +      "type": "boolean"
      +    },
      +    "url": {
      +      "description": "The page that was read.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "url",
      +    "markdown"
      +  ],
      +  "type": "object"
      +}
  2. Changed1 schema field changed
    • addedInput schema / properties / url / description
      Added value: +"Public http or https address of the page to extract; private or local addresses are refused."
  3. First observed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/open-world/non-destructive, but the description adds substantial non-structured context: per-call x402 payment, no charge on failure, 2-5 second latency, and the 8,000-character truncation with a 'truncated' flag. It stops short of describing caching or rate limits, but this is well beyond the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the return contract, then usage routing, then constraints and the payment consequence. No filler and no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists and is not redundantly explained, annotations cover the safety profile, and the description still adds payment model, latency, truncation, and failure behavior. An agent has everything needed to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single well-described 'url' parameter, so the schema already carries the load; the description's mention of clean markdown output and truncation is not parameter syntax. Baseline 3 applies when the schema does the parameter work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (returns/extracts) and resource (main text of one public web page) plus the exact output shape and scope limit (8,000 characters, not summarized). It explicitly contrasts with 'not summarized', which cleanly separates it from the sibling bob_summarize.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Directly instructs when to use it ('when you need the page's own words') and names the alternative plus its selection condition ('use bob_summarize for a short version'). It also lists failure conditions (login-required pages, invalid URLs) so the agent knows when the call will not yield content.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources