Skip to main content
Glama

Fetch Page Content

infobroker_fetch_page
Idempotent

Fetch a URL to extract clean readable text, ranked passages for a question, or the page's last-updated date.

Instructions

Fetch a URL and extract clean content via a renderer (Jina Reader by default; native-HTTP, Wikipedia, Internet Archive, arXiv, and Stack Exchange alternatives). Use when you have a URL and need readable text, passages ranked against a question, or the page's last-updated date. Do NOT use for topic search (use infobroker_search_web) or cross-source claim verification (use infobroker_verify_claims). question returns passages ranked against it, sized by passage_size and capped by max_passages; crawl bounds the same-origin crawl to config caps; extract adds JSON-LD, OpenGraph, and microdata; renderer selects the backend, and native_fetch is the keyless fallback when Jina is throttled or anti-bot challenged. max_length truncates the returned text only; a truncated page is written in full to a temp file. Makes external HTTP calls and needs no API key. Fetched pages auto-index into the knowledge base unless the content policy flags them (see infobroker_manage_kb): flag mode returns without storage, block mode refuses. Unreachable URLs return an [ERROR] envelope with remediation; success returns an [OK] envelope with status, provider, results, and meta.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesURL to fetch: a single URL, or up to five URLs fetched in parallel
crawlNoBounded same-origin crawl: recursively fetch same-origin pages up to config caps (default off)
extractNoReturn structured metadata (JSON-LD, OpenGraph, microdata) alongside the content (default off)
questionNoQuestion to extract ranked passages for, instead of returning the whole page
rendererNoRenderer: jina (default), native_fetch, wikipedia, internet_archive, arxiv, or stack_exchange
max_lengthNoMaximum characters to return (default 50000)
detect_dateNoDetect and report the page's last-updated date (default from config)
max_passagesNoNumber of passages to return (default from config)
passage_sizeNoTarget words per passage (default from config)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.1.2
    • addedInput schema / properties / crawl
      Added value: +{
      +  "default": false,
      +  "description": "Bounded same-origin crawl: recursively fetch same-origin pages up to config caps (default off)",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / extract
      Added value: +{
      +  "default": false,
      +  "description": "Return structured metadata (JSON-LD, OpenGraph, microdata) alongside the content (default off)",
      +  "type": "boolean"
      +}
  2. Changed9 schema fields changedv0.1.1
    • addedInput schema / properties / detect_date
      Added value: +{
      +  "description": "Detect and report the page's last-updated date (default from config)",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / max_length / description
      Added value: +"Maximum characters to return (default 50000)"
    • addedInput schema / properties / max_passages
      Added value: +{
      +  "description": "Number of passages to return (default from config)",
      +  "type": "number"
      +}
    • addedInput schema / properties / passage_size
      Added value: +{
      +  "description": "Target words per passage (default from config)",
      +  "type": "number"
      +}
    • addedInput schema / properties / question
      Added value: +{
      +  "description": "Question to extract ranked passages for, instead of returning the whole page",
      +  "type": "string"
      +}
    • addedInput schema / properties / renderer / description
      Added value: +"Renderer: jina (default), native_fetch, wikipedia, internet_archive, arxiv, or stack_exchange"
    • addedInput schema / properties / url / anyOf
      Added value: +[
      +  {
      +    "description": "URL to fetch",
      +    "type": "string"
      +  },
      +  {
      +    "description": "Multiple URLs to fetch in parallel (max 5)",
      +    "items": {
      +      "type": "string"
      +    },
      +    "maxItems": 5,
      +    "type": "array"
      +  }
      +]
    • changedInput schema / properties / url / description
      Previous value: -"URL to fetch"New value: +"URL to fetch: a single URL, or up to five URLs fetched in parallel"
    • removedInput schema / properties / url / type
      Removed value: -"string"
  3. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false) leave room, and the description fills it substantially: external HTTP calls with no API key, native_fetch as the keyless fallback when Jina is throttled or anti-bot challenged, auto-indexing into the knowledge base with content-policy flag/block semantics, temp-file write on truncation, and the [OK]/[ERROR] envelope shapes. This is well beyond what annotations convey and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then usage/exclusions, then parameter behavior, then envelope/indexing behavior — a sensible ordering with no filler sentences. It is dense and somewhat overstuffed (many semicolon-chained clauses), but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing both the [OK] and [ERROR] envelopes and their fields. Combined with the mutation/indexing side effects and keyless-fallback behavior, an agent has everything needed to invoke this 9-parameter tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics: the question/passage_size/max_passages interaction, max_length truncating only returned text while the full page is written to a temp file, and renderer acting as backend selection with native_fetch as fallback. Some clauses (crawl, extract) largely restate the schema descriptions, keeping it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Fetch a URL and extract clean content via a renderer') and immediately characterizes the renderer backends. It explicitly names the two sibling tools it is not (infobroker_search_web for topic search, infobroker_verify_claims for claim verification), so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the triggering condition ('Use when you have a URL and need readable text, passages ranked against a question, or the page's last-updated date') and gives explicit exclusions with named alternatives ('Do NOT use for topic search... or cross-source claim verification'). Both the when and the when-not are present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.