Skip to main content
Glama
lhswgzy

free-search-mcp-ts

by lhswgzy

Fetch a URL as Markdown

fetch_url
Read-only

Fetch a web page or document and convert it to clean Markdown by removing navigation, ads, and cookie banners; supports HTML, PDF, Office, EPUB, CSV, JSON, and text.

Instructions

Fetch one web page or document and return its main content as clean Markdown, with the navigation, adverts and cookie banners removed. Handles HTML (via a readability pass), PDF, DOCX, XLSX, PPTX, EPUB, ODT, CSV, JSON and plain text. Results are stored in a local SQLite index, so reading the same URL again is instant and the page becomes searchable offline with search_index. Long pages are truncated rather than refused: the response reports a next offset you can pass back to continue. robots.txt is honoured by default and private/loopback addresses are refused as a safety measure.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute http(s) URL to fetch.
formatNoOutput format. "markdown" (default) is compact and readable; "json" returns structured data for post-processing.
offsetNoSkip this many characters into the content. Use the `next offset` from a previous call to page through a long document.
refreshNoIgnore the local cache and re-fetch (default false).
max_charsNoMaximum characters of content to return. Longer content is truncated and a `next offset` is reported so you can continue reading.
use_cacheNoUse the cached copy when it is fresh (default true).
include_linksNoAlso return the outgoing links found in the body, for follow-up fetching.
respect_robotsNoHonour robots.txt for this request (default: the server setting, which is true).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/openWorldHint annotations, it discloses substantial behavior: results cached in a local SQLite index, long pages truncated rather than refused, robots.txt honoured by default, and private/loopback addresses refused as a safety measure. These are exactly the operational traits an agent needs and are not derivable from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core behavior (fetch → clean Markdown) and then layers caching, truncation, and safety in tight successive sentences. It is dense but each sentence carries distinct information; only the long format enumeration is slightly list-like.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema, the description covers return shape (Markdown main content, `next offset` for continuation), caching semantics, and safety constraints, which is everything needed to invoke it correctly. Nothing material is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters are already documented in the schema, and the description adds only the conceptual paging loop around `offset`/`max_chars` plus the supported file types. It reinforces but does not meaningfully extend the schema's parameter documentation, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Fetch one web page or document') and names the exact output ('clean Markdown' with navigation, adverts and cookie banners removed). It is clearly distinguishable from the plural sibling fetch_urls and from parse_document by emphasizing single-URL retrieval of the main content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: reading the same URL again is instant, the page becomes searchable offline with search_index, and the `next offset` is passed back to continue long documents. However, it never explicitly contrasts itself with fetch_urls for batch fetching or with parse_document, leaving the sibling routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.