Skip to main content
Glama
MarekCziba

Website to Markdown MCP

by MarekCziba

🌐 Website to Markdown MCP

PyPI Python CI License: MIT

Turn any web page into clean, structured Markdown β€” purpose-built for AI agents, LLM context and RAG pipelines.

No API key. No account. No pay-per-event. Runs locally over stdio.


✨ Why

LLMs read text, not DOM trees. This server fetches a page, strips the navigation, ads and boilerplate, and hands your agent semantic Markdown it can reason over immediately β€” a browser that speaks the model's language.

  • One command, no config β€” uvx website-to-markdown-mcp and it is running

  • Real extraction β€” trafilatura for article text, BeautifulSoup fallback for everything else

  • Batch + concurrency β€” up to 5 pages in flight at once, one bad URL never kills the batch

  • Honest output β€” guarantees an H1 heading, never leaks raw HTML tags

  • Metadata too β€” title, description, author, Open Graph tags, canonical URL

Related MCP server: fetch-guard

🧰 Tools

Tool

Description

Arguments

fetch_url

Fetch one page and return clean Markdown

url, max_length = 50000

fetch_urls

Fetch several pages concurrently; returns {url, markdown, error} per item

urls, max_length = 50000

extract_metadata

Extract title, description, author, OG tags, canonical

url

πŸš€ Install

# one-off, no installation (recommended)
uvx website-to-markdown-mcp

# or with pip
pip install website-to-markdown-mcp

# or from source
git clone https://github.com/MarekCziba/website-to-markdown-mcp
cd website-to-markdown-mcp
uv sync

πŸ”Œ Connect your MCP client

Claude Desktop / Claude Code β€” claude mcp add:

claude mcp add website-to-markdown -- uvx website-to-markdown-mcp

Cursor, Windsurf, or any JSON-config client:

{
  "mcpServers": {
    "website-to-markdown": {
      "command": "uvx",
      "args": ["website-to-markdown-mcp"]
    }
  }
}

opencode (opencode.json):

{
  "mcp": {
    "website-to-markdown": {
      "type": "local",
      "command": ["uvx", "website-to-markdown-mcp"],
      "enabled": true
    }
  }
}

πŸ’‘ Example session

You:  Summarise https://example.com
Agent: [calls fetch_url]
       # Example Domain

       This domain is for use in documentation examples without needing
       permission. ...

You:  Pull the title and description from three pages at once
Agent: [calls fetch_urls with 3 URLs]

πŸ—οΈ How it works

  1. Your MCP client starts the server over stdio

  2. httpx fetches the page (redirects followed, 30 s timeout, honest User-Agent)

  3. trafilatura extracts the article body as Markdown

  4. If extraction is too thin, boilerplate tags are stripped and markdownify takes over

  5. The document title is re-attached as an # H1 when the extractor dropped it

πŸ§ͺ Development

uv sync                 # install runtime + dev dependencies
uv run pytest           # runs against LIVE urls (example.com, docs.python.org, wikipedia.org)
uv run ruff check src tests
uv run ruff format src tests

The test suite asserts on actual output β€” expected phrases, no leaked HTML tags, max_length enforcement, batch error isolation and URL validation β€” not on β€œit looks fine”.

πŸ“„ License

MIT β€” see LICENSE.


Built by Marek Cziba Β· GitHub Β· LinkedIn Β· marekcziba@gmail.com

Available Tools

3 tools
extract_metadataB

Extract page metadata (title, description, author, Open Graph tags).

Args: url: Absolute http(s) URL of the page to inspect.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say whether the tool fetches the URL over the network, how redirects or unreachable/timeout URLs are handled, whether auth or robots restrictions apply, or whether a head-only request is made. 'Inspect' is the only hint about mechanism.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence naming the tool's output, followed by a compact arg note with no filler. The 'Args:' block is slightly redundant with a single obvious parameter but is not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output is covered by an output schema, so return values need no explanation, and a single required parameter keeps the surface small. The remaining gap is the absent behavioral detail around network access and failure handling, which matters for a URL-fetching tool but is minor given overall simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only types url as a bare string, so the description's 'Absolute http(s) URL of the page to inspect' meaningfully constrains the expected format. For a one-parameter tool this is adequate compensation, though it adds no detail on encoding or query-string handling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Extract') and resource ('page metadata') and enumerates the extracted fields (title, description, author, Open Graph tags). It is distinguishable from fetch_url/fetch_urls, which retrieve content rather than parsed metadata, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by contrast with the sibling fetch tools: an agent can infer this is for metadata rather than full page content. There is no explicit when-to-use, when-not-to-use, or alternative routing statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_urlB

Fetch a single web page and convert it to clean Markdown.

Args: url: Absolute http(s) URL of the page to convert. max_length: Maximum number of characters of Markdown to return.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the core behavior (fetch + convert to Markdown) but omits important traits: whether it follows redirects, handles non-HTML content, requires network access, or what happens on failure/robots.txt. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence purpose followed by concise parameter notes. No filler, though the Args block is somewhat redundant with the schema structure itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. However, with no annotations and a network-facing fetch tool, the absence of any behavioral notes (redirects, errors, encoding) leaves meaningful gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It defines both parameters: url as an absolute http(s) URL, and max_length as the character cap on returned Markdown. This meaningfully fills the gap left by the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (a single web page) with the transformation (to clean Markdown). The word 'single' implicitly distinguishes it from the sibling fetch_urls, though the differentiation is not explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus fetch_urls or extract_metadata. The 'single' qualifier hints at batch alternatives but does not name the condition for choosing this tool over them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_urlsA

Fetch several web pages concurrently and convert each to Markdown.

Args: urls: List of absolute http(s) URLs. max_length: Maximum number of characters of Markdown per page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
max_lengthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose useful traits: concurrency, per-page Markdown conversion, and per-page length capping. However, it says nothing about failure behavior (does one bad URL abort the batch?), timeouts, rate limits, or authentication, which are material for a network-fetch tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence carries the purpose and behavior, followed by a compact Args block. No filler, no repetition of the tool name, and the concurrency fact appears before the parameter details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and both parameters are covered. The remaining gap is batch edge-case behavior: no statement about result ordering, partial failures, or per-URL error reporting, which an agent needs before committing to a multi-URL call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: 'urls' is constrained to absolute http(s) URLs (a real validation rule absent from the schema), and 'max_length' is defined semantically as character count of Markdown per page, clarifying it is not bytes or tokens.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch several web pages'), plus the transformation ('convert each to Markdown') and the concurrency behavior. The plural 'several' implicitly separates it from the singular sibling fetch_url, though the distinction is left to inference rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'several web pages' suggests this is the batch counterpart to fetch_url, and extract_metadata is never mentioned. There is no explicit when-to-use, when-not-to-use, or named alternative, so the agent must infer routing from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedextract_metadata
    • First observedfetch_url
    • First observedfetch_urls

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation4/5

fetch_url and fetch_urls share nearly identical purpose, differing only in single vs. batch input, but their descriptions make the distinction explicit. extract_metadata targets a clearly different output (page metadata) and stands apart cleanly.

Naming Consistency5/5

All three names follow a consistent snake_case verb_noun pattern (fetch_url, fetch_urls, extract_metadata). The plural form for the batch variant is a sensible, predictable convention.

Tool Count4/5

Three tools is on the lean side but well-matched to a narrow single-purpose server (web page to Markdown conversion). Each tool earns its place with no redundancy.

Completeness4/5

Covers the core surface: single fetch, concurrent batch fetch, and metadata extraction, with a sensible max_length control. Missing niceties like custom headers/auth or content-selector options, but no dead ends for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Fast, token-efficient web content extraction tool that converts websites to clean Markdown for AI agents, featuring smart caching, content extraction with Mozilla Readability, and polite crawling capabilities.
    1
    402 npm
    160
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Fetch URLs and return clean, LLM-ready markdown with metadata and layered prompt injection defense. Configurable timeouts, word limits, JS rendering, and link extraction. All-in-one MCP server + CLI.
    1
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.
    3
    6 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Converts any web page URL into clean Markdown for LLM context (Claude, ChatGPT, etc.) with zero external API calls, running entirely locally.
    1
    66 npm
    MIT