Website to Markdown MCP
# π Website to Markdown MCP
[](https://pypi.org/project/website-to-markdown-mcp/)
[](https://www.python.org/)
[](https://github.com/MarekCziba/website-to-markdown-mcp/actions/workflows/ci.yml)
[](LICENSE)
> Turn any web page into clean, structured Markdown β purpose-built for AI agents, LLM context and RAG pipelines.
No API key. No account. No pay-per-event. Runs locally over stdio.
---
## β¨ Why
LLMs read text, not DOM trees. This server fetches a page, strips the navigation,
ads and boilerplate, and hands your agent semantic Markdown it can reason over
immediately β a browser that speaks the model's language.
- **One command, no config** β `uvx website-to-markdown-mcp` and it is running
- **Real extraction** β [trafilatura](https://trafilatura.readthedocs.io/) for article text, BeautifulSoup fallback for everything else
- **Batch + concurrency** β up to 5 pages in flight at once, one bad URL never kills the batch
- **Honest output** β guarantees an H1 heading, never leaks raw HTML tags
- **Metadata too** β title, description, author, Open Graph tags, canonical URL
## π§° Tools
| Tool | Description | Arguments |
|------|-------------|-----------|
| `fetch_url` | Fetch one page and return clean Markdown | `url`, `max_length = 50000` |
| `fetch_urls` | Fetch several pages concurrently; returns `{url, markdown, error}` per item | `urls`, `max_length = 50000` |
| `extract_metadata` | Extract title, description, author, OG tags, canonical | `url` |
## π Install
```bash
# one-off, no installation (recommended)
uvx website-to-markdown-mcp
# or with pip
pip install website-to-markdown-mcp
# or from source
git clone https://github.com/MarekCziba/website-to-markdown-mcp
cd website-to-markdown-mcp
uv sync
```
## π Connect your MCP client
**Claude Desktop / Claude Code** β `claude mcp add`:
```bash
claude mcp add website-to-markdown -- uvx website-to-markdown-mcp
```
**Cursor, Windsurf, or any JSON-config client:**
```json
{
"mcpServers": {
"website-to-markdown": {
"command": "uvx",
"args": ["website-to-markdown-mcp"]
}
}
}
```
**opencode** (`opencode.json`):
```json
{
"mcp": {
"website-to-markdown": {
"type": "local",
"command": ["uvx", "website-to-markdown-mcp"],
"enabled": true
}
}
}
```
## π‘ Example session
```
You: Summarise https://example.com
Agent: [calls fetch_url]
# Example Domain
This domain is for use in documentation examples without needing
permission. ...
You: Pull the title and description from three pages at once
Agent: [calls fetch_urls with 3 URLs]
```
## ποΈ How it works
1. Your MCP client starts the server over stdio
2. `httpx` fetches the page (redirects followed, 30 s timeout, honest User-Agent)
3. `trafilatura` extracts the article body as Markdown
4. If extraction is too thin, boilerplate tags are stripped and `markdownify` takes over
5. The document title is re-attached as an `# H1` when the extractor dropped it
## π§ͺ Development
```bash
uv sync # install runtime + dev dependencies
uv run pytest # runs against LIVE urls (example.com, docs.python.org, wikipedia.org)
uv run ruff check src tests
uv run ruff format src tests
```
The test suite asserts on **actual output** β expected phrases, no leaked HTML tags,
`max_length` enforcement, batch error isolation and URL validation β not on βit looks fineβ.
## π License
MIT β see [LICENSE](LICENSE).
---
Built by **Marek Cziba** Β· [GitHub](https://github.com/MarekCziba) Β· [LinkedIn](https://www.linkedin.com/in/marekcziba) Β· [marekcziba@gmail.com](mailto:marekcziba@gmail.com)
TDQS
Scored across 3 tools
fetch_url and fetch_urls share nearly identical purpose, differing only in single vs. batch input, but their descriptions make the distinction explicit. extract_metadata targets a clearly different output (page metadata) and stands apart cleanly.
All three names follow a consistent snake_case verb_noun pattern (fetch_url, fetch_urls, extract_metadata). The plural form for the batch variant is a sensible, predictable convention.
Three tools is on the lean side but well-matched to a narrow single-purpose server (web page to Markdown conversion). Each tool earns its place with no redundancy.
Covers the core surface: single fetch, concurrent batch fetch, and metadata extraction, with a sensible max_length control. Missing niceties like custom headers/auth or content-selector options, but no dead ends for the stated purpose.