Local Jina Reader MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Local Jina Reader MCPread https://example.com as markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Local Jina Reader MCP
A small stdio MCP adapter for a self-hosted Jina Reader instance.
The Reader service does the browser work and Markdown extraction. This project only exposes one safe, local MCP tool so an Agent can call it when web content is needed.
Architecture
Agent --stdio MCP--> jina-reader-mcp --HTTP--> Jina Reader OSSThe adapter defaults to http://127.0.0.1:3000, uses one in-flight fetch, limits each URL to 8,000 Reader tokens, and rejects obvious local or private targets. It exposes read_url for one page and read_urls for a bounded batch of up to two pages.
Related MCP server: @hauntapi/mcp-server
Run the Reader backend
Docker is required for the upstream Reader image:
docker run --rm -p 127.0.0.1:3000:8081 ghcr.io/jina-ai/reader:ossSmoke-test it:
curl.exe -X POST http://127.0.0.1:3000/ `
-H "Accept: text/markdown" `
-H "Content-Type: application/json" `
-d '{"url":"https://example.com"}'Run the MCP adapter
uv run --with . jina-reader-mcpOr install it into a virtual environment:
uv sync
uv run jina-reader-mcpThe MCP server uses stdio, so configure your Agent to launch jina-reader-mcp from this checkout. The tools are named read_url and read_urls.
Example stdio configuration:
{
"mcpServers": {
"local-jina-reader": {
"command": "uv",
"args": ["run", "--directory", "D:/path/to/jina-reader-mcp-local", "jina-reader-mcp"],
"env": {
"READER_BASE_URL": "http://127.0.0.1:3000"
}
}
}
}Configuration
Variable | Default | Purpose |
|
| Local Reader HTTP base URL |
| empty | Optional Bearer token for a protected Reader endpoint |
|
| Default request timeout, 1-180 |
|
| Default Reader output cap, 500-50000 |
|
| Maximum concurrent fetches, 1-4 |
read_urls returns a JSON object with one result per input URL. A failed URL is reported as an error entry without discarding successful results from the same batch. The batch is capped at two URLs and still passes through the adapter's global concurrency limit.
Search is intentionally not exposed yet: the self-hosted Reader search process needs a populated local index or an external Google/Bing SERP provider, neither of which is available in the default stateless setup. The adapter also does not pretend that X-Max-Tokens is resumable pagination; add a continuation protocol only if a Reader response contract provides a real cursor.
Development
uv sync
uv run python -m unittest discover -s tests -vLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP server (stdio): fetch web pages as clean readable markdown via the AgentForge API
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Web pages to Markdown, metadata and small crawls for agents. Free daily calls, then USDC x402/MPP.
Read any web page as clean Markdown for AI agents: fetch, search, metadata, links. SSRF-safe.
Related MCP Servers
- AlicenseAqualityCmaintenanceFetch URLs and return clean, LLM-ready markdown with metadata and layered prompt injection defense. Configurable timeouts, word limits, JS rendering, and link extraction. All-in-one MCP server + CLI.11MIT
- AlicenseAqualityBmaintenanceWeb extraction MCP server for AI agents. Extract structured data from any URL with built-in Cloudflare bypass, JavaScript rendering, and intelligent parsing. Returns clean markdown or JSON.5794 npm2MIT
- FlicenseNot gradedqualityAmaintenanceEnables AI agents to securely scrape single or batches of public web pages via stdio MCP, with configurable modes, Markdown extraction, and built-in SSRF and resource protections.-
- AlicenseAqualityCmaintenanceGives MCP-capable agents live web access: search the web, scrape pages into Markdown (including JavaScript-heavy and bot-protected sites), and extract named fields as JSON, with job polling, token-aware content offloading, and built-in research guidance. Ships as a self-hostable stdio or HTTP service with spend caps and per-request key support.7MIT