Skip to main content
Glama
vanshulgoyal101

mcp.vanshul.com

mcp.vanshul.com

MCP Registry Endpoint

A public Model Context Protocol (MCP) server, running as a Cloudflare Worker, that lets any AI agent read the live web as clean Markdown. It builds on the same extraction pipeline as the sibling reader/ project (Mozilla Readability + Turndown), exposed over the MCP Streamable HTTP transport so agentic clients can plug straight in.

Endpoint

POST https://mcp.vanshul.com/mcp     # JSON-RPC 2.0 (MCP)
GET  https://mcp.vanshul.com/health  # { ok: true, tools: [...] }

Related MCP server: pagewatch-mcp

Tools

Tool

Input

Returns

fetch_markdown

{ url, max_chars? }

The page's main content as clean Markdown (optionally truncated to max_chars)

search_page

{ url, query, max_matches?, context_chars? }

Only the passages matching query, each with its heading breadcrumb, ranked by relevance — token-efficient alternative to fetch_markdown

fetch_metadata

{ url }

JSON: title, byline, siteName, excerpt, wordCount

extract_links

{ url, limit? }

JSON: all outbound http(s) links + anchor text

Connect from an MCP client

Remote/HTTP-capable clients (Claude Desktop, Cursor, Continue, …):

{
  "mcpServers": {
    "web-reader": { "url": "https://mcp.vanshul.com/mcp" }
  }
}

Stdio-only clients can bridge with mcp-remote:

npx mcp-remote https://mcp.vanshul.com/mcp

Add it to your client

  • Cursor — Settings → MCP → Add new server, or drop this into ~/.cursor/mcp.json:

    { "mcpServers": { "web-reader": { "url": "https://mcp.vanshul.com/mcp" } } }
  • Claude Desktop — add the same block to claude_desktop_config.json (Settings → Developer → Edit Config). If your version is stdio-only, use:

    { "mcpServers": { "web-reader": { "command": "npx", "args": ["mcp-remote", "https://mcp.vanshul.com/mcp"] } } }
  • Continue / VS Code — add web-reader with URL https://mcp.vanshul.com/mcp to your MCP servers config.

Try it with curl

curl -s https://mcp.vanshul.com/mcp \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

curl -s https://mcp.vanshul.com/mcp \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":2,"method":"tools/call",
       "params":{"name":"fetch_markdown","arguments":{"url":"https://example.com"}}}'

# Return only the passages matching a query (token-efficient):
curl -s https://mcp.vanshul.com/mcp \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":3,"method":"tools/call",
       "params":{"name":"search_page","arguments":{"url":"https://example.com","query":"more information"}}}'

Layout

mcp/
├── src/
│   ├── worker.ts     # entry: routes /mcp, /health, rate limit, CORS
│   ├── mcp.ts        # JSON-RPC dispatch + tool definitions
│   ├── extract.ts    # HTML -> Markdown / links (Readability + Turndown)
│   ├── search.ts     # query-focused passage search over extracted Markdown
│   ├── fetcher.ts    # fetch with timeout, size cap, re-validated redirects
│   └── security.ts   # SSRF guard (block private/internal addresses)
├── public/
│   ├── index.html    # landing page (served for non-API paths)
│   ├── og.png        # social share image (1200×630)
│   ├── og.svg        # social image source
│   ├── robots.txt
│   └── sitemap.xml
├── tests/            # vitest: security, extract, search, mcp dispatch, fetcher, worker
├── wrangler.toml
├── package.json
├── tsconfig.json
└── README.md

Develop & deploy

cd mcp
npm install
npm run typecheck
npm test          # vitest — security, extraction, search, MCP dispatch, fetcher, worker
npm run dev        # local worker at http://localhost:8787  (POST /mcp)
npm run deploy     # wrangler deploy

After the first deploy, attach the custom domain mcp.vanshul.com in the Cloudflare dashboard (Workers & Pages → this worker → Settings → Domains), or uncomment the [[routes]] block in wrangler.toml.

Security

  • SSRF-safe: only public http(s) URLs; localhost, private ranges, cloud metadata and every redirect hop are blocked/re-validated.

  • Bounded: 10s fetch timeout, ~3 MB page cap, max 5 redirects, per-IP rate limit.

  • Stateless & private: no page content is stored; extraction is deterministic with no LLM in the loop.

MCP protocol

Implements initialize, ping, tools/list, tools/call and notifications (protocol version 2025-06-18). Tool failures are returned as { content, isError: true } so the agent can read the message and recover; malformed requests use standard JSON-RPC error codes.

Documentation

License

MIT © Vanshul Goyal

Related MCP Connectors

  • MCP server (stdio): fetch web pages as clean readable markdown via the AgentForge API

  • Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.

  • Hosted MCP server for live public-data APIs and Skills for AI agents.

  • Your agent needs the open web — searched by more than one engine, and read as clean markdown rather than raw HTML. **What you can ask for** • "Search this question with two providers and tell me where they disagree." • "Scrape these 40 URLs into markdown, in one batch." • "Crawl this documentation site and give me every page." • "Do deep research on this topic and cite the sources." • "Find the academic papers behind this claim." **How to use it** Point any MCP client at https://mcp.aisa.one/search/mcp and sign in with OAuth — there is no key to create or paste. 30 tools across several independent providers: Tavily and Exa search, answers, contents and agent runs; Firecrawl scrape, batch scrape, crawl, map and search; Perplexity Sonar, Sonar Pro, reasoning and deep research; Oxylabs AI search and LLM jobs; OpenAI and Anthropic web search; and scholarly search. **Why this rather than the source** Several independent indexes behind one account, because one engine's blind spot is not visible from inside it. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Find the page here, then ask the same agent who links to it or how much traffic it gets — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo-serp/mcp for the Google results page itself, https://mcp.aisa.one/seo-serp-other-engines/mcp for Bing, Baidu and Naver.

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that fetches web pages and extracts clean, AI-usable context from them, enabling tools for link discovery, content search, and integrated fetch-and-search operations.
    5
    8 npm
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    MCP server that provides read_page, screenshot, and pdf tools using a real browser, enabling agents to fetch clean markdown, screenshots, and PDFs from any URL.
    5
    16 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A custom MCP server that provides AI agents with tools to fetch web pages, extract readable article text, extract structured data by CSS selector, and check robots.txt permissions.
    MIT