Skip to main content
Glama

@page2ai/mcp — Web to Markdown for LLM Context

npm version CI MCP Registry license MIT

No downloads badge on purpose: npm download counts are dominated by registry mirrors and bots, and this README only shows numbers a human can trust.

One-click webpage-to-Markdown MCP server for ChatGPT, Claude, and Gemini. Preserves code blocks with language hints, tables, and reading structure. Fetch, extract, convert — one tool, one command, entirely on your machine. No API key, no account, no telemetry; the only network requests are to the URLs you pass it.

Companion to the Page2AI Chrome & Firefox extensions; both share the same @page2ai/core extraction library.

Quick start

npx -y page2ai-mcp

This starts the stdio MCP server — it waits silently for a client, so wire it into one of the clients below rather than running it bare. Requires Node.js 20+. No API key. If npx ever serves you a stale cached version, run npx -y page2ai-mcp@latest once.

page2ai-mcp is a thin unscoped wrapper around @page2ai/mcp; since 0.2.5 it follows the latest server automatically via a ^ range.

Hosted endpoint (beta) — no install at all

For clients that only accept a remote MCP server URL (Streamable HTTP):

https://page2ai-mcp-remote.vercel.app/api/mcp

Same tool, same code (see remote/), rate-limited per IP. Privacy trade-off: URLs sent to the hosted endpoint are fetched from that server, so they transit infrastructure operated by the author (on Vercel). Running locally via npx keeps everything on your machine — prefer it when you can.

Claude Code

claude mcp add page2ai -- npx -y page2ai-mcp

Claude Desktop

One-click bundle: download page2ai-mcp-<version>.mcpb from the latest release, then in Claude Desktop open Settings → Extensions → Advanced settings → Install Extension… and pick the file.

Or add to claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\), then fully quit and restart Claude Desktop:

{
  "mcpServers": {
    "page2ai": { "command": "npx", "args": ["-y", "page2ai-mcp"] }
  }
}

Cursor

Install MCP Server

Or add to ~/.cursor/mcp.json (global) / .cursor/mcp.json (project):

{
  "mcpServers": {
    "page2ai": { "command": "npx", "args": ["-y", "page2ai-mcp"] }
  }
}

VS Code (GitHub Copilot)

code --add-mcp '{"name":"page2ai","command":"npx","args":["-y","page2ai-mcp"]}'

Or add to .vscode/mcp.json:

{
  "servers": {
    "page2ai": { "type": "stdio", "command": "npx", "args": ["-y", "page2ai-mcp"] }
  }
}

Codex CLI

codex mcp add page2ai -- npx -y page2ai-mcp

Gemini CLI

gemini mcp add page2ai npx -y page2ai-mcp

Antigravity

agy mcp add page2ai npx -y page2ai-mcp

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json (macOS/Linux) or %USERPROFILE%\.codeium\windsurf\mcp_config.json (Windows):

{
  "mcpServers": {
    "page2ai": { "command": "npx", "args": ["-y", "page2ai-mcp"] }
  }
}

Zed

Add to settings.json (zed: open settings file):

{
  "context_servers": {
    "page2ai": { "command": "npx", "args": ["-y", "page2ai-mcp"] }
  }
}

Other MCP clients

Any client that speaks MCP over stdio works — Cline, Continue, JetBrains AI Assistant, LM Studio and others. Point their MCP server settings at the same command: npx -y page2ai-mcp.

Letting an AI agent install it

If an AI agent (Cline, Claude Code, Codex, …) is doing the setup for you, point it at llms-install.md — a machine-oriented install guide with the exact command per client and a verification step.

Related MCP server: Webustler

Tools

Tool

Description

Read-only

Example

page_to_markdown

Fetch a web page URL and convert to clean Markdown

page_to_markdown(url="https://docs.anthropic.com/en/api/messages")

Example prompts

1. Fetch documentation and answer questions:

"Use page_to_markdown to fetch https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking and summarize the three main use cases for extended thinking."

2. Extract API reference into a code snippet:

"Fetch https://ai.google.dev/gemini-api/docs/thinking with page_to_markdown, then generate a Python code sample using the thinking budget parameter."

3. Compare two documentation pages:

"Fetch both https://docs.anthropic.com/en/docs/prompt-engineering and https://platform.openai.com/docs/guides/prompt-engineering with page_to_markdown, then summarize the differences in approach."

What a call returns

Verbatim output of page_to_markdown(url="https://docs.anthropic.com/en/api/messages"), captured 2026-08-29 on core 0.1.7, truncated:

---
title: "Messages"
source: "https://platform.claude.com/docs/en/api/messages"
captured_at: "2026-08-29T09:01:05.806Z"
extractor: "page2ai-core"
extractor_version: "0.1.7"
extractor_source: "content-negotiation"
---
# Messages

## Create a Message

**POST** `/v1/messages`

Send a structured list of input messages with text and/or image content, and the
model will generate the next message in the conversation.

YAML front matter (title, source, capture timestamp, extractor provenance), then the article as clean Markdown — headings, tables and fenced code preserved, navigation and chrome dropped.

Configuration

No setup required — every argument has a sensible default. page_to_markdown accepts:

Argument

Type

Default

Effect

url

string, required

Absolute http(s) URL to convert. Private ranges, loopback and cloud metadata endpoints are blocked (SSRF guard)

include_images

boolean

true

Keep ![alt](src) image references in the output

include_frontmatter

boolean

true

Prepend the YAML front matter block

timeout_ms

integer

15000

Abort the fetch after N ms (1000–60000)

Planned: a profile argument (auto | docs | marketing | research | dashboard | wordpress-marketing).

Privacy

@page2ai/mcp collects no data, sends no telemetry, and makes no external network calls beyond the URLs you explicitly provide. Note that your AI client (Claude, ChatGPT, …) still sends the extracted Markdown to its own model provider under that client's privacy policy — that part is outside this server's control. See PRIVACY.md for details.

The optional hosted endpoint is different by nature: URLs you ask it to convert are fetched from a server operated by the author (deployed on Vercel), and Vercel keeps its standard operational logs. No accounts, no tracking. If that trade-off matters for your use case, run the server locally — it is the same code.

Protocol revisions

Since 0.2.0 this server answers both MCP protocol revisions on the same stdio connection:

Revision

How a client opens the connection

Served

2026-07-28

server/discover, carrying io.modelcontextprotocol/protocolVersion in _meta

yes

2025-11-25 and earlier

the initialize handshake

yes

serveStdio picks the era from the opening message and pins one server instance for the life of the connection, so clients on older SDKs are unaffected. Note that a message carrying no version claim is treated by the specification as a 2025-era opening; server/discover without that _meta field is therefore answered with -32601 rather than being upgraded.

Verify against a build with the raw JSON-RPC frames, without a client library:

printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"server/discover","params":{"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28"}}}' | node dist/index.js

Known advisories

None. Moving to the v2 SDK removed 90 transitive packages (117 production dependencies down to 26) and with them the @hono/node-server advisory that 0.1.x carried through the monolithic @modelcontextprotocol/sdk; that package pulled the HTTP and OAuth server stack even for a stdio-only server. npm audit is clean as of 0.2.0.

Development

git clone https://github.com/igorsaevets/page2ai-mcp
cd page2ai-mcp
npm install
npm run build
node dist/index.js  # runs stdio MCP server

Test with MCP Inspector:

npx @modelcontextprotocol/inspector node dist/index.js

Support

Pick the channel by what you have, not by what is quickest to type:

You have

Use

A bug, a page that extracts badly, a feature idea

GitHub issues — public and searchable, so the fix helps the next person too

A security vulnerability

Private vulnerability reporting, see SECURITY.md. Do not open a public issue

A usage question

Docs first, then an issue

About the author

Written and maintained by Igor Saevets — AI expert and founder of Page2AI. Full bio: igorsaevets.github.io/page2ai-docs/about/.

These are identity and collaboration links, not the support queue. A bug reported in a DM is a bug nobody else can find later, so anything you want fixed belongs in the table above.

Build provenance

An MCP server runs with the privileges of whatever launched it, so where the tarball came from is a security question, not a formality. Releases from v0.1.2 onward are published from GitHub Actions with npm provenance: each version carries a Sigstore attestation naming the commit and workflow run that produced it, recorded in the public Rekor transparency ledger and shown as a badge on npmjs.com.

Check it before you trust it:

npm audit signatures
npm view @page2ai/mcp --json | jq '.dist.attestations'

License

MIT — see LICENSE. Copyright © 2026 Igor Saevets.

Available Tools

1 tool
page_to_markdownConvert Page to MarkdownA
Read-only

Fetch a web page URL and convert it to clean Markdown optimized for LLM context. Preserves headings, code blocks (with language hints), links, and tables; strips ads, navigation, and cookie banners. For Mintlify docs (docs.anthropic.com, OpenAI platform, Vercel, Stripe, etc.) tries the URL.md convention first for cleanest output. Discovers tab groups statically and emits each panel as a ### Tab: {label} section instead of concatenating (prevents Python+TypeScript examples merging into one broken block). Zero external API calls — parses locally via linkedom.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute HTTP or HTTPS URL of the web page to convert. Blocks private ranges, loopback, and cloud metadata endpoints.
timeout_msNoAbort the fetch after N milliseconds. Range: 1000-60000.
include_imagesNoInclude image references (![alt](src)) in the output.
include_frontmatterNoPrepend a YAML frontmatter block with title, source URL, capture timestamp, language, description, and canonical URL.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description fully discloses the tool's behavior beyond annotations: it performs a fetch and local conversion, is read-only (consistent with readOnlyHint=true), handles Mintlify docs with a URL.md convention, and processes tab groups statically. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet comprehensive, front-loading the core purpose in the first sentence. Every sentence adds distinct value: preservation/stripping details, Mintlify handling, tab group treatment, and technical approach. No redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, no output schema, no sibling tools), the description adequately covers input behavior, special cases, and output format (clean Markdown with preserved elements). It could mention the default parameter values from the schema but is otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage with descriptions for all 4 parameters, meeting the baseline. The description adds value by explaining overall behavior but does not provide additional parameter-specific semantics beyond the schema. Thus, a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Fetch a web page URL and convert it to clean Markdown optimized for LLM context.' It specifies what is preserved (headings, code blocks, links, tables) and what is stripped (ads, navigation, cookie banners), making the action and result unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides good context for when to use this tool (converting web pages for LLM context, especially Mintlify docs) and explains its technical approach (local parsing, no external API calls). However, it does not explicitly state when not to use it or mention alternatives, but the absence of sibling tools reduces the need for such differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The tool's purpose is clearly distinct from any other.

Naming Consistency5/5

A single tool named 'page_to_markdown' follows a clear verb_noun pattern. Consistency is perfect by definition.

Tool Count2/5

A single tool is very minimal for an MCP server. While the tool is focused and well-described, it provides only one operation, which may limit the server's utility.

Completeness5/5

The server fully covers its stated purpose of converting a web page to markdown. It handles the single operation thoroughly, with details on preserving structure and handling special cases.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Converts any webpage into clean, LLM-ready Markdown, removing noise and supporting JavaScript rendering.
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/igorsaevets/page2ai-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server