Pulpie MCP
pulpie-mcp
![]()
An MCP server that gives AI agents their own tools for turning web pages into clean Markdown. It runs the Pulpie content extraction model locally, so tables, code blocks, links, and images come through intact instead of getting summarized away.
The agent can read a page inline, save it as a .md file, crawl a whole docs section into a folder, and check what it already saved before fetching again.
Why I made it
I already had a little local UI for Pulpie that I used to pull docs and articles into Markdown, and it works really well. But I was still the one doing the pulling. Most agent fetch tools run pages through a small model that summarizes or trims them, which is not what you want when the agent needs the actual API reference.
So this gives the agent the same thing I was using. It can grab full documentation whenever it needs it and save it somewhere useful, and I don't have to go do it by hand.
Related MCP server: WebforAI Text Extractor
What it does
Tool | What it's for |
| Read a page right now. The Markdown comes back inline, cut off at |
| Save one page as a |
| Save a documentation section as a folder of |
| List saved docs with their source URL and fetch time, optionally filtered by a search term. |
A few details worth knowing:
Where files go is up to the agent. For project work it passes a folder inside the project, like
/path/to/project/DOCS/reference/uv. For general research it leavesdirectoryempty and the file goes into the global library at~/.pulpie/library/<site>/.Every saved file has frontmatter with
title,source, andfetched, so you (and the agent) can always tell where a file came from and how old it is.Crawling checks the sitemap first. It looks at
robots.txt, then<docs path>/sitemap.xml, then/sitemap.xml. If none of those exist, it follows links from the pages instead. It only crawls pages on the same host and under the starting path.Markdown and plain-text URLs pass straight through. If a URL already serves
text/markdownortext/plain(likellms.txt), you get the file as is.
How it works
There are two pieces:
The MCP server is what your agent launches. It's small and doesn't load the model.
The backend is a local HTTP server that holds the Pulpie model in memory and does the actual fetching and extraction.
You never have to start the backend yourself. The first time a tool needs it, the MCP server starts it in the background and waits for the model to load. Every agent session shares that one backend, so you only ever have one copy of the model on your GPU. After 30 minutes with no requests it shuts itself down and frees the memory, and the next request starts it again.
The model is feyninc/pulpie-orange-small, a 210M parameter encoder. It uses around 420 MB of VRAM. On my RTX 4090 laptop, a cold start (including loading an already downloaded model) took about 10 seconds. After that, a single docs page saved in under a second, and an 8-page crawl of the uv docs took about 3 seconds. Most of that is network time.
Requirements
Python 3.10 or newer
uv(recommended) orpipAn NVIDIA GPU is nice but not required. Pulpie uses CUDA if it can, then Apple MPS, then falls back to CPU.
Heads up: this pulls in PyTorch, so the install is a few GB.
I've only tested this on Linux with an NVIDIA GPU so far. It should work on macOS, and Windows is untested.
Installation
Install it as a tool so the pulpie-mcp command is on your PATH:
uv tool install git+https://github.com/pinkpixel-dev/pulpie-mcpOr from a local clone:
git clone https://github.com/pinkpixel-dev/pulpie-mcp
cd pulpie-mcp
uv tool install .The model downloads from Hugging Face the first time the backend starts.
Adding it to your agent
Claude Code
claude mcp add --scope user pulpie -- pulpie-mcpCodex
In ~/.codex/config.toml:
[mcp_servers.pulpie]
command = "pulpie-mcp"Other MCP clients
Most clients take a JSON config like this:
{
"mcpServers": {
"pulpie": {
"command": "pulpie-mcp"
}
}
}If you add environment variables (see below), put them in the client's env block for this server.
Using it
You don't need special prompts. Just ask for what you want:
"Pull the uv docs into this project's
DOCS/referencefolder.""Read the asyncio queue docs and tell me how
join()works.""Do we already have the FastAPI docs saved somewhere?"
The server tells the agent to save project docs inside the project and general research in the global library, so it usually picks the right spot on its own.
Configuration
Everything is optional.
Variable | Default | What it does |
|
| Where the backend runs. Only local URLs get auto-started. |
|
| Holds the global library and |
|
| Where saved docs go when no directory is given. |
| off | Set to |
|
| Seconds without requests before the backend exits. |
|
| How long to wait for the backend to start. The first run includes the model download. |
Private addresses
By default the server won't fetch localhost, 127.0.0.1, LAN addresses like 192.168.x.x, or cloud metadata addresses. Agents follow links they find on the open web, so I'd rather block those than have a random page send the agent poking around your network. This is checked on every redirect too.
If you want to pull docs from a local dev server, set PULPIE_ALLOW_PRIVATE=1.
Commands
pulpie-mcp # run the MCP server over stdio (your agent does this)
pulpie-mcp serve # run the backend in the foreground, handy for debugging
pulpie-mcp status # show whether the backend is running, and on which device
pulpie-mcp stop # stop the backendIf something goes wrong with the backend, check ~/.pulpie/backend.log.
Limitations
Pages that render entirely in JavaScript come back mostly empty, because the backend fetches the HTML the server sends and doesn't run a browser. Most docs sites (Docusaurus, MkDocs, VitePress, Sphinx, Next.js) render on the server, so this hasn't been an issue for docs so far.
Code blocks are indented, not fenced. Pulpie converts HTML to Markdown with
html2text, which writes code as 4-space indented blocks without a language tag. It's valid Markdown, just not as pretty.No PDFs. Non-HTML pages other than Markdown and plain text are rejected.
Saved pages are untrusted web content. Anything an agent reads from them is data, not instructions.
Model license
This project's code is Apache 2.0, but the Pulpie model it downloads,feyninc/pulpie-orange-small, is licensed CC BY-NC 4.0. That means non-commercial use only. The pulpie Python library itself is Apache 2.0. If you want to use this in a commercial product, check the model license with Feyn first.
pulpie-mcp doesn't bundle or redistribute the model weights. They're downloaded from Hugging Face on your machine.
Development
uv sync
uv run pytestThe integration test runs the real backend and model against a live page, so it's skipped unless you opt in:
PULPIE_INTEGRATION=1 uv run pytest tests/test_integration.pyLicense
Apache 2.0. See LICENSE.
Pulpie is made by Feyn.
Made with 💖 by Pink Pixel
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP server (stdio): fetch web pages as clean readable markdown via the AgentForge API
Turn any public website into an MCP server for agents to search, read and navigate.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceAn MCP server for web content extraction that converts HTML pages into clean, LLM-optimized Markdown using Mozilla's Readability. It supports batch processing, intelligent multi-page crawling, and configurable caching while respecting robots.txt standards.69 npm-
- AlicenseNot gradedqualityNot gradedmaintenanceAn MCP server that extracts clean, structured Markdown content from web page URLs using the WebforAI library. It simplifies feeding web content into AI models by removing HTML noise and intelligently processing tables and links.-
- AlicenseAqualityDmaintenanceMCP server that allows AI agents to fetch and process llms.txt documentation from various sources. Fetch documentation from any HTTPS URL and automatically convert HTML content to readable markdown.211 npm2MIT
- FlicenseNot gradedqualityDmaintenanceMCP server that enables AI agents to search the web and extract clean Markdown content, with support for JavaScript rendering, structured data extraction, and screenshots.1-