mcp-safe-fetch
by PortfolioKB
README.md
# mcp-safe-fetch

An injection-aware content fetcher, exposed over the Model Context Protocol.
Agents that read the open web read content nobody on your team wrote. A page or a post
can carry a fake closing tag followed by "ignore previous instructions and ...". This
server fetches a URL, strips it to text, runs defense-in-depth sanitization, caps the
size, and caches the result. It is small on purpose. It is the reference implementation
of a set of production hardening field notes, not a framework.
## Why it exists
Most MCP servers are written for a demo: one caller, one happy path, no untrusted input,
no cost ceiling. In production the failures cluster in three places, and none of them
show up on day one:
- **Untrusted content.** A single tag stripper feels safe and is not. Unicode variants and
unclosed tags walk straight through one regex.
- **Cost.** An unbounded response body or an oversized context is money spent on noise.
- **Concurrency.** Two tools touching one SQLite file throw `database is locked`.
This server answers all three in code you can read in five minutes.
## What it does
`fetch_clean(url, max_chars=2500)` returns sanitized, size-capped text plus an audit of
what was done. The defense is order, not cleverness:
1. strip the injection wrapper by its literal name, first
2. normalize unicode (NFKC) so homoglyph tags cannot hide
3. drop script and style bodies, then the generic tag strip
4. a second net for common instruction-override phrases
It caps input size with a `[truncated]` marker (if the model has to ignore most of the
input, you are paying for nothing), clamps the per-call cap at the entry to a hard ceiling,
and caches results in SQLite opened with `WAL` and `busy_timeout` so overlapping callers
wait instead of crashing.
## Install
```bash
pip install -e .
```
## Run
As a standalone MCP server (stdio):
```bash
mcp-safe-fetch
```
Register it with an MCP client (for example, Claude Code) by pointing the client at the
`mcp-safe-fetch` command. The server exposes one tool, `fetch_clean`.
## Test
```bash
pip install -e ".[dev]"
pytest
```
## The field notes behind it
The reasoning, with the production incidents that motivated each defense, is written up
here: a short essay on MCP hardening (concurrency, prompt injection, cost) and what breaks
after day 30. The sanitizer in `src/mcp_safe_fetch/sanitize.py` is the exact function from
that write-up.
## License
MIT.
TDQS
A3.6/5.0
Scored across 1 tool
Disambiguation5/5
Only one tool exists, so there is no ambiguity at all. An agent cannot misselect between tools.
Naming Consistency5/5
With a single tool, naming is trivially consistent. The name 'fetch_clean' follows a clear verb_noun pattern.
Tool Count3/5
One tool feels thin for a server dedicated to fetching content. While the tool is well-defined, the server would benefit from additional tools (e.g., batch fetch, status check) to justify its scope.
Completeness4/5
The tool covers the core fetch-and-sanitize operation completely for a single URL. Minor gaps exist (e.g., no way to fetch multiple URLs or configure sanitization levels), but the basic use case is fully served.