Skip to main content
Glama
PortfolioKB

mcp-safe-fetch

by PortfolioKB

mcp-safe-fetch

tests

An injection-aware content fetcher, exposed over the Model Context Protocol.

Agents that read the open web read content nobody on your team wrote. A page or a post can carry a fake closing tag followed by "ignore previous instructions and ...". This server fetches a URL, strips it to text, runs defense-in-depth sanitization, caps the size, and caches the result. It is small on purpose. It is the reference implementation of a set of production hardening field notes, not a framework.

Why it exists

Most MCP servers are written for a demo: one caller, one happy path, no untrusted input, no cost ceiling. In production the failures cluster in three places, and none of them show up on day one:

  • Untrusted content. A single tag stripper feels safe and is not. Unicode variants and unclosed tags walk straight through one regex.

  • Cost. An unbounded response body or an oversized context is money spent on noise.

  • Concurrency. Two tools touching one SQLite file throw database is locked.

This server answers all three in code you can read in five minutes.

Related MCP server: MCP URL Fetcher

What it does

fetch_clean(url, max_chars=2500) returns sanitized, size-capped text plus an audit of what was done. The defense is order, not cleverness:

  1. strip the injection wrapper by its literal name, first

  2. normalize unicode (NFKC) so homoglyph tags cannot hide

  3. drop script and style bodies, then the generic tag strip

  4. a second net for common instruction-override phrases

It caps input size with a [truncated] marker (if the model has to ignore most of the input, you are paying for nothing), clamps the per-call cap at the entry to a hard ceiling, and caches results in SQLite opened with WAL and busy_timeout so overlapping callers wait instead of crashing.

Install

pip install -e .

Run

As a standalone MCP server (stdio):

mcp-safe-fetch

Register it with an MCP client (for example, Claude Code) by pointing the client at the mcp-safe-fetch command. The server exposes one tool, fetch_clean.

Test

pip install -e ".[dev]"
pytest

The field notes behind it

The reasoning, with the production incidents that motivated each defense, is written up here: a short essay on MCP hardening (concurrency, prompt injection, cost) and what breaks after day 30. The sanitizer in src/mcp_safe_fetch/sanitize.py is the exact function from that write-up.

License

MIT.

Available Tools

1 tool
fetch_cleanB

Fetch a URL and return injection-sanitized, size-capped text.

Args: url: an http/https URL to fetch. max_chars: per-fetch character cap (clamped to a hard ceiling).

Returns a dict with the cleaned content and an audit of what was done.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool performs injection sanitization and size capping, and returns a dict with cleaned content and an audit. However, it lacks details on error behavior, rate limits, or side effects, and there are no annotations to supplement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: a single sentence for purpose followed by two bullet-point-like arg descriptions. Every sentence adds value, and the key action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description omits details about the return value structure (e.g., what 'audit' contains), error handling, and supported URL schemes. Given no output schema, this leaves gaps in understanding the tool's full behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by explaining that 'url' must be an http/https URL and 'max_chars' is a per-fetch cap clamped to a hard ceiling. This adds significant meaning beyond the schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches a URL and returns injection-sanitized, size-capped text, which is a specific verb-resource pair. However, since no sibling tools are provided, it cannot be evaluated for differentiation, hence slightly less than perfect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it mention any preconditions or exclusions. It simply describes what the tool does without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.6/5.0
Disambiguation5/5

Only one tool exists, so there is no ambiguity at all. An agent cannot misselect between tools.

Naming Consistency5/5

With a single tool, naming is trivially consistent. The name 'fetch_clean' follows a clear verb_noun pattern.

Tool Count3/5

One tool feels thin for a server dedicated to fetching content. While the tool is well-defined, the server would benefit from additional tools (e.g., batch fetch, status check) to justify its scope.

Completeness4/5

The tool covers the core fetch-and-sanitize operation completely for a single URL. Minor gaps exist (e.g., no way to fetch multiple URLs or configure sanitization levels), but the basic use case is fully served.

Maintenance

ActivityStale
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/PortfolioKB/mcp-safe-fetch'

If you have feedback or need assistance with the MCP directory API, please join our Discord server