Skip to main content
Glama
CormickKneey

Web Fetch MCP Server

by CormickKneey

Web Fetch MCP Server (Python)

基于 zcaceres/fetch-mcp 的 Python 实现,使用 FastMCP 框架。

功能

提供 4 个网页内容获取工具:

工具

描述

fetch_html

获取网页原始 HTML

fetch_json

获取并解析 JSON 数据

fetch_txt

获取纯文本(去除 HTML 标签)

fetch_markdown

获取网页并转换为 Markdown

所有工具支持以下参数:

  • url - 目标 URL(必填)

  • headers - 自定义请求头(可选)

  • max_length - 最大返回字符数,默认 50000(可选)

  • start_index - 起始字符位置,默认 0(可选)

Related MCP server: Fetch MCP Server

使用方式

本地运行

# 安装依赖
uv venv && source .venv/bin/activate
uv pip install -e .

# Stdio 模式
uv run python fetch_mcp_server.py

# SSE 模式(HTTP 服务)
uv run fastmcp run fetch_mcp_server.py:mcp --transport sse --port 8000

Docker 运行

# 构建镜像
docker build -t web-fetch-mcp .

# 运行容器
docker run -d -p 8000:8000 --name web-fetch-mcp web-fetch-mcp

# 或使用 docker-compose
docker-compose up -d

MCP 客户端配置

Streamable HTTP 模式(Docker 默认)

{
  "mcpServers": {
    "fetch": {
      "url": "http://localhost:8000/mcp"
    }
  }
}

SSE 模式

{
  "mcpServers": {
    "fetch": {
      "url": "http://localhost:8000/sse"
    }
  }
}

Stdio 模式

{
  "mcpServers": {
    "fetch": {
      "command": "python",
      "args": ["/path/to/fetch_mcp_server.py"]
    }
  }
}

环境变量

  • DEFAULT_LIMIT - 默认最大返回字符数(默认: 50000)

License

MIT

Available Tools

4 tools
fetch_htmlA

Fetch a website and return its unmodified contents as HTML.

Args: url: URL of the website to fetch headers: Optional headers to include in the request max_length: Maximum number of characters to return (default: 50000) start_index: Start content from this character index (default: 0)

Returns: The raw HTML content of the webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
headersNo
max_lengthNo
start_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It does add useful behavior: the tool returns unmodified raw HTML and supports max_length/start_index for content limits. However, it does not disclose handling of HTTP errors, redirects, timeouts, encodings, or non-HTML responses.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The structure is clean and front-loaded: one purpose sentence, an Args list, and a Returns line. Each parameter earns its place because the schema lacks descriptions. The Returns line is somewhat redundant with the opening sentence, which costs a point for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only fetch tool with four parameters and no annotations, the definition covers the core behavior and every parameter meaning. Because an output schema exists, the return value needs less explanation. The main completeness gap is the lack of guidance on selecting this tool over the sibling fetch_* tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully compensates by explaining all four parameters: the URL to fetch, optional headers, max character limit, and start index. Defaults are also restated clearly, giving the agent enough meaning to invoke the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific action and result: 'Fetch a website and return its unmodified contents as HTML.' This clearly identifies the resource and output format, and the 'unmodified HTML' wording differentiates it from the siblings fetch_json, fetch_txt, and fetch_markdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance is given, and none of the sibling tools are mentioned. The phrase 'return its unmodified contents as HTML' implies it should be chosen when raw HTML is needed, but the agent is left to infer this distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_jsonA

Fetch a JSON file from a URL.

Args: url: URL of the JSON to fetch headers: Optional headers to include in the request max_length: Maximum number of characters to return (default: 50000) start_index: Start content from this character index (default: 0)

Returns: The JSON content as a string

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
headersNo
max_lengthNo
start_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the core behavior (returns the JSON content as a string, honors max_length and start_index) but omits error handling, status codes, content-type validation, and header usage details. These are meaningful gaps for a network tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact docstring with a summary line, Args list, and Returns line. Every sentence adds information, and the summary is front-loaded. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It covers the essential invocation details: the required parameter, optional parameters, and return type. An agent can correctly call it using only this description. It falls short of 5 because it does not address error scenarios or explicitly relate the tool to siblings, but output schema and simple parameter set keep the gap small.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description compensates fully by explaining all four parameters in the Args section, including defaults and the char-based meaning of max_length and start_index. Without this, an agent could not correctly invoke the tool with those options.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line 'Fetch a JSON file from a URL' uses a specific verb and resource, naming the exact content type. This clearly distinguishes it from siblings like fetch_html and fetch_markdown by the format it retrieves.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear context: use this tool when you need to fetch JSON from a URL. It does not explicitly name alternative tools or exclusion criteria, so it is not a 5, but the intended use case is unmistakable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_markdownA

Fetch a website and return the content as Markdown.

Args: url: URL of the website to fetch headers: Optional headers to include in the request max_length: Maximum number of characters to return (default: 50000) start_index: Start content from this character index (default: 0)

Returns: The content of the webpage converted to Markdown format

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
headersNo
max_lengthNo
start_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral disclosure burden. It states the core HTML-to-Markdown conversion behavior and gives useful operational detail through max_length and start_index, implying truncation/pagination. However, it does not disclose error handling, redirects, JavaScript rendering, rate limits, or behavior on fetch failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured with Args and Returns sections. Each parameter is on one line with a concise explanation, and there is no filler or unnecessary repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple four-parameter fetch tool with an output schema present, the description is largely complete: it covers all parameters and the return type. It loses a point because, given four similar siblings, a short routing note on when Markdown is preferred over HTML/JSON/TXT would make it fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the Args block explains all four parameters clearly: URL to fetch, optional headers, maximum character count, and start index. This adequately compensates for the bare schema and clarifies the meaning of each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Fetch') and resource ('a website'), and the phrase 'return the content as Markdown' clearly distinguishes it from sibling tools like fetch_html, fetch_json, and fetch_txt by output format. The tool's job is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when Markdown output is desired, but it does not explicitly state when to prefer this tool over fetch_html, fetch_json, or fetch_txt, nor does it mention any exclusions or alternatives. Sibling tools must be disambiguated mainly by name and inferred format.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_txtA

Fetch a website and return the content as plain text (no HTML).

Args: url: URL of the website to fetch headers: Optional headers to include in the request max_length: Maximum number of characters to return (default: 50000) start_index: Start content from this character index (default: 0)

Returns: The text content of the webpage with HTML tags, scripts, and styles removed

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
headersNo
max_lengthNo
start_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It discloses key processing behavior: HTML tags, scripts, and styles are stripped, and output is limited by max_length and start_index. It does not mention error handling, redirects, or authentication, but for a simple fetch-and-clean tool these omissions are acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact docstring with a clear one-line purpose, structured Args, and a Returns section. Every sentence earns its place, and the main purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with an output schema and no annotations, the description is largely complete: it covers the core behavior, all parameters, and the return type. Minor gaps are the lack of notes on error behavior, character encoding, or whether redirects are followed, but these are not critical for basic use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides zero parameter descriptions, so the description fully compensates by explaining all four parameters: url, headers, max_length, and start_index, including default behavior. This gives an agent the semantic meaning needed to invoke the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Fetch') and resource ('a website') and clearly distinguishes the output format ('plain text (no HTML)') from sibling tools like fetch_html, fetch_json, and fetch_markdown. The phrase 'with HTML tags, scripts, and styles removed' makes the tool's purpose immediately unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the appropriate use case: when you need raw text content of a webpage rather than HTML, JSON, or Markdown. It does not explicitly name alternatives or list exclusions, but the output-format framing gives an agent enough context to select this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedfetch_html
    • First observedfetch_json
    • First observedfetch_markdown
    • First observedfetch_txt

TDQS

A4.4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool is clearly distinguished by its output format: HTML, JSON, plain text, and Markdown. There is no overlap in purpose since the format determines the correct tool to use.

Naming Consistency5/5

All tool names follow the exact same verb_noun pattern: fetch_html, fetch_json, fetch_txt, fetch_markdown. This makes the tool set highly predictable and easy to navigate.

Tool Count5/5

Four tools is well-scoped for a web fetching server, covering the most common output formats without unnecessary bloat. Each tool earns its place and there is no redundancy.

Completeness5/5

The tool surface fully covers the stated purpose of fetching web content, including raw HTML, structured JSON, readable text, and Markdown. There are no obvious dead ends or missing core operations for a fetch-only server.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers