Skip to main content
Glama
jialinhome

page-fetcher-mcp

by jialinhome
README.md
# Page Fetcher MCP Server

一个用于替代 Claude Code 内置 fetch 功能的 MCP (Model Context Protocol) 服务器,提供网页抓取和 API 数据获取能力。

## 功能特性

- **fetch_url**: 从 URL 获取数据(API 端点、JSON 或任何网页内容)
- **scrape_url**: 抓取并提取网页的可读文本内容

## 安装

```bash
npm install
```

## 开发

### 构建项目

```bash
npm run build
```

### 开发模式运行

```bash
npm run dev
```

### 生产模式运行

```bash
npm start
```

## 使用方法

### 在 Claude Code 中配置

将此服务器添加到 Claude Code 的 MCP 配置中:

```json
{
  "mcpServers": {
    "page-fetcher-mcp": {
      "command": "node",
      "args": ["/path/to/page-fetcher-mcp/dist/index.js"],
      "env": {}
    }
  }
}
```

### 可用工具

#### fetch_url

获取任意 URL 的数据(API、JSON、网页内容等)

**参数:**
- `url` (必需): 要获取的 URL
- `headers` (可选): 要包含在请求中的 HTTP 头

**示例:**
```javascript
// 获取 API 数据
await mcp_fetch_call_tool('fetch_url', {
  url: 'https://api.github.com/users/octocat'
});

// 获取 JSON 数据
await mcp_fetch_call_tool('fetch_url', {
  url: 'https://jsonplaceholder.typicode.com/posts/1'
});

// 带自定义头获取
await mcp_fetch_call_tool('fetch_url', {
  url: 'https://api.example.com/data',
  headers: {
    'Authorization': 'Bearer YOUR_TOKEN'
  }
});
```

#### scrape_url

抓取网页并提取可读文本内容

**参数:**
- `url` (必需): 要抓取的 URL
- `selector` (可选): CSS 选择器,用于提取特定元素

**示例:**
```javascript
// 获取整个页面的文本内容
await mcp_fetch_call_tool('scrape_url', {
  url: 'https://example.com'
});

// 使用 CSS 选择器提取特定内容
await mcp_fetch_call_tool('scrape_url', {
  url: 'https://example.com',
  selector: 'h1'  // 只提取所有 h1 标题
});

// 提取文章内容
await mcp_fetch_call_tool('scrape_url', {
  url: 'https://example.com/article',
  selector: '.article-content'
});
```

## 项目结构

```
page-fetcher-mcp/
├── src/
│   └── index.ts          # 主服务器实现
├── dist/                 # 编译输出目录
├── package.json          # 项目配置
├── tsconfig.json         # TypeScript 配置
└── README.md            # 项目文档
```

## 技术栈

- Node.js
- TypeScript
- [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/sdk) - MCP 协议 SDK
- [axios](https://github.com/axios/axios) - HTTP 客户端
- [cheerio](https://github.com/cheeriojs/cheerio) - HTML 解析和爬虫

## 许可证

MIT

TDQS

C2.9/5.0

Scored across 2 tools

Disambiguation3/5

The two tools have broadly similar purposes - both fetch data from URLs. The distinction between 'fetch data' and 'scrape readable text content' is somewhat clear but could easily cause misselection for agents wanting raw HTML vs parsed content. The line between the two is fuzzy.

Naming Consistency4/5

Both tools follow the verb_noun pattern (fetch_url, scrape_url) with consistent snake_case naming. The conventions match, though the verbs are distinct in style rather than sharing a common action term.

Tool Count2/5

With only 2 tools, the server feels extremely thin. Even for a focused purpose like fetching web content, one might expect additional tools for handling different output formats, following redirects, or handling authentication. Two tools barely justify a server.

Completeness2/5

The surface appears incomplete even for the stated purpose. Missing operations like fetching HTML versus JSON versus extracted text as separate concerns, handling redirects, following links, or downloading files. The gap between 'raw fetch' and 'readable text extract' leaves a large middle ground uncovered.

Maintenance

ActivityInactive
ResponsivenessNo issues