Skip to main content
Glama
joyfullyplump

pku-opendata-mcp

README.md
# pku-opendata-mcp

把**北京大学开放研究数据平台**(<https://opendata.pku.edu.cn>,由北京大学图书馆运营的 Dataverse 实例,497 个数据集 / 8611 个数据文件)接成 MCP server,让 WorkBuddy、Claude Desktop、Cursor 等客户端可以直接检索。

零依赖、单文件、纯 Node 18+ 实现(用内建 `fetch`),不需要 `npm install`。

---

## 一、环境要求

- Node.js **18 以上**(用到全局 `fetch` 与 `AbortController`)
- 无需 API 密钥,无需注册,无需 Python 环境

---

## 二、提供 5 个工具

| 工具 | 作用 |
|---|---|
| `pku_platform_stats` | 返回数据集 / 文件总量与接口边界说明 |
| `pku_search_datasets` | 全文检索数据集,返回标题、DOI、发布日期、所属子库、着陆页 |
| `pku_search_files` | 文件级检索,返回 file_id、格式、字节数、md5、所属数据集 DOI |
| `pku_get_dataset` | 用 DOI 精确定位单条数据集记录 + 人类访问页 URL |
| `pku_dataset_files` | 枚举某个 DOI 下的数据文件(平台无此原生接口,本地过滤实现) |

### 调用示例(对话里直接说)

> 检索北大开放数据平台里的老年健康追踪调查数据

> 找一下 autism 相关的数据集,给出 DOI 和下载页

> 10.18170/DVN/VXPXUG 这个数据集有哪些文件,分别多大

---

## 三、已知边界(实测结论,务必知悉)

本平台**只对外开放 `/api/search` 这一个路径**,实测结果:

| 路径 | 状态 |
|---|---|
| `/api/search?q=...&type=dataset\|file` | **200 可用**,关键词检索、分页均有效 |
| `/api/datasets/{id}`、`/versions/latest/files` | 403 |
| `/api/files/{id}/metadata` | 403 |
| `/api/info/*`、`/api/metrics/*` | 403 |
| URL 中含 `:` 的 `persistentId` 写法 | 403(WAF 拦截) |

因此:

- ✅ 检索、定位、拿元数据、拿文件清单 —— 全部可用且免密钥
- ❌ **通过 API 直接下载文件内容** —— 平台未提供,本 server 不尝试绕过
- 受限数据集(Restricted-Use Data)必须在网页端注册后向发布者申请授权,这是平台的访问控制策略

下载请走返回的 DOI 着陆页(`https://doi.org/10.18170/DVN/XXXX`),在浏览器中完成。

---

## 四、安装到 WorkBuddy

### 方式 A:本地目录(最快,推荐先用这个)

编辑 `C:\Users\Administrator\.workbuddy\mcp.json`(**不是** `.mcp.json`,没有点前缀),把本目录的路径填进去:

```json
{
  "mcpServers": {
    "pku-opendata": {
      "command": "C:/Users/Administrator/.workbuddy/binaries/node/versions/22.22.2-3/node.exe",
      "args": [
        "C:/Users/Administrator/WorkBuddy/2026-09-26-23-37-43/pku-opendata-mcp/index.js"
      ]
    }
  }
}
```

> 如果本机已有其他 MCP server,不要覆盖文件——只把 `pku-opendata` 这一条**合并**进已有的 `mcpServers` 对象里。
>
> Windows 下建议写 `node.exe` 的完整绝对路径。如果 `node` 已在系统 PATH 中,也可以直接写 `"command": "node"`。

### 方式 B:npm 全局安装

在本目录执行:

```bash
npm install -g .
```

然后把命令名写进配置:

```json
{
  "mcpServers": {
    "pku-opendata": {
      "command": "pku-opendata-mcp"
    }
  }
}
```

### 方式 C:发布到 npm 后一行接入(面向其他 WorkBuddy 用户)

维护者执行:

```bash
npm publish --access public
```

使用者只需:

```json
{
  "mcpServers": {
    "pku-opendata": {
      "command": "npx",
      "args": ["-y", "pku-opendata-mcp"]
    }
  }
}
```

### 方式 D:直接从 GitHub 分发

```json
{
  "mcpServers": {
    "pku-opendata": {
      "command": "npx",
      "args": ["-y", "github:<你的账号>/pku-opendata-mcp"]
    }
  }
}
```

---

## 五、启用步骤

写入 `mcp.json` 后,MCP server **不会自动生效**,还需要一步:

1. 完全退出并重启 WorkBuddy
2. 打开连接器管理页
3. 右上角找到自定义连接器入口,对新出现的 `pku-opendata` 点 **信任**
4. 之后在对话里即可用自然语言调用

---

## 六、其他客户端

Claude Desktop(`claude_desktop_config.json`):

```json
{
  "mcpServers": {
    "pku-opendata": {
      "command": "node",
      "args": ["/absolute/path/to/pku-opendata-mcp/index.js"]
    }
  }
}
```

Cursor(`.cursor/mcp.json`)与 mcp.json 格式一致,直接套用即可。

---

## 七、本地自检

仓库里带了回归 fixture,无需启动客户端即可验证协议层:

```bash
node index.js < test-input.jsonl > test-output.jsonl 2>test-err.log
```

`index.js` 只向 **stdout** 输出 JSON-RPC 报文,所有日志走 stderr,不会污染协议。

---

## 八、发布前的邮箱隐私检查

**npm 会把维护者邮箱公开,且无法撤回。**

由于 npm 侧的暴露风险高于 GitHub,发布前请逐项确认:

| 检查项 | 要求 |
|---|---|
| `package.json` 的 `author` | 保持空字符串,**务必不要填邮箱** |
| `contributors` 字段 | 不要出现 email |
| npm 账号主邮箱 | 使用专用别名邮箱,不要用私人常用邮箱 |
| 本机 `git config user.email` | 用 GitHub noreply 地址 |
| GitHub 设置 | 勾选 Keep my email addresses private + Block command line pushes that expose my email |

设置本机提交邮箱(`<ID>` 见 GitHub → Settings → Emails):

```bash
git config --global user.email "<ID>+<用户名>@users.noreply.github.com"
```

> npm 即使改用别名邮箱,历史版本的 maintainers 元数据里旧邮箱也会永久留存。首次发布前就把邮箱换好,代价最小。

## 九、许可

MIT。数据来源为北京大学开放研究数据平台,使用时请按其条款规范引用各数据集的 DOI。

TDQS

A3.9/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct aspect of the platform: statistics, dataset search, file search, dataset resolution by DOI, and file enumeration for a dataset. No overlaps or ambiguous boundaries; an agent can clearly choose the right tool for each action.

Naming Consistency3/5

All tools share a consistent 'pku_' prefix, but the naming pattern is mixed: 'search_datasets' and 'get_dataset' follow verb_noun, while 'platform_stats' and 'dataset_files' are noun_noun. This inconsistency could cause slight confusion, though the prefix helps.

Tool Count5/5

Five tools is well-scoped for an open data platform server, covering the essential operations without redundancy. Each tool serves a clear purpose and earns its place.

Completeness4/5

The tool surface covers core workflows: searching datasets, searching files, resolving DOIs, and listing files for a dataset. A minor gap is the absence of direct file download or detailed metadata retrieval, but the landing page URLs and file metadata returned are sufficient for most use cases.

Maintenance

ActivityMaintained
ResponsivenessNo issues