Skip to main content
Glama
joyfullyplump

pku-opendata-mcp

pku-opendata-mcp

把北京大学开放研究数据平台(https://opendata.pku.edu.cn,由北京大学图书馆运营的 Dataverse 实例,497 个数据集 / 8611 个数据文件)接成 MCP server,让 WorkBuddy、Claude Desktop、Cursor 等客户端可以直接检索。

零依赖、单文件、纯 Node 18+ 实现(用内建 fetch),不需要 npm install。


一、环境要求

  • Node.js 18 以上(用到全局 fetch 与 AbortController)

  • 无需 API 密钥,无需注册,无需 Python 环境


Related MCP server: DataCite MCP Server

二、提供 5 个工具

工具

作用

pku_platform_stats

返回数据集 / 文件总量与接口边界说明

pku_search_datasets

全文检索数据集,返回标题、DOI、发布日期、所属子库、着陆页

pku_search_files

文件级检索,返回 file_id、格式、字节数、md5、所属数据集 DOI

pku_get_dataset

用 DOI 精确定位单条数据集记录 + 人类访问页 URL

pku_dataset_files

枚举某个 DOI 下的数据文件(平台无此原生接口,本地过滤实现)

调用示例(对话里直接说)

检索北大开放数据平台里的老年健康追踪调查数据

找一下 autism 相关的数据集,给出 DOI 和下载页

10.18170/DVN/VXPXUG 这个数据集有哪些文件,分别多大


三、已知边界(实测结论,务必知悉)

本平台只对外开放 /api/search 这一个路径,实测结果:

路径

状态

/api/search?q=...&type=dataset|file

200 可用,关键词检索、分页均有效

/api/datasets/{id}、/versions/latest/files

403

/api/files/{id}/metadata

403

/api/info/*、/api/metrics/*

403

URL 中含 : 的 persistentId 写法

403(WAF 拦截)

因此:

  • ✅ 检索、定位、拿元数据、拿文件清单 —— 全部可用且免密钥

  • ❌ 通过 API 直接下载文件内容 —— 平台未提供,本 server 不尝试绕过

  • 受限数据集(Restricted-Use Data)必须在网页端注册后向发布者申请授权,这是平台的访问控制策略

下载请走返回的 DOI 着陆页(https://doi.org/10.18170/DVN/XXXX),在浏览器中完成。


四、安装到 WorkBuddy

方式 A:本地目录(最快,推荐先用这个)

编辑 C:\Users\Administrator\.workbuddy\mcp.json(不是 .mcp.json,没有点前缀),把本目录的路径填进去:

{
  "mcpServers": {
    "pku-opendata": {
      "command": "C:/Users/Administrator/.workbuddy/binaries/node/versions/22.22.2-3/node.exe",
      "args": [
        "C:/Users/Administrator/WorkBuddy/2026-09-26-23-37-43/pku-opendata-mcp/index.js"
      ]
    }
  }
}

如果本机已有其他 MCP server,不要覆盖文件——只把 pku-opendata 这一条合并进已有的 mcpServers 对象里。

Windows 下建议写 node.exe 的完整绝对路径。如果 node 已在系统 PATH 中,也可以直接写 "command": "node"。

方式 B:npm 全局安装

在本目录执行:

npm install -g .

然后把命令名写进配置:

{
  "mcpServers": {
    "pku-opendata": {
      "command": "pku-opendata-mcp"
    }
  }
}

方式 C:发布到 npm 后一行接入(面向其他 WorkBuddy 用户)

维护者执行:

npm publish --access public

使用者只需:

{
  "mcpServers": {
    "pku-opendata": {
      "command": "npx",
      "args": ["-y", "pku-opendata-mcp"]
    }
  }
}

方式 D:直接从 GitHub 分发

{
  "mcpServers": {
    "pku-opendata": {
      "command": "npx",
      "args": ["-y", "github:<你的账号>/pku-opendata-mcp"]
    }
  }
}

五、启用步骤

写入 mcp.json 后,MCP server 不会自动生效,还需要一步:

  1. 完全退出并重启 WorkBuddy

  2. 打开连接器管理页

  3. 右上角找到自定义连接器入口,对新出现的 pku-opendata 点 信任

  4. 之后在对话里即可用自然语言调用


六、其他客户端

Claude Desktop(claude_desktop_config.json):

{
  "mcpServers": {
    "pku-opendata": {
      "command": "node",
      "args": ["/absolute/path/to/pku-opendata-mcp/index.js"]
    }
  }
}

Cursor(.cursor/mcp.json)与 mcp.json 格式一致,直接套用即可。


七、本地自检

仓库里带了回归 fixture,无需启动客户端即可验证协议层:

node index.js < test-input.jsonl > test-output.jsonl 2>test-err.log

index.js 只向 stdout 输出 JSON-RPC 报文,所有日志走 stderr,不会污染协议。


八、发布前的邮箱隐私检查

npm 会把维护者邮箱公开,且无法撤回。

由于 npm 侧的暴露风险高于 GitHub,发布前请逐项确认:

检查项

要求

package.json 的 author

保持空字符串,务必不要填邮箱

contributors 字段

不要出现 email

npm 账号主邮箱

使用专用别名邮箱,不要用私人常用邮箱

本机 git config user.email

用 GitHub noreply 地址

GitHub 设置

勾选 Keep my email addresses private + Block command line pushes that expose my email

设置本机提交邮箱(<ID> 见 GitHub → Settings → Emails):

git config --global user.email "<ID>+<用户名>@users.noreply.github.com"

npm 即使改用别名邮箱,历史版本的 maintainers 元数据里旧邮箱也会永久留存。首次发布前就把邮箱换好,代价最小。

九、许可

MIT。数据来源为北京大学开放研究数据平台,使用时请按其条款规范引用各数据集的 DOI。

Available Tools

5 tools
pku_dataset_filesA

Enumerate files belonging to one dataset by its DOI, since the platform does not expose a per-dataset file listing API.

ParametersJSON Schema
NameRequiredDescriptionDefault
per_pageNoMax files to return. Default 20.
identifierYesDataset DOI, e.g. 10.18170/DVN/VXPXUG.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden of behavioral disclosure. It reveals only the core listing behavior and says nothing about pagination, returned file metadata, error cases, rate limits, or whether this is a best-effort workaround with caveats. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly written sentence that front-loads the action and resource, then adds only the useful platform-context rationale. Every word earns its place, and the description is neither bloated nor underspecified enough to feel like a placeholder.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with complete schema coverage, the description is minimally viable: an agent knows to provide a dataset DOI and can expect a file list. But with no output schema and no annotations, the lack of any statement about response shape, pagination defaults, or how the per_page parameter affects results leaves the agent with some uncertainty.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so per the baseline this dimension should receive 3 even without additional parameter detail in the description. The description does clarify that 'identifier' is a DOI by saying 'by its DOI', but that repeats the schema's own example, so no substantial meaning is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Enumerate'), a concrete resource ('files belonging to one dataset'), and a precise selection method ('by its DOI'). This clearly distinguishes the tool from sibling pku_search_files, which would search rather than list an entire dataset's files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrasing implies the tool is for retrieving a dataset's files when you have its DOI, and the rationale 'since the platform does not expose a per-dataset file listing API' gives useful context. However, it never explicitly names an alternative or states when not to use it, leaving the agent to infer the boundary against pku_search_files.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pku_get_datasetA

Resolve a DOI to the exact dataset record and give the human-facing landing page URL used to request access or download.

ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYesDOI such as 10.18170/DVN/VXPXUG, or a doi.org URL containing it.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It meaningfully discloses that the result is a landing page for requesting access/download rather than a direct file or full dataset payload. It does not discuss error cases or access restrictions, but for a simple read-only resolver this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the action ('Resolve a DOI'), states the target resource, and specifies the output type, making every phrase informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter resolver with no output schema, the description sufficiently states the input (DOI) and the output (dataset record and landing page URL). It lacks explicit return-field details and error behavior, but the simple scope and sibling context make this a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage of the single identifier parameter, including an example DOI and doi.org URL format. The description adds no new parameter-level meaning beyond referring to DOI resolution, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Resolve a DOI') and resource ('dataset record'), and clearly states the deliverable: a human-facing landing page URL for access or download. It distinguishes itself from search siblings by emphasizing exact DOI resolution, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The DOI input implies this tool is for when a DOI is already known, but the description gives no explicit when-to-use or when-not-to-use guidance. It does not mention alternatives like pku_search_datasets for finding datasets without a DOI.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pku_platform_statsA

Report dataset and file counts for the Peking University Open Research Data Platform, plus its API access limits.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It discloses that the tool reports counts and access limits, which suggests a read-only operation, but it does not describe side effects, authentication requirements, or how access limits are represented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every phrase contributes to scope: counts, platform identity, and API access limits.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter schema, low complexity, and absence of an output schema, the description is largely complete. It could be slightly more explicit about the meaning or format of 'API access limits', but this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is complete by default. With no parameters to document, the baseline of 4 applies and the description adds no unnecessary param detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Report') and a clear resource: dataset and file counts for the Peking University Open Research Data Platform, plus API access limits. This distinguishes it from sibling search/get tools, which are focused on individual datasets and files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: this tool is for platform-level statistics rather than per-dataset searches or retrievals. However, it does not explicitly mention alternatives or state when not to use it, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pku_search_datasetsA

Full-text search across the 497 published datasets of Peking University Open Research Data Platform. Returns title, DOI, release date, source dataverse and landing page URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoSearch terms, e.g. "longevity". Use * to list everything. Default: *
startNoZero-based offset for paging. Default 0.
per_pageNoResults per page, 1-100. Default 20.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the return fields (title, DOI, release date, source dataverse, landing page URL) and the scope (497 published datasets), which goes beyond the schema. It does not mention pagination behavior, but the schema covers start/per_page, and it does not describe side effects (none expected for a read-only search). The disclosure of concrete output fields is valuable and transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero redundancy. The primary purpose is front-loaded, and the return fields are listed compactly. Every word earns its place, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple search tool with 100% schema coverage and no output schema, the description covers the essential information: what it searches, the scope, and the return fields. It does not explicitly explain how to page through results, but the schema provides start/per_page. It also does not mention sorting or default behavior, but those are minor gaps for a straightforward search. Overall, it is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameters are already well-documented in the input schema. The description does not add any extra meaning beyond what the schema provides (e.g., query is the search term, start is offset, per_page is limit). It neither clarifies nor expands on the parameters, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Full-text search') and resource ('datasets') with a precise scope ('497 published datasets of Peking University Open Research Data Platform'). It clearly distinguishes itself from sibling tools like pku_search_files (search files) and pku_get_dataset (retrieve a specific dataset), leaving no ambiguity about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the use case: searching across datasets. It does not explicitly mention alternatives or exclusions, but the verb 'search' and the resource 'datasets' make it obvious this is the tool for dataset search, not file search or dataset retrieval. The context is clear enough for an agent to select correctly without extra guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pku_search_filesA

Search individual data files on the Peking University Open Research Data Platform. Returns file id, format, byte size, md5 and the owning dataset DOI.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoSearch terms, or * for all files. Default: *
startNoZero-based offset for paging. Default 0.
per_pageNoResults per page, 1-100. Default 20.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose the output shape (id, format, byte size, md5, DOI). It does not mention paging behavior, wildcard query semantics, or any operational constraints, but the tool appears simple and non-destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the purpose and the return values with no filler. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple search tool with three optional parameters and full schema coverage, the description is sufficient to understand what it returns and how it differs from dataset-level search. It could name sibling tools or give a usage example, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully describes all three parameters with defaults and ranges, so the baseline is 3. The description adds no parameter-specific meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search'), a specific resource ('individual data files'), and the platform, and it names the returned fields. This clearly distinguishes it from dataset-level search and retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'individual data files' implies this is the file-level search tool and gives some selection context. However, it does not explicitly tell the agent when to use this tool over pku_search_datasets or pku_dataset_files, nor does it name any alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv1.0.0
    • First observedpku_dataset_files
    • First observedpku_get_dataset
    • First observedpku_platform_stats
    • First observedpku_search_datasets
    • First observedpku_search_files

TDQS

A3.9/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct aspect of the platform: statistics, dataset search, file search, dataset resolution by DOI, and file enumeration for a dataset. No overlaps or ambiguous boundaries; an agent can clearly choose the right tool for each action.

Naming Consistency3/5

All tools share a consistent 'pku_' prefix, but the naming pattern is mixed: 'search_datasets' and 'get_dataset' follow verb_noun, while 'platform_stats' and 'dataset_files' are noun_noun. This inconsistency could cause slight confusion, though the prefix helps.

Tool Count5/5

Five tools is well-scoped for an open data platform server, covering the essential operations without redundancy. Each tool serves a clear purpose and earns its place.

Completeness4/5

The tool surface covers core workflows: searching datasets, searching files, resolving DOIs, and listing files for a dataset. A minor gap is the absence of direct file download or detailed metadata retrieval, but the landing page URLs and file metadata returned are sufficient for most use cases.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables natural language search and discovery of open-access scientific datasets through the EOSC Data Commons OpenSearch service. Provides tools to search datasets and retrieve file metadata using LLM-assisted queries.
    15
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Provides read-only access to DataCite's index of 125M+ research DOIs via natural language queries, enabling searching, metadata retrieval, citation formatting, and relationship exploration.
    9
    378 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Searches and fetches research datasets across Zenodo, DataCite (Dryad/Figshare/Dataverse/OSF), NCBI omics archives (GEO/SRA/BioProject), and the literature (PubMed/OpenAIRE) through one normalized model — deduplicating by DOI, expanding organism queries with NCBI Taxonomy synonyms, and bridging papers to the datasets they produced. Resolves citations and open-access full text, and downloads files.
    6
    310 PyPI
    4
    MIT