pku-opendata-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pku-opendata-mcpRetrieve the file list and sizes for dataset 10.18170/DVN/VXPXUG"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
pku-opendata-mcp
把北京大学开放研究数据平台(https://opendata.pku.edu.cn,由北京大学图书馆运营的 Dataverse 实例,497 个数据集 / 8611 个数据文件)接成 MCP server,让 WorkBuddy、Claude Desktop、Cursor 等客户端可以直接检索。
零依赖、单文件、纯 Node 18+ 实现(用内建 fetch),不需要 npm install。
一、环境要求
Node.js 18 以上(用到全局
fetch与AbortController)无需 API 密钥,无需注册,无需 Python 环境
Related MCP server: DataCite MCP Server
二、提供 5 个工具
工具 | 作用 |
| 返回数据集 / 文件总量与接口边界说明 |
| 全文检索数据集,返回标题、DOI、发布日期、所属子库、着陆页 |
| 文件级检索,返回 file_id、格式、字节数、md5、所属数据集 DOI |
| 用 DOI 精确定位单条数据集记录 + 人类访问页 URL |
| 枚举某个 DOI 下的数据文件(平台无此原生接口,本地过滤实现) |
调用示例(对话里直接说)
检索北大开放数据平台里的老年健康追踪调查数据
找一下 autism 相关的数据集,给出 DOI 和下载页
10.18170/DVN/VXPXUG 这个数据集有哪些文件,分别多大
三、已知边界(实测结论,务必知悉)
本平台只对外开放 /api/search 这一个路径,实测结果:
路径 | 状态 |
| 200 可用,关键词检索、分页均有效 |
| 403 |
| 403 |
| 403 |
URL 中含 | 403(WAF 拦截) |
因此:
✅ 检索、定位、拿元数据、拿文件清单 —— 全部可用且免密钥
❌ 通过 API 直接下载文件内容 —— 平台未提供,本 server 不尝试绕过
受限数据集(Restricted-Use Data)必须在网页端注册后向发布者申请授权,这是平台的访问控制策略
下载请走返回的 DOI 着陆页(https://doi.org/10.18170/DVN/XXXX),在浏览器中完成。
四、安装到 WorkBuddy
方式 A:本地目录(最快,推荐先用这个)
编辑 C:\Users\Administrator\.workbuddy\mcp.json(不是 .mcp.json,没有点前缀),把本目录的路径填进去:
{
"mcpServers": {
"pku-opendata": {
"command": "C:/Users/Administrator/.workbuddy/binaries/node/versions/22.22.2-3/node.exe",
"args": [
"C:/Users/Administrator/WorkBuddy/2026-09-26-23-37-43/pku-opendata-mcp/index.js"
]
}
}
}如果本机已有其他 MCP server,不要覆盖文件——只把
pku-opendata这一条合并进已有的mcpServers对象里。Windows 下建议写
node.exe的完整绝对路径。如果node已在系统 PATH 中,也可以直接写"command": "node"。
方式 B:npm 全局安装
在本目录执行:
npm install -g .然后把命令名写进配置:
{
"mcpServers": {
"pku-opendata": {
"command": "pku-opendata-mcp"
}
}
}方式 C:发布到 npm 后一行接入(面向其他 WorkBuddy 用户)
维护者执行:
npm publish --access public使用者只需:
{
"mcpServers": {
"pku-opendata": {
"command": "npx",
"args": ["-y", "pku-opendata-mcp"]
}
}
}方式 D:直接从 GitHub 分发
{
"mcpServers": {
"pku-opendata": {
"command": "npx",
"args": ["-y", "github:<你的账号>/pku-opendata-mcp"]
}
}
}五、启用步骤
写入 mcp.json 后,MCP server 不会自动生效,还需要一步:
完全退出并重启 WorkBuddy
打开连接器管理页
右上角找到自定义连接器入口,对新出现的
pku-opendata点 信任之后在对话里即可用自然语言调用
六、其他客户端
Claude Desktop(claude_desktop_config.json):
{
"mcpServers": {
"pku-opendata": {
"command": "node",
"args": ["/absolute/path/to/pku-opendata-mcp/index.js"]
}
}
}Cursor(.cursor/mcp.json)与 mcp.json 格式一致,直接套用即可。
七、本地自检
仓库里带了回归 fixture,无需启动客户端即可验证协议层:
node index.js < test-input.jsonl > test-output.jsonl 2>test-err.logindex.js 只向 stdout 输出 JSON-RPC 报文,所有日志走 stderr,不会污染协议。
八、发布前的邮箱隐私检查
npm 会把维护者邮箱公开,且无法撤回。
由于 npm 侧的暴露风险高于 GitHub,发布前请逐项确认:
检查项 | 要求 |
| 保持空字符串,务必不要填邮箱 |
| 不要出现 email |
npm 账号主邮箱 | 使用专用别名邮箱,不要用私人常用邮箱 |
本机 | 用 GitHub noreply 地址 |
GitHub 设置 | 勾选 Keep my email addresses private + Block command line pushes that expose my email |
设置本机提交邮箱(<ID> 见 GitHub → Settings → Emails):
git config --global user.email "<ID>+<用户名>@users.noreply.github.com"npm 即使改用别名邮箱,历史版本的 maintainers 元数据里旧邮箱也会永久留存。首次发布前就把邮箱换好,代价最小。
九、许可
MIT。数据来源为北京大学开放研究数据平台,使用时请按其条款规范引用各数据集的 DOI。
Available Tools
5 toolspku_dataset_filesA
Enumerate files belonging to one dataset by its DOI, since the platform does not expose a per-dataset file listing API.
| Name | Required | Description | Default |
|---|---|---|---|
| per_page | No | Max files to return. Default 20. | |
| identifier | Yes | Dataset DOI, e.g. 10.18170/DVN/VXPXUG. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden of behavioral disclosure. It reveals only the core listing behavior and says nothing about pagination, returned file metadata, error cases, rate limits, or whether this is a best-effort workaround with caveats. This is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tightly written sentence that front-loads the action and resource, then adds only the useful platform-context rationale. Every word earns its place, and the description is neither bloated nor underspecified enough to feel like a placeholder.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with complete schema coverage, the description is minimally viable: an agent knows to provide a dataset DOI and can expect a file list. But with no output schema and no annotations, the lack of any statement about response shape, pagination defaults, or how the per_page parameter affects results leaves the agent with some uncertainty.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so per the baseline this dimension should receive 3 even without additional parameter detail in the description. The description does clarify that 'identifier' is a DOI by saying 'by its DOI', but that repeats the schema's own example, so no substantial meaning is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Enumerate'), a concrete resource ('files belonging to one dataset'), and a precise selection method ('by its DOI'). This clearly distinguishes the tool from sibling pku_search_files, which would search rather than list an entire dataset's files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrasing implies the tool is for retrieving a dataset's files when you have its DOI, and the rationale 'since the platform does not expose a per-dataset file listing API' gives useful context. However, it never explicitly names an alternative or states when not to use it, leaving the agent to infer the boundary against pku_search_files.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pku_get_datasetA
Resolve a DOI to the exact dataset record and give the human-facing landing page URL used to request access or download.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | DOI such as 10.18170/DVN/VXPXUG, or a doi.org URL containing it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It meaningfully discloses that the result is a landing page for requesting access/download rather than a direct file or full dataset payload. It does not discuss error cases or access restrictions, but for a simple read-only resolver this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the action ('Resolve a DOI'), states the target resource, and specifies the output type, making every phrase informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter resolver with no output schema, the description sufficiently states the input (DOI) and the output (dataset record and landing page URL). It lacks explicit return-field details and error behavior, but the simple scope and sibling context make this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage of the single identifier parameter, including an example DOI and doi.org URL format. The description adds no new parameter-level meaning beyond referring to DOI resolution, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Resolve a DOI') and resource ('dataset record'), and clearly states the deliverable: a human-facing landing page URL for access or download. It distinguishes itself from search siblings by emphasizing exact DOI resolution, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The DOI input implies this tool is for when a DOI is already known, but the description gives no explicit when-to-use or when-not-to-use guidance. It does not mention alternatives like pku_search_datasets for finding datasets without a DOI.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pku_platform_statsA
Report dataset and file counts for the Peking University Open Research Data Platform, plus its API access limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses that the tool reports counts and access limits, which suggests a read-only operation, but it does not describe side effects, authentication requirements, or how access limits are represented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every phrase contributes to scope: counts, platform identity, and API access limits.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter schema, low complexity, and absence of an output schema, the description is largely complete. It could be slightly more explicit about the meaning or format of 'API access limits', but this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is complete by default. With no parameters to document, the baseline of 4 applies and the description adds no unnecessary param detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Report') and a clear resource: dataset and file counts for the Peking University Open Research Data Platform, plus API access limits. This distinguishes it from sibling search/get tools, which are focused on individual datasets and files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: this tool is for platform-level statistics rather than per-dataset searches or retrievals. However, it does not explicitly mention alternatives or state when not to use it, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pku_search_datasetsA
Full-text search across the 497 published datasets of Peking University Open Research Data Platform. Returns title, DOI, release date, source dataverse and landing page URL.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search terms, e.g. "longevity". Use * to list everything. Default: * | |
| start | No | Zero-based offset for paging. Default 0. | |
| per_page | No | Results per page, 1-100. Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the return fields (title, DOI, release date, source dataverse, landing page URL) and the scope (497 published datasets), which goes beyond the schema. It does not mention pagination behavior, but the schema covers start/per_page, and it does not describe side effects (none expected for a read-only search). The disclosure of concrete output fields is valuable and transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero redundancy. The primary purpose is front-loaded, and the return fields are listed compactly. Every word earns its place, making it efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with 100% schema coverage and no output schema, the description covers the essential information: what it searches, the scope, and the return fields. It does not explicitly explain how to page through results, but the schema provides start/per_page. It also does not mention sorting or default behavior, but those are minor gaps for a straightforward search. Overall, it is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the parameters are already well-documented in the input schema. The description does not add any extra meaning beyond what the schema provides (e.g., query is the search term, start is offset, per_page is limit). It neither clarifies nor expands on the parameters, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Full-text search') and resource ('datasets') with a precise scope ('497 published datasets of Peking University Open Research Data Platform'). It clearly distinguishes itself from sibling tools like pku_search_files (search files) and pku_get_dataset (retrieve a specific dataset), leaving no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the use case: searching across datasets. It does not explicitly mention alternatives or exclusions, but the verb 'search' and the resource 'datasets' make it obvious this is the tool for dataset search, not file search or dataset retrieval. The context is clear enough for an agent to select correctly without extra guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pku_search_filesA
Search individual data files on the Peking University Open Research Data Platform. Returns file id, format, byte size, md5 and the owning dataset DOI.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search terms, or * for all files. Default: * | |
| start | No | Zero-based offset for paging. Default 0. | |
| per_page | No | Results per page, 1-100. Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the output shape (id, format, byte size, md5, DOI). It does not mention paging behavior, wildcard query semantics, or any operational constraints, but the tool appears simple and non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the purpose and the return values with no filler. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with three optional parameters and full schema coverage, the description is sufficient to understand what it returns and how it differs from dataset-level search. It could name sibling tools or give a usage example, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully describes all three parameters with defaults and ranges, so the baseline is 3. The description adds no parameter-specific meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search'), a specific resource ('individual data files'), and the platform, and it names the returned fields. This clearly distinguishes it from dataset-level search and retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'individual data files' implies this is the file-level search tool and gives some selection context. However, it does not explicitly tell the agent when to use this tool over pku_search_datasets or pku_dataset_files, nor does it name any alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
pku_dataset_files - First observed
pku_get_dataset - First observed
pku_platform_stats - First observed
pku_search_datasets - First observed
pku_search_files
TDQS
Scored across 5 tools
Each tool targets a distinct aspect of the platform: statistics, dataset search, file search, dataset resolution by DOI, and file enumeration for a dataset. No overlaps or ambiguous boundaries; an agent can clearly choose the right tool for each action.
All tools share a consistent 'pku_' prefix, but the naming pattern is mixed: 'search_datasets' and 'get_dataset' follow verb_noun, while 'platform_stats' and 'dataset_files' are noun_noun. This inconsistency could cause slight confusion, though the prefix helps.
Five tools is well-scoped for an open data platform server, covering the essential operations without redundancy. Each tool serves a clear purpose and earns its place.
The tool surface covers core workflows: searching datasets, searching files, resolving DOIs, and listing files for a dataset. A minor gap is the absence of direct file download or detailed metadata retrieval, but the landing page URLs and file metadata returned are sufficient for most use cases.
Maintenance
Related MCP Connectors
Search, sample and query open reproducible datasets published as immutable Parquet with schemas.
Discover, resolve, and query official Brazilian economic data with semantic search and provenance.
Scholarly search: OpenAlex, Crossref, arXiv, OpenCitations and PubMed in one endpoint.
공공데이터포털(data.go.kr) 10.9만 데이터셋을 말로 찾고 컬럼·첫 행까지 바로 확인. 키 없이. Search data.go.kr easily.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables natural language search and discovery of open-access scientific datasets through the EOSC Data Commons OpenSearch service. Provides tools to search datasets and retrieve file metadata using LLM-assisted queries.15MIT
- AlicenseAqualityCmaintenanceProvides read-only access to DataCite's index of 125M+ research DOIs via natural language queries, enabling searching, metadata retrieval, citation formatting, and relationship exploration.9378 npmMIT
- AlicenseAqualityAmaintenanceSearches and fetches research datasets across Zenodo, DataCite (Dryad/Figshare/Dataverse/OSF), NCBI omics archives (GEO/SRA/BioProject), and the literature (PubMed/OpenAIRE) through one normalized model — deduplicating by DOI, expanding organism queries with NCBI Taxonomy synonyms, and bridging papers to the datasets they produced. Resolves citations and open-access full text, and downloads files.6310 PyPI4MIT
- AlicenseNot gradedqualityDmaintenanceSearches and retrieves scholarly metadata from the CrossRef REST API, covering over 150 million records across all disciplines, without requiring an API key.MIT