Zotero MCP
Zotero MCP
面向本地 Zotero 库的只读 MCP 服务器。可浏览分类、查看论文元数据,并通过 FastMCP 工具从 PDF 中提取全文。
要求
Zotero 桌面版已同步到
~/Zotero/zotero.sqlite(Linux 默认路径)Python 3.13+
Related MCP server: zotero-mcp
配置
这是一个 MCP 服务器。你需要在编码代理的 MCP 配置中注册它,之后代理即可使用其工具。
大多数代理接受类似的配置。例如,在 Opencode 中,你将其添加到 opencode.json:
{
"mcp": {
"zotero-mcp": {
"command": [
"uvx",
"--from",
"git+https://github.com/404Simon/zotero-mcp",
"zotero-mcp"
],
"enabled": true,
"type": "local"
}
}
}完成此设置后,代理会自动发现工具,你可以直接向它提问,例如:
“我的库中有哪些关于 RAG 的论文?最新一篇的摘要是什么?”
工具
list_library
以格式化树形结构列出所有分类和论文。可选按分类名称、论文标题或作者过滤。
参数 | 类型 | 描述 |
|
| 按分类名称、论文标题或作者过滤(不区分大小写) |
每篇论文行都包含其条目键,显示在 [KEY] 方括号中。可将该键用于 paper_details 或 paper_text。过滤后的输出仅显示匹配的子树,包括匹配的分类及其祖先。
├── AI (2 papers)
│ [N3G6XKB9] [preprint] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2021) - Patrick Lewis
│ [PPJJCMXJ] [book] Grundkurs Künstliche Intelligenz: eine praxisorientierte Einführung (2021) - Wolfgang Ertel
├── Bachelorarbeit (42 papers)
│ ├── GraalVM (4 papers)
│ │ [DSUERN67] [book] Supercharge your applications with GraalVM ... (2021) - A. B. Vijay Kumar
│ └── Java Performance (1 papers)
│ [Q3ANMGC3] [conferencePaper] Applying Optimizations for Dynamically-typed Languages to Java (2017) - Matthias Grimmer
│ [MTSF327R] [book] Pro Spring Boot 3: An Authoritative Guide with Best Practices (2024) - Felipe Gutierrez
├── Studienarbeit (49 papers)
│ [T3ZCDWC7] [preprint] MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (2024) - Yubo Wang
├── T3000 (4 papers)
│ [DTZF77Y4] [webpage] Conventional Commits (n.d.) - Unknown
├── TheGreenEpoch (8 papers)
│ [YINRQ63P] [preprint] Distributed LLM Pretraining During Renewable Curtailment Windows (2026) - Philipp Wiesner
└── VesSkel (19 papers)
[ZUXQUGHW] [journalArticle] Open-source analysis and visualization of segmented vasculature datasets with VesselVio (2022) - Jacob R. Bumgarnerpaper_details
通过条目键获取论文的完整元数据。可从 list_library 输出(显示为 [KEY])或 search_papers 结果中获取该键。
参数 | 类型 | 描述 |
|
| Zotero 条目键(在 list_library 输出或 search_papers 结果中显示为 |
返回标题、类型、键、添加/修改日期、作者、所有字段元数据、所属分类以及 PDF 附件信息:
Title: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Type: preprint
Key: T3ZCDWC7
Authors: Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, ...
date: 2024-11-06
DOI: 10.48550/arXiv.2406.01574
url: http://arxiv.org/abs/2406.01574
abstractNote: In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) ...
Collections: Studienarbeit
PDF: Wang et al. - 2024 - MMLU-Pro A More Robust and Challenging Multi-Task Language Understanding Benchmark.pdfsearch_papers
按标题或作者搜索所有论文。返回包含条目键的结构化结果。
参数 | 类型 | 描述 |
|
| 搜索词(不区分大小写,匹配标题和作者) |
[{"key": "ZUXQUGHW",
"title": "Open-source analysis and visualization of segmented vasculature datasets with VesselVio",
"type": "journalArticle", "year": "2022",
"first_author": "Jacob R. Bumgarner",
"authors": ["Jacob R. Bumgarner", "Randy J. Nelson"],
"url": "https://linkinghub.elsevier.com/retrieve/pii/S2667237522000443"},
{"key": "6DU4XPQQ",
"title": "Robust Vessel Segmentation in Fundus Images",
"type": "journalArticle", "year": "2013",
"first_author": "A. Budai",
"authors": ["A. Budai", "R. Bock", "A. Maier", "J. Hornegger", "G. Michelson"],
"url": "http://www.hindawi.com/journals/ijbi/2013/154860/"}]paper_text
使用 PyMuPDF(作为 Python 依赖打包,无需系统工具)提取论文 PDF 的全文。需要 PDF 附件存储在 ~/Zotero/storage/ 中。
参数 | 类型 | 描述 |
|
| Zotero 条目键(在 list_library 输出或 search_papers 结果中显示为 |
返回原始提取文本,以标题、作者和摘要开头:
MMLU-Pro: A More Robust and Challenging
Multi-Task Language Understanding Benchmark
1Yubo Wang∗, 1Xueguang Ma∗, 1Ge Zhang, 1Yuansheng Ni, 1Abhranil Chandra, ...
1University of Waterloo, 2University of Toronto, 3Carnegie Mellon University
Abstract
In the age of large-scale language models, benchmarks like the Massive Multitask
Language Understanding (MMLU) have been pivotal in pushing the boundaries
of what AI can achieve in language comprehension and reasoning across diverse
domains. ...文件结构
src/
main.py # FastMCP server, tool definitions
zotero.py # SQLite queries, data models, formatting数据库直接从 ~/Zotero/zotero.sqlite 读取。PDF 从 ~/Zotero/storage/ 解析。无需 API 密钥。
Available Tools
4 toolslist_libraryA
List collections and papers in the Zotero library. Optionally filter by collection name, paper title, or author. Each paper line includes its item key in [KEY] brackets — use that key in paper_details or paper_text.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries full burden. It discloses that each paper line includes its item key in [KEY] brackets and that this key is used in other tools — a useful behavioral detail. However, it does not mention pagination, sorting, exact output scope, or whether the list includes collections only as names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action, and the second sentence provides crucially useful downstream context about item keys. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (not shown) and a small sibling set, the description is largely complete: it covers the main list functionality, optional filters, and the connection to paper_details/paper_text via the item key. It lacks a few edge details like default sort order, but is sufficient for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a nullable string 'query' with 0% coverage, so the description must compensate. It explains that the parameter filters by collection name, paper title, or author, but is vague about the exact format — whether it's a single free-text search or separate fields. It adds some meaning but not enough for precise invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('List') and resource ('collections and papers in the Zotero library'), and it distinguishes itself from siblings: paper_details and paper_text focus on individual items, while search_papers implies searching rather than listing all.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says the list can be optionally filtered by collection name, paper title, or author, which indicates a browsing/listing use case. It also points to paper_details and paper_text for downstream access, but does not explicitly state when not to use this tool versus search_papers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_detailsA
Get full metadata for a paper by its item key. Obtain the item_key from list_library output (shown as [KEY] before each paper) or from search_papers results.
| Name | Required | Description | Default |
|---|---|---|---|
| item_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavior. It communicates that the operation is a read ('Get') and focuses on metadata, which is the primary behavior. However, it does not disclose potential error conditions, permission requirements, or any side effects, leaving some room for ambiguity in edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two clear sentences. The first sentence immediately states the tool's purpose, and the second provides necessary usage guidance. There is no redundant information or filler, making it exceptionally concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one parameter and an output schema exists, the description covers all necessary context: the purpose, the source of the key, and the fact that the output is full metadata. The output schema handles return value specifics, so the description is complete for this low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must add meaning to the parameter. It does this effectively by explaining what item_key is ('shown as [KEY]') and exactly where to get it from (list_library or search_papers). This is far more informative than the schema's bare 'string' type, fully compensating for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Get full metadata for a paper', clearly specifying the action (get) and the resource (paper metadata). It distinguishes itself from siblings like list_library and paper_text by focusing on metadata retrieval. It also mentions how to obtain the required item_key, reinforcing the tool's specific role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states where the item_key comes from ('Obtain the item_key from list_library output... or from search_papers results'), which gives clear context for when to use this tool. It doesn't explicitly exclude alternatives, but the reference to source tools implies a workflow. No alternatives are named directly, so it's not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_textA
Extract full text from a paper's PDF using pdftotext. Obtain the item_key from list_library output (shown as [KEY] before each paper) or from search_papers results.
| Name | Required | Description | Default |
|---|---|---|---|
| item_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations provided, so the description carries the full burden. It does not disclose side effects, performance implications, error handling, or explicitly confirm it is read-only. Mentioning 'using pdftotext' is an implementation detail, not behavioral transparency. The extract action implies reading, but the description lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the action, and the second provides essential input sourcing. It is well-structured and front-loaded with the primary purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, output schema present), the description covers purpose and parameter sourcing adequately. It does not mention limitations like scanned PDFs or potential errors, but the presence of an output schema likely handles return values. Slight gap in edge-case behaviors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter and 0% schema description coverage, the description fully compensates by explaining item_key as a paper identifier and providing specific instructions on where to find it (list_library output prefixed with [KEY], or search_papers results). This adds meaning beyond the bare string type in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extracts full text from a paper's PDF, using the specific verb 'Extract' and identifying the resource. This distinguishes it from siblings like list_library (listing papers), paper_details (metadata), and search_papers (searching), as it's the only one that retrieves full text content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by instructing how to obtain the required item_key from list_library or search_papers results. This implies the tool is used when full text is needed and even mentions prerequisite tools, though it does not explicitly contrast with alternatives or provide exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersA
Search papers by title or author and return structured results with item keys. Each result includes key, title, type, year, first_author, authors, and url. Use the key in paper_details or paper_text to get full metadata or PDF text.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of transparency. It discloses the result structure and the need for keys, but it does not mention matching behavior (exact vs. fuzzy), pagination, result limits, or error conditions, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action, uses three concise sentences, and includes no redundant details—every sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema covers return values, the description sufficiently covers purpose, result fields, and downstream tool usage. A brief mention of how it differs from list_library would make it more complete, but it is adequate as-is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero coverage for the single 'query' parameter, but the description clarifies that it is a title/author search term, adding semantic meaning beyond the bare schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search papers by title or author') and the specific resource, and differentiates from siblings by mentioning the returned fields and how to use the key with paper_details or paper_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use the tool (searching by title/author) and explains the follow-up usage of keys with related tools, though it does not explicitly mention when not to use it or contrast with list_library.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The tools are mostly distinct: list_library for browsing, search_papers for targeted search, paper_details for metadata, and paper_text for PDF content. However, list_library and search_papers both filter by title/author, which could cause some confusion.
The naming mixes verb_noun patterns (list_library, search_papers) with noun_noun patterns (paper_details, paper_text). While the paper_* prefix helps, the overall convention is not fully consistent.
With only 4 tools, the server is well-scoped for its purpose of accessing a Zotero library. Each tool is essential and there is no unnecessary bloat.
The tool set covers the core read workflow: discovering papers (list/search), retrieving full metadata, and extracting PDF text. Missing write operations and a dedicated collection endpoint are minor gaps that agents can work around.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Remote MCP server for full read/write access to a Zotero library
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
PubMed MCP — wraps the NCBI E-utilities API (biomedical literature, free, no auth)
Multi-engine scholarly research server for search, traversal, full text, and reading lists.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceRead-only MCP server for browsing, searching, and exporting a Zotero library from AI assistants.
- AlicenseAqualityCmaintenanceRead-only MCP server that lets Claude or any MCP client search and retrieve metadata, notes, full text, citations, and BibTeX from your local Zotero library via its built-in API.11MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that connects AI assistants to your Zotero library, enabling full-text PDF extraction and metadata search.MIT
- AlicenseNot gradedqualityBmaintenanceLocal read-only MCP server for Zotero libraries, enabling search, retrieval, and full-text access via the Zotero Web API.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/404Simon/zotero-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server