Zotero MCP
Zotero MCP
Servidor MCP de solo lectura para tu biblioteca local de Zotero. Navega por colecciones, inspecciona metadatos de artículos y extrae el texto completo de los PDF, todo mediante herramientas FastMCP.
Requisitos
Zotero de escritorio sincronizado con
~/Zotero/zotero.sqlite(por defecto en Linux)Python 3.13+
Related MCP server: zotero-mcp
Configuración
Este es un servidor MCP. Lo registras en la configuración MCP de tu agente de codificación, y luego el agente puede usar sus herramientas.
La mayoría de los agentes aceptan una configuración similar. Por ejemplo, en Opencode lo añades al opencode.json:
{
"mcp": {
"zotero-mcp": {
"command": [
"uvx",
"--from",
"git+https://github.com/404Simon/zotero-mcp",
"zotero-mcp"
],
"enabled": true,
"type": "local"
}
}
}Después de esta configuración, el agente descubre las herramientas automáticamente y puedes simplemente preguntarle cosas como:
"¿Qué artículos sobre RAG hay en mi biblioteca y cuál es el resumen del más reciente?"
Herramientas
list_library
Lista todas las colecciones y artículos como un árbol formateado. Opcionalmente filtra por nombre de colección, título de artículo o autor.
Argumento | Tipo | Descripción |
|
| Filtra por nombre de colección, título de artículo o autor (no distingue mayúsculas de minúsculas) |
Cada línea de artículo incluye su clave de elemento entre corchetes [KEY]. Usa esa clave con paper_details o paper_text. La salida filtrada solo muestra el subárbol coincidente, con las colecciones coincidentes y sus ancestros.
├── AI (2 papers)
│ [N3G6XKB9] [preprint] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2021) - Patrick Lewis
│ [PPJJCMXJ] [book] Grundkurs Künstliche Intelligenz: eine praxisorientierte Einführung (2021) - Wolfgang Ertel
├── Bachelorarbeit (42 papers)
│ ├── GraalVM (4 papers)
│ │ [DSUERN67] [book] Supercharge your applications with GraalVM ... (2021) - A. B. Vijay Kumar
│ └── Java Performance (1 papers)
│ [Q3ANMGC3] [conferencePaper] Applying Optimizations for Dynamically-typed Languages to Java (2017) - Matthias Grimmer
│ [MTSF327R] [book] Pro Spring Boot 3: An Authoritative Guide with Best Practices (2024) - Felipe Gutierrez
├── Studienarbeit (49 papers)
│ [T3ZCDWC7] [preprint] MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (2024) - Yubo Wang
├── T3000 (4 papers)
│ [DTZF77Y4] [webpage] Conventional Commits (n.d.) - Unknown
├── TheGreenEpoch (8 papers)
│ [YINRQ63P] [preprint] Distributed LLM Pretraining During Renewable Curtailment Windows (2026) - Philipp Wiesner
└── VesSkel (19 papers)
[ZUXQUGHW] [journalArticle] Open-source analysis and visualization of segmented vasculature datasets with VesselVio (2022) - Jacob R. Bumgarnerpaper_details
Obtén los metadatos completos de un artículo mediante su clave de elemento. Obtén la clave de la salida de list_library (mostrada como [KEY]) o de los resultados de search_papers.
Argumento | Tipo | Descripción |
|
| Clave de elemento de Zotero (mostrada como |
Devuelve título, tipo, clave, fechas de añadido/modificado, autores, todos los metadatos de campos, membresías de colecciones e información del adjunto PDF:
Title: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Type: preprint
Key: T3ZCDWC7
Authors: Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, ...
date: 2024-11-06
DOI: 10.48550/arXiv.2406.01574
url: http://arxiv.org/abs/2406.01574
abstractNote: In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) ...
Collections: Studienarbeit
PDF: Wang et al. - 2024 - MMLU-Pro A More Robust and Challenging Multi-Task Language Understanding Benchmark.pdfsearch_papers
Busca todos los artículos por título o autor. Devuelve resultados estructurados con claves de elemento.
Argumento | Tipo | Descripción |
|
| Término de búsqueda (no distingue mayúsculas de minúsculas, coincide con título y autor) |
[{"key": "ZUXQUGHW",
"title": "Open-source analysis and visualization of segmented vasculature datasets with VesselVio",
"type": "journalArticle", "year": "2022",
"first_author": "Jacob R. Bumgarner",
"authors": ["Jacob R. Bumgarner", "Randy J. Nelson"],
"url": "https://linkinghub.elsevier.com/retrieve/pii/S2667237522000443"},
{"key": "6DU4XPQQ",
"title": "Robust Vessel Segmentation in Fundus Images",
"type": "journalArticle", "year": "2013",
"first_author": "A. Budai",
"authors": ["A. Budai", "R. Bock", "A. Maier", "J. Hornegger", "G. Michelson"],
"url": "http://www.hindawi.com/journals/ijbi/2013/154860/"}]paper_text
Extrae el texto completo del PDF de un artículo usando PyMuPDF (incluido como dependencia de Python, no se necesitan herramientas del sistema). Requiere un adjunto PDF almacenado en ~/Zotero/storage/.
Argumento | Tipo | Descripción |
|
| Clave de elemento de Zotero (mostrada como |
Devuelve el texto extraído en bruto, comenzando con el título, los autores y el resumen:
MMLU-Pro: A More Robust and Challenging
Multi-Task Language Understanding Benchmark
1Yubo Wang∗, 1Xueguang Ma∗, 1Ge Zhang, 1Yuansheng Ni, 1Abhranil Chandra, ...
1University of Waterloo, 2University of Toronto, 3Carnegie Mellon University
Abstract
In the age of large-scale language models, benchmarks like the Massive Multitask
Language Understanding (MMLU) have been pivotal in pushing the boundaries
of what AI can achieve in language comprehension and reasoning across diverse
domains. ...Estructura de archivos
src/
main.py # FastMCP server, tool definitions
zotero.py # SQLite queries, data models, formattingLa base de datos se lee directamente de ~/Zotero/zotero.sqlite. Los PDF se resuelven desde ~/Zotero/storage/. No se necesita clave API.
Available Tools
4 toolslist_libraryA
List collections and papers in the Zotero library. Optionally filter by collection name, paper title, or author. Each paper line includes its item key in [KEY] brackets — use that key in paper_details or paper_text.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries full burden. It discloses that each paper line includes its item key in [KEY] brackets and that this key is used in other tools — a useful behavioral detail. However, it does not mention pagination, sorting, exact output scope, or whether the list includes collections only as names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action, and the second sentence provides crucially useful downstream context about item keys. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (not shown) and a small sibling set, the description is largely complete: it covers the main list functionality, optional filters, and the connection to paper_details/paper_text via the item key. It lacks a few edge details like default sort order, but is sufficient for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a nullable string 'query' with 0% coverage, so the description must compensate. It explains that the parameter filters by collection name, paper title, or author, but is vague about the exact format — whether it's a single free-text search or separate fields. It adds some meaning but not enough for precise invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('List') and resource ('collections and papers in the Zotero library'), and it distinguishes itself from siblings: paper_details and paper_text focus on individual items, while search_papers implies searching rather than listing all.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says the list can be optionally filtered by collection name, paper title, or author, which indicates a browsing/listing use case. It also points to paper_details and paper_text for downstream access, but does not explicitly state when not to use this tool versus search_papers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_detailsA
Get full metadata for a paper by its item key. Obtain the item_key from list_library output (shown as [KEY] before each paper) or from search_papers results.
| Name | Required | Description | Default |
|---|---|---|---|
| item_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavior. It communicates that the operation is a read ('Get') and focuses on metadata, which is the primary behavior. However, it does not disclose potential error conditions, permission requirements, or any side effects, leaving some room for ambiguity in edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two clear sentences. The first sentence immediately states the tool's purpose, and the second provides necessary usage guidance. There is no redundant information or filler, making it exceptionally concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one parameter and an output schema exists, the description covers all necessary context: the purpose, the source of the key, and the fact that the output is full metadata. The output schema handles return value specifics, so the description is complete for this low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must add meaning to the parameter. It does this effectively by explaining what item_key is ('shown as [KEY]') and exactly where to get it from (list_library or search_papers). This is far more informative than the schema's bare 'string' type, fully compensating for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Get full metadata for a paper', clearly specifying the action (get) and the resource (paper metadata). It distinguishes itself from siblings like list_library and paper_text by focusing on metadata retrieval. It also mentions how to obtain the required item_key, reinforcing the tool's specific role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states where the item_key comes from ('Obtain the item_key from list_library output... or from search_papers results'), which gives clear context for when to use this tool. It doesn't explicitly exclude alternatives, but the reference to source tools implies a workflow. No alternatives are named directly, so it's not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_textA
Extract full text from a paper's PDF using pdftotext. Obtain the item_key from list_library output (shown as [KEY] before each paper) or from search_papers results.
| Name | Required | Description | Default |
|---|---|---|---|
| item_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations provided, so the description carries the full burden. It does not disclose side effects, performance implications, error handling, or explicitly confirm it is read-only. Mentioning 'using pdftotext' is an implementation detail, not behavioral transparency. The extract action implies reading, but the description lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the action, and the second provides essential input sourcing. It is well-structured and front-loaded with the primary purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, output schema present), the description covers purpose and parameter sourcing adequately. It does not mention limitations like scanned PDFs or potential errors, but the presence of an output schema likely handles return values. Slight gap in edge-case behaviors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter and 0% schema description coverage, the description fully compensates by explaining item_key as a paper identifier and providing specific instructions on where to find it (list_library output prefixed with [KEY], or search_papers results). This adds meaning beyond the bare string type in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extracts full text from a paper's PDF, using the specific verb 'Extract' and identifying the resource. This distinguishes it from siblings like list_library (listing papers), paper_details (metadata), and search_papers (searching), as it's the only one that retrieves full text content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by instructing how to obtain the required item_key from list_library or search_papers results. This implies the tool is used when full text is needed and even mentions prerequisite tools, though it does not explicitly contrast with alternatives or provide exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersA
Search papers by title or author and return structured results with item keys. Each result includes key, title, type, year, first_author, authors, and url. Use the key in paper_details or paper_text to get full metadata or PDF text.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of transparency. It discloses the result structure and the need for keys, but it does not mention matching behavior (exact vs. fuzzy), pagination, result limits, or error conditions, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action, uses three concise sentences, and includes no redundant details—every sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema covers return values, the description sufficiently covers purpose, result fields, and downstream tool usage. A brief mention of how it differs from list_library would make it more complete, but it is adequate as-is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero coverage for the single 'query' parameter, but the description clarifies that it is a title/author search term, adding semantic meaning beyond the bare schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search papers by title or author') and the specific resource, and differentiates from siblings by mentioning the returned fields and how to use the key with paper_details or paper_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use the tool (searching by title/author) and explains the follow-up usage of keys with related tools, though it does not explicitly mention when not to use it or contrast with list_library.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The tools are mostly distinct: list_library for browsing, search_papers for targeted search, paper_details for metadata, and paper_text for PDF content. However, list_library and search_papers both filter by title/author, which could cause some confusion.
The naming mixes verb_noun patterns (list_library, search_papers) with noun_noun patterns (paper_details, paper_text). While the paper_* prefix helps, the overall convention is not fully consistent.
With only 4 tools, the server is well-scoped for its purpose of accessing a Zotero library. Each tool is essential and there is no unnecessary bloat.
The tool set covers the core read workflow: discovering papers (list/search), retrieving full metadata, and extracting PDF text. Missing write operations and a dedicated collection endpoint are minor gaps that agents can work around.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Remote MCP server for full read/write access to a Zotero library
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
PubMed MCP — wraps the NCBI E-utilities API (biomedical literature, free, no auth)
Multi-engine scholarly research server for search, traversal, full text, and reading lists.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceRead-only MCP server for browsing, searching, and exporting a Zotero library from AI assistants.
- AlicenseAqualityCmaintenanceRead-only MCP server that lets Claude or any MCP client search and retrieve metadata, notes, full text, citations, and BibTeX from your local Zotero library via its built-in API.11MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that connects AI assistants to your Zotero library, enabling full-text PDF extraction and metadata search.MIT
- AlicenseNot gradedqualityBmaintenanceLocal read-only MCP server for Zotero libraries, enabling search, retrieval, and full-text access via the Zotero Web API.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/404Simon/zotero-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server