Skip to main content
Glama

zotero-mcp

A Model Context Protocol server for Zotero. It lets Claude Code, Claude Desktop and any other MCP client search your library, read items, take notes, file things into collections and find papers by meaning rather than by exact words.

This is a ground-up rewrite of the earlier zotero-mcp server with a small, typed tool surface, bounded responses, real error reporting and a semantic search index that is chunk-level, incremental and reranked.

What it does

Read

Tool

Purpose

search_library

Keyword search across the library, paged, with type filter

get_item

One item's full record, children, tags and collections; flags items in the Trash

list_collections / collection_items

Browse the collection tree and its contents

list_tags / library_stats

Tags and a library overview

server_health

Is Zotero reachable, is the index built, which optional features are on

Find by meaning

Tool

Purpose

semantic_search

Vector search over titles, abstracts and PDF fulltext, reranked by a cross-encoder. Each hit shows the passage that matched, whether it came from the metadata or the fulltext, and how many passages agreed. Filter by item type, year range or collection.

Write (only when a web API key with write permission is configured)

Tool

Purpose

create_item / update_item / delete_item

Create, edit and trash items. Delete is always "move to Trash", never permanent.

modify_tags / modify_collections

Add or remove tags and collection memberships

create_note

Attach a note to an item, or create a standalone one

Every write tool defaults to dry_run=true and returns a preview of the change. The client must pass dry_run=false to apply it. Updates carry Zotero's version header, so a concurrent edit in the desktop app produces a clear conflict error rather than a silent overwrite.

Related MCP server: zotero-mcp

Install

Requires Python 3.10+ and uv.

uv tool install 'zotero-mcp-next[semantic,pdf] @ git+https://github.com/v-i-n-a-y/zotero-mcp-v2'

The semantic extra pulls in ChromaDB and sentence-transformers for semantic search; pdf adds PyMuPDF and ebooklib for fulltext extraction. Omit both for a lean install with keyword search only.

Configure

The server reads environment variables (or a JSON file via ZOTERO_MCP_CONFIG).

Variable

Meaning

ZOTERO_MCP_MODE

local, web, hybrid or auto (default). See below.

ZOTERO_API_KEY

Web API key from https://www.zotero.org/settings/keys

ZOTERO_LIBRARY_ID

Your numeric user ID (shown on the same page)

ZOTERO_LIBRARY_TYPE

user (default) or group

ZOTERO_MCP_INDEX_SCHEDULE

daily (default), weekly, startup or manual

ZOTERO_MCP_RERANK

false to skip cross-encoder reranking

ZOTERO_MCP_EMBEDDING_MODEL

Sentence-transformers model; default BAAI/bge-small-en-v1.5

ZOTERO_MCP_DB_PATH

Where the index lives; default ~/.config/zotero-mcp/chroma

Modes. local talks to the running Zotero desktop app (port 23119): fast, no key needed, read-only. web talks to api.zotero.org and needs a key. hybrid reads locally and writes through the web API, which is the best everyday setting when the desktop app is open. auto picks local if the app is reachable, otherwise web.

Claude Code

claude mcp add zotero -s user \
  -e ZOTERO_MCP_MODE=hybrid \
  -e ZOTERO_API_KEY=... \
  -e ZOTERO_LIBRARY_ID=... \
  -- zotero-mcp

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "zotero": {
      "command": "zotero-mcp",
      "env": {
        "ZOTERO_MCP_MODE": "hybrid",
        "ZOTERO_API_KEY": "...",
        "ZOTERO_LIBRARY_ID": "..."
      }
    }
  }
}

Build the index once from a terminal; it needs the desktop app open (or web mode) to read fulltext:

zotero-mcp index build     # first run: roughly an hour per thousand PDFs
zotero-mcp index status

After that the running server refreshes the index on the schedule above, re-embedding only items whose Zotero version changed. Extracted fulltext is cached beside the index, so switching embedding models rebuilds in minutes.

How a query is answered: the vector index returns a pool of candidate chunks (metadata chunk plus fulltext chunks per item, with trailing reference lists stripped), a cross-encoder rescores the pool against the query, citation-dense chunks are penalised, and chunks collapse to one hit per item with a small bonus for items that matched in several places. Everything runs locally; on Apple silicon the models use Metal.

CLI

zotero-mcp                 # serve over stdio (what MCP clients run)
zotero-mcp health          # can the configured library be reached?
zotero-mcp index build     # build or incrementally refresh the semantic index
zotero-mcp index status
zotero-mcp index clear
zotero-mcp version

Design

  • Every response is bounded. Lists are paged with opaque cursors; long text is truncated with an explicit marker.

  • Failures raise MCP tool errors with a stable code (not_found, invalid_input, auth, write_conflict, backend_unavailable, unsupported) and a hint, never a success whose text starts with "Error:".

  • Tools return both markdown for humans and structured content for clients that can use it.

  • Logging goes to stderr; stdout carries only the protocol.

More in docs/design.md.

Development

uv sync --extra dev
uv run pytest
uv run ruff check src tests

Licence

MIT.

Available Tools

8 tools
collection_itemsCollection ItemsA

List the top-level items in one collection.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items per page.
cursorNoThe 'next_cursor' from a previous call.
collection_keyYesThe collection key from `list_collections`.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It communicates that the tool is a read-only list operation and that it only returns top-level items, which is useful. It does not disclose pagination behavior, ordering, or whether the result contains full item details, but the schema partially covers cursor and limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core action and scope. There is no redundant wording, and it earns its place without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, complete parameter schema, and no output schema, the description provides adequate context for basic invocation. It could be slightly stronger by noting pagination expectations or what 'top-level' excludes, but nothing essential is missing for a straightforward listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter meanings are already documented: collection_key identifies the collection, limit controls page size, and cursor continues pagination. The description adds little beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'List', and a specific resource, 'top-level items in one collection,' so its purpose is clear. It does not explicitly contrast itself with sibling tools like search_library or get_item, but the phrasing is precise enough to avoid serious ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when you want the top-level contents of a specific collection, and use the collection_key from list_collections as the schema notes. However, there is no explicit guidance on when to prefer this over search_library, semantic_search, or get_item.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_itemGet ItemA

Fetch the full record for one item, including its attachments and note count.

ParametersJSON Schema
NameRequiredDescriptionDefault
item_keyYesThe 8-character Zotero item key from a search result.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral burden. 'Fetch' clearly indicates a read operation, and the mention of attachments and note count reveals what the response includes. It doesn't cover error cases or auth, but for a simple single-item lookup the key behavioral traits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence that leads with the action and resource, then adds useful output detail. Every word contributes meaning; no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only lookup tool, the description combined with the schema covers what the tool does, what input it needs, and what the response contains (full record, attachments, note count). There is no output schema, but the description gives enough context for correct invocation. Minor gaps like not-found behavior are acceptable given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with item_key already described as an 8-character Zotero item key from a search result. The description adds no further parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Fetch') and resource ('full record for one item'), and adds concrete return details (attachments and note count). This clearly distinguishes get_item from siblings like collection_items or search_library, which operate on sets of items.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes clear context: use this tool when you need the full details of a single item identified by an item_key. It doesn't explicitly name alternatives or exclusions, but the purpose is specific enough that an agent can infer when it applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_statsLibrary StatsA

Report a quantitative overview of the active library.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations to set a baseline, so the description must carry the full burden. It states the tool reports a 'quantitative overview' but does not explicitly confirm it is read-only, what side effects (if any) might occur, or what specific data points it returns. The ambiguity about its operational behavior limits transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence with no extraneous words. It efficiently communicates the core purpose without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema or parameters, the description provides sufficient context for a simple stats tool. The phrase 'quantitative overview' is somewhat generic but adequately conveys that the tool returns summary counts or metrics about the active library, which is enough for typical usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline score is 4. The description does not need to explain parameter semantics, and no parameter-related information is missing or misleading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('Report') and a specific object ('quantitative overview of the active library'), which distinguishes it from sibling tools like get_item, search_library, and list_collections. It unambiguously conveys that this tool provides aggregate statistics about the library.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives, nor does it mention any conditions or prerequisites. While the name and object imply a stats-gathering role, the absence of any guidance on selecting it over list_collections or search_library leaves the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_collectionsList CollectionsA

List the library's collections as slash-joined paths (Parent/Child).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax collections per page.
cursorNoThe 'next_cursor' from a previous call.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It communicates a read-style listing operation and the return format ('slash-joined paths (Parent/Child)'), which adds hierarchy semantics not present in the schema. It does not explicitly say the operation is non-destructive or describe pagination behavior, though the cursor parameter hints at it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One clean, front-loaded sentence states the action, resource, and output shape. There is no fluff, and the useful detail about slash-joined paths earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple: no required parameters, no nested objects, and the schema covers the pagination mechanics. The description covers what is returned and its format. A minor gap is the lack of explicit alternatives or when-not-to-use guidance, but this does not materially hurt invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters already have clear descriptions: 'limit' is max collections per page and 'cursor' is the next_cursor from a previous call. The tool description adds no parameter-level meaning but does not need to, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('List') and a clear resource ('the library's collections'), and it adds the output form ('slash-joined paths (Parent/Child)'). This differentiates list_collections from siblings like list_tags or collection_items even without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: retrieve the library's collections. However, it gives no explicit guidance about when to choose this tool over alternatives or when not to use it. Sibling names provide context, but the description itself does not route the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tagsList TagsC

List tags used in the library.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax tags to return.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description must carry the burden of behavioral disclosure. It only says tags are listed; it does not mention ordering, deduplication, whether counts are returned, pagination behavior, or any side effects. This is minimal transparency for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence with no filler or repetition. It is front-loaded with the core purpose, though it could sacrifice some brevity for additional behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with one optional parameter, this is minimally viable: an agent knows what endpoint to call and that limit is available. However, with no output schema and no annotations, details like result format, tag count semantics, and ordering are missing, leaving small but real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of the single parameter (limit) with a clear description, so the baseline is 3. The tool description adds no parameter-specific information, but the schema already provides adequate meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List') and the resource ('tags used in the library'), making it distinguishable from sibling tools like list_collections. It is not a tautology and communicates the core function effectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as list_collections or search_library. The context signals imply a library domain, but the description does not state exclusions or preferred scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_librarySearch LibraryA

Search the active Zotero library and return a page of matching items.

Each result carries the item key you pass to get_item, plus author, year, and a short abstract preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return this page.
queryNoWords to match anywhere in the item (title, author, abstract, tags, fulltext). Empty lists recent items.
cursorNoThe 'next_cursor' from a previous call, to fetch the next page. Keep every other argument identical.
item_typeNoRestrict to one Zotero item type, e.g. 'journalArticle', 'book'.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It usefully reveals that results are paged, that the search applies to the active library, and that each result contains a key, author, year, and abstract preview. It does not cover pagination mechanics, but the schema already documents the cursor parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences: the first states the primary action and result shape, the second explains how the result connects to get_item. There is no filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with the fully documented schema, gives an agent enough to invoke the tool and understand the basic output. It explains the output fields, active-library scope, and follow-up path, though sort order and semantic-search alternatives are not discussed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description adds no parameter-level detail beyond general search and paging behavior, which matches the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: search the active Zotero library and return matching items. It is clear and actionable, though it does not explicitly differentiate itself from the sibling semantic_search tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a search-then-fetch workflow by telling the caller to pass the returned item key to get_item. However, it does not state when to prefer this over semantic_search or other listing tools, nor does it provide explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

server_healthServer HealthA

Check that the server can actually reach the configured Zotero library.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full burden of transparency. It indicates a read-only check, but does not disclose what happens on success or failure, nor any side effects or return format. The behavior is partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence with no redundant words. It is concise and directly conveys the tool's purpose without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a simple health check, but it does not specify what the tool returns (e.g., boolean, status code) or what constitutes a successful check. This omission is minor given the straightforward nature of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so there is nothing to add beyond the empty schema. The description does not need to explain any inputs, making this dimension fully satisfied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks server connectivity to the Zotero library. It uses the specific verb 'check' and the resource 'server... Zotero library', distinguishing it from the sibling tools which perform retrieval or list operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose implies usage when needing to verify connectivity, though it doesn't explicitly state when to use it versus other tools. Since no alternative is directly applicable for health checks, the context is clear but could benefit from explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedcollection_items
    • First observedget_item
    • First observedlibrary_stats
    • First observedlist_collections
    • First observedlist_tags
    • First observedsearch_library
    • First observedsemantic_search
    • First observedserver_health

TDQS

A3.7/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: single-item retrieval, keyword search, semantic search, collection browsing, tag listing, stats, and health check. Even search_library and semantic_search are explicitly differentiated as exact-word vs. meaning-based complement.

Naming Consistency3/5

Naming mixes verb-prefixed tools (get_item, search_library, list_collections, list_tags) with noun-phrase tools (collection_items, library_stats, server_health, semantic_search). The pattern is readable but not consistently applied.

Tool Count5/5

Eight tools is well-scoped for a Zotero library exploration and search server. Every tool addresses a necessary capability without redundancy or bloat.

Completeness4/5

The tool set covers the core read/query workflow: item lookup, two search modes, collection/tag navigation, stats, and diagnostics. The only notable gap is collection-specific detail beyond top-level items, but within the apparent read-only scope this is minor.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    MCP server for interacting with a Zotero library via the local API. Enables searching, retrieving, creating, updating, and deleting Zotero items, managing collections and tags, and generating citations.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that gives AI agents read access to a Zotero library, enabling search, item metadata, collections, tags, notes, and full-text search of attached PDFs.
    MIT