Literature Evidence MCP
Reads live Zotero metadata from a configured collection and its descendants, providing collection/item discovery and metadata lookup for literature evidence.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Literature Evidence MCPSearch my papers for 'deep learning' and show the top 3 evidence excerpts with page numbers."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Literature Evidence MCP
Small, versioned evidence from your local papers—without loading the whole library into an AI conversation. Optional modules connect Zotero, writing projects, fixed-commit source code and experiment records.
中文说明 · Configuration · Tools · Release guide
Status: 0.5.0a1, an unreleased alpha candidate. Python 3.11+, macOS and Linux. Windows is not supported in this release (POSIX file locking/process management). No hosted account or API key is needed for local file access.
What it does
Search local names and cached full text; retrieve text with PDF page numbers and SHA-256 hashes. Use original page images for equations and figures with Poppler.
Return compact indexes and bounded excerpts: 4,000 text characters by default, at most 12,000 per read. Explicit continuations preserve the evidence version.
Read live Zotero metadata in one configured collection and its descendants.
Optionally read allowlisted writing files, save agreed review notes, inspect code at a fixed commit and exchange experiment plans/results.
The server does not modify papers, Zotero data, manuscript text or source code. MCP calls never launch an experiment. Enabled write tools save notes/plans only. Search is lexical, PDFs have no OCR, and conversation history is managed by your client. A cache avoids repeated extraction, not accumulated tool-result history.
Related MCP server: Keepygaga RAG
Try the bundled example
git clone https://github.com/phy-zhangzl/literature-evidence-mcp.git
cd literature-evidence-mcpFrom this source checkout, with uv installed:
uv sync --locked
uv run --locked literature-mcp --config config.example.json doctor
uv run --locked literature-mcp --config config.example.json serveserve defaults to stdio and waits for an MCP client. The example uses only the
synthetic document in examples/papers and writes its cache to .state/cache.
It does not use a personal Zotero library, a tunnel or an existing service config.
Install and configure
Build a wheel from a reviewed source checkout, or use a wheel attached to a future GitHub release. This candidate has not been uploaded to PyPI; do not assume the package name on a registry belongs to this project.
uv build
python3 -m venv .venv-runtime
.venv-runtime/bin/python -m pip install dist/local_literature_mcp-0.5.0a1-py3-none-any.whl
mkdir -p "$HOME/Papers"
.venv-runtime/bin/literature-mcp init --papers "$HOME/Papers"
.venv-runtime/bin/literature-mcp doctorinit creates ~/.config/literature-mcp/config.json and never overwrites an existing
file. XDG_CONFIG_HOME, LITERATURE_MCP_CONFIG or the global --config argument
can select another location. Relative paths are relative to the config file.
Use a separate config, cache, experiment store and runtime environment for each
installation. Migration from the local predecessor.
Configure your MCP client with this stdio server entry, substituting absolute paths for your runtime and config:
{
"mcpServers": {
"literature-product": {
"command": "/absolute/path/to/.venv-runtime/bin/literature-mcp",
"args": ["--config", "/absolute/path/to/config.json", "serve"]
}
}
}Client configuration formats vary; this is a common stdio entry, not an automatic
client installer. Optional local HTTP: add --transport streamable-http after
serve. It listens on 127.0.0.1, at /mcp, with the configured port. It has no
public authentication layer; do not expose it directly to the Internet.
Optional capabilities
Configuration | Tools made available |
Only | Library status, file listing, search, fetch, page images |
| Collection/item discovery and metadata sections |
| Writing context and append-only feedback |
| Fixed-commit code listing, search, reads and version context |
| Experiment index, artifact discovery and evidence reads |
Code repository + experiment store | Save agreed experiment plans |
Poppler is optional for text and required for page images. Install with
brew install poppler on macOS or your distribution's poppler-utils package on
Linux. Git is needed only for code tools. Zotero requires its desktop local API to
be enabled. See configuration and data boundaries.
The optional foamCase adapter is for existing scientific workflows. File-only users need neither it nor a simulation environment.
Development and release
uv run --locked pytest -q
uv run --locked python scripts/smoke_test.py
uv run --locked python scripts/smoke_test.py --transport http
uv build
uv run --locked python scripts/check_release.pyTests use synthetic fixtures and temporary stores. See CONTRIBUTING.md. CI checks macOS/Linux with Python 3.11–3.13; a configured workflow is not a claim that remote CI has already run. The tag workflow creates a draft GitHub prerelease; it does not publish to PyPI. See release guide.
License
MIT. Papers and research data remain under their own licenses and are not included in this repository.
Available Tools
5 toolsfetchARead-onlyIdempotent
Read a source excerpt by file:, zotero:, project: or feedback: ID. Defaults to one PDF page and 4000 text characters; max_chars is 100..12000. Select at most 50 pages. Continue only as needed, retaining page range and expected_sha256. Use get_writing_context for project IDs. Parent Zotero items auto-select only when one readable PDF exists.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| offset | No | ||
| end_page | No | ||
| max_chars | No | ||
| start_page | No | ||
| expected_sha256 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds substantial behavior beyond that: default page count, default character count, max_chars bounds, the 50-page limit, continuation guidance, and the auto-selection caveat for parent Zotero items. This is rich, useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each adding distinct value: purpose, defaults and limits, continuation guidance, and an important caveat. The most identifying information is front-loaded and there is no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values do not need explanation. The description covers core behavior, constraints, and an important sibling alternative. The only real gap is the underspecified offset and expected_sha256 semantics, but defaults make the tool callable without those details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify the ID scheme, max_chars range and default, page-selection limits, and the importance of retaining expected_sha256. However, offset and the exact meaning of expected_sha256 are not fully explained, so it is strong but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Read a source excerpt by file:, zotero:, project: or feedback: ID.' This clearly distinguishes it from siblings like list_files, search, and read_page_image, and names the ID formats it accepts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing guidance: 'Use get_writing_context for project IDs.' It also provides operational guidance about continuing only as needed and retaining page range and expected_sha256. It does not explicitly state when to prefer fetch over siblings, but the purpose distinction is strong enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
library_statusARead-onlyIdempotent
Check library health, PDF/full-text coverage and enabled capabilities. Does not load writing/research context or experiment records.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive; the description adds value beyond that by specifying what the tool actually inspects and by disclaiming that it will not load writing/research context or experiment records. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences with no filler. The first states what it does; the second prevents a common misuse. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with an output schema and strong annotations, the description covers purpose, scope, and an important exclusion. No additional information is necessary to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema has 100% coverage (empty properties), so there is nothing for the description to clarify. The baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('Check') and names the exact scope: library health, PDF/full-text coverage, and enabled capabilities. It also draws a boundary by saying it does not load writing/research context or experiment records, making the purpose distinct from siblings like fetch and search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it for health/coverage/capability checks. The explicit 'Does not load...' sentence provides a when-not condition, but no sibling alternative is named, so it stops short of a full routing guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_filesARead-onlyIdempotent
Browse the local literature folder. folder is relative; results contain file: IDs for fetch. Hidden files and paths outside the configured folder are excluded.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| folder | No | ||
| offset | No | ||
| recursive | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds non-obvious context: folder paths are relative, hidden files are excluded, paths outside the configured folder are excluded, and results contain file: IDs for downstream fetch. This is genuinely useful behavioral disclosure beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary purpose is front-loaded, and the secondary details about relative paths, exclusions, and file IDs are compact and meaningful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description communicates the core browsing behavior and folder constraints, and an output schema exists to document return values. However, it leaves limit, offset, and recursive semantics undocumented, and it lacks explicit guidance for choosing between list_files and the search/fetch siblings. These are meaningful completeness gaps for an agent calling the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only clarifies the folder parameter ('folder is relative'), while limit, offset, and recursive receive no explanation. For a tool with four parameters but zero schema descriptions, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: 'Browse the local literature folder.' It clarifies folder scope, notes that results contain file: IDs for fetch, and excludes hidden files and paths outside the configured folder. This makes the tool's purpose distinct from siblings like search and fetch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when browsing the local literature folder, and it tells agents that file IDs in results are intended for fetch. However, it does not explicitly contrast with search or other siblings, nor state when not to use list_files.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_page_imageARead-onlyIdempotent
Inspect an original PDF page as an image, especially for plots, tables, equations or scanned pages. page is 1-based. Requires locally installed Poppler.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| page | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only, idempotent, and non-destructive, so the description only needs to add operational context. It adds the Poppler prerequisite and notes that page numbering is 1-based, which is useful, but it does not describe the return format or failure behavior. This matches the lower bar for annotation-backed tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. The core purpose is front-loaded, and the indexing detail and prerequisite are separated cleanly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool the description is mostly sufficient, and the Poppler prerequisite is helpful. However, with no output schema and an unexplained 'id' parameter, the agent is left to infer where the id comes from and what form the returned image takes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It clarifies that 'page' is 1-based, but the required 'id' parameter is never explained, leaving ambiguity about what identifier should be passed and where it comes from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Inspect an original PDF page as an image') and a clear resource, with concrete use cases like plots, tables, equations, and scanned pages. It is clear about what the tool does, but it does not explicitly differentiate from siblings such as fetch or search, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'especially for plots, tables, equations or scanned pages' gives a clear context for when this tool is appropriate, implying it is for visual content where text-based extraction may fail. It does not state exclusions or explicitly name alternatives, but the guidance is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchARead-onlyIdempotent
Search live Zotero metadata, local file names and cached PDF/text content. Returns IDs for fetch, snippets and indexing coverage. Uncached PDFs are not full-text searched; query terms are matched lexically on the same page.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds meaningful behavioral limits beyond those annotations: uncached PDFs are not full-text searched, and query terms are matched lexically on the same page. This is valuable operational context for setting agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The core scope and return value are front-loaded, and the important limitation about uncached PDFs is placed at the end without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the search surface, return contents, and a key limitation, and an output schema exists to document return structure. However, it does not explain parameter behavior or provide guidance for choosing among siblings, leaving the definition slightly incomplete for a tool with low schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the two parameters, but it does not explain query or limit semantics. The phrase 'query terms' hints at the query parameter, and limit likely controls result count, but neither is explicitly defined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search') and distinct resources: live Zotero metadata, local file names, and cached PDF/text content. It also differentiates from siblings by noting it returns IDs for fetch, snippets, and indexing coverage, which separates it from list_files, fetch, and read_page_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use clear: search across metadata, file names, and cached content, returning IDs that could be passed to fetch. It does not explicitly name alternatives or state when not to use it, so it stops short of a 5, but the context is clear enough for an agent to select it sensibly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
fetch - First observed
library_status - First observed
list_files - First observed
read_page_image - First observed
search
TDQS
Scored across 5 tools
Each tool maps to a distinct function: health/status, file enumeration, search across metadata and full text, targeted excerpt retrieval, and visual page inspection. The only close pair, fetch and read_page_image, is clearly differentiated as text versus image access.
Names are consistently lowercase snake_case but mix conventions: list_files and read_page_image follow verb_noun, while library_status is noun_noun and search/fetch are bare verbs. The set remains readable but does not follow one predictable naming pattern.
Five tools is appropriately sized for a literature retrieval server; each covers a distinct need without redundancy. The count is well within the ideal 3-15 range and feels neither sparse nor bloated.
The surface covers the core evidence workflow: locate sources via search/list, retrieve text via fetch, and inspect visual content via read_page_image. Minor gaps exist, such as no dedicated full-metadata display or comparison tool, but these are workable.
Maintenance
Related MCP Connectors
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Cited, versioned knowledge for agents: retrieve sourced passages and propose owner-approved fixes.
Retrieve citation-ready technical context and coordinate evidence-backed work between AI agents.
Related MCP Servers
- AlicenseBqualityBmaintenanceEnables local agents to search and retrieve cited evidence from PDFs and Markdown notes, including page-specific passages and rendered page images.6GPL 3.0
- AlicenseAqualityCmaintenanceEnables local-first hybrid knowledge retrieval from authorized Markdown and plain-text files, combining full-text and vector search with reranking and traceable source references via a single search tool.1MIT
- AlicenseAqualityBmaintenanceEnables AI writing assistants to retrieve citable evidence from local PDFs and verify draft citations against their sources locally, providing verifiable support for claims.21MIT
- FlicenseBqualityCmaintenanceEnables evidence-first analysis of research papers and slide decks by converting local PDFs, DOCX, PPT, and PPTX files into a cached, token-efficient index for batched searching and selective visual verification.42-