Skip to main content
Glama
phy-zhangzl

Literature Evidence MCP

by phy-zhangzl

Literature Evidence MCP

Small, versioned evidence from your local papers—without loading the whole library into an AI conversation. Optional modules connect Zotero, writing projects, fixed-commit source code and experiment records.

中文说明 · Configuration · Tools · Release guide

Status: 0.5.0a1, an unreleased alpha candidate. Python 3.11+, macOS and Linux. Windows is not supported in this release (POSIX file locking/process management). No hosted account or API key is needed for local file access.

What it does

  • Search local names and cached full text; retrieve text with PDF page numbers and SHA-256 hashes. Use original page images for equations and figures with Poppler.

  • Return compact indexes and bounded excerpts: 4,000 text characters by default, at most 12,000 per read. Explicit continuations preserve the evidence version.

  • Read live Zotero metadata in one configured collection and its descendants.

  • Optionally read allowlisted writing files, save agreed review notes, inspect code at a fixed commit and exchange experiment plans/results.

The server does not modify papers, Zotero data, manuscript text or source code. MCP calls never launch an experiment. Enabled write tools save notes/plans only. Search is lexical, PDFs have no OCR, and conversation history is managed by your client. A cache avoids repeated extraction, not accumulated tool-result history.

Related MCP server: Keepygaga RAG

Try the bundled example

git clone https://github.com/phy-zhangzl/literature-evidence-mcp.git
cd literature-evidence-mcp

From this source checkout, with uv installed:

uv sync --locked
uv run --locked literature-mcp --config config.example.json doctor
uv run --locked literature-mcp --config config.example.json serve

serve defaults to stdio and waits for an MCP client. The example uses only the synthetic document in examples/papers and writes its cache to .state/cache. It does not use a personal Zotero library, a tunnel or an existing service config.

Install and configure

Build a wheel from a reviewed source checkout, or use a wheel attached to a future GitHub release. This candidate has not been uploaded to PyPI; do not assume the package name on a registry belongs to this project.

uv build
python3 -m venv .venv-runtime
.venv-runtime/bin/python -m pip install dist/local_literature_mcp-0.5.0a1-py3-none-any.whl
mkdir -p "$HOME/Papers"
.venv-runtime/bin/literature-mcp init --papers "$HOME/Papers"
.venv-runtime/bin/literature-mcp doctor

init creates ~/.config/literature-mcp/config.json and never overwrites an existing file. XDG_CONFIG_HOME, LITERATURE_MCP_CONFIG or the global --config argument can select another location. Relative paths are relative to the config file. Use a separate config, cache, experiment store and runtime environment for each installation. Migration from the local predecessor.

Configure your MCP client with this stdio server entry, substituting absolute paths for your runtime and config:

{
  "mcpServers": {
    "literature-product": {
      "command": "/absolute/path/to/.venv-runtime/bin/literature-mcp",
      "args": ["--config", "/absolute/path/to/config.json", "serve"]
    }
  }
}

Client configuration formats vary; this is a common stdio entry, not an automatic client installer. Optional local HTTP: add --transport streamable-http after serve. It listens on 127.0.0.1, at /mcp, with the configured port. It has no public authentication layer; do not expose it directly to the Internet.

Optional capabilities

Configuration

Tools made available

Only papers_dir

Library status, file listing, search, fetch, page images

zotero_collection

Collection/item discovery and metadata sections

project_dir

Writing context and append-only feedback

code_repository

Fixed-commit code listing, search, reads and version context

experiments_dir

Experiment index, artifact discovery and evidence reads

Code repository + experiment store

Save agreed experiment plans

Poppler is optional for text and required for page images. Install with brew install poppler on macOS or your distribution's poppler-utils package on Linux. Git is needed only for code tools. Zotero requires its desktop local API to be enabled. See configuration and data boundaries.

The optional foamCase adapter is for existing scientific workflows. File-only users need neither it nor a simulation environment.

Development and release

uv run --locked pytest -q
uv run --locked python scripts/smoke_test.py
uv run --locked python scripts/smoke_test.py --transport http
uv build
uv run --locked python scripts/check_release.py

Tests use synthetic fixtures and temporary stores. See CONTRIBUTING.md. CI checks macOS/Linux with Python 3.11–3.13; a configured workflow is not a claim that remote CI has already run. The tag workflow creates a draft GitHub prerelease; it does not publish to PyPI. See release guide.

License

MIT. Papers and research data remain under their own licenses and are not included in this repository.

Available Tools

5 tools
fetchA
Read-onlyIdempotent

Read a source excerpt by file:, zotero:, project: or feedback: ID. Defaults to one PDF page and 4000 text characters; max_chars is 100..12000. Select at most 50 pages. Continue only as needed, retaining page range and expected_sha256. Use get_writing_context for project IDs. Parent Zotero items auto-select only when one readable PDF exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
offsetNo
end_pageNo
max_charsNo
start_pageNo
expected_sha256No

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds substantial behavior beyond that: default page count, default character count, max_chars bounds, the 50-page limit, continuation guidance, and the auto-selection caveat for parent Zotero items. This is rich, useful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each adding distinct value: purpose, defaults and limits, continuation guidance, and an important caveat. The most identifying information is front-loaded and there is no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values do not need explanation. The description covers core behavior, constraints, and an important sibling alternative. The only real gap is the underspecified offset and expected_sha256 semantics, but defaults make the tool callable without those details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does clarify the ID scheme, max_chars range and default, page-selection limits, and the importance of retaining expected_sha256. However, offset and the exact meaning of expected_sha256 are not fully explained, so it is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Read a source excerpt by file:, zotero:, project: or feedback: ID.' This clearly distinguishes it from siblings like list_files, search, and read_page_image, and names the ID formats it accepts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing guidance: 'Use get_writing_context for project IDs.' It also provides operational guidance about continuing only as needed and retaining page range and expected_sha256. It does not explicitly state when to prefer fetch over siblings, but the purpose distinction is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_statusA
Read-onlyIdempotent

Check library health, PDF/full-text coverage and enabled capabilities. Does not load writing/research context or experiment records.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it read-only, idempotent, and non-destructive; the description adds value beyond that by specifying what the tool actually inspects and by disclaiming that it will not load writing/research context or experiment records. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with no filler. The first states what it does; the second prevents a common misuse. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool with an output schema and strong annotations, the description covers purpose, scope, and an important exclusion. No additional information is necessary to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and the schema has 100% coverage (empty properties), so there is nothing for the description to clarify. The baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('Check') and names the exact scope: library health, PDF/full-text coverage, and enabled capabilities. It also draws a boundary by saying it does not load writing/research context or experiment records, making the purpose distinct from siblings like fetch and search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use it for health/coverage/capability checks. The explicit 'Does not load...' sentence provides a when-not condition, but no sibling alternative is named, so it stops short of a full routing guide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_filesA
Read-onlyIdempotent

Browse the local literature folder. folder is relative; results contain file: IDs for fetch. Hidden files and paths outside the configured folder are excluded.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
folderNo
offsetNo
recursiveNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds non-obvious context: folder paths are relative, hidden files are excluded, paths outside the configured folder are excluded, and results contain file: IDs for downstream fetch. This is genuinely useful behavioral disclosure beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary purpose is front-loaded, and the secondary details about relative paths, exclusions, and file IDs are compact and meaningful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description communicates the core browsing behavior and folder constraints, and an output schema exists to document return values. However, it leaves limit, offset, and recursive semantics undocumented, and it lacks explicit guidance for choosing between list_files and the search/fetch siblings. These are meaningful completeness gaps for an agent calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only clarifies the folder parameter ('folder is relative'), while limit, offset, and recursive receive no explanation. For a tool with four parameters but zero schema descriptions, this is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'Browse the local literature folder.' It clarifies folder scope, notes that results contain file: IDs for fetch, and excludes hidden files and paths outside the configured folder. This makes the tool's purpose distinct from siblings like search and fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when browsing the local literature folder, and it tells agents that file IDs in results are intended for fetch. However, it does not explicitly contrast with search or other siblings, nor state when not to use list_files.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_page_imageA
Read-onlyIdempotent

Inspect an original PDF page as an image, especially for plots, tables, equations or scanned pages. page is 1-based. Requires locally installed Poppler.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
pageYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only, idempotent, and non-destructive, so the description only needs to add operational context. It adds the Poppler prerequisite and notes that page numbering is 1-based, which is useful, but it does not describe the return format or failure behavior. This matches the lower bar for annotation-backed tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. The core purpose is front-loaded, and the indexing detail and prerequisite are separated cleanly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool the description is mostly sufficient, and the Poppler prerequisite is helpful. However, with no output schema and an unexplained 'id' parameter, the agent is left to infer where the id comes from and what form the returned image takes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undocumented parameters. It clarifies that 'page' is 1-based, but the required 'id' parameter is never explained, leaving ambiguity about what identifier should be passed and where it comes from.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Inspect an original PDF page as an image') and a clear resource, with concrete use cases like plots, tables, equations, and scanned pages. It is clear about what the tool does, but it does not explicitly differentiate from siblings such as fetch or search, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'especially for plots, tables, equations or scanned pages' gives a clear context for when this tool is appropriate, implying it is for visual content where text-based extraction may fail. It does not state exclusions or explicitly name alternatives, but the guidance is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedfetch
    • First observedlibrary_status
    • First observedlist_files
    • First observedread_page_image
    • First observedsearch

TDQS

A4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool maps to a distinct function: health/status, file enumeration, search across metadata and full text, targeted excerpt retrieval, and visual page inspection. The only close pair, fetch and read_page_image, is clearly differentiated as text versus image access.

Naming Consistency3/5

Names are consistently lowercase snake_case but mix conventions: list_files and read_page_image follow verb_noun, while library_status is noun_noun and search/fetch are bare verbs. The set remains readable but does not follow one predictable naming pattern.

Tool Count5/5

Five tools is appropriately sized for a literature retrieval server; each covers a distinct need without redundancy. The count is well within the ideal 3-15 range and feels neither sparse nor bloated.

Completeness4/5

The surface covers the core evidence workflow: locate sources via search/list, retrieve text via fetch, and inspect visual content via read_page_image. Minor gaps exist, such as no dedicated full-metadata display or comparison tool, but these are workable.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    Enables local agents to search and retrieve cited evidence from PDFs and Markdown notes, including page-specific passages and rendered page images.
    6
    GPL 3.0
  • A
    license
    A
    quality
    C
    maintenance
    Enables local-first hybrid knowledge retrieval from authorized Markdown and plain-text files, combining full-text and vector search with reranking and traceable source references via a single search tool.
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI writing assistants to retrieve citable evidence from local PDFs and verify draft citations against their sources locally, providing verifiable support for claims.
    2
    1
    MIT
  • F
    license
    B
    quality
    C
    maintenance
    Enables evidence-first analysis of research papers and slide decks by converting local PDFs, DOCX, PPT, and PPTX files into a cached, token-efficient index for batched searching and selective visual verification.
    4
    2
    -