Skip to main content
Glama
Oxidane-bot

Paper Download MCP Server

by Oxidane-bot

Paper Download MCP Server

English | 简体中文

MCP server for downloading academic papers by DOI, arXiv ID, or URL.

What You Get

  • paper_download: Download one or more papers (1-50 per call)

  • paper_get_metadata: Get paper metadata without downloading

  • Optional PDF-to-Markdown conversion via to_markdown

Related MCP server: rust-research-mcp

Quick Start (MCP Clients)

Before configuration, make sure uvx is available:

uvx --version

Claude Code

Add as a project-scoped MCP server:

claude mcp add --transport stdio --scope project --env PAPER_DOWNLOAD_EMAIL=your-email@university.edu paper-download -- uvx paper-download-mcp

This writes .mcp.json in the current project. Equivalent config:

{
  "mcpServers": {
    "paper-download": {
      "command": "uvx",
      "args": ["paper-download-mcp"],
      "env": {
        "PAPER_DOWNLOAD_EMAIL": "your-email@university.edu"
      }
    }
  }
}

Codex

Add with CLI:

codex mcp add paper-download --env PAPER_DOWNLOAD_EMAIL=your-email@university.edu -- uvx paper-download-mcp

Equivalent ~/.codex/config.toml snippet:

[mcp_servers.paper-download]
command = "uvx"
args = ["paper-download-mcp"]

[mcp_servers.paper-download.env]
PAPER_DOWNLOAD_EMAIL = "your-email@university.edu"

Claude Desktop

Edit Claude Desktop MCP config:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\\Claude\\claude_desktop_config.json

{
  "mcpServers": {
    "paper-download": {
      "command": "uvx",
      "args": ["paper-download-mcp"],
      "env": {
        "PAPER_DOWNLOAD_EMAIL": "your-email@university.edu"
      }
    }
  }
}

Restart Claude Desktop after editing the file.

Configuration

Required

  • PAPER_DOWNLOAD_EMAIL: Required for Unpaywall API usage.

Optional (Advanced)

  • PAPER_DOWNLOAD_OUTPUT_DIR: Global fallback output directory.

In most cases, you do not need PAPER_DOWNLOAD_OUTPUT_DIR. Prefer passing output_dir in the paper_download tool call when you want a specific location.

Legacy env vars are still supported for compatibility:

  • SCIHUB_CLI_EMAIL

  • SCIHUB_OUTPUT_DIR

Tools

paper_download

Download papers with configurable concurrency (default parallel=10). If parallel=1, papers are processed sequentially with a 2-second delay between items. OA-first routing uses OpenAlex and Unpaywall first; CORE is disabled by default in MCP runtime.

Parameters:

  • identifiers (required): list[str], 1-50 items

  • output_dir (optional): target directory (default uses runtime fallback: PAPER_DOWNLOAD_OUTPUT_DIR or ./downloads)

  • parallel (optional): concurrent workers, 1-50 (default 10)

  • to_markdown (optional): convert PDF to Markdown (false by default)

  • md_output_dir (optional): Markdown directory (default <output_dir>/md)

Examples:

paper_download(["10.1038/nature12373"])
paper_download(["10.1038/nature12373", "2301.00001"], output_dir="/path/to/papers")
paper_download(["10.1038/nature12373", "10.1126/science.169.3946.635"], parallel=10)
paper_download(["10.1038/nature12373"], to_markdown=true)

paper_get_metadata

Get metadata quickly (no PDF download).

Parameters:

  • identifier (required): DOI, arXiv ID, or URL

Example:

paper_get_metadata("10.1038/nature12373")

Troubleshooting

PAPER_DOWNLOAD_EMAIL environment variable is required

Set PAPER_DOWNLOAD_EMAIL in your MCP server config.

uvx: command not found

Install uv, then re-run the MCP configuration.

Download path errors

Pass a writable directory with output_dir, for example:

paper_download(["10.1038/nature12373"], output_dir="/absolute/path")

This tool can access papers from multiple sources, including Unpaywall and Sci-Hub. You are responsible for complying with copyright and local laws in your jurisdiction.

License

MIT. See LICENSE.

Available Tools

3 tools
paper_batch_downloadA
Download multiple papers sequentially (1-50 max, 2s delay).

Prioritizes open access sources (Unpaywall, arXiv, CORE) before Sci-Hub.

Args:
    identifiers: List of DOIs, arXiv IDs, or URLs
    output_dir: Save directory (default: './downloads')

Returns:
    Markdown summary with statistics, successes, and failures

Examples:
    paper_batch_download(["10.1038/nature12373", "2301.00001"])
    paper_batch_download(dois, "/papers")
ParametersJSON Schema
NameRequiredDescriptionDefault
identifiersYes
output_dirNo./downloads

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden and discloses key behavioral traits: sequential processing, 1-50 max limit, 2s delay, source prioritization (open access before Sci-Hub), and output format (markdown summary). It does not cover error handling or authentication needs, but provides substantial operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with core functionality, followed by structured sections for args, returns, and examples. Every sentence earns its place: the first sentence defines purpose and constraints, subsequent sections efficiently document usage without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (batch download with constraints), no annotations, and an output schema present (so return values need not be explained), the description is complete. It covers purpose, usage, parameters, behavior, and examples, providing sufficient context for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains both parameters: 'identifiers' as a list of DOIs, arXiv IDs, or URLs, and 'output_dir' as a save directory with default. This adds meaningful semantics beyond the bare schema, though it could detail format constraints for identifiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'download' and resource 'multiple papers' with specific scope 'sequentially (1-50 max, 2s delay)'. It distinguishes from sibling tools by specifying batch processing versus paper_download (likely single) and paper_metadata (information retrieval).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: for downloading multiple papers (1-50) with a delay. It implies an alternative to paper_download for batch operations but does not explicitly state when-not-to-use or compare with paper_metadata. The prioritization of sources offers some guidance on behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_downloadA
Download academic paper by DOI, arXiv ID, or URL.

Prioritizes open access sources (Unpaywall, arXiv, CORE) before Sci-Hub.
Sources: Unpaywall (OA), arXiv (OA), CORE (OA), Sci-Hub (last resort)

Args:
    identifier: DOI, arXiv ID, or URL
    output_dir: Save directory (default: './downloads')

Returns:
    Markdown with file path, metadata, source, or error message

Examples:
    paper_download("10.1038/nature12373")  # DOI
    paper_download("2301.00001")  # arXiv ID
    paper_download("https://arxiv.org/abs/2301.00001")  # URL
ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYes
output_dirNo./downloads

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well by disclosing key behavioral traits: it specifies the prioritization of sources (open access first, Sci-Hub as last resort), the return format (Markdown with file path, metadata, source, or error), and includes examples. It does not mention rate limits or authentication needs, but covers the core behavior adequately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the purpose, followed by prioritized sources, args, returns, and examples. Each section adds value without redundancy, and the structure is clear and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (2 parameters, no annotations, but with output schema), the description is complete enough. It explains the tool's purpose, usage, behavior, parameters, and return values, and the output schema likely covers return details, so no gaps remain for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics: 'identifier' is explained as 'DOI, arXiv ID, or URL' with examples, and 'output_dir' is described as 'Save directory (default: './downloads')'. This clarifies beyond the schema's basic types, though it could detail format constraints (e.g., URL patterns).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Download academic paper') and the resources involved ('by DOI, arXiv ID, or URL'). It distinguishes from sibling tools like 'paper_batch_download' (which handles multiple papers) and 'paper_metadata' (which retrieves metadata only) by focusing on single-paper downloading with file output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use this tool (for downloading papers via identifiers) and mentions prioritization of open access sources, which guides usage. However, it does not explicitly state when not to use it or directly compare to alternatives like 'paper_batch_download' for multiple papers, though the distinction is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_metadataA
Get paper metadata without downloading (fast, <1s).

Sources: Unpaywall, Crossref, arXiv APIs
Returns: title, authors, year, journal, OA status, available sources

Args:
    identifier: DOI, arXiv ID, or URL

Returns:
    JSON with metadata fields

Examples:
    paper_metadata("10.1038/nature12373")  # DOI
    paper_metadata("2301.00001")  # arXiv ID
ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key traits: the operation is read-only (implied by 'Get'), performance characteristics ('fast, <1s'), data sources ('Unpaywall, Crossref, arXiv APIs'), and what information is returned. It doesn't mention error handling or rate limits, but covers most essential aspects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly structured and front-loaded: the first sentence states the core purpose, followed by sources, returns, args, and examples. Every sentence earns its place with no wasted words, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, no annotations, and the presence of an output schema, the description provides excellent completeness. It covers purpose, usage context, behavioral traits, parameter semantics with examples, and mentions the return format. The output schema will handle return value details, so the description doesn't need to duplicate that information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must fully compensate. It clearly explains the 'identifier' parameter with specific examples of valid formats (DOI, arXiv ID, URL) and provides concrete usage examples, adding substantial value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('paper metadata'), and distinguishes it from sibling tools by emphasizing it's 'without downloading' (unlike 'paper_download' and 'paper_batch_download'). The phrase 'fast, <1s' adds useful context about performance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('without downloading') and implicitly contrasts with download-focused siblings. However, it doesn't explicitly state when NOT to use it or name specific alternatives, which prevents a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no ambiguity: paper_batch_download handles multiple papers sequentially, paper_download handles single paper downloads, and paper_metadata retrieves metadata without downloading. The descriptions clearly differentiate between batch processing, single downloads, and metadata-only operations.

Naming Consistency5/5

All three tools follow a perfect verb_noun pattern with consistent snake_case naming: paper_batch_download, paper_download, and paper_metadata. The naming convention is predictable and readable throughout the entire tool set.

Tool Count3/5

With only 3 tools, the set feels somewhat thin for a paper download server that could benefit from additional operations like search, citation management, or format conversion. While the core functionality is covered, the scope could be expanded to provide more comprehensive paper management capabilities.

Completeness4/5

The tool set covers the essential operations for paper downloading and metadata retrieval well, with clear separation between batch and single operations. A minor gap exists in search functionality (finding papers by keyword/topic) and citation-related operations, but the core download workflow is complete and agents can work effectively with what's provided.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    An MCP server for searching and downloading academic papers from multiple sources including arXiv, PubMed, bioRxiv, and Sci-Hub, designed for seamless integration with large language models like Claude Desktop.
    795
    57
    2,515
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server for academic research that enables paper search across 14 sources, PDF download with multi-provider fallback, metadata extraction, and bibliography generation.
    2
    GPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for searching, downloading, and reading academic papers from multiple sources such as arXiv, Google Scholar, and Elsevier.
    6
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    MCP server for downloading academic papers from DOI or title, resolving references, and generating citations. Supports batch downloads, multiple mirrors, and optional Unpaywall integration.
    2
    11
    1
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Oxidane-bot/paper-download-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server