Paper Download MCP Server
This server downloads academic papers and retrieves their metadata using DOI, arXiv ID, or URL identifiers.
Download papers (
paper_download): Download 1–50 papers per call with configurable concurrency (default 10 parallel workers), optional PDF-to-Markdown conversion, and customizable output directoriesRetrieve paper metadata (
paper_get_metadata): Fetch metadata (title, authors, year, journal, OA status, available sources) without downloading the PDF, using Unpaywall, Crossref, and arXiv APIsSmart source prioritization: Automatically tries open-access sources (Unpaywall, arXiv, CORE) before falling back to Sci-Hub as a last resort
Flexible input: Accepts DOI, arXiv ID, or direct URLs
Configuration: Requires an email via the
PAPER_DOWNLOAD_EMAILenvironment variable for the Unpaywall API
Enables the automatic identification and downloading of academic preprints directly from the arXiv repository.
Allows for the retrieval of academic papers and associated metadata using Digital Object Identifiers (DOIs) from various scholarly sources.
Supports downloading academic research articles by integrating with PubMed Central (PMC) for open access content.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Paper Download MCP ServerDownload the paper with DOI 10.1038/nature12373"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Paper Download MCP Server
English | 简体中文
MCP server for downloading academic papers by DOI, arXiv ID, or URL.
What You Get
paper_download: Download one or more papers (1-50 per call)paper_get_metadata: Get paper metadata without downloadingOptional PDF-to-Markdown conversion via
to_markdown
Related MCP server: rust-research-mcp
Quick Start (MCP Clients)
Before configuration, make sure uvx is available:
uvx --versionClaude Code
Add as a project-scoped MCP server:
claude mcp add --transport stdio --scope project --env PAPER_DOWNLOAD_EMAIL=your-email@university.edu paper-download -- uvx paper-download-mcpThis writes .mcp.json in the current project. Equivalent config:
{
"mcpServers": {
"paper-download": {
"command": "uvx",
"args": ["paper-download-mcp"],
"env": {
"PAPER_DOWNLOAD_EMAIL": "your-email@university.edu"
}
}
}
}Codex
Add with CLI:
codex mcp add paper-download --env PAPER_DOWNLOAD_EMAIL=your-email@university.edu -- uvx paper-download-mcpEquivalent ~/.codex/config.toml snippet:
[mcp_servers.paper-download]
command = "uvx"
args = ["paper-download-mcp"]
[mcp_servers.paper-download.env]
PAPER_DOWNLOAD_EMAIL = "your-email@university.edu"Claude Desktop
Edit Claude Desktop MCP config:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\\Claude\\claude_desktop_config.json
{
"mcpServers": {
"paper-download": {
"command": "uvx",
"args": ["paper-download-mcp"],
"env": {
"PAPER_DOWNLOAD_EMAIL": "your-email@university.edu"
}
}
}
}Restart Claude Desktop after editing the file.
Configuration
Required
PAPER_DOWNLOAD_EMAIL: Required for Unpaywall API usage.
Optional (Advanced)
PAPER_DOWNLOAD_OUTPUT_DIR: Global fallback output directory.
In most cases, you do not need PAPER_DOWNLOAD_OUTPUT_DIR. Prefer passing output_dir in the paper_download tool call when you want a specific location.
Legacy env vars are still supported for compatibility:
SCIHUB_CLI_EMAILSCIHUB_OUTPUT_DIR
Tools
paper_download
Download papers with configurable concurrency (default parallel=10).
If parallel=1, papers are processed sequentially with a 2-second delay between items.
OA-first routing uses OpenAlex and Unpaywall first; CORE is disabled by default in MCP runtime.
Parameters:
identifiers(required):list[str], 1-50 itemsoutput_dir(optional): target directory (default uses runtime fallback:PAPER_DOWNLOAD_OUTPUT_DIRor./downloads)parallel(optional): concurrent workers,1-50(default10)to_markdown(optional): convert PDF to Markdown (falseby default)md_output_dir(optional): Markdown directory (default<output_dir>/md)
Examples:
paper_download(["10.1038/nature12373"])
paper_download(["10.1038/nature12373", "2301.00001"], output_dir="/path/to/papers")
paper_download(["10.1038/nature12373", "10.1126/science.169.3946.635"], parallel=10)
paper_download(["10.1038/nature12373"], to_markdown=true)paper_get_metadata
Get metadata quickly (no PDF download).
Parameters:
identifier(required): DOI, arXiv ID, or URL
Example:
paper_get_metadata("10.1038/nature12373")Troubleshooting
PAPER_DOWNLOAD_EMAIL environment variable is required
Set PAPER_DOWNLOAD_EMAIL in your MCP server config.
uvx: command not found
Install uv, then re-run the MCP configuration.
Download path errors
Pass a writable directory with output_dir, for example:
paper_download(["10.1038/nature12373"], output_dir="/absolute/path")Legal Notice
This tool can access papers from multiple sources, including Unpaywall and Sci-Hub. You are responsible for complying with copyright and local laws in your jurisdiction.
License
MIT. See LICENSE.
Available Tools
3 toolspaper_batch_downloadA
Download multiple papers sequentially (1-50 max, 2s delay).
Prioritizes open access sources (Unpaywall, arXiv, CORE) before Sci-Hub.
Args:
identifiers: List of DOIs, arXiv IDs, or URLs
output_dir: Save directory (default: './downloads')
Returns:
Markdown summary with statistics, successes, and failures
Examples:
paper_batch_download(["10.1038/nature12373", "2301.00001"])
paper_batch_download(dois, "/papers")
| Name | Required | Description | Default |
|---|---|---|---|
| identifiers | Yes | ||
| output_dir | No | ./downloads |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and discloses key behavioral traits: sequential processing, 1-50 max limit, 2s delay, source prioritization (open access before Sci-Hub), and output format (markdown summary). It does not cover error handling or authentication needs, but provides substantial operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with core functionality, followed by structured sections for args, returns, and examples. Every sentence earns its place: the first sentence defines purpose and constraints, subsequent sections efficiently document usage without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (batch download with constraints), no annotations, and an output schema present (so return values need not be explained), the description is complete. It covers purpose, usage, parameters, behavior, and examples, providing sufficient context for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains both parameters: 'identifiers' as a list of DOIs, arXiv IDs, or URLs, and 'output_dir' as a save directory with default. This adds meaningful semantics beyond the bare schema, though it could detail format constraints for identifiers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'download' and resource 'multiple papers' with specific scope 'sequentially (1-50 max, 2s delay)'. It distinguishes from sibling tools by specifying batch processing versus paper_download (likely single) and paper_metadata (information retrieval).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: for downloading multiple papers (1-50) with a delay. It implies an alternative to paper_download for batch operations but does not explicitly state when-not-to-use or compare with paper_metadata. The prioritization of sources offers some guidance on behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_downloadA
Download academic paper by DOI, arXiv ID, or URL.
Prioritizes open access sources (Unpaywall, arXiv, CORE) before Sci-Hub.
Sources: Unpaywall (OA), arXiv (OA), CORE (OA), Sci-Hub (last resort)
Args:
identifier: DOI, arXiv ID, or URL
output_dir: Save directory (default: './downloads')
Returns:
Markdown with file path, metadata, source, or error message
Examples:
paper_download("10.1038/nature12373") # DOI
paper_download("2301.00001") # arXiv ID
paper_download("https://arxiv.org/abs/2301.00001") # URL
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | ||
| output_dir | No | ./downloads |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well by disclosing key behavioral traits: it specifies the prioritization of sources (open access first, Sci-Hub as last resort), the return format (Markdown with file path, metadata, source, or error), and includes examples. It does not mention rate limits or authentication needs, but covers the core behavior adequately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded: the first sentence states the purpose, followed by prioritized sources, args, returns, and examples. Each section adds value without redundancy, and the structure is clear and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (2 parameters, no annotations, but with output schema), the description is complete enough. It explains the tool's purpose, usage, behavior, parameters, and return values, and the output schema likely covers return details, so no gaps remain for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics: 'identifier' is explained as 'DOI, arXiv ID, or URL' with examples, and 'output_dir' is described as 'Save directory (default: './downloads')'. This clarifies beyond the schema's basic types, though it could detail format constraints (e.g., URL patterns).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Download academic paper') and the resources involved ('by DOI, arXiv ID, or URL'). It distinguishes from sibling tools like 'paper_batch_download' (which handles multiple papers) and 'paper_metadata' (which retrieves metadata only) by focusing on single-paper downloading with file output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool (for downloading papers via identifiers) and mentions prioritization of open access sources, which guides usage. However, it does not explicitly state when not to use it or directly compare to alternatives like 'paper_batch_download' for multiple papers, though the distinction is implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_metadataA
Get paper metadata without downloading (fast, <1s).
Sources: Unpaywall, Crossref, arXiv APIs
Returns: title, authors, year, journal, OA status, available sources
Args:
identifier: DOI, arXiv ID, or URL
Returns:
JSON with metadata fields
Examples:
paper_metadata("10.1038/nature12373") # DOI
paper_metadata("2301.00001") # arXiv ID
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key traits: the operation is read-only (implied by 'Get'), performance characteristics ('fast, <1s'), data sources ('Unpaywall, Crossref, arXiv APIs'), and what information is returned. It doesn't mention error handling or rate limits, but covers most essential aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly structured and front-loaded: the first sentence states the core purpose, followed by sources, returns, args, and examples. Every sentence earns its place with no wasted words, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, no annotations, and the presence of an output schema, the description provides excellent completeness. It covers purpose, usage context, behavioral traits, parameter semantics with examples, and mentions the return format. The output schema will handle return value details, so the description doesn't need to duplicate that information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must fully compensate. It clearly explains the 'identifier' parameter with specific examples of valid formats (DOI, arXiv ID, URL) and provides concrete usage examples, adding substantial value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Get') and resource ('paper metadata'), and distinguishes it from sibling tools by emphasizing it's 'without downloading' (unlike 'paper_download' and 'paper_batch_download'). The phrase 'fast, <1s' adds useful context about performance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('without downloading') and implicitly contrasts with download-focused siblings. However, it doesn't explicitly state when NOT to use it or name specific alternatives, which prevents a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose with no ambiguity: paper_batch_download handles multiple papers sequentially, paper_download handles single paper downloads, and paper_metadata retrieves metadata without downloading. The descriptions clearly differentiate between batch processing, single downloads, and metadata-only operations.
All three tools follow a perfect verb_noun pattern with consistent snake_case naming: paper_batch_download, paper_download, and paper_metadata. The naming convention is predictable and readable throughout the entire tool set.
With only 3 tools, the set feels somewhat thin for a paper download server that could benefit from additional operations like search, citation management, or format conversion. While the core functionality is covered, the scope could be expanded to provide more comprehensive paper management capabilities.
The tool set covers the essential operations for paper downloading and metadata retrieval well, with clear separation between batch and single operations. A minor gap exists in search functionality (finding papers by keyword/topic) and citation-related operations, but the core download workflow is complete and agents can work effectively with what's provided.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Academic research MCP server for paper search, citation checks, graphs, and deep research.
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
Search and download academic papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic…
MCP server for Altmetric APIs - track research attention across news, policy, social media, and more
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server for searching and downloading academic papers from multiple sources including arXiv, PubMed, bioRxiv, and Sci-Hub, designed for seamless integration with large language models like Claude Desktop.795572,515MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server for academic research that enables paper search across 14 sources, PDF download with multi-provider fallback, metadata extraction, and bibliography generation.2GPL 3.0
- AlicenseNot gradedqualityCmaintenanceMCP server for searching, downloading, and reading academic papers from multiple sources such as arXiv, Google Scholar, and Elsevier.6MIT
- AlicenseAqualityAmaintenanceMCP server for downloading academic papers from DOI or title, resolving references, and generating citations. Supports batch downloads, multiple mirrors, and optional Unpaywall integration.2111MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Oxidane-bot/paper-download-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server