Skip to main content
Glama

arXiv MCP Server

PyPI version Python License: MIT

A Model Context Protocol (MCP) server that enables LLMs to search, download, and read arXiv papers. Gives AI assistants direct access to scientific literature.

Features

  • Search papers - Search by title, keywords, author, or arXiv ID

  • Read full text - Download PDFs and extract text automatically

  • Section extraction - Get specific sections (abstract, introduction, methods, conclusion)

  • Local caching - Downloaded papers are cached locally for fast re-access

  • Zero configuration - Works out of the box with sensible defaults

Related MCP server: arXiv MCP Server

Getting Started

Prerequisites

This MCP server uses uvx to run. First, install uv:

# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Or using Homebrew
brew install uv

After installation, restart your terminal.

Installation

Install the arXiv MCP server with your client.

Standard config works in most tools:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}
amp mcp add arxiv -- uvx arxiv-paper-mcp-server
claude mcp add arxiv-server -- uvx arxiv-paper-mcp-server

Add to your claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}
codex mcp add arxiv -- uvx arxiv-paper-mcp-server

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}

Add to Factory MCP settings:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}
gemini mcp add arxiv -- uvx arxiv-paper-mcp-server

Run goose configure, then add to ~/.config/goose/config.yaml:

extensions:
  arxiv:
    command: uvx
    args:
      - arxiv-paper-mcp-server

Add to Kiro MCP settings:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}

Add to LM Studio MCP settings:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}
opencode mcp add arxiv -- uvx arxiv-paper-mcp-server

Add to Qodo Gen MCP configuration:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}

Add to .vscode/mcp.json in your workspace:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}

Add to Warp MCP settings:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}

Add to ~/.windsurf/mcp.json:

{
  "mcpServers": {
    "arxiv": {
      "command": "uvx",
      "args": ["arxiv-paper-mcp-server"]
    }
  }
}
pip install arxiv-paper-mcp-server
arxiv-mcp-server

Tools

Tool

Description

search

Search arXiv papers by title, keywords, or arXiv ID (e.g., 2401.12345)

get_paper

Download and read the full text of a paper, with optional section filtering

list_downloaded_papers

List all locally cached papers

Tool Details

search(query, max_results=10)

Search for papers on arXiv. Supports:

  • Keywords: "transformer attention mechanism"

  • Paper ID: "2401.12345" or "arXiv:2401.12345"

  • Author: "Yann LeCun"

Returns paper ID, title, authors, publication date, and abstract preview.

get_paper(paper_id, section="all")

Download and extract text from a paper.

Section

Description

all

Full paper text (default)

abstract

Abstract only

introduction

Introduction section

method

Methods/Approach section

conclusion

Conclusion/Discussion section

list_downloaded_papers()

List all papers that have been downloaded and cached locally.

Configuration

Environment Variable

Description

Default

ARXIV_STORAGE_DIR

Directory for downloaded papers

~/.arxiv-mcp/papers

Usage Examples

Search for papers:

User: Find recent papers about prompt compression

Claude: [Uses search("prompt compression", max_results=5)]
Found 5 papers:
- 2504.16574: PIS: Linking Importance Sampling...
- ...

Read a specific paper:

User: Read the introduction of paper 2401.12345

Claude: [Uses get_paper("2401.12345", section="introduction")]
[Returns the introduction section]

Review cached papers:

User: What papers have I downloaded?

Claude: [Uses list_downloaded_papers()]
You have 3 papers cached locally:
- 2401.12345: Paper Title...

Development

# Clone the repository
git clone https://github.com/AnnaSuSu/arxiv-mcp.git
cd arxiv-mcp

# Install dependencies
uv sync

# Run server locally
uv run arxiv-mcp-server

Requirements

  • Python 3.10+

  • Dependencies: mcp, arxiv, pymupdf

License

MIT License - see LICENSE for details.

Available Tools

3 tools
get_paperA

Get the full text of an arXiv paper.

Args:
    paper_id: arXiv paper ID (e.g., "2401.12345")
    section: Which section to return: "all", "abstract", "introduction", "method", "conclusion"

Returns:
    The paper text (full or specified section)
ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes
sectionNoall

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It describes the basic behavior of retrieving text and specifies return values, but lacks details on error handling, rate limits, or authentication needs. It does not contradict annotations, as none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core purpose followed by structured Arg and Return sections. Every sentence adds value without redundancy, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, no annotations, and an output schema (which handles return values), the description is mostly complete. It covers purpose, parameters, and returns, but could improve by adding behavioral context like error cases or usage guidelines relative to siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% description coverage. It explains 'paper_id' with an example format and 'section' with allowed values and default, compensating well for the schema's lack of documentation. However, it doesn't detail constraints like paper ID validation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and resource 'full text of an arXiv paper', making the purpose specific and actionable. It distinguishes from sibling tools like 'list_downloaded_papers' (which lists) and 'search' (which searches) by focusing on retrieving specific paper content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you need paper text, but does not explicitly state when to use this tool versus alternatives like 'search' for finding papers or 'list_downloaded_papers' for checking availability. No exclusions or prerequisites are mentioned, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_downloaded_papersB

List all locally downloaded papers.

Returns:
    List of downloaded papers with their metadata
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. While it states it 'List[s] all locally downloaded papers' and mentions the return format, it doesn't address important behavioral aspects like whether this requires file system access, how it handles large collections, what metadata fields are included, or if there are any limitations on what 'locally downloaded' means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with two brief sentences that get straight to the point. The first sentence states the purpose, the second describes the return format. There's no unnecessary verbiage, though the structure could be slightly improved by integrating the return information more seamlessly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has zero parameters, an output schema exists, and no annotations, the description provides basic but incomplete context. It states what the tool does and the return type, but doesn't explain important contextual details like what constitutes 'locally downloaded' (file paths? database records?), what metadata fields are included, or how this differs from sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters with 100% schema description coverage, so the baseline is 4. The description appropriately doesn't waste space discussing non-existent parameters, though it could theoretically mention that no filtering options are available (which would be redundant with the empty schema).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('List') and resource ('locally downloaded papers'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'search' or 'get_paper' - it doesn't specify that this only shows already-downloaded content versus searching for new papers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention that 'search' might be for finding new papers online, or that 'get_paper' might retrieve specific papers. There's no context about prerequisites or when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • First observedget_paper
    • First observedlist_downloaded_papers
    • First observedsearch

TDQS

A3.5/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: get_paper retrieves full text or sections of a specific paper, list_downloaded_papers shows locally stored papers, and search finds papers based on queries. The descriptions make it easy to differentiate between retrieving, listing local content, and searching.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (get_paper, list_downloaded_papers, search), using snake_case throughout. The naming is predictable and readable, with no deviations or mixed conventions.

Tool Count3/5

With only 3 tools, the server feels thin for an arXiv domain that might benefit from more operations like filtering searches, managing downloads, or accessing metadata. While the tools cover basic needs, the count is borderline low for comprehensive arXiv interaction.

Completeness3/5

The tools cover core arXiv operations (search, retrieve, list local), but there are notable gaps such as no ability to download papers, update local storage, or access advanced metadata. Agents can work around this by using existing tools, but the surface is incomplete for full paper management workflows.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables searching, downloading, and managing academic papers from arXiv.org through natural language interactions. Provides tools for paper discovery, PDF downloads, and local paper collection management.
    4
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables arXiv paper search, PDF download, text extraction, and context chunking for LLM pipelines, along with advanced features like citation graphs and reproducibility scoring.
    2
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables LLM agents to search arXiv, download papers, parse PDFs into structured sections, and extract key findings using client-side LLMs. Features persistent caching and layout-aware PDF extraction.
    6
    MIT