Skip to main content
Glama
YI-TING-EE13

ArXiv Insight MCP Server

by YI-TING-EE13

ArXiv Insight MCP Server

An MCP server for academic-paper demos. It exposes arXiv search, citation, PDF/text retrieval, a recent-search resource, and reusable review prompts.

What It Provides

Tools:

  • health_check: local health and cache status, no network required.

  • search_arxiv: search arXiv by topic, category, sort order, year range, and offset.

  • get_paper_fulltext: download and extract paper text, using local cache.

  • download_pdf: download a PDF into the local downloads/ directory only.

  • get_bibtex: generate BibTeX for an arXiv paper ID.

  • extract_section: extract a named section from a cached or downloaded paper.

Resources:

  • papers://recent: IDs and titles from the most recent search_arxiv result.

Prompts:

  • review_paper: standard review workflow for one paper.

  • compare_papers: comparison workflow for multiple papers.

Related MCP server: MCP Research Server (arXiv)

Requirements

  • Python 3.12+

  • uv

  • Network access to arXiv for search, PDF, and BibTeX tools

The project currently uses MCP Python SDK v1.x through mcp[cli]>=1.25.0. Before refreshing dependencies, consider adding a <2 upper bound while SDK v2 remains a separate pre-release line.

Install

uv sync --frozen

Run

uv run main.py

The server uses stdio transport by default, which is the expected mode for local MCP clients.

MCP Client Configuration

From the sibling mcp-client or chainlit-mcp-client repo, use:

{
  "mcpServers": {
    "arxiv-insight": {
      "command": "uv",
      "args": ["--directory", "../mcp-server", "run", "main.py"]
    }
  }
}

For absolute Windows paths, escape backslashes in JSON.

Test

Local smoke tests do not contact arXiv or require extra dev dependencies:

uv run python scripts/smoke_test.py

Manual MCP Inspector flow:

npx -y @modelcontextprotocol/inspector

In the Inspector, configure a stdio server with command uv and args:

--directory C:\absolute\path\to\mcp-server run main.py

First call health_check, then search_arxiv with a small max_results such as 3.

Demo Baseline

This server has been verified through the sibling ../mcp-client Chainlit demo on Windows with LM Studio Local Server:

  • Endpoint: http://localhost:1234/v1

  • Model identifier: google/gemma-4-e4b

  • Resource shown in Chainlit: papers://recent

  • Successful MCP tool calls: health_check, then search_arxiv

Data And Security Notes

  • paper_cache/, downloads/, metadata_db.json, and Log.txt are local runtime artifacts.

  • download_pdf validates the arXiv ID and restricts writes to downloads/.

  • arXiv calls are rate-limited to a 3-second interval through the shared client state.

  • Tool errors are returned as text for demo readability; production servers should consider structured error output and stricter auditing.

Troubleshooting

  • If uv run python scripts/smoke_test.py cannot import dependencies, run uv sync --frozen first.

  • If arXiv tools fail but health_check works, check network connectivity and arXiv availability.

  • If PDF extraction is slow, repeat the same paper ID; cached text should be reused from paper_cache/.

Available Tools

6 tools
download_pdfA

Download a paper's PDF to a local directory. Securely restricts downloads to the project's 'downloads' folder and its subdirectories.

ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes
save_dirNodownloads

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses a security constraint (downloads restricted to the 'downloads' folder and subdirectories), which is important behavioral context for a file-writing tool. It does not mention permissions, rate limits, or failure modes, but the key safety-relevant fact is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no fluff, and the core action plus the security scoping constraint are front-loaded. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a 2-param, 1-required tool, the description covers the action and the key constraint. It could be more complete by noting paper_id expectations or the default save directory, but it is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that save_dir is a local directory and that downloads are locked to the project's 'downloads' folder, which adds meaning beyond the raw schema prop names. It still doesn't explain paper_id format or that save_dir defaults to 'downloads', but it partially covers the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Download) and resource (a paper's PDF) with a clear outcome (to a local directory). Differentiates itself from siblings like get_bibtex or get_paper_fulltext by specifying it's a PDF download rather than metadata or text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage (when you need the actual PDF file locally) but does not explicitly state when to choose this over get_paper_fulltext or other retrieval siblings. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_sectionC

Extract a specific section (e.g., Abstract, Introduction, Conclusion) from a paper.

ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes
section_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose whether this is a read-only operation, how missing/unnamed sections are handled, or what happens when section_name doesn't match a known heading. With an output schema present the return format is partly covered, but behavioral gaps remain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core action front-loaded and the clarifying examples inline. Nothing is padded, though there is very little content overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter extraction tool with an output schema, the description covers the essential action but omits edge-case behavior (section not found, paper not found) and valid section_name values beyond three examples. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It gives useful examples for section_name (Abstract, Introduction, Conclusion), which is genuinely additive, but says nothing about paper_id format or whether section matching is fuzzy or exact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Extract) and resource (a specific section of a paper), and the parenthetical examples make the scope concrete. It doesn't explicitly differentiate itself from the sibling get_paper_fulltext, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given, and the obvious alternative (get_paper_fulltext, which presumably returns the whole paper) is never mentioned. The agent must infer that this is the targeted alternative to full-text retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_bibtexB

Get the BibTeX citation for a specific arXiv paper.

ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, but for a simple read retrieval the verb 'Get' plus the output schema adequately convey the operation. It does not state auth needs, error behavior, or whether the citation is returned as raw text, so it stops short of rich disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, return values need not be described, and the tool is simple with one parameter. However, the description omits the expected ID format and any read-only/auth context, leaving minor gaps for a tool of this simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is one required parameter (paper_id), but the description's phrase 'a specific arXiv paper' lightly signals that the parameter identifies an arXiv paper. It adds marginal meaning beyond the schema and does not specify the ID format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get the BibTeX citation' for 'a specific arXiv paper'. This clearly distinguishes it from siblings like get_paper_fulltext or search_arxiv, though it does not explicitly name a sibling to contrast with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use or when-not-to-use guidance, no prerequisites, and no mention of alternatives. An agent must infer from the name that this is the citation-format retrieval option, which is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_paper_fulltextB

Retrieve the full text of a paper. Checks local cache first.

ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose one meaningful behavioral trait: a local cache lookup precedes retrieval. However, it says nothing about what 'full text' means (extracted plaintext vs. PDF), what happens on a cache miss, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and followed by the cache detail. No wasted words, though it is arguably under-specified rather than truly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. But given zero annotation coverage and an undocumented identifier format, the description should say more about what is returned and how to handle cache misses or unavailable full text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One required parameter, paper_id, with 0% schema description coverage. The description does not state the expected identifier format (arXiv ID, DOI, internal key), leaving an ambiguity that the schema also fails to resolve. This does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Retrieve the full text of a paper.' This is clearly distinct from siblings like get_bibtex and download_pdf, which fetch other representations. It stops short of explicitly contrasting itself with those siblings, so a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance about when to use this tool versus download_pdf or extract_section, which serve overlapping needs. 'Checks local cache first' hints at behavior but not selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkA

Return local server health and cache status without contacting arXiv.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does add a genuinely useful behavioral trait — that the check performs no arXiv/network calls — which signals a cheap, self-contained read. However, it says nothing about what the health output reports (e.g., healthy vs. degraded), permission requirements, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The scope qualifier is placed at the end where it usefully bounds behavior without delaying the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument diagnostic with an output schema available to describe return values, the description covers the essential purpose and the key behavioral constraint. It would be marginally stronger if it hinted at what the health/cache status indicates, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline this scores 4. There is nothing further the description could add on parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resources ('local server health and cache status'), and the qualifier 'without contacting arXiv' sharply separates it from the arXiv-fetching siblings. It does not name a sibling explicitly, but the local/diagnostic scope makes the distinction clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: an agent infers this is the diagnostic entry point for checking server/cache state. There is no explicit when-to-use framing, no prerequisites, and no named alternative, so the guidance is only adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_arxivA
Search for papers on arXiv and cache metadata for Resources.

Args:
    topic: The search query. Supports advanced prefixes like 'ti:' (title), 'au:' (author), 'abs:' (abstract).
    max_results: Maximum number of results to return (default: 100, max: 300).
    offset: The index of the first result to return (for pagination).
    category: Optional category filter (e.g., 'cs.AI').
    sort_by: Sort order. Options: 'relevance' (default), 'submitted', 'updated'.
    start_year: Filter by submission year (start). Set to 0 to ignore.
    end_year: Filter by submission year (end). Set to 0 to ignore.
ParametersJSON Schema
NameRequiredDescriptionDefault
topicYes
offsetNo
sort_byNorelevance
categoryNo
end_yearNo
start_yearNo
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does add one genuinely useful disclosure beyond the schema — that searching caches metadata as Resources — but says nothing about rate limits, auth, pagination limits beyond offset, or what the cache write implies for subsequent calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in one sentence, followed by a tight Args list where each line adds distinct information. It is slightly list-heavy but no line is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and all seven parameters are covered. The only meaningful gap is the absence of any note about how the cached Resources relate to the search results or what downstream tools expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate fully and it does: it documents the 'ti:/au:/abs:' prefixes for topic, the default and max of max_results, offset as pagination index, an example category, the three sort_by options with the default, and the '0 to ignore' convention for both year filters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource ('Search for papers on arXiv') plus a disclosed side effect ('cache metadata for Resources'). It is clearly the discovery tool among siblings like get_paper_fulltext and download_pdf, though it never explicitly names an alternative to route against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool name and by the presence of clearly downstream siblings (get_bibtex, download_pdf), but the description never states when to search here versus fetching fulltext or downloading a PDF, and gives no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observeddownload_pdf
    • First observedextract_section
    • First observedget_bibtex
    • First observedget_paper_fulltext
    • First observedhealth_check
    • First observedsearch_arxiv

TDQS

A3.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct action and output: server health, search, citation retrieval, full-text retrieval, PDF download, and section extraction. Although full-text and section extraction are related, their scopes are clearly differentiated by the descriptions.

Naming Consistency4/5

Most tools follow a consistent snake_case verb_noun pattern (search_arxiv, get_bibtex, get_paper_fulltext, download_pdf, extract_section). health_check deviates slightly by using a noun-based name, but it is a standard exception and remains readable.

Tool Count5/5

Six tools are well-scoped for an arXiv search and retrieval server. Each tool covers a distinct part of the workflow without unnecessary duplication or bloat.

Completeness4/5

The surface covers the main read-only lifecycle: search, metadata caching, citation generation, full-text retrieval, section extraction, PDF download, and health checking. A direct get-paper-metadata-by-ID tool is absent, but search results and cached Resources likely fill that gap for most workflows.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables searching and retrieving academic papers from arXiv by various criteria including title, author, and category, with support for extracting full text content from PDFs.
    4
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables searching and retrieving academic papers from arXiv by topic, allowing users to discover research papers and extract their metadata including titles, authors, abstracts, and PDF links through natural language queries.
    -
  • A
    license
    A
    quality
    B
    maintenance
    Enables searching and retrieving academic papers from arXiv with support for advanced filtering by author, category, and date, plus full paper content extraction.
    6
    15
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables searching academic papers on arXiv and retrieving detailed information such as title, authors, summary, and PDF link.
    1
    6
    MIT