mcp-simple-arxiv
mcp-simple-arxiv
API를 통해 arXiv 논문에 대한 액세스를 제공하는 MCP 서버입니다.
특징
이 서버를 사용하면 LLM 클라이언트(예: Claude Desktop)에서 다음 작업을 수행할 수 있습니다.
제목 및 초록 내용으로 arXiv에서 과학 논문 검색
논문 메타데이터 및 초록 받기
사용 가능한 논문 형식(PDF/HTML)에 대한 링크에 액세스하세요.
서버는 arXiv의 API 가이드라인에 따라 적절한 속도 제한을 구현합니다(3초마다 최대 1개 요청).
Related MCP server: Research Paper Agent
설치
Smithery를 통해 설치
Smithery를 통해 Claude Desktop에 Simple Arxiv를 자동으로 설치하는 방법:
지엑스피1
수동 설치
pip install mcp-simple-arxivClaude Desktop과 함께 사용
claude_desktop_config.json 에 다음 구성을 추가하세요.
(맥 OS)
{
"mcpServers": {
"simple-arxiv": {
"command": "python",
"args": ["-m", "mcp_simple_arxiv"]
}
}
}(Windows 버전):
{
"mcpServers": {
"simple-arxiv": {
"command": "C:\\Users\\YOUR_USERNAME\\AppData\\Local\\Programs\\Python\\Python311\\python.exe",
"args": [
"-m",
"mcp_simple_arxiv"
]
}
}
}Claude Desktop을 다시 시작하면 다음 기능을 사용할 수 있습니다.
논문 검색
다음과 같은 쿼리를 사용하여 Claude에게 논문 검색을 요청할 수 있습니다.
Can you search arXiv for recent papers about large language models?검색을 통해 다음을 포함한 일치하는 논문에 대한 기본 정보가 반환됩니다.
논문 제목
저자
arXiv ID
출판일
논문 세부 정보 얻기
종이 신분증을 받으면 자세한 내용을 문의할 수 있습니다.
Can you show me the details for paper 2103.08220?그러면 다음이 반환됩니다.
전체 논문 제목
저자
출판 및 업데이트 날짜
저널 참조(가능한 경우)
논문 초록
사용 가능한 형식(PDF/HTML)에 대한 링크
개발
개발을 위해 설치하려면:
git clone https://github.com/andybrandt/mcp-simple-arxiv
cd mcp-simple-arxiv
pip install -e .arXiv API 가이드라인
이 서버는 arXiv API 사용 지침을 따릅니다.
3초당 최대 1개의 요청으로 속도 제한
한 번에 하나의 연결
적절한 오류 처리 및 재시도 논리
특허
MIT
Available Tools
4 toolsget_paper_dataGet arXiv Paper DataBRead-only
Get detailed information about a specific paper including abstract and available formats.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, indicating a safe, read-only operation with open-world assumptions. The description adds some context by specifying the type of information retrieved ('abstract and available formats'), but does not disclose additional behavioral traits like rate limits, error handling, or authentication needs. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Get detailed information about a specific paper including abstract and available formats.' It is front-loaded with the core purpose and includes key details without unnecessary words, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one parameter) and the presence of annotations (readOnlyHint, openWorldHint) and an output schema (which handles return values), the description is reasonably complete. It specifies the information retrieved, but lacks usage guidelines and parameter details, which are minor gaps in this context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the parameter 'paper_id' is undocumented in the schema. The description does not add any semantic details about this parameter, such as expected format (e.g., arXiv ID like '2401.12345') or examples. With one parameter and no schema documentation, the description fails to compensate, resulting in a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get detailed information about a specific paper including abstract and available formats.' It specifies the verb ('Get'), resource ('paper'), and scope ('detailed information'), but does not explicitly differentiate it from sibling tools like 'search_papers' or 'list_categories', which prevents a score of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention sibling tools like 'search_papers' for broader searches or 'list_categories' for category listings, nor does it specify prerequisites or exclusions, such as requiring a specific paper ID format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_categoriesList arXiv CategoriesARead-only
List all available arXiv categories and how to use them in search.
| Name | Required | Description | Default |
|---|---|---|---|
| primary_category | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, indicating this is a safe read operation with a closed set of results. The description adds that it lists categories and how to use them in search, which provides some behavioral context beyond annotations, but doesn't detail aspects like response format, pagination, or error handling. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that efficiently conveys the core purpose and an additional use case ('how to use them in search'). It's front-loaded with the main action and avoids unnecessary words, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (read-only, 1 optional parameter), annotations covering safety, and the presence of an output schema (which handles return values), the description is mostly complete. It explains what the tool does and adds search usage context. However, it misses explaining the optional parameter, which is a minor gap in an otherwise adequate description for this context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and 1 parameter, the description carries the burden of explaining parameters. It doesn't mention the 'primary_category' parameter at all, which is a gap. However, since there's only one optional parameter with a default null value, and the tool likely functions without it (listing all categories), the description's omission is less critical. Baseline for 0 parameters would be 4, but with 1 undocumented parameter, it's slightly penalized.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List all available arXiv categories') and the resource ('arXiv categories'), making the purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'update_categories' or 'search_papers', which would require a 5. The phrase 'and how to use them in search' adds useful context but doesn't fully distinguish it from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'search_papers' or 'update_categories'. It mentions 'how to use them in search' which implies a connection to searching, but doesn't specify if this should be used before searching, for filtering, or as a reference. No explicit when/when-not instructions or named alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersSearch arXiv PapersARead-only
Search for papers on arXiv by title and abstract content.
You can use advanced search syntax:
Search in title: ti:"search terms"
Search in abstract: abs:"search terms"
Search by author: au:"author name"
Combine terms with: AND, OR, ANDNOT
Filter by category: cat:cs.AI (use list_categories tool to see available categories)
Examples:
"machine learning" (searches all fields)
ti:"neural networks" AND cat:cs.AI (title with category)
au:bengio AND ti:"deep learning" (author and title)
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and openWorldHint=true, indicating a safe, read-only operation with open-world data. The description adds valuable behavioral context beyond annotations by detailing advanced search syntax, examples, and the ability to filter by category, which helps the agent understand how to construct effective queries. However, it does not mention rate limits, pagination, or error handling, leaving some behavioral aspects uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with the core purpose, followed by syntax details and examples. Every sentence earns its place by providing essential information without redundancy, making it efficient and easy to parse for an AI agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (search with advanced syntax), low schema coverage (0%), and the presence of an output schema (which handles return values), the description is complete enough. It covers purpose, usage, parameters, and examples, leaving no critical gaps for the agent to operate effectively in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description fully compensates by explaining the semantics of the 'query' parameter in detail (e.g., advanced syntax, examples) and implies the use of 'max_results' through examples like limiting results. It adds significant meaning beyond the bare schema, making parameter usage clear and actionable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Search for papers on arXiv') and resource ('by title and abstract content'), distinguishing it from sibling tools like 'get_paper_data' (likely for retrieving specific papers), 'list_categories' (for listing categories), and 'update_categories' (for updating categories). It provides a verb+resource+scope combination that is precise and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides usage guidance by naming an alternative tool ('use list_categories tool to see available categories') and includes examples that illustrate when to use specific syntax (e.g., for title, abstract, author, category filtering). It effectively guides the agent on how to apply the tool in different contexts without being misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_categoriesUpdate arXiv CategoriesA
Update the stored category taxonomy by fetching the latest version from arxiv.org
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=false (indicating mutation) and openWorldHint=true (suggesting external data fetching). The description adds valuable behavioral context: it specifies that this tool fetches from arxiv.org (external source) and updates stored data. This goes beyond annotations by revealing the external dependency and the 'fetch-and-update' nature of the operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that efficiently conveys the entire operation: action (update), target (stored category taxonomy), method (fetching), and source (arxiv.org). No wasted words, perfectly front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with annotations covering mutation/external access and an output schema present, the description provides exactly what's needed: it explains what the tool does, where it gets data, and what it updates. The presence of an output schema means return values don't need explanation in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage (empty schema). The description doesn't need to explain parameters, but it implicitly confirms there are no required inputs by describing a self-contained fetch operation. This is appropriate for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('update'), the resource ('stored category taxonomy'), and the source ('fetching the latest version from arxiv.org'). It distinguishes from siblings like 'list_categories' (which presumably reads existing data) and 'get_paper_data'/'search_papers' (which work with papers rather than taxonomy).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: when you need to refresh the locally stored taxonomy with the latest from arxiv.org. It doesn't explicitly state when NOT to use it or name alternatives, but the context is clear enough for an agent to infer this is for maintenance/update operations rather than regular queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- Changed
get_paper_data3 fields changed- removed
Input schema / properties / paper_id / titleRemoved value: -"Paper Id" - removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / titleRemoved value: -"_WrappedResult"
- Changed
list_categories3 fields changed- removed
Input schema / properties / primary_category / titleRemoved value: -"Primary Category" - removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / titleRemoved value: -"_WrappedResult"
- Changed
search_papers4 fields changed- removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / titleRemoved value: -"_WrappedResult"
- Changed
update_categories2 fields changed- removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / titleRemoved value: -"_WrappedResult"
4 tool updates
- First observed
get_paper_data - First observed
list_categories - First observed
search_papers - First observed
update_categories
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: get_paper_data retrieves specific paper details, list_categories shows available categories, search_papers performs searches, and update_categories refreshes category data. There is no overlap in functionality that would cause confusion.
All tools follow a consistent verb_noun naming pattern (get_paper_data, list_categories, search_papers, update_categories). The naming is uniform and predictable throughout the set.
With 4 tools, this server is well-scoped for arXiv paper access. Each tool serves a clear purpose (retrieval, listing, searching, updating) without being too sparse or bloated, making it efficient for the domain.
The tools cover core arXiv operations: searching, retrieving paper details, and managing categories. A minor gap is the lack of tools for downloading papers or handling user-specific features like saved papers, but the basic workflow is well-supported.
Maintenance
Related MCP Connectors
Academic research MCP server for paper search, citation checks, graphs, and deep research.
An MCP server that provides congressional transcripts
An MCP server that gives your AI access to the source code and docs of all public github repos
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceThis MCP server enables users to search for scientific papers on arXiv and retrieve detailed metadata for specific papers. It provides tools to perform search queries and fetch in-depth information using paper IDs.3Apache 2.0
- FlicenseNot gradedqualityDmaintenanceA remote MCP server for searching arXiv papers, extracting paper details, and generating structured prompts for LLM agents.1-
- FlicenseAqualityDmaintenanceA streamlined MCP server that connects AI assistants to arXiv's vast collection of academic papers, enabling search, retrieval, and analysis of research papers.71-
- AlicenseAqualityCmaintenanceMCP server for searching and retrieving arXiv papers with full-text PDF extraction.52MIT