paper-search
Searches English academic papers from arXiv API and retrieves PDFs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@paper-searchsearch for papers on reinforcement learning"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
paper-search MCP
日本語・英語の学術論文を検索し、本文の取得可能性を判定し、公開 PDF をローカル収集する MCP サーバー。Claude Code などの LLM エージェントから利用することを想定しています。
英語論文: arXiv(API キー不要)
日本語論文: CiNii Research(API キー不要。appid は任意)
本文解決: arXiv PDF / J-STAGE 公開論文 / 検索結果に含まれる既知 URL(機関リポジトリ等)
データソースごとの API 差異はサーバー内部で吸収し、LLM には統一された 4 つの Tool だけを公開します。
提供 Tool
Tool | 役割 |
| 日英複数ソースの並列検索。正規化・重複除去・ランキング済みの候補を返す |
| 複数論文の詳細と本文取得可能性(available / restricted 等)を一括調査 |
| 公開 PDF を一括ダウンロードしローカル保存(SHA-256 記録、重複防止) |
| 収集済みファイルのローカルパス等を返す |
すべてバッチ入力(paper_ids[])で、一部の失敗があっても成功分は返します(部分成功)。
Related MCP server: paper-mcp
セットアップ
要件: uv(Python 3.12+ は uv が自動管理)
git clone <this-repo>
cd mcp-paper-search
uv syncClaude Code への登録
プロジェクトルート(または利用したいワークスペース)の .mcp.json に追加:
{
"mcpServers": {
"paper-search": {
"command": "uv",
"args": ["run", "--directory", "/path/to/mcp-paper-search", "paper-search-mcp"],
"env": {
"PAPER_SEARCH_DATA_DIR": "/path/to/mcp-paper-search/data"
}
}
}
}環境変数
変数 | 必須 | 説明 |
| 任意 | PDF・収集記録の保存先(デフォルト |
| 任意 | CiNii Research の appid。なしでも動作するが、継続利用には NII へのアプリケーション登録を推奨 |
使い方の例
Claude Code に次のように依頼すると、search → inspect → collect のワークフローが実行されます。
組込みソフトウェア開発における AI エージェント活用について、
日本語と英語の論文を探して、関連度の高いものを PDF で収集して。収集した PDF は data/papers/{年}/{paper-id}.pdf に保存されます。
開発
uv run pytest # offline テスト(実 API は叩かない)
uv run pytest -m live # 実 API との整合検証(arXiv / CiNii)
uv run ruff check src tests # lint
uv run mypy # 型チェック(strict)設計ドキュメントは docs/plan.md、開発時の規約は CLAUDE.md を参照してください。
実装状況
✅ Phase 1–7: ドメインモデル / arXiv・CiNii 検索 / 検索統合(重複除去・ランキング)/ arXiv・J-STAGE 本文解決 / PDF 収集 / MCP サーバー
⬜ Phase 8: 機関リポジトリ・DOI Resolver
⬜ Phase 9: SQLite 永続化(現状は検索結果がプロセス内メモリのため、サーバー再起動後は再検索が必要)
⬜ Skills(paper-discovery / literature-survey / paper-collection)
制約・ポリシー
公開・合法的に取得可能な本文のみ収集します(認証突破・Shadow Library 非対応)
各 API にはホスト単位の rate limit(arXiv 3 秒 / CiNii 1 秒 / J-STAGE 2 秒)を設けています
ダウンロード対象 URL はサーバーが解決した候補のみに限定し、SSRF 検証を行います
Available Tools
4 toolscollect_papersA
複数論文の本文 PDF を一括取得しローカル保存する。
公開・合法的に取得可能な本文のみ収集する。部分成功を許容し、 per-item の status(collected / already_collected / unavailable / restricted / failed)と理由を返す。
| Name | Required | Description | Default |
|---|---|---|---|
| paper_ids | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses batch operation, local saving, legal constraints, partial success, and per-item statuses with a specific list of values. This provides substantial behavioral context beyond a vague 'collect papers' claim.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action, followed by critical constraints and return details. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one parameter) and the description covers the main behavioral aspects: what it does, constraints, partial success, and output statuses. It relies on the output schema for detailed return structure, which is appropriate. Minor gaps include no mention of storage location or how saved files are retrieved, but sibling get_collected_papers likely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single 'paper_ids' parameter. While the description does not add parameter-specific details, the parameter name is self-explanatory (IDs of papers to collect) and the tool purpose makes its meaning clear. The lack of elaboration is acceptable but not exemplary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's action (batch-fetch and locally save full-text PDFs) and resource (multiple papers' full-text PDFs). It distinguishes itself from sibling tools like search_papers, inspect_papers, and get_collected_papers by focusing on the downloading/saving action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for collecting PDFs but does not explicitly state when to choose this tool over alternatives. It notes legal constraints (only publicly/legally obtainable texts) but lacks clear 'when-to-use' or 'when-not-to-use' guidance relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_collected_papersA
収集済み論文のローカル情報(パス・SHA-256)を返す。
paper_ids を省略すると収集済み論文の一覧を返す。
| Name | Required | Description | Default |
|---|---|---|---|
| paper_ids | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It usefully discloses that omitting paper_ids returns a list and that the output includes path and SHA-256, but it does not explicitly state read-only safety, error conditions, or required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the primary purpose, followed by an important conditional usage. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one optional parameter and an output schema is indicated, so the description need not enumerate return values. It covers the main behavior and the conditional list mode; only minor gaps such as error handling and permission requirements remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that paper_ids is optional and that omitting it returns a list, adding meaning beyond the schema. However, it does not describe the expected string format of IDs or what happens if IDs are provided but not found in the collection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns local information (path, SHA-256) of collected papers, with a specific verb (返す) and resource (収集済み論文のローカル情報). It also distinguishes itself from siblings by focusing on already-collected papers' local metadata and adding the list behavior when paper_ids is omitted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call this when you need local path/hash of collected papers. It explains the conditional behavior of omitting paper_ids, but it does not explicitly mention when to prefer this over sibling tools like search_papers/inspect_papers, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_papersA
複数論文の詳細メタデータと本文取得可能性を一括調査する。
search_papers が返した paper_id を渡す。ファイル保存は行わない。 fulltext.status: available / metadata_only / restricted / not_found / unknown。
| Name | Required | Description | Default |
|---|---|---|---|
| paper_ids | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that no files are saved and enumerates the possible fulltext.status values (available, metadata_only, restricted, not_found, unknown), giving the agent concrete knowledge of what to expect from the tool's output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences, each adding a distinct piece of information: purpose, input source and side effect disclaimer, and output status vocabulary. There is no redundancy or filler content; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (single parameter, no annotations) and the presence of an output schema, the description is complete: it explains the input origin, the operation, and the meaning of status values. An agent has all necessary context to invoke the tool correctly and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines paper_ids as an array of strings, with 0% description coverage. The description compensates by explaining that these IDs come from search_papers, adding critical semantic context about the parameter's origin and intended use. It does not provide format examples, but the schema already covers the type structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Investigate detailed metadata and full-text availability for multiple papers at once,' which clearly identifies the tool's function and resource. It distinguishes itself from siblings: search_papers finds papers, collect_papers saves them, and get_collected_papers retrieves saved ones, while this tool inspects details without saving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: 'Pass paper_id returned by search_papers,' establishing a clear workflow after searching. It also warns 'Does not save files,' which implicitly differentiates it from collect_papers and tells the agent when not to use this tool for saving purposes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersA
日英の学術論文を検索する。
languages(ja / en)に応じて arXiv・CiNii Research を検索し、 重複除去・ランキング済みの論文候補を返す。本文取得は行わない。 categories は分野フィルタ(例: cs.SE、cs.AI。arXiv のみ有効)。 1 データソースの障害は warnings として返し、検索全体は失敗しない。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| offset | No | ||
| sources | No | ||
| year_to | No | ||
| languages | No | ||
| year_from | No | ||
| categories | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses several behaviors: deduplication, ranking, no full-text fetch, and graceful handling of individual source failures via warnings. This is substantial but not exhaustive (e.g., auth needs, rate limits).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with key search behavior front-loaded and examples in parentheses. It is slightly unstructured but efficient and earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter search tool with no annotations, the description covers core behavior (source failure, dedupe, ranking) but lacks detailed parameter guidance and does not mention pagination or output shape (though output schema exists). Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage. The description only explains 'categories' (arXiv-only filter with examples) and 'languages' (ja/en), but omits semantics for query, limit, offset, sources, year_from, and year_to. Many parameter meanings are left to inference from names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches Japanese/English academic papers across arXiv and CiNii, returning deduplicated/ranked candidates. This distinguishes it from sibling tools like inspect_papers or collect_papers, which imply different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides context that search depends on languages and notes it does not fetch full text, implying it is not for full-text retrieval. However, it never explicitly names alternative tools or provides when/when-not criteria relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
collect_papers - First observed
get_collected_papers - First observed
inspect_papers - First observed
search_papers
TDQS
Scored across 4 tools
Each tool has a clearly distinct role: searching, inspecting metadata, collecting full-text, and listing collected papers. There is no overlap or ambiguity.
All tools follow a consistent verb_noun pattern: search_papers, inspect_papers, collect_papers, get_collected_papers. Naming is uniform and predictable.
With 4 tools, the set is well-scoped for the paper search and management workflow. Each tool serves a necessary step and none is redundant.
The workflow covers search, inspection, collection, and listing. A minor gap is the lack of a delete/remove operation for collected papers, but the core lifecycle is well covered.
Maintenance
Related MCP Connectors
Search and download academic papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic…
Find academic papers across major sources like arXiv, PubMed, bioRxiv, and more. Download PDFs whe…
Academic literature search, retrieval, and private library management on top of OpenAlex.
Federated search of books and papers, BibTeX/RIS citations, open-access retrieval and reading.
Related MCP Servers
- FlicenseAqualityCmaintenanceEnables searching, downloading, and reading academic papers from multiple platforms including arXiv, Semantic Scholar, PubMed, bioRxiv, medRxiv, IACR, Google Scholar, RePEc/IDEAS, and Sci-Hub with PDF to Markdown conversion.297-
- AlicenseAqualityCmaintenanceEnables retrieval of academic paper metadata, PDFs, full text, citations, and references by title via Semantic Scholar, arXiv, and other sources.61MIT
- AlicenseNot gradedqualityDmaintenanceEnables searching, downloading, and exporting academic papers from 20+ scholarly sources including arXiv, PubMed, and Semantic Scholar. Supports multi-source concurrent search, citation network tracing, and export to CSV, RIS, and BibTeX.1MIT
- FlicenseNot gradedqualityDmaintenanceAggregates academic paper search from multiple databases (OpenAlex, Semantic Scholar, etc.) with PDF storage and full-text search capabilities.1-