literature-bot-mcp
Provides tools for searching arXiv papers, retrieving paper metadata, downloading and parsing PDFs, and retrieving relevant text evidence from prepared papers for question answering with page citations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@literature-bot-mcp分析 arXiv 论文 2508.19294v2 的主要方法,并给出页码依据"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
文献搜索与问答机器人
一个在本地运行的文献助手:搜索 arXiv 论文、读取 PDF,并根据检索到的正文证据回答问题。通过 Streamlit 提供聊天界面,通过 MCP 将文献能力封装为 Tools、Resources 和 Prompts。
适合查找研究资料、了解论文方法和核对回答出处。当前版本面向本地使用,不是已经部署的公共在线服务。
功能
根据研究主题搜索 arXiv,查看标题、作者、摘要和论文链接。
下载并解析论文 PDF,检索与问题相关的正文片段。
使用支持工具调用的 OpenAI 兼容接口生成回答。
提供「搜索文献」「分析论文」「总结论文」三个快捷入口。
为涉及正文的关键结论提供页码引用;在回答下方展开对应页的提取文本,或打开原始 PDF。
提供可独立使用的本地 MCP Server,支持 Tools、Resources 和 Prompts。
Related MCP server: paperstack
快速开始
以下命令用于 Windows 的 Anaconda Prompt。需要已经安装 Anaconda 或 Miniconda,并能够访问模型服务及 arXiv。
1. 获取项目
在 Anaconda Prompt 中,进入你希望存放项目的位置,然后执行:
git clone https://github.com/Remus-cloud/local-mcp-literature-assistant.git
cd local-mcp-literature-assistant后续安装、配置和启动命令,都在这个项目根目录执行。
如果没有安装 Git,也可以在仓库页面选择 Code → Download ZIP,解压后在 Anaconda Prompt 中进入解压后的项目目录。
2. 创建环境并安装项目
conda create -n literature-bot python=3.12 -y
conda activate literature-bot
cd /d F:\literature-bot
python -m pip install -e .如果已创建同名环境,跳过第一条命令。安装命令会根据项目配置自动安装运行所需的软件包,无需逐个安装。
3. 配置模型接口
首次配置时执行:
copy .env.example .env
notepad .env如果已有 .env,不要重新复制覆盖,直接编辑现有文件。填写下面三个配置项,并保存文件:
DASHSCOPE_API_KEY=替换为你的API密钥
DASHSCOPE_BASE_URL=https://你的服务地址/v1
DASHSCOPE_MODEL=替换为服务商提供的模型ID这些变量名是当前代码沿用的名称,不代表必须使用阿里云。程序使用 OpenAI Python SDK 的 Chat Completions 接口,即 client.chat.completions.create(...)。
接口及模型需要支持:
OpenAI 兼容的 Chat Completions 请求和响应格式。
tools、tool_choice="auto"以及响应中的tool_calls。将工具执行结果以
tool消息返回给模型。
BASE_URL 填写服务商给出的接口基础地址,不要自行追加 /chat/completions。模型 ID 以服务商文档为准,仅支持普通文本聊天的模型不一定能运行本项目。
当前版本的兼容性注意事项: 模型请求仍包含 extra_body={"enable_thinking": False}。这是服务商扩展参数,不属于通用 OpenAI 接口字段。如果目标服务不接受该参数,需要先删除或按服务商条件传入这一参数。相关调用位于 src/literature_bot/mcp_chatbot.py 和 src/literature_bot/llm_service.py;仅替换 .env 尚不能保证兼容所有服务商。
不要把真实密钥提交到 GitHub。项目的 .gitignore 已忽略 .env,公开配置示例中只应保留占位值。
4. 启动聊天界面
在项目根目录执行:
python -m streamlit run src/literature_bot/streamlit_app.py --server.address 127.0.0.1 --server.port 8501 --browser.gatherUsageStats false浏览器打开 http://localhost:8501。如果没有自动打开,手动访问该地址即可。
使用时保持 Anaconda Prompt 窗口开启;停止程序时,在该窗口按 Ctrl+C。
以后再次启动只需:
conda activate literature-bot
cd /d F:\literature-bot
python -m streamlit run src/literature_bot/streamlit_app.py --server.address 127.0.0.1 --server.port 8501 --browser.gatherUsageStats false如何使用
可以点击快捷入口填写任务,也可以直接在页面底部提问,例如:
帮我搜索 5 篇关于 multimodal object detection 的 arXiv 论文。
请分析 arXiv 论文 2508.19294v2 的主要方法,并给出页码依据。
总结 arXiv 论文 2508.19294v2 的研究问题、主要贡献和局限。涉及具体论文时,建议提供 arXiv ID,而不是仅引用上一条消息中的「第二篇」。当前每次提问都是独立任务,页面显示之前的消息,不意味着模型具有跨问题的对话记忆。
回答包含有效页码引用时,可以在回答下方展开「查看原文」核对对应页的提取文本。PDF 链接中的页码是 PDF 文件页序号,不一定等于论文印刷页码;浏览器是否直接跳转到指定页,取决于其 PDF 阅读器。
MCP 接口
程序的主要调用关系为:
Streamlit 界面 → Chatbot(模型与工具调用循环)→ MCP Client
↓ 本地 stdio
MCP Server
↓
arXiv / PDF / 正文检索聊天界面会通过 MCP Client 启动本地 Server 子进程,正常使用时无需另外启动 Server,也无需运行 MCP Inspector。MCP Inspector 用于开发和验证接口,不是聊天界面的运行桥梁。
Tools
名称 | 用途 |
| 搜索 arXiv 论文 |
| 获取论文详情 |
| 下载、解析 PDF 并准备检索 |
| 检索与问题相关的正文证据 |
Resources
URI | 内容 |
| 服务使用指南 |
| 论文元数据 |
| 已准备论文的指定页文本 |
读取页面 Resource 前,需要在同一个 Server 会话中调用 literature_prepare_paper。页面 Resource 本身不会自动下载论文。
Prompts
search_literature:搜索文献任务模板。analyze_paper:围绕指定问题分析论文。summarize_paper:总结论文。
Prompt 是任务模板,不是另一个模型,也不会仅因获取模板就自动执行工具。
如需将 Server 接入其他 MCP 客户端,在安装项目的 Conda 环境中可使用以下启动命令:
literature-bot-mcp传输方式为 stdio。该命令等待 MCP 客户端通信,不会启动网页或出现聊天输入框。其他客户端的启动配置应使用此环境中 Python 的绝对路径,并以 -m literature_bot.mcp_integration.server 为参数,将工作目录设为项目根目录。
命令行使用
除了网页,也可以在同一环境、同一项目目录运行:
literature-bot-chat或只提问一次:
literature-bot-chat --question "帮我搜索 3 篇关于 multimodal object detection 的论文"项目结构
literature-bot/
├── src/literature_bot/ # 文献业务逻辑、Chatbot 和界面
│ ├── mcp_integration/ # MCP Server、Client、Tools、Resources、Prompts
│ ├── mcp_chatbot.py # 模型与 MCP 工具调用循环
│ ├── streamlit_app.py # 网页聊天入口
│ └── citations.py # 页码引用处理
├── notebooks/ # 开发与分步验证用的 Notebook
├── tests/ # 自动化测试
├── data/papers/ # 运行时下载的论文 PDF 缓存
├── .env.example # 不含密钥的配置示例
├── .env # 本地配置,不提交到 Git
├── .gitignore
├── pyproject.toml # 项目安装配置
└── README.md使用网页界面不需要先运行 Notebook。data/papers/ 是论文文件缓存,不是聊天记录。
运行测试
在项目根目录、已激活的 Conda 环境中执行:
python -m pip install -e ".[test]"
python -m pytest自动化测试不能替代真实服务验证。更换模型服务后,还应通过一次实际文献搜索与正文问答,确认工具调用和网络连接正常。
限制与数据说明
当前文献来源为 arXiv,不覆盖全部学术数据库。
当前版本不提供扫描 PDF 的 OCR;公式、表格和多栏排版的文本提取可能不完整。
正文问答基于检索出的证据片段,不保证已逐页审阅整篇论文;模型输出仍需人工核对。
不持久化聊天记录;页面消息用于当前浏览器会话展示。
问题和相关论文片段会发送到所配置的模型服务。调用可能产生费用,数据处理政策由相应服务商决定。
当前 MCP 使用本地 stdio,没有实现供公网访问的 HTTP 服务、登录或多用户隔离。上传到 GitHub 不等于已经部署为在线服务。
常见问题
提示找不到模块或命令
确认已执行 conda activate literature-bot,并在包含 pyproject.toml 的目录运行 python -m pip install -e .。
提示缺少配置或认证失败
检查项目根目录中的 .env、密钥、接口基础地址和模型 ID。不要在 Issue 或截图中公开密钥。
模型能聊天,但不能完成文献任务
确认接口与模型支持工具调用。如果报错指出 enable_thinking 不被支持,请按上面的兼容性说明处理。
arXiv 搜索或 PDF 下载失败
检查本机网络、代理和 arXiv 是否可访问。网络限制、服务端限流或暂时不可用都可能导致失败,稍后重试。
8501 端口已被占用
将启动命令中的 --server.port 8501 改为 --server.port 8502,然后访问 http://localhost:8502。
"# local-mcp-literature-assistant"
Available Tools
4 toolsliterature_get_paper_detailsGet arXiv Paper DetailsARead-onlyIdempotent
Return detailed metadata for one arXiv paper.
Use this for title, authors, abstract, categories, dates, and URLs. If the paper has not been seen in this session, the server resolves the ID through arXiv. It does not download the PDF.
| Name | Required | Description | Default |
|---|---|---|---|
| arxiv_id | Yes | arXiv ID such as '2508.19294v2'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| paper | Yes | |
| prepared_in_memory | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, openWorld and non-destructive, so the safety profile is free. The description adds genuinely new behavior: the server performs a live arXiv resolution when the paper was not seen this session, and it does not download the PDF — a useful scope boundary and a latency/network hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, then field list, then behavioral caveats. No filler, though the field enumeration is somewhat redundant with the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers the notable non-obvious behavior (session-aware arXiv resolution, no PDF). The only gap is not situating itself against the three sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is documented with a concrete example ('2508.19294v2'), so the schema carries the semantics. The description adds nothing about ID format, versioning, or invalid-ID behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return detailed metadata for one arXiv paper') and enumerates the fields returned (title, authors, abstract, categories, dates, URLs). This distinguishes it implicitly from the search sibling by scope ('one paper'), though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this for title, authors, abstract, categories, dates, and URLs' gives a context of use, but there is no when-not guidance and no routing to literature_search_papers or literature_retrieve_paper_evidence, which an agent choosing among four siblings would benefit from.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
literature_prepare_paperPrepare Paper for RetrievalAIdempotent
Download, extract, and chunk one paper for evidence retrieval.
The PDF is cached in the configured local papers directory and is not overwritten. Extracted pages and chunks stay only in server memory. Repeating the call is safe and reuses the cache.
| Name | Required | Description | Default |
|---|---|---|---|
| arxiv_id | Yes | arXiv ID returned by literature_search_papers. |
Output Schema
| Name | Required | Description |
|---|---|---|
| title | Yes | |
| arxiv_id | Yes | |
| pdf_path | Yes | |
| page_count | Yes | |
| chunk_count | Yes | |
| nonempty_page_count | Yes | |
| reused_memory_cache | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that the PDF is cached in a configured local directory and never overwritten, that extracted pages/chunks live only in server memory, and that repeat calls are safe and reuse the cache. This substantiates the idempotentHint=true and readOnlyHint=false (disk cache write) annotations with concrete mechanics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the action front-loaded and secondary behavioral facts (cache location, memory-only chunks, idempotency) following. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, output-schema-backed preparation tool, the description covers the action, the caching side effect, the memory footprint, and repeat-call safety. An agent has everything needed to invoke it correctly and predict its side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single arxiv_id parameter is well documented in the schema ('arXiv ID returned by literature_search_papers'). The description adds no syntax, format, or validation detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific three-step verb chain (download, extract, chunk) on a specific resource (one paper) with a stated purpose (evidence retrieval). An agent can distinguish this preparation step from literature_search_papers and literature_retrieve_paper_evidence without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for evidence retrieval' implies this is a prerequisite for literature_retrieve_paper_evidence, giving implied usage. However, no sibling is named explicitly, and there is no statement of when not to call it (e.g. already-prepared papers) or which tool must precede it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
literature_retrieve_paper_evidenceRetrieve Evidence from a PaperAIdempotent
Retrieve page-cited evidence chunks from one paper.
Use this before answering a question about paper content. The query must be concise English keywords because retrieval uses local BM25 rather than a model. If necessary, the paper is prepared automatically. Return the page numbers with any answer built from these results.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Concise English keywords. Translate non-English questions before calling. | |
| top_k | No | Return between 1 and 8 chunks. | |
| arxiv_id | Yes | arXiv ID identifying the paper to search. |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | |
| title | Yes | |
| arxiv_id | Yes | |
| evidence | Yes | |
| auto_prepared | Yes | |
| evidence_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real mechanics beyond the annotations: retrieval is local BM25 rather than a model, the paper may be prepared as a side effect, and callers should carry page numbers into the answer. The auto-preparation side effect is consistent with readOnlyHint=false (a read tool that can trigger a write), so there is no contradiction, though failure modes (paper not found, no hits) go unmentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the action and the pre-condition, then the query-format constraint, then the side effect and the citation obligation. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations present and an output schema covering the return shape, the description supplies the missing operational context: when to call it, how to phrase the query, and the auto-preparation behavior. It omits edge cases such as an unknown arxiv_id or an empty result set, which keeps it just short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description earns credit by explaining *why* the query must be short English keywords (BM25, not a model), which is the rationale the schema omits. The top_k parameter is not addressed in prose beyond the schema's own range documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: retrieving page-cited evidence chunks, scoped to 'one paper', which implicitly separates it from the corpus-wide literature_search_papers sibling. It never names that sibling, so the differentiation is inferable rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this before answering a question about paper content" gives a clear trigger condition, and "the paper is prepared automatically" tells the agent it need not call literature_prepare_paper first. No explicit when-not or named alternative is given, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
literature_search_papersSearch arXiv PapersARead-onlyIdempotent
Search arXiv and return concise paper metadata.
Use this first when the user wants literature recommendations or has not supplied an arXiv ID. It does not download PDFs. The returned arXiv IDs can be passed to the other literature tools.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | English arXiv search query. | |
| max_results | No | Return between 1 and 10 papers. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| query | Yes | |
| papers | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered structurally. The description still earns credit for adding non-annotation context: no PDF download, concise metadata output, and that the returned arXiv IDs are reusable inputs for the sibling tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. The core capability is front-loaded and the routing/exclusion notes follow in priority order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description still sketches what comes back (IDs, metadata) and how to chain it. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both query and max_results documented in the schema (including bounds and default), so the baseline is 3. The description adds no syntax, language, or formatting detail for the query parameter beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Search arXiv") plus the return shape ("concise paper metadata"). It also positions itself against siblings by telling the agent to use it first when no arXiv ID is supplied, so an agent can distinguish it from literature_get_paper_details and friends without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rule: "Use this first when the user wants literature recommendations or has not supplied an arXiv ID," plus a clear negative ("It does not download PDFs") and a hand-off path (returned IDs pass to the other literature tools). The only gap is that the alternatives are referred to generically rather than named as literature_get_paper_details / literature_prepare_paper.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.0- First observed
literature_get_paper_details - First observed
literature_prepare_paper - First observed
literature_retrieve_paper_evidence - First observed
literature_search_papers
TDQS
Scored across 4 tools
Each tool targets a distinct stage: search, metadata details, preparation/download, and evidence retrieval. The only mild overlap is prepare_paper vs retrieve_paper_evidence, since retrieve auto-prepares when needed, but descriptions clarify the distinction well.
All four tools follow a strict literature_verb_noun pattern (search_papers, get_paper_details, prepare_paper, retrieve_paper_evidence), making the namespace and conventions fully predictable.
Four tools is slightly lean but well-matched to a focused paper search-and-retrieval server. Each tool earns its place with a distinct role in the pipeline.
The surface covers a coherent lifecycle: search, inspect metadata, prepare, and retrieve cited evidence. Minor gaps exist (no listing/clearing of the local cache), but core literature workflows are fully supported.
Maintenance
Related MCP Connectors
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Search arXiv, fetch paper metadata, and read full-text content.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Extract papers from ArXiv — titles, abstracts, authors, categories & PDF links. Monitor new AI, phys
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables searching, downloading, and managing academic papers from arXiv.org through natural language interactions. Provides tools for paper discovery, PDF downloads, and local paper collection management.41MIT
- AlicenseNot gradedqualityDmaintenanceEnables arXiv paper search, PDF download, text extraction, and context chunking for LLM pipelines, along with advanced features like citation graphs and reproducibility scoring.2MIT
- AlicenseNot gradedqualityDmaintenanceEnables searching arXiv, fetching metadata, reading papers as section-aware Markdown, listing recent papers, and downloading PDFs via five MCP tools.16 npm2MIT
- FlicenseNot gradedqualityCmaintenanceA local, rule-based MCP server for searching and analyzing academic papers from arXiv. Enables paper search, ranking, smart summarization, keyword extraction, and citation generation without API keys or LLM calls.-