arxiv-zh
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@arxiv-zhTranslate https://arxiv.org/abs/1706.03762 to Simplified Chinese and compile the PDF."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
arxivZH MCP
把 arXiv、DOI 或论文网页链接对应的 LaTeX 源码批量译为简体中文,并用 XeLaTeX 编译出中文 PDF。
固定使用 DeepSeek 官方 deepseek-flash;多篇论文并行处理,单篇内部串行,支持术语表、结构校验、断点续跑和编译 QA。
输出固定在 <work_dir>/arxivZH/:
<work_dir>/arxivZH/
<arxiv_id>v<version>/ # 中文 .tex 源码、glossary.json、manifest.json、编译日志与页面预览
PDFzh/<arxiv_id>v<version>.zh.pdf
.jobs/<job_id>.json # 任务记录(进度、用量、逐篇结果)无 TeX 源码的论文会跳过(回退查 arXiv),不做 PDF 文本翻译。
环境要求
macOS(密钥存于 macOS Keychain,其他系统需改
arxiv_zh/network.py中的密钥读取)Python >= 3.11、uv、网络可访问
api.deepseek.comMacTeX(ctex / Fandol 字体 / XeLaTeX / latexmk)、Poppler(pdftoppm / pdftotext)
一个 DeepSeek API 密钥(仅查询模型列表不收费,翻译按 token 计费)
Related MCP server: Paper Search MCP Server
安装
方式一:一键安装(macOS + Codex CLI)
git clone https://github.com/FjkFjkFjk314/arxiv-zh-mcp.git
cd arxiv-zh-mcp
python3 install.pyinstall.py 会:复制运行时到 ~/.local/share/arxiv-zh-mcp/ 并 uv sync;
把 skill 装到 ~/.agents/skills/arxiv-zh/;向 Codex 注册 MCP(不改动其他 MCP 配置)。
它不读写任何 API 密钥。
方式二:手动安装(任意 MCP 客户端)
git clone https://github.com/FjkFjkFjk314/arxiv-zh-mcp.git
cd arxiv-zh-mcp
uv sync # 创建 .venv 并安装依赖
chmod +x run.sh setup-keychain.command然后按客户端注册 stdio MCP,启动命令为仓库内的 run.sh 绝对路径:
Codex:
codex mcp add arxiv-zh -- /绝对路径/arxiv-zh-mcp/run.shClaude Code:
claude mcp add arxiv-zh /绝对路径/arxiv-zh-mcp/run.shClaude Desktop / Cursor(JSON 配置):
{
"mcpServers": {
"arxiv-zh": {
"command": "/绝对路径/arxiv-zh-mcp/run.sh"
}
}
}注册后新开一个会话加载工具。
配置 API 密钥
macOS 上双击(或终端执行)安装目录里的 setup-keychain.command,
在隐藏输入中粘贴 DeepSeek 密钥,保存到 Keychain 的 arxiv-zh-deepseek / api-key 条目。
密钥不作为 CLI/MCP 参数,不写入任何配置或日志。配置前可先跑:
.venv/bin/python cli.py status # 本机依赖与 Keychain 检查
.venv/bin/python cli.py status --check-api # 附加查询模型权限(不产生翻译请求)MCP 工具
工具 | 说明 |
| 检查依赖与密钥配置; |
| 提交 1–50 个链接,后台并行翻译、排队编译,返回 |
| 查询进度、逐篇结果、PDF 路径、编译日志、页面预览;不发起翻译 |
| 续跑失败/中断任务,复用已验证译文缓存;可能产生新费用 |
关键约定(写进给 agent 的 skill,见 skill/arxiv-zh/SKILL.md):
work_dir必须显式传用户当前工作目录的绝对路径,不能用 MCP 安装目录。提交后每 15–30 秒轮询同一
job_id;不因等待或超时而重新提交。完成后要实际查看返回的
qa.previews页面图像;visual_review: pending不算版式验收。失败先读错误与日志,再对同一
job_id续跑;中断块可能重复计费,不要无上限重试。
给其他 agent 的一句话用法
翻译论文:先
arxiv_zh_status()确认就绪,再translate_papers(links=[...], work_dir="<当前工作目录绝对路径>"), 保存返回的job_id并每 15–30 秒调get_translation()轮询,完成后打开返回的 PDF 路径与预览图验收。
并行限制
按 2026-09-23 官方并发定义,
deepseek-flash额度为 2500 个同时在途请求(不是 RPM),本机取 60%,即 1500。同一系统用户的所有 arxivZH 工作目录共用 API 请求槽位;API 请求完成或异常退出会释放槽位,重试退避和等待编译不占用 API 槽位。其他程序或其他电脑使用同一账号的请求不在本机控制范围内。
每批最多 50 篇、每篇同时最多 1 个 API 请求,所以单批实际最多 50 路;同一工作目录仍只允许一个活动任务。上限不会预先发起空请求。
论文术语表、译文缓存和用量彼此独立;结果保持输入顺序,
progress.papers提供按输入序号区分的进度,progress.finished提供已结束条目数。PDF 编译在本机串行排队,不占 API 槽位。模型额度为经核对的固定配置;不会自动抓取网页更改上限或更换模型。
可通过 arxiv_zh_status().concurrency 查看当前配置。一次提交全部待译论文即可启用论文级并行,无需逐篇创建任务。链接解析按输入顺序进行,已解析论文随即进入处理;arXiv 下载继续遵守跨任务约 3 秒的请求间隔。
实测范围(2026-09-23):已用已安装运行时并行翻译两份短 TeX 样例,记录到真实 DeepSeek HTTP 请求重叠;6 次生成请求均返回 HTTP 200,两份中文 PDF 均编译成功并通过页面检查。此次验证了 2 路实际并发,未对 50 路或 1500 路进行线上满载测试。详见 验证记录。
使用 Skill
skill/arxiv-zh/ 是完整的 agent skill(SKILL.md + agents 配置),按你的 agent 的 skill 机制放置即可:
通用约定:
cp -R skill/arxiv-zh ~/.agents/skills/arxiv-zh(项目级则放<项目>/.agents/skills/)Codex / Kimi Code 等识别
~/.agents/skills/的客户端放这里即可
之后对 agent 说:
使用 arxiv-zh 翻译 https://arxiv.org/abs/1706.03762,放到当前工作目录。
CLI(不依赖 MCP,推荐作为兜底通道)
MCP 工具未加载时可直接用 CLI,功能等价:
.venv/bin/python cli.py status [--check-api]
.venv/bin/python cli.py translate --work-dir /绝对/工作目录 <链接...> [--main-tex 相对路径]
.venv/bin/python cli.py get --work-dir /绝对/工作目录 <job_id>
.venv/bin/python cli.py resume --work-dir /绝对/工作目录 <job_id>测试
uv sync --dev
.venv/bin/python -m pytest -q # 离线测试包含模拟 API 的真实 XeLaTeX 编译流水线、断点续跑、公式/引用保护、恶意压缩包防护、 DOI/标题匹配歧义、论文级并行、跨进程限流、异常退出后释放槽位和 MCP stdio 握手。 2026-09-23 共 46 项测试通过。离线测试不证明 DeepSeek 账户或翻译质量; 真实在线验证记录见 VERIFICATION.md。
行为与边界
官方模型固定为
deepseek-flash,Chat Completions 非思考模式。密钥未配置或模型不可用时明确报错,不替换模型。术语表先行:按论文上下文采用国内学术界通行译名,全文一致;数学、代码、作者姓名、参考文献著录、专名与图片保持原样。数学环境和图片内的英文不翻译。自定义宏作为 TeX 代码整体保护。
不保留英文 TeX 或原始压缩包;完成的文件原位替换为中文,
.bib/.bbl/.sty/.cls和图片保留,无original/副本。分块保护结构、公式、路径和引用,校验标记顺序与数值,失败最多重试一次;通过的块按内容哈希缓存,文件以哈希 + 原子写入记进度;检测到外部修改时停止覆盖。
网络断连不自动重复付费请求;手动续跑可能重做未确认的当前块。
一批最多 50 篇,论文之间并行,单篇内部的文件与分块串行;arXiv 请求跨任务间隔约 3 秒;单篇失败不阻塞其他篇;相同链接列表重复提交返回原任务。
DOI 优先精确匹配 arXiv,标题回退要求高相似度并核对作者;歧义即跳过。普通 PDF 链接只尝试文献元数据解析,不做 PDF 文本翻译。
解包拒绝路径穿越、符号链接和超限文件;编译禁用 shell escape 与用户 latexmkrc,限制 TeX 文件读写。仍应使用可信论文源;本工具不是通用 TeX 沙箱。
成功 PDF 必须通过中文文本、缺字、未解析引用检查,并生成页面预览与溢出警告。
参考
许可
MIT,见 LICENSE。
Available Tools
4 toolsarxiv_zh_statusARead-only
检查本机依赖和 Keychain 配置。check_api=true 查询官方模型列表,不进行付费翻译,不返回密钥。
| Name | Required | Description | Default |
|---|---|---|---|
| check_api | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/open-world/non-destructive profile, and the description adds meaningful extra context: no paid translation is performed and secrets are not returned. That is genuinely useful safety disclosure beyond the annotations, though it says nothing about output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the primary purpose and then the flag behavior and side-effect guarantees. No filler, though the terse telegraphic style borders on under-explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple diagnostic tool with no output schema and one boolean parameter, the description covers purpose, the parameter's effect, and the side-effect boundaries. Adequate for correct invocation, with only return-shape details missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one parameter at 0% schema coverage, the description must carry the load, and it does explain that check_api=true queries the official model list while the default leaves it off. It does not describe latency or cost implications of that call, but the core semantics are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource pair: checks local dependencies and Keychain configuration. This clearly separates it from the translation siblings (translate_papers, get_translation, resume_translation) without needing to name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage as a preflight/diagnostic status check, and it clarifies what check_api=true does. However, it never states when an agent should call this vs. jumping into translation, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_translationARead-only
查询已有任务进度、逐篇结果、PDF 路径、编译日志和页面预览。不会发起翻译。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| work_dir | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint and destructiveHint annotations already establish a safe read-only profile, but the description adds useful context by enumerating the returned data types (progress, per-paper results, PDF paths, logs, previews) and explicitly clarifying that the tool does not initiate translation. It does not cover pagination, authentication, or failure modes, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the queryable outputs and followed by the single behavioral exclusion. Every phrase earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully lists what data can be retrieved, and annotations cover the safety profile. However, the complete absence of parameter guidance for two required inputs (work_dir and job_id) leaves a notable gap for an agent trying to construct a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain either required parameter. There is a vague implicit connection to '已有任务' (existing task) that hints at job_id, but work_dir is completely unaddressed, leaving an agent without guidance on parameter meaning or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (查询) and lists the exact resources returned: task progress, per-paper results, PDF paths, compilation logs, and page previews. It also distinguishes from translation-starting tools with '不会发起翻译', so an agent can tell what this tool does and does not do without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use this tool (query existing task information) and when not to use it (will not initiate translation), which implicitly routes the agent to translate_papers for starting work. It does not explicitly name alternatives like resume_translation or arxiv_zh_status, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_translationAIdempotent
续跑失败或中断的同一任务,复用通过校验的译文缓存。可能产生新的 API 费用。
已运行的任务只返回进度。可为单篇失败论文补充 main_tex;不切换模型、不自动覆盖用户改动。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| main_tex | No | ||
| work_dir | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds context beyond annotations: may incur new API costs, does not switch models, does not automatically overwrite user changes, and already-running tasks return progress only. These behavior disclosures are genuinely useful; annotations already cover the read-only/idempotent profile, which the cache-reuse framing is consistent with.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, each carrying distinct information (purpose, cost, running-task behavior, parameter caveat, non-changes). Front-loaded with the core action. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description does well on cost, safety-relevant non-changes, and running-task behavior. The remaining gap is parameter documentation (0% schema coverage), but overall the agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It partially explains main_tex ('可为单篇失败论文补充 main_tex' - supplement main_tex for a single failed paper) but says nothing about job_id or work_dir, leaving two of three parameters undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: resume a failed/interrupted translation task, reusing the validated cache. It is clearly a continuation operation distinct from translate_papers, though it does not name the sibling tools explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use condition (failed or interrupted same task) and an implicit when-not: '已运行的任务只返回进度' tells the agent running tasks only report progress. It stops short of naming an alternative tool for a fresh start (translate_papers).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translate_papersAIdempotent
提交 1–50 个论文链接,会调用付费 DeepSeek API。后台批量翻译并编译,返回 job_id。
work_dir 为用户当前工作目录绝对路径。main_tex 仅在单篇论文有多个主文件时指定相对路径。 同一目录同一链接列表不会重复启动。部分失败继续处理其他论文。
| Name | Required | Description | Default |
|---|---|---|---|
| links | Yes | ||
| main_tex | No | ||
| work_dir | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=true), yet the description adds genuinely new behavioral facts: it invokes a paid DeepSeek API (cost), runs as a background async batch (returns job_id), deduplicates identical submissions, and continues on partial failure. This exceeds what annotations alone provide, though it omits rough cost/scale expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and return value come first, followed by parameter notes and then behavioral caveats — a sensible front-loading order. Every sentence carries information, with only minor density from mixed-language technical terms.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description states the return (job_id) and the async model, and the presence of resume_translation/arxiv_zh_status implies a pollable job lifecycle. It is close to complete for an async submission tool; a note on how to retrieve results would close the loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description carries the burden and largely does: work_dir is defined as the absolute path of the current working directory, and main_tex as a relative path used only when a single paper has multiple main files. links is only implied by '1–50 paper links' rather than defined explicitly, leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (submit 1–50 paper links for batch translation and compilation) and an outcome (returns job_id), which distinguishes it from siblings like get_translation and resume_translation. It doesn't explicitly name those siblings, so sibling differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys operational context — a 1–50 link batch, deduplication on identical directory/link lists, and continued processing on partial failure. However, it never explains when to choose this over resume_translation or how it relates to arxiv_zh_status/get_translation, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
arxiv_zh_status - First observed
get_translation - First observed
resume_translation - First observed
translate_papers
TDQS
Scored across 4 tools
Each tool has a distinct primary role: environment check, start translation, query progress, and resume failed tasks. Minor overlap exists because resume_translation on an already-running task only returns progress, similar to get_translation, but the intended use cases are clearly differentiated.
Three tools follow a verb_noun snake_case pattern (translate_papers, get_translation, resume_translation), while arxiv_zh_status uses a server-prefixed noun phrase. This is a minor deviation but the set remains readable and consistently cased.
Four tools map cleanly to the essential operations of the translation workflow—setup check, submission, monitoring, and recovery. The count is well-scoped with no redundant tools.
The surface covers the core lifecycle: check environment, submit translation, query results, and resume interrupted jobs. Minor gaps include no cancel/delete operation and no way to list all existing jobs if the job_id is lost, but these are workable in most cases.
Maintenance
Related MCP Connectors
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Persistent AI LaTeX workspace: edit and compile multi-file projects, export publication-ready PDFs.
Edit your Overleaf LaTeX projects from Claude and ChatGPT; every change is a real Git commit.
A hosted LaTeX editor your assistant can actually use. Search 1,019 free templates, create and edit projects, and compile them to PDF on a real TeX Live farm, getting back the PDF or the compile log when a build fails. Thirteen tools behind OAuth 2.1, with nothing to install and nothing to run locally.
Related MCP Servers
- FlicenseBqualityNot gradedmaintenanceEnables extraction of mathematical content from TeX papers and conversion to Lean code through a structured intermediate representation. Supports project scaffolding, entity management, and task tracking for mathematical formalization workflows.14-
- FlicenseAqualityBmaintenanceEnables searching, downloading, and reading academic papers from multiple platforms including arXiv, Semantic Scholar, PubMed, bioRxiv, medRxiv, IACR, Google Scholar, RePEc/IDEAS, and Sci-Hub with PDF to Markdown conversion.297-
- FlicenseAqualityFmaintenanceEnables AI assistants to create, edit, and validate LaTeX documents through a standardized protocol with support for multiple document classes and package management. It provides tools for document structure analysis and file organization to streamline the generation of professional academic papers.819-
- AlicenseNot gradedqualityDmaintenanceEnables searching, downloading, and exporting academic papers from 20+ scholarly sources including arXiv, PubMed, and Semantic Scholar. Supports multi-source concurrent search, citation network tracing, and export to CSV, RIS, and BibTeX.1MIT