Skip to main content
Glama

arxivZH MCP

把 arXiv、DOI 或论文网页链接对应的 LaTeX 源码批量译为简体中文,并用 XeLaTeX 编译出中文 PDF。 固定使用 DeepSeek 官方 deepseek-flash;多篇论文并行处理,单篇内部串行,支持术语表、结构校验、断点续跑和编译 QA。

输出固定在 <work_dir>/arxivZH/:

<work_dir>/arxivZH/
  <arxiv_id>v<version>/    # 中文 .tex 源码、glossary.json、manifest.json、编译日志与页面预览
  PDFzh/<arxiv_id>v<version>.zh.pdf
  .jobs/<job_id>.json     # 任务记录(进度、用量、逐篇结果)

无 TeX 源码的论文会跳过(回退查 arXiv),不做 PDF 文本翻译。

环境要求

  • macOS(密钥存于 macOS Keychain,其他系统需改 arxiv_zh/network.py 中的密钥读取)

  • Python >= 3.11、uv、网络可访问 api.deepseek.com

  • MacTeX(ctex / Fandol 字体 / XeLaTeX / latexmk)、Poppler(pdftoppm / pdftotext)

  • 一个 DeepSeek API 密钥(仅查询模型列表不收费,翻译按 token 计费)

Related MCP server: Paper Search MCP Server

安装

方式一:一键安装(macOS + Codex CLI)

git clone https://github.com/FjkFjkFjk314/arxiv-zh-mcp.git
cd arxiv-zh-mcp
python3 install.py

install.py 会:复制运行时到 ~/.local/share/arxiv-zh-mcp/ 并 uv sync; 把 skill 装到 ~/.agents/skills/arxiv-zh/;向 Codex 注册 MCP(不改动其他 MCP 配置)。 它不读写任何 API 密钥。

方式二:手动安装(任意 MCP 客户端)

git clone https://github.com/FjkFjkFjk314/arxiv-zh-mcp.git
cd arxiv-zh-mcp
uv sync                          # 创建 .venv 并安装依赖
chmod +x run.sh setup-keychain.command

然后按客户端注册 stdio MCP,启动命令为仓库内的 run.sh 绝对路径:

  • Codex:codex mcp add arxiv-zh -- /绝对路径/arxiv-zh-mcp/run.sh

  • Claude Code:claude mcp add arxiv-zh /绝对路径/arxiv-zh-mcp/run.sh

  • Claude Desktop / Cursor(JSON 配置):

{
  "mcpServers": {
    "arxiv-zh": {
      "command": "/绝对路径/arxiv-zh-mcp/run.sh"
    }
  }
}

注册后新开一个会话加载工具。

配置 API 密钥

macOS 上双击(或终端执行)安装目录里的 setup-keychain.command, 在隐藏输入中粘贴 DeepSeek 密钥,保存到 Keychain 的 arxiv-zh-deepseek / api-key 条目。 密钥不作为 CLI/MCP 参数,不写入任何配置或日志。配置前可先跑:

.venv/bin/python cli.py status            # 本机依赖与 Keychain 检查
.venv/bin/python cli.py status --check-api  # 附加查询模型权限(不产生翻译请求)

MCP 工具

工具

说明

arxiv_zh_status(check_api=false)

检查依赖与密钥配置;check_api=true 查模型列表,不收费

translate_papers(links, work_dir, main_tex=null)

提交 1–50 个链接,后台并行翻译、排队编译,返回 job_id。会产生 API 费用

get_translation(work_dir, job_id)

查询进度、逐篇结果、PDF 路径、编译日志、页面预览;不发起翻译

resume_translation(work_dir, job_id, main_tex=null)

续跑失败/中断任务,复用已验证译文缓存;可能产生新费用

关键约定(写进给 agent 的 skill,见 skill/arxiv-zh/SKILL.md):

  • work_dir 必须显式传用户当前工作目录的绝对路径,不能用 MCP 安装目录。

  • 提交后每 15–30 秒轮询同一 job_id;不因等待或超时而重新提交。

  • 完成后要实际查看返回的 qa.previews 页面图像;visual_review: pending 不算版式验收。

  • 失败先读错误与日志,再对同一 job_id 续跑;中断块可能重复计费,不要无上限重试。

给其他 agent 的一句话用法

翻译论文:先 arxiv_zh_status() 确认就绪,再 translate_papers(links=[...], work_dir="<当前工作目录绝对路径>"), 保存返回的 job_id 并每 15–30 秒调 get_translation() 轮询,完成后打开返回的 PDF 路径与预览图验收。

并行限制

  • 按 2026-09-23 官方并发定义,deepseek-flash 额度为 2500 个同时在途请求(不是 RPM),本机取 60%,即 1500。

  • 同一系统用户的所有 arxivZH 工作目录共用 API 请求槽位;API 请求完成或异常退出会释放槽位,重试退避和等待编译不占用 API 槽位。其他程序或其他电脑使用同一账号的请求不在本机控制范围内。

  • 每批最多 50 篇、每篇同时最多 1 个 API 请求,所以单批实际最多 50 路;同一工作目录仍只允许一个活动任务。上限不会预先发起空请求。

  • 论文术语表、译文缓存和用量彼此独立;结果保持输入顺序,progress.papers 提供按输入序号区分的进度,progress.finished 提供已结束条目数。

  • PDF 编译在本机串行排队,不占 API 槽位。模型额度为经核对的固定配置;不会自动抓取网页更改上限或更换模型。

可通过 arxiv_zh_status().concurrency 查看当前配置。一次提交全部待译论文即可启用论文级并行,无需逐篇创建任务。链接解析按输入顺序进行,已解析论文随即进入处理;arXiv 下载继续遵守跨任务约 3 秒的请求间隔。

实测范围(2026-09-23):已用已安装运行时并行翻译两份短 TeX 样例,记录到真实 DeepSeek HTTP 请求重叠;6 次生成请求均返回 HTTP 200,两份中文 PDF 均编译成功并通过页面检查。此次验证了 2 路实际并发,未对 50 路或 1500 路进行线上满载测试。详见 验证记录。

使用 Skill

skill/arxiv-zh/ 是完整的 agent skill(SKILL.md + agents 配置),按你的 agent 的 skill 机制放置即可:

  • 通用约定:cp -R skill/arxiv-zh ~/.agents/skills/arxiv-zh(项目级则放 <项目>/.agents/skills/)

  • Codex / Kimi Code 等识别 ~/.agents/skills/ 的客户端放这里即可

之后对 agent 说:

使用 arxiv-zh 翻译 https://arxiv.org/abs/1706.03762,放到当前工作目录。

CLI(不依赖 MCP,推荐作为兜底通道)

MCP 工具未加载时可直接用 CLI,功能等价:

.venv/bin/python cli.py status [--check-api]
.venv/bin/python cli.py translate --work-dir /绝对/工作目录 <链接...> [--main-tex 相对路径]
.venv/bin/python cli.py get --work-dir /绝对/工作目录 <job_id>
.venv/bin/python cli.py resume --work-dir /绝对/工作目录 <job_id>

测试

uv sync --dev
.venv/bin/python -m pytest -q        # 离线测试

包含模拟 API 的真实 XeLaTeX 编译流水线、断点续跑、公式/引用保护、恶意压缩包防护、 DOI/标题匹配歧义、论文级并行、跨进程限流、异常退出后释放槽位和 MCP stdio 握手。 2026-09-23 共 46 项测试通过。离线测试不证明 DeepSeek 账户或翻译质量; 真实在线验证记录见 VERIFICATION.md。

行为与边界

  • 官方模型固定为 deepseek-flash,Chat Completions 非思考模式。密钥未配置或模型不可用时明确报错,不替换模型。

  • 术语表先行:按论文上下文采用国内学术界通行译名,全文一致;数学、代码、作者姓名、参考文献著录、专名与图片保持原样。数学环境和图片内的英文不翻译。自定义宏作为 TeX 代码整体保护。

  • 不保留英文 TeX 或原始压缩包;完成的文件原位替换为中文,.bib/.bbl/.sty/.cls 和图片保留,无 original/ 副本。

  • 分块保护结构、公式、路径和引用,校验标记顺序与数值,失败最多重试一次;通过的块按内容哈希缓存,文件以哈希 + 原子写入记进度;检测到外部修改时停止覆盖。

  • 网络断连不自动重复付费请求;手动续跑可能重做未确认的当前块。

  • 一批最多 50 篇,论文之间并行,单篇内部的文件与分块串行;arXiv 请求跨任务间隔约 3 秒;单篇失败不阻塞其他篇;相同链接列表重复提交返回原任务。

  • DOI 优先精确匹配 arXiv,标题回退要求高相似度并核对作者;歧义即跳过。普通 PDF 链接只尝试文献元数据解析,不做 PDF 文本翻译。

  • 解包拒绝路径穿越、符号链接和超限文件;编译禁用 shell escape 与用户 latexmkrc,限制 TeX 文件读写。仍应使用可信论文源;本工具不是通用 TeX 沙箱。

  • 成功 PDF 必须通过中文文本、缺字、未解析引用检查,并生成页面预览与溢出警告。

参考

许可

MIT,见 LICENSE。

Available Tools

4 tools
arxiv_zh_statusA
Read-only

检查本机依赖和 Keychain 配置。check_api=true 查询官方模型列表,不进行付费翻译,不返回密钥。

ParametersJSON Schema
NameRequiredDescriptionDefault
check_apiNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/open-world/non-destructive profile, and the description adds meaningful extra context: no paid translation is performed and secrets are not returned. That is genuinely useful safety disclosure beyond the annotations, though it says nothing about output shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the primary purpose and then the flag behavior and side-effect guarantees. No filler, though the terse telegraphic style borders on under-explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple diagnostic tool with no output schema and one boolean parameter, the description covers purpose, the parameter's effect, and the side-effect boundaries. Adequate for correct invocation, with only return-shape details missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With one parameter at 0% schema coverage, the description must carry the load, and it does explain that check_api=true queries the official model list while the default leaves it off. It does not describe latency or cost implications of that call, but the core semantics are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource pair: checks local dependencies and Keychain configuration. This clearly separates it from the translation siblings (translate_papers, get_translation, resume_translation) without needing to name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implied usage as a preflight/diagnostic status check, and it clarifies what check_api=true does. However, it never states when an agent should call this vs. jumping into translation, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_translationA
Read-only

查询已有任务进度、逐篇结果、PDF 路径、编译日志和页面预览。不会发起翻译。

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
work_dirYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint and destructiveHint annotations already establish a safe read-only profile, but the description adds useful context by enumerating the returned data types (progress, per-paper results, PDF paths, logs, previews) and explicitly clarifying that the tool does not initiate translation. It does not cover pagination, authentication, or failure modes, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the queryable outputs and followed by the single behavioral exclusion. Every phrase earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully lists what data can be retrieved, and annotations cover the safety profile. However, the complete absence of parameter guidance for two required inputs (work_dir and job_id) leaves a notable gap for an agent trying to construct a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not explain either required parameter. There is a vague implicit connection to '已有任务' (existing task) that hints at job_id, but work_dir is completely unaddressed, leaving an agent without guidance on parameter meaning or format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (查询) and lists the exact resources returned: task progress, per-paper results, PDF paths, compilation logs, and page previews. It also distinguishes from translation-starting tools with '不会发起翻译', so an agent can tell what this tool does and does not do without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use this tool (query existing task information) and when not to use it (will not initiate translation), which implicitly routes the agent to translate_papers for starting work. It does not explicitly name alternatives like resume_translation or arxiv_zh_status, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_translationA
Idempotent

续跑失败或中断的同一任务,复用通过校验的译文缓存。可能产生新的 API 费用。

已运行的任务只返回进度。可为单篇失败论文补充 main_tex;不切换模型、不自动覆盖用户改动。

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
main_texNo
work_dirYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds context beyond annotations: may incur new API costs, does not switch models, does not automatically overwrite user changes, and already-running tasks return progress only. These behavior disclosures are genuinely useful; annotations already cover the read-only/idempotent profile, which the cache-reuse framing is consistent with.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, each carrying distinct information (purpose, cost, running-task behavior, parameter caveat, non-changes). Front-loaded with the core action. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description does well on cost, safety-relevant non-changes, and running-task behavior. The remaining gap is parameter documentation (0% schema coverage), but overall the agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It partially explains main_tex ('可为单篇失败论文补充 main_tex' - supplement main_tex for a single failed paper) but says nothing about job_id or work_dir, leaving two of three parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: resume a failed/interrupted translation task, reusing the validated cache. It is clearly a continuation operation distinct from translate_papers, though it does not name the sibling tools explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use condition (failed or interrupted same task) and an implicit when-not: '已运行的任务只返回进度' tells the agent running tasks only report progress. It stops short of naming an alternative tool for a fresh start (translate_papers).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

translate_papersA
Idempotent

提交 1–50 个论文链接,会调用付费 DeepSeek API。后台批量翻译并编译,返回 job_id。

work_dir 为用户当前工作目录绝对路径。main_tex 仅在单篇论文有多个主文件时指定相对路径。 同一目录同一链接列表不会重复启动。部分失败继续处理其他论文。

ParametersJSON Schema
NameRequiredDescriptionDefault
linksYes
main_texNo
work_dirYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=true), yet the description adds genuinely new behavioral facts: it invokes a paid DeepSeek API (cost), runs as a background async batch (returns job_id), deduplicates identical submissions, and continues on partial failure. This exceeds what annotations alone provide, though it omits rough cost/scale expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and return value come first, followed by parameter notes and then behavioral caveats — a sensible front-loading order. Every sentence carries information, with only minor density from mixed-language technical terms.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description states the return (job_id) and the async model, and the presence of resume_translation/arxiv_zh_status implies a pollable job lifecycle. It is close to complete for an async submission tool; a note on how to retrieve results would close the loop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description carries the burden and largely does: work_dir is defined as the absolute path of the current working directory, and main_tex as a relative path used only when a single paper has multiple main files. links is only implied by '1–50 paper links' rather than defined explicitly, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (submit 1–50 paper links for batch translation and compilation) and an outcome (returns job_id), which distinguishes it from siblings like get_translation and resume_translation. It doesn't explicitly name those siblings, so sibling differentiation is implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It conveys operational context — a 1–50 link batch, deduplication on identical directory/link lists, and continued processing on partial failure. However, it never explains when to choose this over resume_translation or how it relates to arxiv_zh_status/get_translation, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedarxiv_zh_status
    • First observedget_translation
    • First observedresume_translation
    • First observedtranslate_papers

TDQS

A3.9/5.0

Scored across 4 tools

Disambiguation4/5

Each tool has a distinct primary role: environment check, start translation, query progress, and resume failed tasks. Minor overlap exists because resume_translation on an already-running task only returns progress, similar to get_translation, but the intended use cases are clearly differentiated.

Naming Consistency4/5

Three tools follow a verb_noun snake_case pattern (translate_papers, get_translation, resume_translation), while arxiv_zh_status uses a server-prefixed noun phrase. This is a minor deviation but the set remains readable and consistently cased.

Tool Count5/5

Four tools map cleanly to the essential operations of the translation workflow—setup check, submission, monitoring, and recovery. The count is well-scoped with no redundant tools.

Completeness4/5

The surface covers the core lifecycle: check environment, submit translation, query results, and resume interrupted jobs. Minor gaps include no cancel/delete operation and no way to list all existing jobs if the job_id is lost, but these are workable in most cases.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    Not graded
    maintenance
    Enables extraction of mathematical content from TeX papers and conversion to Lean code through a structured intermediate representation. Supports project scaffolding, entity management, and task tracking for mathematical formalization workflows.
    14
    -
  • F
    license
    A
    quality
    B
    maintenance
    Enables searching, downloading, and reading academic papers from multiple platforms including arXiv, Semantic Scholar, PubMed, bioRxiv, medRxiv, IACR, Google Scholar, RePEc/IDEAS, and Sci-Hub with PDF to Markdown conversion.
    29
    7
    -
  • F
    license
    A
    quality
    F
    maintenance
    Enables AI assistants to create, edit, and validate LaTeX documents through a standardized protocol with support for multiple document classes and package management. It provides tools for document structure analysis and file organization to streamline the generation of professional academic papers.
    8
    19
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables searching, downloading, and exporting academic papers from 20+ scholarly sources including arXiv, PubMed, and Semantic Scholar. Supports multi-source concurrent search, citation network tracing, and export to CSV, RIS, and BibTeX.
    1
    MIT