Scholar Library
Enables searching arXiv, resolving arXiv identifiers, and downloading open-access PDFs for literature research projects.
Uses Cloudflare HTTPS DNS to verify public domain names when resolving synthetic DNS/proxy addresses, sending only the domain.
Resolves DOI identifiers to bibliographic metadata and supports citation/reference workflows.
Supports optional OpenAI-compatible embedding services for semantic and hybrid literature retrieval, with vectors stored locally.
Enables searching PubMed and resolving PMIDs as part of biomedical literature retrieval.
Exchanges bibliographic records with Zotero through BibTeX, RIS, and CSL-JSON formats.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Scholar LibraryCheck workspace, create a project on RAG citation errors, and search arXiv plus Crossref for open papers"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Scholar Library · 文献研究插件
面向高校学生与研究员的本地文献库,通过标准MCP接入AI宿主。宿主负责理解与写作;插件负责收集、版本、原件、TextIn任务、证据检索、引用完整性检查、关系、导出和恢复。
状态:0.1.0 开发版,持续更新、完善与迭代中。 能力及验收边界见验收记录。历史开发记录报告 TextIn 直连 API 完成30篇中英文论文、374页解析,以及字段/表格抽取实测;原始材料、响应与本机诊断库不公开,不能仅凭该摘要独立复核;广泛材料质量、人工语义指标及第二种图形AI宿主验收尚未完成,不能视为计划全部验收通过。
环境与兼容性
Python 3.12+、uv、Git,以及支持本地 stdio MCP 的 AI 客户端。MCP 是客户端调用本地工具的协议,不是另一个聊天模型。
macOS 已有本机运行证据;Linux 使用相同 POSIX 文件锁接口,但尚需独立实机验收。
Windows 原生 Python 暂不支持。可在 WSL2/Linux 内安装并让客户端启动 WSL 中的服务;这一组合尚未实机验收,不提供 Windows 原生安装包。
首次安装会下载 Python/依赖,需要网络。API 服务的额度、价格和产品权限由各服务商控制,开源代码不包含免费额度或凭据。
Related MCP server: surveyHelper
安装与运行
git clone https://github.com/SAKURAfan1023/scholar-library.git
cd scholar-library
uv sync --locked --python 3.12
uv run python scripts/configure_local.py
uv run python scripts/smoke_mcp.py最后一步在临时合成库中验证 MCP,不上传论文、不调用 OCR。将生成的 .mcp.json 内容导入客户端的 MCP 配置;直接作为 Codex 本地插件使用时,仓库中的 manifest 和 Skill 也一并保留。宿主各自的插件注册界面可能不同,配置导入后必须实际调用 workspace_status 确认连接。
手工启动服务用于排障:
uv run scholar-library以上命令启动stdio MCP服务,等待客户端请求;不会自动上传。.mcp.example.json 提供无密钥格式示例,configure_local.py 生成的 .mcp.json 是本机专用文件,不进入 Git;移动源码后重新生成。正式安装使用python3 scripts/build_release.py构建独立安装包,解压后运行python3 install_release.py --target <新的安装目录>,将生成的mcp.json交给支持本地stdio MCP的客户端。安装会下载锁定依赖,不是全离线安装包。
默认库为~/.local/share/scholar-library,可用环境变量SCHOLAR_DATA_DIR指定。运行环境、源码与文献库相互独立;真实材料不进入Git或插件安装包。
TextIn凭据通过系统钥匙串配置:
uv run scholar-configure TEXTIN_APP_ID
uv run scholar-configure TEXTIN_SECRET_CODE同名环境变量可覆盖钥匙串,适合受控部署;不要把值写入MCP配置或分享给AI。OpenAlex可配置OPENALEX_API_KEY,可选语义索引配置EMBEDDING_API_KEY。服务商是否开通产品仍须真实验证。
第一次使用
创建专用材料目录,将要研究的PDF/DOCX/图片/文本放入其中。
向AI说明研究问题、输入目录、哪些服务允许联网、可用页数和请求预算。
AI调用workspace_status、create_project,建立授权项目,再检索或导入指定材料。
阅读、追问、对比后将带引用主张保存为research记录,按需要导出。
示例提示:
使用文献插件研究“检索增强生成如何降低文献问答中的错误引用”。先检查工作区,建立研究项目。优先检索Crossref和arXiv,列出候选与来源状态。只下载我选中的开放PDF。全文解析前核对项目上传授权和页数预算。回答必须带原文片段和PDF页序,区分作者结论、推断和证据不足。
已有授权在项目内持续有效,无需每次调用重复确认。host_text控制向宿主返回正文,但无法控制宿主接收后的保存策略。TextIn页范围限制处理页,上传的仍是整个文件。
能力
Crossref/OpenAlex/arXiv/PubMed检索及逐来源状态;DOI/PMID/arXiv解析;公开PDF下载。
BibTeX/RIS/CSL-JSON与Zotero交换题录,付费数据库通过授权本地全文导入。
独立文献版本、来源快照、阅读和筛选理由,重复与冲突候选。
免费离线PDF文字层解析、原页图核对;TextIn xParse任务续接、分批论文解析、原始缓存、页图读取;单独字段/表格抽取。
中英文关键词检索、可选OpenAI兼容embedding混合检索;向量本地保存。
带精确摘录的回答/卡片/综述,跨论文关系,方法评价;生成产物不污染原文检索。
Markdown、JSON、CSV、BibTeX、RIS、CSL-JSON、HTML、PDF、DOCX,GB/T 7714与APA参考文献。
带哈希的备份/恢复、删除影响预览、版本冲突控制。
边界
单人本地库,允许多个新服务进程连接;不放网络盘。备份需关闭其他连接;Windows尚未支持POSIX文件锁。
单文件100MiB、PDF最多2000页、每次TextIn处理1–100页;这是本地保护值,不代表服务商必然支持。服务商拒绝时保留任务及失败状态。
request_budget累计计入网络检索、下载、解析提交、查询和向量请求,page_budget计入解析/抽取预留页数。未知提交不释放预留量,避免超支;不报告无法核实的货币费用。本地MD/TXT/DOCX解析不做OCR。DOCX使用段落/表格位置而不是稳定页码,文本框和复杂布局可能未提取。PDF可用parse_local_pdf离线提取文字层进入索引;扫描页/图片需TextIn。文字层完整不代表视觉内容完整,表格、公式、双栏顺序须看原页复核。
OpenAlex标志与Crossref撤稿通知都是来源信息,不是论文质量保证;元数据源不全、错误或过期均可能影响结论。
下载和embedding拒绝内网地址并固定实际连接IP;遇到本机代理的198.18/15合成DNS地址时,会通过Cloudflare HTTPS DNS核实公开域名(仅发送域名,另计一次请求预算),不会关闭内网保护。
表格CSV保存表格文本;复杂单元格结构在解析JSON中保留。HTML使用离线KaTeX;PDF/DOCX保留TeX源码,暂不转换为原生数学对象。
删除清理题录、证据和索引并标记相关产物过期;输入文件、旧导出、备份、原始对象文件及服务商副本保留,不能称为彻底擦除。
不提供后台调研/通知、团队服务、付费墙绕过或自动安装模型;宿主不调用工具时不会自主运行研究。
开发验证
uv run pytest -q
uv run ruff check src tests scripts
uv run python scripts/smoke_mcp.py
uv run python scripts/benchmark.py
uv run python scripts/probe_metadata.py前四项使用独立合成测试库,不调用付费接口。最后一项只联网查询公开题录并保存来源状态,不上传本机论文。可用scripts/evaluate.py对人工标注证据问题计算Recall@10;人工语义支持率不能用合成测试替代。标注格式及指标边界见评估说明。
题录交换保留常用出版字段,原始导入文件和未映射字段可追溯;两条ACL官方题录的三格式往返已实测。复杂定制字段与不同管理器兼容性仍需针对样本验证。
直接调用TextIn API
uv run python scripts/textin_acceptance.py submit VERSION_ID --start 1 --end 2提交现有版本;poll JOB_ID只查询已有任务;extract VERSION_ID --key '全部作者(按原文顺序)'抽取字段,--table-header可重复指定表头。这些命令使用HTTP API,不通过MCP,凭据读取系统钥匙串。每次请求先落盘记录,已知任务不重复上传。
独立安装目录同时提供textin_acceptance.py,可用该目录的runtime/bin/python -I执行。用户明确授权无限预算时page_budget/request_budget支持null;未授权时默认0。
配置与数据链路
配置 | 用途 | 默认/要求 |
| SQLite、原件、证据与导出所在的本地库 |
|
| TextIn 解析与抽取 | 按需配置;本地文字层解析不需要 |
| OpenAlex 检索 | 按服务商当前要求配置 |
| 可选语义向量服务 | 非必需;使用前确认正文传输授权 |
项目 | 可读取的材料目录 | 建项目时显式指定 |
项目 policy / budgets | 联网、上传、正文返回与请求/页数预算 | 默认拒绝未授权动作,不把旧测试授权带入新项目 |
flowchart LR
User[用户授权与研究问题] --> Host[AI 宿主]
Host --> MCP[本地 MCP 工具]
MCP --> Store[独立 SQLite 库与原件]
MCP --> Evidence[证据片段与引用校验]
MCP -->|按项目授权和预算| Provider[题录检索 / TextIn / 可选向量服务]
Evidence --> Export[研究记录与多格式导出]AI 推理在宿主侧进行;本插件负责存储、工具执行和证据约束,不要求另外填一个通用聊天模型 Key。host_text 开启后宿主会接收正文,仍须考虑宿主自身的数据政策。
常见问题与更新
服务启动后没有聊天窗口:stdio 服务在等待 MCP 客户端连接;先用
smoke_mcp.py核对,再在宿主调用workspace_status。找不到凭据:钥匙串和客户端所用系统账户应一致;Linux 需要可用的 keyring 后端,也可从受控启动环境注入同名变量。
检索/上传被拒绝:先核对当前项目授权与预算。不要通过关闭安全检查或自动重试来绕过失败。
找不到之前的记录:核对
SCHOLAR_DATA_DIR和当前运行版本;更新源码不等于已重启宿主中的旧服务。解析成功但理解错误:先回看原页和具体摘录,按 评估说明 保存待人工核对状态。
源码方式升级:先保存本地改动,再 git pull --ff-only、uv sync --locked、重新生成 .mcp.json 并在客户端重载服务。独立安装包应安装到新目录,验证后切换配置。数据始终与源码/运行环境分开保存,重要库先备份。
目录与验证证据
src/ 是核心实现,skills/ 是 Agent 工作流,tests/ 是自动回归,scripts/ 含构建、自检和显式联网实验,docs/ 保存设计与验证边界。output/、数据目录、凭据和本机安装清单不公开。
当前公开版本复现结果见 发布验证;历史验收摘要 和 完成度审计 是带日期的开发记录,不代表读者设备、实时外部接口或全部语义质量通过。
维护、贡献与使用声明
持续更新、完善与迭代中。禁止恶意转载与滥用。 转载须保留版权和许可声明,请注明原仓库及修改内容;不得冒充作者或官方、夹带恶意代码、盗取凭据、泄露个人资料或伪造研究/学习证据。
代码按 MIT License 开源。上述反滥用声明不额外撤销 MIT 授予的合法使用、修改、转载或商用权利;第三方材料和服务不随代码一并授权。完整说明见 USE_POLICY.md、第三方声明。
欢迎通过仓库 Issues 反馈 Bug 或建议。请包含版本、系统、脱敏复现步骤、预期与实际结果;不要上传密钥或真实私密材料。提交代码前阅读 贡献指南,安全问题见 SECURITY.md。
Available Tools
38 toolsadd_workA
保存题录(title必填;doi/pmid/arxiv、authors、year、journal、url等)。一致标识和标题才自动复用;冲突保留独立记录。
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | Yes | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish a non-destructive, closed-world write, so the bar is lower; the description adds genuinely non-obvious behavior beyond that: records with matching identifiers and title are auto-reused, while conflicting ones are kept as separate records. This directly affects data integrity and is not derivable from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler; the required-field constraint and the supported identifiers are front-loaded, and the dedup caveat follows. Density is high without hurting readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema and an entirely undescribed nested metadata object, the description covers purpose, the required field, common fields, and dedup behavior. It omits what project_id means, what a successful call returns, and any permission expectations, so it is adequate but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and metadata is a free-form nested object with no property definitions, so the description's field list (title required, doi/pmid/arxiv, authors, year, journal, url) is the only documentation of metadata contents. However, project_id is left entirely unexplained in both schema and description, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action and resource ('保存题录') and enumerates the metadata fields it accepts (title required; doi/pmid/arxiv, authors, year, journal, url), so an agent knows exactly what kind of record is created. It does not explicitly differentiate itself from sibling update_work, which prevents a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the dedup rule ('一致标识和标题才自动复用;冲突保留独立记录') tells the agent what happens on repeated submissions, which helps decide whether to call this versus update_work, but no alternative is named and no exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
backup_libraryA
备份整个本地文献库和原件;要求其他连接已关闭,备份可能包含多个项目,不含凭据。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, destructive=false, openWorld=false. The description adds genuinely new behavioral facts: it spans multiple projects and excludes credentials, and it imposes a concurrency constraint. It does not say where the backup is written or its format, but the additions go meaningfully beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence packed with three useful facts, front-loaded on the verb and resource. The semicolon-joined clauses make it slightly dense but no clause is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description covers scope (whole library and originals), preconditions, content breadth (multiple projects), and an important exclusion (credentials). The main omission is the destination/location of the backup, which an agent might need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is nothing for the description to clarify; baseline 4 applies. The description does not introduce any parameters, and none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource – backing up the entire local literature library plus original files – and immediately scopes it ('entire local library', not a single project or record). This clearly separates it from siblings like export_research, preview_deletion, and restore_library.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition ('other connections must be closed'), which is actionable usage context. However, it does not name when to choose this over export_research or restore_library, so alternatives are left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_evidence_packB
构建有界证据包和关联阅读线索。默认本地关键词;semantic=true会将问题发送至项目配置的embedding服务,需要授权与已建索引。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| semantic | No | ||
| project_id | Yes | ||
| char_budget | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses the most important behavioral trait: semantic=true sends the query to an external embedding service, which requires authorization and a pre-built index. This is real context an agent needs to avoid surprise failures. It does not explain why the tool is flagged readOnlyHint=false (what, if anything, is persisted), leaving one annotation unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core capability front-loaded and the semantic-mode caveat second. Nothing is wasted, though the compactness comes partly at the cost of parameter coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter, no-output-schema tool with 0% schema coverage, the description is thin: it omits the return shape of an evidence pack, how limit and char_budget bound the result, and the auth/index prerequisites are stated only as a consequence of the semantic flag rather than as prerequisites for calling the tool at all.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, so the description carries the full burden. It explains only semantic and hints at boundedness via char_budget; limit, project_id, query format, and the meaning of the char budget/limit interaction are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb+resource: building bounded evidence packs plus related reading clues. It is clear what the tool produces, but it never differentiates itself from the nearby read_evidence sibling or from search_library/index_semantic_search, so an agent has to infer its place in the family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It draws one useful decision boundary: default is local keyword search, while semantic=true routes to the project embedding service. That implies when to set the flag, but there is no explicit when-to-use-this-tool-versus-alternatives guidance and no warning about failure when authorization or an index is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connect_papersA
保存文献联系;支持/反对关系必须提供两端原文证据和具体条件说明。similar仅是线索。
| Name | Required | Description | Default |
|---|---|---|---|
| relation | Yes | ||
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (write, non-destructive, closed-world), and the description adds real behavioral rules beyond them: evidence and condition requirements for supports/contradicts, and the weak status of 'similar'. It does not cover permissions or the shape of the response, but it meaningfully extends the annotation surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences with the core constraint (evidence required) front-loaded and the caveat about 'similar' appended. It is well-sized and wastes little, though the terseness leaves gaps that a slightly fuller description could have closed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a nested relation object, eight relation kinds, and no output schema, the description covers the most important constraint (evidence for support/contradiction) but omits guidance on version_id/target_version and the remaining kinds, so it is only adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the parameter burden. It explains the evidence_ids/explanation requirements for supports/contradicts and the semantics of 'similar', but it says nothing about version_id or target_version and leaves five other relation kinds unaddressed, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (save/store) and resource (literature relations between papers), which lets an agent distinguish it from siblings like add_work or verify_work. It does not name a specific sibling to contrast against, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage rules by mandating that supports/contradicts relations carry evidence from both ends and concrete conditions, and it flags 'similar' as only a clue. However, it gives no explicit when-to-use versus when-not-to-use guidance relative to alternatives, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_documentB
将已有解析导出为Markdown、结构JSON和表格文本CSV;不重新上传,不承诺Word版式还原。
| Name | Required | Description | Default |
|---|---|---|---|
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so safety is partly covered. The description adds two useful behavioral facts — no re-upload of source, no guarantee of Word layout fidelity — but says nothing about where output is written or whether it overwrites existing artifacts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the capabilities front-loaded and the two disclaimers trailing; no filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does list the produced formats, which is the most important missing structured info. However, for a conversion tool it omits where/how results are retrieved, overwrite behavior, and any prerequisite on the parse state, leaving gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter version_id is undocumented in both schema and description. The phrase '已有解析' loosely implies version_id refers to an existing parse version, but format, sourcing, and valid values are unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and target ('export an existing parse') plus the three output formats (Markdown, structured JSON, table CSV), which distinguishes it from parse_document and read_parse_result. It is clear what the tool produces, though it does not explicitly name the sibling it replaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'不重新上传' implies this runs on an already-parsed version rather than re-parsing, hinting at when to use it over parse_document, but no alternative tool is named and no explicit precondition (e.g. parse job must have succeeded) is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectA
建立单人研究项目;input_root是用户明确指定的输入目录。policy须反映用户授权,默认全部关闭;预算是累计页数和请求上限。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| policy | Yes | ||
| question | Yes | ||
| input_root | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so they convey little beyond 'safe-ish write'. The description adds real behavioral context: policy flags default to off and must mirror user consent, and budgets are cumulative caps on pages and requests. It still doesn't say what the call returns or what happens if a project already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the purpose front-loaded and constraints packed into the remainder; no filler. Dense but readable, though the semicolon-chained clauses make it slightly harder to scan than a fully structured version.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only minimal annotations, the description does cover the creation semantics and policy/budget meaning, but omits the return value, duplicate-project behavior, and what 'name' and 'question' should contain. Adequate but with clear gaps for a 4-parameter mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 4 required params including a nested Policy object, so the description must carry the load. It explains input_root (user-designated input directory), policy (must reflect authorization, all off by default), and touches on the budget fields, but never mentions 'name' or 'question', leaving two required params unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('建立单人研究项目' – create a single-person research project), which is clearly distinct from siblings like update_project_policy or save_research. It does not explicitly name or contrast with any sibling, which keeps it out of the 5 range.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives two conditional constraints – input_root must be a directory the user explicitly designated, and policy must reflect user authorization with all flags defaulting off – which implies prerequisites. However, it never states when to call this versus update_project_policy or how it relates to the rest of the workflow, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_workADestructive
仅在用户明确要求删除并查看最新预览后调用,使用预览digest防止影响范围变化。
| Name | Required | Description | Default |
|---|---|---|---|
| work_id | Yes | ||
| inventory_digest | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true/readOnlyHint=false, but the description adds non-obvious behavior: the call is gated on human confirmation and the digest acts as a staleness guard so the blast radius cannot drift from the preview. It still doesn't say what is destroyed, whether it is recoverable, or what a digest mismatch returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the gating condition front-loaded and no filler. It is tight, though the single clause bundles two distinct requirements (explicit user request and preview review) that could be easier to parse if split.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, 2-required-param tool with no output schema, the description covers the critical safety workflow (preview plus digest) but omits the consequences of success (what exactly is removed, whether backup_library should be used first) and the failure mode when the digest is stale.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning; it usefully explains that 'inventory_digest' comes from the preview and exists to detect scope changes, which the schema alone does not convey. work_id is left entirely to the name, so only half the parameters are meaningfully covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb '删除' (delete) appears only inside a precondition clause, so the tool's actual operation is inferred from the name rather than stated (e.g., it never says a work record and its dependent artifacts are removed). Siblings like preview_deletion, backup_library and restore_library are not named or contrasted, so differentiation is left to the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit, testable precondition: call only after the user explicitly asks for deletion AND has seen the latest preview. That is real when-to-use guidance rather than an implied context. It stops short of naming preview_deletion as the required preceding tool or stating what to do when no preview exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_fulltextA
从用户选定的公开HTTPS地址下载PDF,最多100MiB;检查每次重定向,不处理账号、付费墙或登录页。
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| work_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond the annotations (which only flag openWorldHint=true, readOnlyHint=false) by disclosing concrete operational traits: a 100MiB size cap, per-redirect validation, and the refusal to handle authentication, paywalls, or login pages. These are real constraints an agent must plan around; only the return/failure behavior is left unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tightly packed sentence front-loads the core action and follows with constraints. No filler or redundancy; every clause adds an actionable fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a network-fetch tool with no output schema, the description covers the safety-relevant constraints well, but leaves unanswered what gets returned, what happens on failure/redirect exhaustion, and how the downloaded file is bound to work_id. Adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It partially covers 'url' (must be HTTPS, public, size-capped) but says nothing about what 'work_id' is for or how the two parameters relate. With both parameters required and undocumented, a key gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: downloading a PDF from a public HTTPS URL. It scopes the source clearly (user-selected, public, HTTPS) which distinguishes it from sibling file-ingest tools like parse_local_pdf or import_document, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The negative scope ('不处理账号、付费墙或登录页') gives an implicit when-not condition, and '用户选定的公开HTTPS地址' implies the intended trigger. But there is no positive routing guidance telling the agent when to choose this over import_document or parse_document.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_researchB
导出MD/JSON/CSV/BibTeX/RIS/CSL-JSON/HTML/PDF/DOCX及证据清单;拒绝过期产物。style=gb7714或apa;原全文不打包。
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | gb7714 | |
| project_id | Yes | ||
| artifact_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds genuinely useful behavior beyond them: it refuses stale artifacts (拒绝过期产物) and discloses that full text is not packaged (原全文不打包), which prevents a false expectation about export contents. It still omits error/auth/return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with formats front-loaded and constraints trailing, with no filler. It is appropriately sized for the tool's scope, though the run-on semicolon structure compresses several distinct facts (formats, staleness, style, packaging) into one breath.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-format export tool with no output schema and 0% schema coverage, the description covers formats, style, and packaging scope but leaves key gaps: it never explains what artifact_ids are or what happens (error vs partial) when artifacts are stale. Adequate but with clear holes an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters. The description compensates only for 'style' by giving the allowed values (gb7714/apa), which is helpful since the schema has no enum. project_id and artifact_ids receive no meaning at all (e.g., what an artifact_id refers to, id format), leaving most parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (导出/export) and resource, and enumerates the concrete output formats (MD/JSON/CSV/BibTeX/RIS/CSL-JSON/HTML/PDF/DOCX), so the agent knows exactly what it produces. It does not differentiate itself from adjacent tools like convert_document or build_evidence_pack, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or named alternative; the agent must infer that this is the batch-export path versus convert_document for single documents. The only usage-flavored hints are the style options and the staleness rejection, which are not framed as selection conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_fieldsB
调用TextIn字段/表格抽取,会上传整个PDF或图片并预留全部页数;结果需复核,不成为原始证据索引。相同请求复用任务。
| Name | Required | Description | Default |
|---|---|---|---|
| keys | Yes | ||
| version_id | Yes | ||
| table_headers | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses meaningful behavior: it uploads the ENTIRE PDF/image and reserves all pages, results require manual review and do not serve as the original evidence index, and identical requests reuse the same task. These are real operational traits (whole-document upload, idempotency, output trust level) that annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, dense sentence that front-loads the action (call TextIn extraction) and then packs behavior/caveats. It is efficient, though the compound clauses make it somewhat heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-read-only, open-world extraction tool with no output schema and zero parameter documentation, the description adequately covers behavioral expectations (upload scope, review requirement, idempotency) but leaves the three required inputs unexplained. It is viable but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three required parameters. The description loosely hints that keys map to extracted fields and table_headers to table columns, but version_id (presumably the document version to upload) is never explained, and no format or constraint details are provided. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation (TextIn field/table extraction) and resource (PDF/image document), so an agent can tell it is the tool that triggers extraction. It does not, however, explicitly contrast itself with siblings like read_extraction, which reads the extracted output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (invoke when you need fields/tables pulled from a document) and the caveat that results need review plus the idempotent 'same request reuses task' note give some context. There is no explicit when-to-use/when-not or named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_duplicate_candidatesARead-only
查找同标识冲突或标题相似的候选,不自动合并不同版本或附件。最多比较最近1000条,候选上限100。
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly=true, destructive=false, openWorld=false), and the description adds behavioral value beyond them: it states the tool does not auto-merge, caps comparison to the 1000 most recent records, and limits candidates to 100. These operational limits are genuinely useful context not present in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the core purpose and followed by scope and limits. No filler sentences; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only finder with no output schema, the description covers what is found and the comparison caps, but does not explain what a 'candidate' result contains or how the caller should act on it. Adequate but leaves the return/next-step semantics to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the single required parameter project_id, adding no meaning, scope, or format for it. Although the parameter is fairly self-evident from its name, the description fails to compensate for the zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource (find duplicate candidates) and defines what qualifies: identifier conflicts or similar titles. It also clarifies the boundary that it does not auto-merge versions/attachments, which distinguishes it from a hypothetical merge tool. It does not name a sibling explicitly, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the stated purpose (finding duplicates) but there is no explicit when-to-use guidance and no alternatives named (e.g., resolve_identifier, verify_work, list_records). The 'does not auto-merge' clause gives a boundary but is not a routing instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_parse_jobB
默认只读本地任务;poll=true查询已有远端任务,至少间隔3秒。结果保存并合并各批页面,不重新上传。
| Name | Required | Description | Default |
|---|---|---|---|
| poll | No | ||
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true, readOnlyHint=false and destructiveHint=false. The description usefully adds the 3-second polling throttle, that results are saved and merged across page batches, and that nothing is re-uploaded. It does not explain what 'readOnlyHint=false' writes or what triggers a state change, but given annotation coverage this is a reasonable addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses front-load the default behavior, then the poll alternative, then the merge/no-re-upload detail. No filler, though the clauses are dense and could be slightly clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotation clarity on the write side, the description covers mode selection, polling cadence, and merge semantics but omits what job_id identifies, what happens when a remote job is absent, and what the saved/merged result contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. It explains poll=true versus the default local read well, but job_id is left entirely undocumented, so the description only compensates for half of the two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that it reads a local parse job by default and queries existing remote jobs when poll=true, and that results are merged. However, it never plainly says it retrieves parse-job status/results, so the core resource is only implied; an agent must infer the purpose from the name and sibling tools like read_parse_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete condition for poll=true (query an existing remote job) and a rate constraint (at least 3 seconds apart), which is implied rather than stated as guidance. It does not say when to use this versus retry_parse or read_parse_result, so alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_documentB
本地导入PDF/DOCX/PNG/JPG/MD/TXT,保留不可变原件;不上传。每个新文件版本单独保存。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| label | No | manuscript | |
| work_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations giving only readOnlyHint=false and destructiveHint=false, the description adds real behavioral context: the original is kept immutable, nothing is uploaded (local-only), and every new file version is stored separately. That versioning and immutability detail is valuable beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A tight multi-clause sentence with the verb and format list front-loaded and no filler. Every clause conveys a distinct fact, though the structure is slightly telegraphic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the behavioral profile (immutable original, local-only, per-version storage) is reasonably covered. However, with three undocumented parameters and no usage routing, the definition is not complete enough to call correctly without inspecting the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across all three parameters (work_id, path, label), so the description carries the full burden. It only mentions file formats, leaving work_id, path, and the label default of 'manuscript' entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (import) and resource (document) plus the supported formats (PDF/DOCX/PNG/JPG/MD/TXT), and specifies 'local'. This distinguishes it from upload-oriented siblings, though it never names an alternative tool explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or named alternative among the many siblings (import_references, download_fulltext, convert_document). 'Does not upload' implies a local-only context, but the agent must infer when to pick this over parse_local_pdf or import_references.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_referencesA
导入授权目录内的BibTeX、RIS或CSL-JSON题录;附件须另行导入,不爬取付费数据库。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false but say nothing about scope, so the description carries real added value: it clarifies this imports records only (attachments excluded) and that it does not crawl paywalled databases, and that inputs are confined to an authorized directory. It still omits behavior on duplicates or existing-library conflicts, which matters for an import mutation, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence front-loads the action and formats, then appends the two key constraints. Every clause earns its place with no filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no output schema and no extra structured guidance, the description covers formats, the attachment exclusion, the source-directory restriction, and the no-crawling limit. It nonetheless omits duplicate handling, expected outcome (what gets created), and the role of project_id, so an agent could still invoke it without knowing the full behavioral surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters, so the description must do all the work – and it largely does not. It constrains 'path' to an authorized directory, which is useful, but says nothing about the format/meaning of project_id or what path forms are accepted (file vs. directory, relative vs. absolute). This leaves the parameters materially under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (导入/import) and resource (题录/bibliographic records) and enumerates the accepted formats (BibTeX, RIS, CSL-JSON), which distinguishes it from sibling importers such as import_document or parse_document. It also implies a distinct scope by noting attachments must be imported separately. It stops short of explicitly naming an alternative tool, so the differentiation is inferential rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one real prerequisite – the path must be inside an authorized directory – and one routing hint ('附件须另行导入'), which points the agent toward a separate attachment-import flow. However, there is no explicit when-to-use vs. when-not guidance relative to siblings like import_document, nor any mention of preconditions like project_id being an existing project. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_semantic_searchB
将最多32个原文片段发送到已配置的embedding服务并本地保存向量;重调续接,模型改变后使用新索引。
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly=false, openWorld=true, destructive=false), it discloses the batch cap of 32 chunks, local vector persistence, resumption of interrupted runs, and the model-change/new-index rule. It still omits cost, auth, and what happens to a pre-existing index, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A compact, front-loaded sentence that leads with the core action and appends the resumption and model-change caveats. Dense but every clause carries operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an indexing tool with no output schema and a single parameter, the description covers batching, persistence, resumption, and re-indexing triggers. It leaves the project_id parameter and the fate of an existing index unexplained, so it is adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (project_id) with 0% schema description coverage, and the description says nothing about it. With low coverage the description is expected to compensate, and it does not, though the parameter name is largely self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action (send up to 32 source chunks to the configured embedding service and store vectors locally), making the indexing purpose clear. It does not explicitly contrast with siblings like search_literature or search_library, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives two usage conditions: re-invocation resumes an interrupted job, and a new index should be used after the model changes. However, it names no alternative tools and gives no when-not guidance, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_input_filesARead-only
仅列输入目录第一层的普通文件名和大小,不读内容、不递归、不上传。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered structurally. The description adds real value beyond that by clarifying that content is not read, subdirectories are not traversed, and nothing is uploaded — explicit non-behaviors an agent cannot infer from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with zero padding. The scope constraint (first-level only) is front-loaded and every clause ('不读内容、不递归、不上传') earns its place by ruling out a distinct behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with no output schema, the description covers the essentials of what is returned (names and sizes) and what is excluded. The gap is that paging parameters with defaults (limit=20, offset=0) are never explained, leaving an agent unaware that results are paginated and can be iterated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, so the description must carry the load and it does not. limit, offset, and project_id are entirely undocumented — no mention of pagination semantics or what project_id scopes to, even though the description references an 'input directory' without tying it to project_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('仅列输入目录第一层的普通文件名和大小' — list first-level regular file names and sizes in the input directory) and scopes it precisely to the top level. This distinguishes it from document-reading siblings like read_document or parse_document, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The negative constraints ('不读内容、不递归、不上传' — no content read, no recursion, no upload) implicitly tell the agent when this is appropriate versus a read/parse tool. However, no alternative tool is named and no positive trigger condition is stated, so usage is only implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recordsBRead-only
分页找回文献、版本、任务、笔记、关系、检索和核验记录。job列表不重提;内容记录要求host_text授权。
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| limit | No | ||
| offset | No | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds a genuinely useful behavioral constraint (content records need host_text authorization) and a note about job lists, but the phrasing is cryptic and says nothing about pagination semantics or result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the primary operation front-loaded; little waste. The second sentence's phrasing ('job列表不重提') is ambiguous enough to slightly undercut clarity, but the overall size is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, paginated list tool with no output schema and no parameter descriptions, the definition leaves pagination behavior, return shape, and the meaning of limit/offset unexplained. Annotations cover the safety profile, but the description is only minimally sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description carries the burden. It lists record kinds that partially mirror the enum, but says nothing about limit, offset, or project_id semantics, nor about how pagination is bounded. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (paginated retrieval) and enumerates the record resources (works, versions, jobs, notes, relations, searches, assessments). An agent can tell this is the list/collection reader, though it does not explicitly name the sibling it replaces (e.g., read_record for single records).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one conditional rule (content records require host_text authorization) and one exclusion note ('job列表不重提'), which hints at usage boundaries. However, it never contrasts itself with read_record or search_library, so the agent must infer when to page records vs fetch a single one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_documentA
按项目授权向TextIn上传原件并解析指定页范围;每批最多100页。页范围只限制处理,上传的是整个文件。相同版本和范围复用原任务,包括失败/未知提交。
| Name | Required | Description | Default |
|---|---|---|---|
| end_page | No | ||
| start_page | No | ||
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true and destructiveHint=false, so the safety profile is partly covered. The description adds genuinely useful behavior beyond that: the entire file is uploaded even when only a page range is processed, and identical version+range reuses the original task including failed/unknown submissions (idempotent caching).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action and followed by the constraints and reuse rule. Dense but nearly every clause carries distinct information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should signal what happens next, and the reuse-of-task behavior partially does. It still never indicates whether a job is created that must be polled (get_parse_job/retry_parse siblings suggest async), which is the main remaining gap for a parsing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning, and it does clarify that start/end page restrict processing only (not upload) and that version+range jointly identify a reusable task. It does not explain where version_id comes from or its relation to import/reference records, leaving one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource and mechanism: it uploads the original file to TextIn and parses a specified page range. This implicitly separates it from parse_local_pdf, though no sibling is named explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives operational context (max 100 pages per batch, page range only limits processing, whole file is uploaded) and idempotency behavior, which helps an agent decide how to call it. However, it never states when to choose this over parse_local_pdf, get_parse_job, or retry_parse, so when-versus-alternative guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_local_pdfA
免费离线解析PDF文字层,无上传;扫描页需TextIn OCR。公式、表格、阅读顺序须看原页复核。切换解析器会使旧引用过期。
| Name | Required | Description | Default |
|---|---|---|---|
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint=false, destructiveHint=false, openWorldHint=false) and the description carries substantive behavior: free, offline, no upload, OCR required for scanned pages, and a real side effect (switching parsers expires old citations). Missing explanation of why readOnlyHint is false or what the call produces, but it adds notable context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense clauses, front-loaded with the core purpose and no wasted filler. Every clause (offline/no-upload, OCR caveat, verification caveat, citation invalidation) carries information, though the density borders on terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, OCR limitation, verification caveats, and a side effect, which is solid for a simple parse tool. However, it omits any guidance on the required version_id and gives no hint of the result shape, which an agent needs given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter 'version_id' has 0% schema description coverage, and the description never mentions it or explains what a version identifier is. With low coverage the description should compensate, but it adds no parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('解析PDF文字层' = parse PDF text layer) with explicit scope qualifiers (offline, no upload, text-layer only). It is clearly distinguishable from the generic sibling 'parse_document' by emphasizing local/offline operation and text-layer limitations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions: use for text-layer PDFs offline, but scanned pages need TextIn OCR (routing the agent elsewhere). It also warns that switching parsers invalidates old citations. It stops short of naming the exact OCR sibling tool, so the alternative is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_deletionARead-only
预览删除文献及关联索引的影响;混合研究产物标记过期,输入原件、旧导出、备份和服务商数据保留。
| Name | Required | Description | Default |
|---|---|---|---|
| work_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), but the description adds real behavioral detail beyond them: mixed research artifacts are marked stale, while originals, old exports, backups, and provider data are retained. This consequence-disclosure is genuinely useful and not derivable from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, followed by consequence and retention details. It is a single dense sentence with no filler, though the trailing retention list is somewhat list-like rather than prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and read-only annotations, the definition provides enough context to call it (what is affected, what is preserved). The main gap is the unexplained work_id format, but the impact-preview contract itself is well communicated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required work_id parameter, so the description must compensate and it does not: it references '文献' (work) only loosely and gives no format, identifier type, or example for work_id. The one parameter remains semantically opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (预览/preview) and resource (删除文献及关联索引/deleting a work and its associated indexes), which is precisely separable from the sibling delete_work that performs the actual deletion. An agent can tell it is an impact-assessment preview rather than the destructive operation itself without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The framing strongly implies this should be run before delete_work, but the description never explicitly states when to use it, when not to, or names the delete_work alternative. Usage is left to inference rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_documentBRead-only
按原文顺序分页读取当前解析片段、页序和质量信息,不调用OCR。无页码的DOCX/文本用段落定位。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavior not in annotations: pagination, in-order traversal, no OCR invocation, and paragraph-based positioning for pageless DOCX/text. It stays silent on error cases and result-size limits, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core capability and the OCR exclusion front-loaded. Slightly dense/ambiguous phrasing (当前解析片段) costs it the top mark, but nothing needs to be cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, 0% param coverage, and many closely related siblings (read_page, read_parse_result, read_evidence, download_fulltext), the description leaves the agent unable to tell which read tool applies or what version_id refers to. It should route to siblings and clarify inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with three parameters (version_id required, limit, offset). The description explains none of them — version_id's meaning and scope in particular is left entirely to inference, and limit/offset are only weakly implied by 分页. It does not compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (读取) and resource (当前解析片段、页序、质量信息), and states the mode (分页、按原文顺序、不调用OCR). This distinguishes it from parse_document/parse_local_pdf style tools in the same family, though it does not name a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 不调用OCR implies this tool reads already-parsed content and that a parse/OCR tool is needed otherwise, and 无页码的DOCX/文本用段落定位 hints at a fallback path. But there is no explicit when-to-use vs read_page, read_parse_result, download_fulltext, or any prerequisite statement — usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_evidenceARead-only
读取准确引用片段及版本、页序、原始坐标和识别疑点。单片段不证明完整上下文。
| Name | Required | Description | Default |
|---|---|---|---|
| evidence_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuine behavioral content about the return payload (version, page order, raw coordinates, recognition doubts) and a scope caveat about single snippets not constituting full context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler: the first enumerates the returned facets, the second states the key limitation. Efficient, though terse enough that some useful context is omitted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return values and does so by listing the fields returned. For a single-parameter read-only lookup it is nearly complete, missing only a hint about what evidence_id refers to or where such IDs originate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions evidence_id, so it adds nothing beyond the schema. However, with a single self-explanatory required identifier, the gap is minor and the baseline of 3 fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: read the accurate citation snippet, and enumerates what comes with it (version, page order, raw coordinates, recognition doubts). This distinguishes it from generic siblings like read_page/read_document, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is the caveat '单片段不证明完整上下文' (a single snippet does not prove the full context), which warns about interpretation but does not say when to use this tool versus read_page, read_document, or build_evidence_pack. No preconditions or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_extractionBRead-only
有界读取已缓存的TextIn原始抽取JSON文本;保留上游位置字段,无可靠定位的字段不得作为核验通过的证据。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| job_id | Yes | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description contributes meaningfully beyond that by disclosing the bounded/paginated nature of the read and the constraint on position fields, but it does not explain return format, pagination behavior, or cache freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact statement with the core operation front-loaded and the evidence caveat appended; there is no filler. Density is high, though the evidence clause bundles two ideas that could be structured more clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With read-only annotations present and no output schema to explain, the description is adequate on operation and safety, but incomplete on selection context and on the three undocumented parameters. An agent knows roughly what it reads but not clearly when to prefer it over sibling readers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it largely does not. '有界读取' loosely implies the limit/offset window and '已缓存' implies job_id targets a cached job, but the defaults (limit=16000, offset=0), units, and job_id source are left entirely to the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (读/read) and a specific resource (已缓存的TextIn原始抽取JSON文本), and the 'bounded' qualifier narrows the operation. It does not, however, distinguish itself explicitly from close siblings like read_parse_result, read_evidence, or read_document, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a downstream consumption rule ('fields without reliable positioning must not serve as verification-passing evidence'), but says nothing about when to choose this tool over read_parse_result/read_evidence/read_document. There is no when-to-use or when-not-to-use guidance for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_pageARead-only
返回原始PDF页或图片的真实MCP图像,用于核对公式、图表、双栏及OCR误差。page为从1开始的PDF页序,不是印刷页码。
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| version_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds meaningful behavior beyond them: the return is a real image, not extracted text, which is the key trait an agent needs. It does not mention size limits, resolution, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded purpose followed by the critical parameter clarification. No filler, though it is slightly terser than ideal given the undocumented version_id.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully discloses that the return is an image, which an agent needs to interpret the result. However, it omits what version_id refers to and how to obtain it, leaving one of two required parameters undocumented in both schema and description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the full burden. It does an excellent job on 'page' (1-based PDF page order, not printed page number), which is a real trap, but leaves 'version_id' entirely unexplained, so half the parameters remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('returns the raw PDF page or image as a real MCP image') and clearly distinguishes itself from text-returning siblings like read_document by emphasizing it returns visual page content. The verification use cases (formulas, charts, two-column, OCR errors) sharpen the purpose, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context — use it to verify formulas, charts, two-column layouts and OCR errors that text extraction would miss. It does not name an alternative tool to prefer instead, so it stops short of the explicit when-not guidance required for a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_parse_resultARead-only
有界读取已缓存的完整xParse原始JSON,包含上游返回的表格结构、字符细节和页元数据;不上传或重提。按next_offset续接。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| job_id | Yes | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/openWorldHint=false/destructiveHint=false, so safety is covered. The description adds real context beyond that: results are cached rather than re-fetched, no upload or resubmit occurs, and the read is bounded with offset-based continuation. It still says nothing about what happens if the job is incomplete or the cache is expired.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the action and resource, then adds the boundaries and the continuation rule. Dense but every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description does tell the agent the shape of the returned payload (tables, character detail, page metadata) and how to page through it. It omits any mention of the required job_id's meaning and any failure/empty-cache behavior, so it is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. '有界读取' and '按next_offset续接' do convey the limit/offset paging semantics and that next_offset is the continuation token, but job_id (the single required parameter) is left unexplained in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: a bounded read of the already-cached full xParse raw JSON, including tables, character detail and page metadata. The clause '不上传或重提' (does not upload or resubmit) implicitly separates it from parse_document/retry_parse, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the correct context (consuming results already produced by an xParse job) and gives a continuation rule via next_offset. It never states when to prefer this over get_parse_job or read_document, nor any preconditions such as 'only after job completion', leaving the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_recordARead-only
分页读取已保存研究、方法评价、来源、关系或检索记录的完整JSON。先list_records找ID;按next_offset续接。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| record_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds meaningful behavioral context: this is a paginated read returning full JSON, and continuation should use next_offset. It does not mention rate limits or authentication, but that is not critical given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two compact sentences, front-loaded with the core action and followed by the essential prerequisite and pagination instruction. Every phrase is useful and there is no redundant text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety and no output schema, the description provides enough to invoke the tool: record types, prerequisite lookup via list_records, pagination continuation, and that full JSON is returned. The missing explanation of limit/offset defaults is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies record_id usage ('先list_records找ID') and pagination continuation ('按next_offset续接'), covering two of three parameters partially. It does not explain the limit parameter or its default of 16000, leaving a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('读取' / read) and resource ('完整JSON' of saved research, method reviews, sources, relations, or retrieval records). It distinguishes from list_records by instructing the agent to find IDs there first, but it does not explicitly differentiate from other read_* siblings like read_evidence or read_extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: use list_records first to find the ID, then call read_record, and use next_offset to continue pagination. It does not state when not to use this tool or name alternatives beyond list_records.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_identifierA
解析DOI、pmid:或arxiv:标识符并保存题录;不获取全文。
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, destructiveHint=false; the description is consistent, confirming it performs a network lookup and persists a record ('保存题录'). It adds the scope boundary of not fetching full text, which is useful behavioral context beyond the annotations. It does not cover failure behavior or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the action and the accepted identifier forms front-loaded, followed by the scope limit. Efficient, though the parameter meaning for project_id is simply absent rather than concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-required-param mutation tool with no output schema and 0% schema coverage, the description covers the identifier formats and the save scope but omits what project_id is and what the tool returns on success or on an unresolvable identifier.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does add real value for 'identifier' by naming the accepted prefixes (DOI, pmid:, arxiv:), but 'project_id' is left completely unexplained in both schema and description, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (解析/resolve) plus the resource (DOI/pmid:/arxiv: identifiers) and its effect (保存题录/save bibliographic record). The closing clause '不获取全文' explicitly separates it from download_fulltext and other full-text siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '不获取全文' clause tells the agent the condition under which this tool stops short of full-text retrieval, effectively routing full-text needs elsewhere. It does not name a specific alternative tool (e.g. download_fulltext) by name, but the use context is clear for a known-identifier resolve.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_libraryA
从用户指定可信备份恢复至全新目录,校验哈希和数据库;不覆盖现有库,恢复后云服务和正文输出权限全部关闭。
| Name | Required | Description | Default |
|---|---|---|---|
| archive_path | Yes | ||
| new_directory | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false and readOnlyHint=false, and the description corroborates with '不覆盖现有库'. Beyond that it adds real behavioral value: hash/database verification during restore and the notable side effect that cloud services and full-text output permissions are disabled post-restore.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the action front-loaded and clauses separated by semicolons. Efficient with no filler, though the permission caveat and non-overwrite guarantee are packed in without further structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with annotations covering the safety profile and no output schema, the description covers the essential what, the non-destructive guarantee, verification steps, and the permission side effect. Return value details are the only notable omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It conceptually maps the two params (archive_path = trusted backup source, new_directory = the fresh target directory), but adds no format, path, or validation detail for either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (恢复/restore) plus resource (library) and scope (from a user-specified trusted backup into a brand-new directory). This clearly distinguishes it from the inverse sibling backup_library and from preview_deletion/delete_work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the scenario (recovering from a trusted backup) and notes it will not overwrite the existing library, which is useful context. But it names no alternative tool and gives no explicit when-to-use vs when-not-to-use guidance relative to backup_library or the deletion tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_parse_observationB
用户解决限流后恢复已知任务的查询;不会上传或重提,下一次get_parse_job(poll=true)才联网。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| user_note | Yes | ||
| expected_revision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=false, destructiveHint=false, so the agent already knows this mutates local state without being destructive or going online. The description usefully confirms the no-upload/no-resubmit boundary and that network activity is deferred to get_parse_job(poll=true), but it never says what local state it actually changes, which is the key gap for a non-read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the trigger condition and followed by the side-effect boundary. Nothing is wasted, though the extreme terseness leaves several terms undefined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating tool with no output schema and no annotations covering semantics, the description positions the tool correctly in the parse workflow and clarifies the network boundary, but omits any explanation of the required parameters or the resulting state, which an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all three parameters (job_id, expected_revision, user_note) are undocumented in both schema and description. expected_revision in particular implies an optimistic-concurrency precondition and user_note implies a required annotation, but neither is explained, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (resume the query/observation of a known task) and explicitly scopes it against the sibling get_parse_job by stating that no upload or resubmission occurs. It is clear enough to distinguish from retry_parse and get_parse_job, though the notion of an 'observation' state is left somewhat abstract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition (after the user has resolved rate limiting) and points to the alternative that performs the actual network call: the next get_parse_job(poll=true). This routes the agent well, but it does not state when this tool should NOT be used or what happens if the rate limit is still active.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retry_parseA
仅按用户明确重试决定调用。失败/提交未知可能已计费;本次另预留页数,不能保证上游只计费一次。先get_parse_job读取revision。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| user_note | Yes | ||
| expected_revision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state that it is a non-read, non-destructive, open-world operation. The description adds critical behavioral context the annotations cannot: a failed or unknown submission may already have been billed, this call reserves additional page quota, and upstream cannot guarantee single billing. For a cost-incurring action this is unusually valuable disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the most important constraint (only on explicit user retry), followed by cost warning and prerequisite. Dense but every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description needn't explain returns, and it covers the gating rule, cost risk, and prerequisite thoroughly. The main gap is that two of three required parameters remain undefined, but for a guarded action tool the operational guidance is otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It does explain how to obtain expected_revision (read it via get_parse_job), which is the highest-risk parameter, but it says nothing about job_id or user_note semantics despite both being required. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name and text make clear this re-runs a parse job ('重试' / retry), and it explicitly distinguishes itself from the sibling get_parse_job by telling the agent to read the revision there first. It does read more like a usage/policy note than a statement of what the operation does (e.g., it never says it re-enqueues a failed parse), but the verb+resource is inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit gating condition ('仅按用户明确重试决定调用' – only call on a user's explicit retry decision) and a concrete prerequisite ('先get_parse_job读取revision'), naming the sibling to use first. This is exactly the when-to-use / when-not-to-use guidance an agent needs on a costly, non-idempotent action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_method_reviewB
保存宿主的方法评价及局限,必须关联该文献原文;这是AI评价而非已确认的来源身份或客观质量分数。
| Name | Required | Description | Default |
|---|---|---|---|
| work_id | Yes | ||
| rationale | Yes | ||
| limitations | Yes | ||
| evidence_ids | Yes | ||
| research_type | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a write (readOnlyHint=false) that is non-destructive, so the safety profile is covered. The description usefully adds that the stored content is an AI-generated evaluation rather than a verified fact, which is meaningful semantic context, but it says nothing about overwrite behavior, auth requirements, or what happens on repeat saves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a semicolon separating the action from the caveat; no filler. It is appropriately compact, though the caveat could be tightened into the primary clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-required-parameter write tool with 0% schema coverage and no output schema, the description is too thin: it never defines the parameter contract (required evidence linkage is only mentioned in prose) nor the response behavior. The agent knows the semantic intent but not how to populate the fields correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 required parameters, so the schema provides no per-parameter help. The description only loosely gestures at 'rationale' and 'limitations' concept, leaving work_id, evidence_ids, and research_type completely unexplained in both places. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (保存/save) and resource (方法评价及局限/method review and limitations), and adds scope constraints (must link to the source document; it is an AI evaluation, not a confirmed identity or quality score). This lets an agent distinguish it conceptually from verification-style siblings like verify_work, though no sibling is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause '这是AI评价而非已确认的来源身份或客观质量分数' implicitly separates it from verification/identity tools, and '必须关联该文献原文' states a precondition. However, there is no explicit when-to-use vs when-not framing and no alternative tool is named, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_researchA
保存宿主生成的阅读卡片/回答/综述。author主张必须引原文;引用须同项目、当前解析、精确摘录。此工具仅校验引用完整性,语义支持需宿主/用户检查。产物不进入原文索引。
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| title | Yes | ||
| claims | Yes | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=false. The description adds real behavioral context: it only validates citation completeness, semantic support must be checked by host/user, and artifacts do not enter the original text index. These are non-obvious traits that annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the action and resource before the constraints. No filler, though the citation-rule sentence packs several conditions without separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so return behavior is unclear (does it return an id or validation report?). The description covers validation scope and index behavior adequately for a write tool, but omits enough parameter meaning that an agent could mis-populate claims or notes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 params, so the description must carry the load. It does add meaningful citation semantics (same project, current parse, exact excerpt; author claims must cite source) but says nothing about project_id, title, notes, or the Claim kind/review/support enums, leaving half the model undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb (保存/save) and the resource (阅读卡片/回答/综述 – host-generated reading cards, answers, summaries). This clearly distinguishes it from siblings like save_method_review or export_research, though it never explicitly names those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (persisting host-generated artifacts) and lays out citation rules, but gives no explicit when-to-use-vs-alternative guidance such as 'use save_method_review for method notes'. Usage is inferable but not routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_libraryBRead-only
本地中英文关键词检索,限定可选版本和字符预算;仅召回当前原文证据,摘要和笔记不作为原始证据。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| project_id | Yes | ||
| char_budget | No | ||
| version_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description adds that only current original text evidence is returned and that summaries/notes are excluded, which is useful result-scope context, but it does not describe pagination, ranking, or truncation behavior for char_budget.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the search type and evidence scope. It contains no filler, though the semicolon-joined clauses could be separated for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Considering five parameters, no schema descriptions, and no output schema, the description covers the core purpose and evidence constraints but leaves usage context, required parameter roles (beyond schema), and result handling unspecified. An agent can invoke it, but with notable uncertainty around project scoping and result limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains that version_ids and char_budget are optional limits; it does not cover project_id, query, or limit. Query is indirectly suggested by '关键词检索', but required project scoping and result paging are left unclarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation: a local Chinese/English keyword search over a project's current original text evidence, with optional version and character-budget constraints. It does not explicitly name or compare against sibling tools such as search_literature or index_semantic_search, so it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are provided. The phrase '本地' implies local use and the evidence scope hints at a distinction from summary-based retrieval, but an agent must infer when to choose this tool over search_literature or index_semantic_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_literatureA
在crossref/openalex/arxiv/pubmed检索,分别报告失败及成功。保存检索式与结果,不自动纳入或下载;会消耗项目请求预算。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| sources | Yes | ||
| project_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, destructiveHint=false; the non-readOnly hint is consistent with the stated 'saves query and results'. The description adds genuinely new behavioral context beyond the annotations: per-source failure reporting, no automatic inclusion/download, and that the call '会消耗项目请求预算' (consumes the project request budget) — a cost/rate signal the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the sources and the search action, followed by output behavior and cost. No filler, though the clauses are compressed to the point of terseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, 0%-coverage tool with no output schema, the description covers source values, side effects, and budget cost, but omits what the return payload actually looks like (per-source result lists? stored record IDs?) and the meaning of limit/project_id. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with no enums, so the description carries extra weight. It usefully enumerates the valid source identifiers (crossref/openalex/arxiv/pubmed) that the 'sources' array accepts, which the schema does not. It says nothing about limit, query, or project_id, leaving half the parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (检索/search) and enumerates the four external resources searched (crossref/openalex/arxiv/pubmed), making it distinguishable from the local-library sibling search_library. It also specifies the reporting shape (per-source success/failure). Strong, though it never explicitly contrasts itself with search_library.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies a discovery-only role by stating '不自动纳入或下载' (does not auto-include or download), which tells the agent this is not import_references or download_fulltext. However, it never names those alternatives or states the prerequisite/context for choosing this tool, leaving the routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_project_policyA
按用户授权修改项目服务与预算;先读取当前revision。预算不能低于已预留用量。
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| project_id | Yes | ||
| expected_revision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state readOnlyHint=false, destructiveHint=false, openWorldHint=false. The description adds genuinely new behavioral context: the operation requires user authorization, requires reading the current revision beforehand (optimistic-concurrency semantics behind expected_revision), and enforces a validation rule that budget cannot go below already-reserved usage. These are real constraints an agent could not infer from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the action and its authorization precondition, followed by the constraint. No filler; every phrase carries operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and 0% parameter description coverage, the description covers authorization, concurrency, and one validation rule, which is a reasonable core. It still omits failure behavior on revision mismatch, whether unmentioned policy fields are preserved or reset, and the meaning of the service-level fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for project_id, policy, and expected_revision. It partially compensates by clarifying the budget lower-bound constraint and the meaning of expected_revision (the revision that must be read first), but it says nothing about project_id or the individual policy fields (textin, metadata, host_text, embeddings), leaving several parameters opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource in Chinese: 修改 (modify) 项目服务与预算 (project services and budget), which maps onto the boolean service toggles and the integer budget fields in the schema. It is clear what the tool does, though it does not name or contrast with any sibling tool (e.g. create_project or workspace_status).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a precondition (按用户授权 / requires user authorization) and an ordering rule (先读取当前revision / read the current revision first), which is real usage guidance. However, it names no alternative tool and gives no explicit when-not guidance, so the routing value is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_workB
修订题录和阅读/纳入状态,保留历史;excluded必须提供理由。selection:unread/reading/included/excluded/read。
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| work_id | Yes | ||
| metadata | Yes | ||
| selection | Yes | ||
| expected_revision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the safety profile is known. The description adds that history is preserved, which is useful behavioral context, but it omits the concurrency/optimistic-locking semantics of expected_revision and what happens on a revision mismatch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with the core mutation scoping front-loaded and the status enum appended. No filler, though the terse semicolon style trades clarity for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 5 required params, a nested metadata object, 0% schema coverage, and no output schema, the description is under-specified. It never explains the metadata shape or the expected_revision concurrency contract an agent must satisfy to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the burden. It documents the selection enum (unread/reading/included/excluded/read) and the reason requirement for 'excluded', but leaves work_id, the free-form metadata object, and especially expected_revision unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (修订/revise) and resource (题录和阅读/纳入状态), and the phrase '保留历史' distinguishes it from destructive siblings like delete_work. An agent can tell it edits an existing record rather than creating one, though it doesn't explicitly name any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete usage rule ('excluded必须提供理由') but no guidance on when to prefer this over verify_work, add_work, or delete_work. Usage is implied rather than contrasted against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_workA
联网核对DOI与标题及Crossref更新通知,保留证据和查询时间;未查到撤稿不等于可靠。
| Name | Required | Description | Default |
|---|---|---|---|
| work_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true, readOnlyHint=false and destructiveHint=false, and the description is consistent with them, adding that evidence and query time are retained (explaining the write behavior implied by readOnlyHint=false) and that absence of a retraction does not imply reliability. That limitation caveat is genuinely useful behavior/interpretation context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence plus a semicolon-delimited caveat, with the core action front-loaded and zero filler. It is efficient, though the terse phrasing leaves selection criteria compressed rather than explicit.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a 1-param tool, the description covers the network/source behavior and a key interpretive caveat but omits parameter format and any usage routing. It is adequate but leaves real gaps for an agent deciding how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single required work_id, so the schema contributes nothing. The description only indirectly implies what work_id is (a DOI/title-bearing work identifier) without stating format, accepted identifier types, or whether an internal ID or DOI string is expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: online verification of a work's DOI and title against Crossref. The scope (DOI/title cross-check, retraction status) distinguishes it from generic siblings like resolve_identifier or search_literature. It names its external source, which sharpens the purpose further, though it does not explicitly contrast itself with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied – call it to verify a work's DOI/title and pick up Crossref update notices – but there is no explicit when-to-use vs. when-to-use-something-else, and no alternatives named among the many verification/search siblings. The retraction caveat is interpretive guidance, not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_statusARead-only
检查安装版本、本地库、已配置凭据(不返回秘密)、项目;不发起联网请求。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false. The description adds useful context beyond those annotations: it explicitly states it does not return secrets and does not make network requests, which reinforces safety and offline behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It efficiently lists the checked items and appends important behavioral constraints in the same compact structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only diagnostic tool with no output schema, the description is complete enough: it names what it checks, confirms no secrets are returned, and states no network calls are made. Annotations cover the safety profile, and no return-value details are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters (empty schema), so the baseline is 4. The description does not need to document parameter meanings, and it correctly focuses on what the tool inspects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (检查/check) and enumerates the resources it inspects: installed version, local library, configured credentials, projects. This clearly distinguishes it from mutation and read siblings, though it does not explicitly name an alternative tool for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not state when to use this tool versus alternatives, nor any preconditions or exclusions. It only restates what the tool checks, leaving the agent to infer its diagnostic purpose from the purpose statement alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
38 tool updates
v0.1.0- First observed
add_work - First observed
backup_library - First observed
build_evidence_pack - First observed
connect_papers - First observed
convert_document - First observed
create_project - First observed
delete_work - First observed
download_fulltext - First observed
export_research - First observed
extract_fields - First observed
find_duplicate_candidates - First observed
get_parse_job - First observed
import_document - First observed
import_references - First observed
index_semantic_search - First observed
list_input_files - First observed
list_records - First observed
parse_document - First observed
parse_local_pdf - First observed
preview_deletion - First observed
read_document - First observed
read_evidence - First observed
read_extraction - First observed
read_page - First observed
read_parse_result - First observed
read_record - First observed
resolve_identifier - First observed
restore_library - First observed
resume_parse_observation - First observed
retry_parse - First observed
save_method_review - First observed
save_research - First observed
search_library - First observed
search_literature - First observed
update_project_policy - First observed
update_work - First observed
verify_work - First observed
workspace_status
TDQS
Scored across 38 tools
Tools are mostly well separated by resource and action (e.g., search_literature vs search_library, parse_local_pdf vs parse_document). Some adjacent read/extract/parse-job tools (read_extraction vs read_parse_result, get_parse_job vs retry_parse vs resume_parse_observation) require careful reading, but descriptions clarify boundaries.
All names use snake_case and nearly all follow a verb_noun or action_object pattern. workspace_status is a minor noun-only exception, but the overall naming convention is predictable and readable.
38 tools is above the typical 3-15 band and exceeds the 25+ heavy threshold. The parsing, extraction, and job-management pipeline contains many granular tools that could be consolidated for easier agent selection.
The surface covers project setup, import, parsing, search, evidence building, export, deletion, backup/restore, and more. Minor lifecycle gaps exist (e.g., no explicit project-level delete/list beyond workspace_status), but major research workflows are well covered.
Maintenance
Related MCP Connectors
PubMed MCP — wraps the NCBI E-utilities API (biomedical literature, free, no auth)
Auditable MCP server for PubMed, Europe PMC, ClinicalTrials.gov, and bioRxiv/medRxiv queries
Multi-engine scholarly research server for search, traversal, full text, and reading lists.
Academic literature search, retrieval, and private library management on top of OpenAlex.
Related MCP Servers
- FlicenseCqualityDmaintenanceExposes a local biomedical literature pipeline as MCP tools for automated research workflows. Enables literature search, open-access paper retrieval, and draft generation for biomedical and pathology domains through standard MCP clients.6-
- AlicenseNot gradedqualityCmaintenanceA local-first MCP server that analyzes research papers, maps citation graphs, and surfaces insights with verbatim-verified contradictions, all while keeping data private on your machine.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables MCP clients to interact with a local-first research knowledge workbench, supporting literature search, evidence-grounded Q&A, and reference export.2AGPL 3.0
- AlicenseBqualityAmaintenanceEnables provenance-first scholarly retrieval, paper ingestion, and reproducible research workflows by searching academic and developer sources, extracting source-located facts and claims, and preserving evidence and provider uncertainty for MCP clients.371Apache 2.0