PaperGraph MCP
This server provides an MCP interface for PaperGraph, an evidence-first math paper reading workflow. It lets you load papers (arXiv, local LaTeX, PDFs), build theorem graphs, inspect proof/citation evidence, create reading sessions/queues, export Markdown reports/plans, and resolve external references via scholarly metadata search and bounded reference expansion.
Load and manage papers: open a persistent SQLite workspace; add local LaTeX projects, arXiv sources, or born-digital PDFs; list/get stored papers.
Map and inspect results: get a Paper Map, list/search theorem-like results, retrieve result text/proofs, source spans, dependencies, dependency diagnostics, citations, and external result mentions.
Read proofs with evidence: export reading bundles or focused result reading contexts, get source slices, and compute deterministic local reading paths.
Track reading state: create/list/get reading sessions, record checkpoints, add notes, export session summaries, create/list/get reading queues, and apply queues to sessions.
Export durable artifacts: generate deterministic Markdown Reading Reports, Cross-Paper Reading Plans, and Workspace Starter artifacts (START_HERE.md, manifest, reports/plans).
Resolve blocked references: search Crossref/OpenAlex/arXiv metadata for blocked citations, list searches/resolutions, apply candidates, or resolve to DOI/URL/PDF/arXiv targets.
Plan external imports: generate import plans for a result, queue, or whole paper.
Bounded reference expansion: create/advance/pause/cancel/export reference expansion runs with approved budgets, auto-import unique strong candidates, and record decisions for ambiguous edges.
Validate raw arXiv requests: normalize bare IDs, URLs, Markdown links, and prose; choose safe next actions or ask the user when ambiguous.
Run diagnostics: check environment and launch reproducibility via
get_environment_diagnostics.
Allows loading arXiv papers by identifier, downloading sources from arXiv's e-print endpoint, caching them, and building a dependency graph of theorem-like environments.
Provides tools for parsing local LaTeX papers, recursively following \input and \include commands, and exploring theorems, dependencies, and usages in the resulting document graph.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PaperGraph MCPload arXiv paper 2401.12345 and list its theorems"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
PaperGraph MCP
Read math papers with evidence, not guesses.
PaperGraph v1.1.5 is the stable evidence-first reading workflow for math papers: start with a Paper Map, inspect source-backed proof and citation evidence, use Evidence Triage to understand sparse extraction, search scholarly metadata for blocked references, then export Markdown Reading Reports or Cross-Paper Reading Plans.
It helps AI agents turn arXiv papers, local LaTeX projects, and born-digital PDFs into a local theorem-centered workspace so a researcher can inspect where every claim came from.
English
What PaperGraph Helps You Do
Start with a Paper Map | Trace proof evidence | Resolve references | Save reading artifacts |
Identify main-result candidates, result structure, proof-path evidence, and external reading risks before choosing where to read. | Inspect proof-local references, cited stops, source slices, and dependency diagnostics with explicit evidence. | Search Crossref, OpenAlex, and arXiv metadata for blocked references, then apply only a chosen candidate through the reference closure workflow. | Export deterministic Markdown reports and cross-paper reading plans that can live in Git, notes, or handoff sessions. |
PaperGraph v1.1.3 adds Scholarly Reference Resolver: low-risk online metadata search for blocked references, deterministic candidate ranking, and explicit boundary messages when the trail stops at ambiguous or non-importable records.
v1.1.4 introduced bounded reference expansion: approve a finite policy, then advance a saved run. Defaults are depth 2 and 10 new papers; unique strong importable identities can be selected automatically, while ambiguous references await review and independent branches continue. Runs retain budgets, evidence, decisions and recovery history. Only arXiv sources and explicitly supplied local PDFs are importable. It does not bypass paywalls or perform unlimited crawling.
v1.1.5 improves reference identity quality: traceable bibliography hints, conservative DOI/arXiv normalization, explicit conflicts, provider outcomes and saved matching explanations. New runs use unique_strong_v2; existing unique_strong_v1 tasks keep their legacy resolver. Back up workspaces before upgrading to schema 9; older versions cannot open them. Scores are not probabilities, and metadata agreement is not independent verification. See the release preparation notes and offline quality corpus.
See the expansion walkthrough, offline JSON, reference tree, and client verification matrix. This branch prepares v1.1.5; the pinned tag commands below become available only after that release is published. Until then, use uv run papergraph-mcp from this checkout.
Why Researchers Use It
Need | How PaperGraph behaves |
"Do not invent dependencies." | PaperGraph reports evidence-backed links and explains empty results as extraction limits, not mathematical facts. |
"Show me the exact source." | Results, proofs, dependencies, and citations carry source spans that can be sliced back out of the original paper. |
"Let me review external papers first." | External references become import plans. Cross-paper plans separate selected-paper citation evidence from unresolved outside risks. |
"Keep my reading state." | Workspaces store queues, sessions, checkpoints, notes, blocked targets, and open questions locally. |
PaperGraph does not verify proofs, perform semantic theorem matching, or claim that similarly worded results are equivalent.
Quick Start
Install uv, then verify the pinned GitHub release without cloning:
uvx --from git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5 papergraph-mcp --version
uvx --from git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5 papergraph-mcp doctorPinning the v1.1.5 tag keeps MCP client installations reproducible.
Add PaperGraph to an MCP client that accepts JSON-style stdio configuration:
{
"mcpServers": {
"papergraph": {
"command": "uvx",
"args": ["--from", "git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5", "papergraph-mcp"]
}
}
}Restart the MCP client after changing its configuration. The server uses stdio, so running the command without --help or --version waits quietly for an MCP client connection.
Ask your agent to set it up
Give a coding agent this request:
I use an MCP-capable agent/client. Clone https://github.com/lotchuazzz-crypto/papergraph-mcp and help me configure PaperGraph for it. After cloning, read .agents/skills/setting-up-papergraph/SKILL.md and follow it.
Compatible agents can follow the repository-local setting-up-papergraph skill. The agent should show you a reusable PaperGraph prompt, explain why uv is needed, and ask before installing software, changing client configuration, or restarting the client.
If you do not use an MCP-capable client yet, PaperGraph can still be run from the CLI with the pinned uvx --from ... papergraph-mcp doctor command above and the workspace commands below.
If your agent clones into a directory that already exists, ask it to run git fetch --tags origin before treating the checkout as current. Existing clones can otherwise remain pinned to an old local origin/main.
For a complete first run, follow First PaperGraph Workspace. The Workspace Starter commands plan-starter-project and bootstrap-reading-project create START_HERE.md, papergraph-starter-manifest.json, Reading Reports, and a Cross-Paper Reading Plan from explicit paper inputs. For the stable surface, see PaperGraph v1 Core Contract and the v1 Release Checklist. Example outputs are available as a Reading Report, a Cross-Paper Reading Plan, a Reference Search, a Reference Resolution, a Starter Summary, and a Starter Manifest.
For raw user requests, prefer load_arxiv_request(input=...) or papergraph-mcp load-arxiv-request "...". These high-level entry points validate bare IDs, URLs, Markdown links, and prose before loading. To inspect the decision without loading, call validate_arxiv_request or papergraph-mcp validate-arxiv-request "...". If validation returns action: ask_user_to_choose, ask the user to choose; detecting a conflict and then continuing is a failure. Use load_arxiv_paper only after the user has provided one already-disambiguated arXiv ID.
A Typical Reading Flow
flowchart LR
Paper[Paper] --> Results[Extract results]
Results --> Evidence[Inspect proof evidence]
Evidence --> Path[Build reading path]
Path --> Queue[Create reading queue]
Queue --> Imports[Review external import plan]
Queue --> Session[Resume reading session]Load a paper from arXiv, local LaTeX, or PDF.
List theorem-like results and choose a target theorem.
Inspect the theorem statement, proof evidence, source slice, and dependency diagnostics.
Generate a reading queue from local proof evidence.
Export a single-paper Reading Report when you want a durable Markdown handoff.
For a few related papers, export a Cross-Paper Reading Plan to see selected-paper citation evidence and remaining risks.
Save checkpoints and notes so the next reading session starts from known state.
What PaperGraph Does Not Do
It does | It does not |
Extract and store evidence from papers. | Prove the paper is correct. |
Follow explicit labels, proof-local references, and citation evidence. | Guess hidden mathematical prerequisites. |
Build reviewable reading queues and import plans. | Automatically crawl the literature. |
Keep local reading state in SQLite. | Upload private manuscripts or PDFs. |
Related MCP server: arxiv-reader-mcp
中文
PaperGraph 能帮你做什么
先看 Paper Map | 追踪证明证据 | 保存阅读产物 | 跨论文规划 |
在选择阅读目标前,先看到 main-result candidates、结果结构、proof-path evidence 和 external reading risks。 | 查看 proof-local references、citation stops、source slices 和 dependency diagnostics,并保留证据来源。 | 导出确定性的 Markdown report,方便放进 Git、笔记或交接会话。 | 对一组显式给定的小规模相关论文,导出跨论文阅读计划、选中论文之间的 citation evidence 和剩余风险。 |
PaperGraph v1.1.5 是稳定的 evidence-first 数学论文阅读工作流:先看 Paper Map,再检查 proof 和 citation 证据,用 Evidence Triage 理解稀疏抽取结果,对被阻塞的外部引用做 scholarly metadata 搜索,然后把选定候选闭环到用户确认的 arXiv、本地 PDF、DOI、URL 或出版信息。
v1.1.4 引入有界引用扩展:默认最多追踪 2 层、导入 10 篇新论文。v1.1.5 进一步改善引用身份匹配:保留解析证据,审慎规范化 DOI/arXiv,明确冲突、服务状态和选择理由。新任务默认 unique_strong_v2,已有 unique_strong_v1 任务保持旧行为。升级到 schema 9 前请备份 workspace,旧版本无法打开新 schema。分数不是概率,多来源元数据一致也不代表独立验证。只支持 arXiv 源码和明确提供的本地 PDF,不绕过付费墙。见完整操作示例及客户端验证状态。当前为发布准备;v1.1.5 标签发布前请使用 uv run papergraph-mcp。
为什么适合数学论文阅读
研究者关心的问题 | PaperGraph 的回答 |
不要猜依赖。 | 只报告有证据的链接;空依赖结果解释为抽取限制,而不是数学事实。 |
我要看到原文位置。 | result、proof、dependency、citation 都尽量保留 source span,可回到原文片段。 |
外部论文先让我审。 | 外部引用先变成 import plan;跨论文计划会区分选中论文之间的 citation evidence 和仍在外部的 unresolved risks。 |
阅读项目要能继续。 | workspace 在本地保存 queue、session、checkpoint、note、blocked target 和 open question。 |
PaperGraph does not verify proofs,也不做 semantic theorem matching;它不会声称两个措辞相似的结果数学上等价。
快速开始
先安装 uv,然后验证固定版本:
uvx --from git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5 papergraph-mcp --version
uvx --from git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5 papergraph-mcp doctor如果你的 MCP client 使用 JSON 风格的 stdio server 配置,可以添加:
{
"mcpServers": {
"papergraph": {
"command": "uvx",
"args": ["--from", "git+https://github.com/lotchuazzz-crypto/papergraph-mcp.git@v1.1.5", "papergraph-mcp"]
}
}
}修改配置后重启 MCP client。这个 server 使用 stdio,所以不带 --help 或 --version 直接运行时,会安静等待 MCP client 连接。
让 agent 帮你设置
你可以把这段话发给 coding agent:
我使用的是支持 MCP 的 agent/client。请克隆 https://github.com/lotchuazzz-crypto/papergraph-mcp,并帮我把 PaperGraph 配置进去。克隆后请先阅读 .agents/skills/setting-up-papergraph/SKILL.md 并按它执行。
支持本仓库 skill 的 agent 会读取 setting-up-papergraph,展示可复用提示词,解释为什么需要 uv,并在安装软件、修改客户端配置或重启客户端前询问你。
如果你暂时没有支持 MCP 的 client,也可以先用 CLI:运行上方固定版本的 uvx --from ... papergraph-mcp doctor 和下方 workspace 命令。
如果目标目录已经存在,请让 agent 先运行 git fetch --tags origin,再判断仓库是否是最新。否则已有 clone 可能仍停留在旧的本地 origin/main。
第一次完整使用可以跟着 First PaperGraph Workspace 走。稳定承诺见 PaperGraph v1 Core Contract,发布前检查见 v1 Release Checklist。示例输出见 Reading Report、Cross-Paper Reading Plan、Reference Search 和 Reference Resolution。
普通用户请求优先走 load_arxiv_request(input=...) 或 papergraph-mcp load-arxiv-request "..."。这些入口会在加载前验证 bare IDs、URLs、Markdown links 和自然语言描述。若验证返回 action: ask_user_to_choose,必须让用户选择;detecting a conflict and then continuing is a failure。Use load_arxiv_paper only after 用户已经给出单一、无歧义的 arXiv ID。
典型阅读流程
从 arXiv、本地 LaTeX 或 PDF 加载论文。
列出 theorem-like results,选择目标定理。
查看 theorem statement、proof evidence、source slice 和 dependency diagnostics。
根据本地 proof evidence 生成 reading queue。
需要持久交接时,导出单篇 Reading Report。
面对几篇相关论文时,导出 Cross-Paper Reading Plan,查看选中论文之间的 citation evidence 和剩余风险。
保存 checkpoints 和 notes,下次继续读时不必从头开始。
Reference
Core Workflows
Workflow | Main tools |
Load papers |
|
Start a project |
|
Map papers |
|
Inspect results |
|
Read a proof |
|
Resume reading |
|
Plan reading |
|
Resolve references |
|
Original single-paper tools: get_environment_diagnostics, validate_arxiv_request, load_arxiv_request, validate_arxiv_input, load_paper, load_arxiv_paper, list_theorems, get_theorem, get_dependencies, get_dependency_diagnostics, and where_used.
Complete workspace tool index: open_workspace, workspace_add_local_paper, workspace_add_arxiv_paper, workspace_list_papers, workspace_get_paper, workspace_search_theorems, workspace_get_dependencies, workspace_get_dependency_diagnostics, workspace_get_citations, workspace_add_pdf_paper, workspace_get_paper_map, workspace_export_paper_reading_report, workspace_export_cross_paper_reading_plan, workspace_plan_starter_project, workspace_bootstrap_reading_project, workspace_list_results, workspace_get_result, workspace_get_result_proof, workspace_get_proof_dependencies, workspace_get_external_result_mentions, workspace_get_evidence, workspace_export_reading_bundle, workspace_export_result_reading_context, workspace_get_source_slice, workspace_get_result_reading_path, workspace_create_reading_session, workspace_list_reading_sessions, workspace_get_reading_session, workspace_record_reading_checkpoint, workspace_add_reading_note, workspace_export_reading_session_summary, workspace_create_reading_queue, workspace_list_reading_queues, workspace_get_reading_queue, workspace_apply_reading_queue_to_session, workspace_plan_external_imports_for_result, workspace_plan_external_imports_for_queue, workspace_plan_external_imports_for_paper, workspace_search_external_reference, workspace_list_external_reference_searches, workspace_resolve_external_reference_candidate, workspace_resolve_external_reference, workspace_list_external_reference_resolutions.
Most workspace operations are available from the CLI with --workspace:
papergraph-mcp validate-arxiv-request "[math/0307200](https://arxiv.org/abs/2609.01574)"
papergraph-mcp plan-starter-project --workspace .\papergraph.sqlite3 --artifact-dir .\papergraph-starter --pdf .\paper-a.pdf=local:paper-a
papergraph-mcp bootstrap-reading-project --workspace .\papergraph.sqlite3 --artifact-dir .\papergraph-starter --pdf .\paper-a.pdf=local:paper-a --no-queue --no-session
papergraph-mcp get-paper-map --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp export-paper-reading-report --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp export-paper-reading-report --workspace .\papergraph.sqlite3 --paper-id local:paper-a --output report.md
papergraph-mcp export-cross-paper-reading-plan --workspace .\papergraph.sqlite3 --paper-id local:paper-a --paper-id arxiv:2401.12345 --output cross-paper-plan.md
papergraph-mcp export-reading-bundle --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp export-result-reading-context --workspace .\papergraph.sqlite3 --result-id local:paper-a::thm:main
papergraph-mcp get-source-slice --workspace .\papergraph.sqlite3 --result-id local:paper-a::thm:main
papergraph-mcp get-result-reading-path --workspace .\papergraph.sqlite3 --result-id local:paper-a::thm:main
papergraph-mcp create-reading-session --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp record-reading-checkpoint --workspace .\papergraph.sqlite3 --session-id SESSION --target-kind result --target-id local:paper-a::thm:main --status reviewed
papergraph-mcp add-reading-note --workspace .\papergraph.sqlite3 --session-id SESSION --text "Need to check the cited fixed point theorem."
papergraph-mcp export-reading-session-summary --workspace .\papergraph.sqlite3 --session-id SESSION
papergraph-mcp create-reading-queue --workspace .\papergraph.sqlite3 --result-id local:paper-a::thm:main
papergraph-mcp list-reading-queues --workspace .\papergraph.sqlite3
papergraph-mcp get-reading-queue --workspace .\papergraph.sqlite3 --queue-id QUEUE
papergraph-mcp apply-reading-queue-to-session --workspace .\papergraph.sqlite3 --queue-id QUEUE --session-id SESSION
papergraph-mcp plan-external-imports-for-result --workspace .\papergraph.sqlite3 --result-id local:paper-a::thm:main
papergraph-mcp plan-external-imports-for-queue --workspace .\papergraph.sqlite3 --queue-id QUEUE
papergraph-mcp plan-external-imports-for-paper --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp search-external-reference --workspace .\papergraph.sqlite3 --paper-id local:paper-a --blocked-id BLOCKED
papergraph-mcp list-external-reference-searches --workspace .\papergraph.sqlite3 --paper-id local:paper-a
papergraph-mcp resolve-external-reference-candidate --workspace .\papergraph.sqlite3 --paper-id local:paper-a --blocked-id BLOCKED --candidate-id CANDIDATE --import-target
papergraph-mcp resolve-external-reference --workspace .\papergraph.sqlite3 --paper-id local:paper-a --blocked-id BLOCKED --doi 10.1000/example --title "Published target"
papergraph-mcp resolve-external-reference --workspace .\papergraph.sqlite3 --paper-id local:paper-a --blocked-id BLOCKED --pdf .\reference.pdf=local:reference
papergraph-mcp list-external-reference-resolutions --workspace .\papergraph.sqlite3 --paper-id local:paper-aFor a compact single-paper check with an already-disambiguated ID, call load_arxiv_paper(arxiv_id="math/0307200"). For ordinary user text, call load_arxiv_request(input="math/0307200"). PaperGraph selects main.tex; a representative first response has "path": "main.tex", "cached": false, and "nodes": 7.
PaperGraph v0.4.4 dependency traversal uses statement_explicit_latex_refs_only: it follows explicit LaTeX references such as \ref, \eqref, \autoref, \cref, and \Cref inside theorem-like statements. An empty dependency result means PaperGraph found no resolvable theorem-label references under that rule. It is not evidence that the theorem has no mathematical dependencies.
Proof dependency extraction is evidence-scoped. PaperGraph looks inside TeX proof environments, direct proof continuations, and short text immediately following a theorem-like result, including evidence tied to the immediately preceding result. It reports explicit references, simple inferred local references, and unresolved mentions separately. It does not infer unstated mathematical prerequisites.
Kind metadata is intentionally explicit:
raw_kind: what the source extractor found.display_kind: the user-facing type label.normalized_kind: the stable grouping key used by tools.
The repository includes a small fixture under tests/fixtures/workspace_tex_project/. A typical local demo imports paper_a, paper_b, and paper_c, then searches for fixed point:
workspace_search_theorems("fixed point")returnslocal:paper-a::thm:main,local:paper-b::thm:main, andlocal:paper-c::thm:main.workspace_get_citations("local:paper-a", direction="outgoing", include_unresolved=True)reports citation keysabsent,missing, andpaper-b.The
paper-bcitation has cited arXiv ID2401.12346, but it does not resolve tolocal:paper-b; the row keepstarget_paper_id: null.To create a resolved target, the cited arXiv ID is imported with
workspace_add_arxiv_paper. A local paper with a similar bibliography entry is not enough; citation resolution is based on explicit cited arXiv ID evidence.
PaperGraph only constructs remote downloads from arXiv's fixed e-print endpoint; arbitrary URLs are not accepted for downloads. Scholarly reference search queries public metadata services and records candidates before any resolution/import is applied. It limits compressed responses to 100 MiB, expanded content to 500 MiB, and archives to 10,000 members. Absolute paths, parent traversal, symbolic links, hard links, devices, FIFOs, and other special archive members are rejected.
Workspaces are ordinary local SQLite files. Local PDFs remain local. Extracted PDF text, source spans, and proof evidence are written only to the workspace you choose. Do not commit databases, private manuscripts, cache data, credentials, tokens, generated distributions, or raw local logs.
PDF extraction is best for born-digital PDFs; scanned PDFs or OCR-heavy files may produce sparse text and missing evidence. Complex projects may need an explicit main_file; the parser is not a full TeX engine.
v1.1.3 adds Scholarly Reference Resolver: metadata search across Crossref, OpenAlex, and arXiv for blocked references, deterministic candidate ranking, candidate apply through Reference Import Closure, and explicit boundaries for ambiguous, old, paywalled, or metadata-only literature.
v1.1.2 adds Reference Import Closure: user-confirmed arXiv and local PDF imports for blocked references, DOI, URL, and published metadata records for non-importable targets, and regenerated reading reports after successful imports.
v1.1.1 adds Evidence Triage for first-use Reading Reports and Starter artifacts, with candidate-start labels, sparse dependency status, external blocker next actions, and unchanged evidence boundaries.
v1.1.0 adds Workspace Starter planning and bootstrap commands for
START_HERE.md,papergraph-starter-manifest.json, Reading Reports, and Cross-Paper Reading Plans from explicit paper inputs.v1.0.0 released the stable PaperGraph core: Paper Map, Reading Report, Cross-Paper Reading Plan, first-workspace onboarding, and the v1 evidence contract.
v0.13.0 added the v1.0 readiness pass with stable core contract docs, first-workspace walkthroughs, and example Reading Report/Cross-Paper Reading Plan artifacts.
v0.12.0 added Cross-Paper Reading Plan, a deterministic Markdown artifact for explicit paper sets with recommended sequence, selected-paper citation evidence, external risks, and evidence boundaries.
v0.11.0 added Reading Report Export, a deterministic Markdown artifact with Paper Map context, main-result candidates, reading route, external risks, and evidence boundaries.
v0.10.0 added Paper Map, an evidence-first first-load overview with main-result candidates, structure, reading route, and external-risk evidence.
v0.9.3 added external import review summaries.
v0.9.2 improved cited-result mention extraction.
v0.9.1 added proof-adjacent dependency evidence.
v0.9.0 introduced reading queues, sessions, and external import planning.
v0.4.0 introduced cross-paper SQLite workspaces with
workspace_add_arxiv_paper,workspace_search_theorems, andworkspace_get_citations. Resolution remains explicit, not semantic.
uv sync
uv run pytest -q -p no:cacheproviderThe automated suite uses synthetic archives, projects, bibliography entries, and PDFs. It does not require the live arXiv service.
Contributing And License
Bug reports, research-reading workflows, reproducible fixtures, and PRs are welcome. Please read Contributing before submitting changes. PaperGraph is released under the MIT License.
Available Tools
63 toolsget_dependenciesC
Return theorem-like nodes referenced by the given theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It conveys a read-only 'return' operation but omits the meaning of the recursive parameter, the direction of traversal, and any side effects or error conditions. The phrase 'theorem-like nodes referenced by' does clarify the edge direction, but significant behavioral variability is left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundant content. It states the action and resource immediately and uses every word effectively, even though brevity sacrifices behavioral detail addressed in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and can explain return values, the description omits the meaning of the recursive flag and does not distinguish this tool from workspace_get_dependencies. These gaps are important for correct invocation, and the schema's zero parameter descriptions do not fill them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare parameter names. It only clarifies theorem_id as 'the given theorem'; the recursive parameter remains completely unexplained despite having a default value. This is insufficient compensation for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and clearly identifies the resource ('theorem-like nodes referenced by the given theorem'). It is unambiguous about the operation, though it does not explicitly differentiate itself from the sibling workspace_get_dependencies tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as workspace_get_dependencies or get_dependency_diagnostics. The presence of a workspace-scoped sibling implies a possible selection criterion, but that context is not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dependency_diagnosticsB
Explain how dependencies were extracted for one theorem-like node.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Explain' implies a read-only operation, but the description does not disclose any behavioral details such as whether recursive traversal is performed, whether results depend on workspace state, or what the output structure looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler or redundancy. It front-loads the core purpose and earns its place, even though other dimensions suffer from missing details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% schema description coverage, the description is too thin. It does not explain what the diagnostics output contains, how recursive affects behavior, or when to choose this tool over related siblings. An agent would likely need to inspect the tool implementation to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds minimal context by saying 'one theorem-like node,' which loosely maps to theorem_id, but it says nothing about the recursive parameter or its meaning. The schema's type and default values are all the agent has to work with.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Explain') and the resource ('how dependencies were extracted for one theorem-like node'). This distinguishes it from sibling tools like get_dependencies, which presumably returns the dependencies themselves rather than explaining their extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as get_dependencies, where_used, or workspace_get_dependency_diagnostics. The description implies a diagnostic context but does not state when it should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_environment_diagnosticsB
Return PaperGraph version and reproducible launch diagnostics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the high-level return content but says nothing about side effects, error behavior, what 'reproducible launch diagnostics' includes, or whether any checks are performed. The risk is low for a 0-parameter read-only tool, but the disclosure is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and resource with zero filler. Every word earns its place, and the length is appropriate for a parameterless tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and annotations, the description is the agent's only source of information. It covers the essentials for making the call, but 'reproducible launch diagnostics' is vague and the description does not distinguish this tool from get_dependency_diagnostics or explain what the returned diagnostics enable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema trivially covers 100%, so there is nothing for the description to explain. Per the baseline for 0-parameter tools, a 4 is appropriate since no semantic gap exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') with a clear resource ('PaperGraph version and reproducible launch diagnostics'), stating exactly what the tool produces. While it doesn't explicitly name a sibling alternative, no other sibling tool covers environment diagnostics, so it is implicitly distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to invoke this tool, no mention of alternatives such as the closely named get_dependency_diagnostics, and no exclusions. The agent must infer that this is for launch/environment troubleshooting entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_theoremB
Return the full text and metadata for one theorem-like node.
| Name | Required | Description | Default |
|---|---|---|---|
| theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. 'Return' clearly signals a read-only operation and the description discloses what is returned, but it does not address behavior for missing or invalid IDs, access restrictions, or output shape beyond 'full text and metadata'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler or redundancy; every word adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers the core return value and scope, but with no output schema and no annotations it leaves unresolved how the ID is supplied and what happens on edge cases. It is minimally viable but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions theorem_id or how it identifies the node. The parameter name is fairly self-explanatory, but the description does not compensate for the lack of schema-level documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('one theorem-like node'), and clarifies the payload ('full text and metadata'). It is distinct from listing or searching siblings, though it does not name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one theorem-like node' implies this is for fetching a single item by ID rather than listing or searching, but there is no explicit when-to-use guidance or mention of alternatives such as list_theorems or workspace_search_theorems.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_theoremsC
List theorem-like environments in the currently loaded paper.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations and the description gives no information about side effects, read-only behavior, permissions, or any impact on the workspace. The user is left to infer that listing is non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no redundancy or unnecessary detail. It is well-structured and easy to read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too brief to provide complete context. It does not clarify what 'theorem-like environments' includes (e.g., theorems, lemmas, corollaries) or how the 'kind' parameter influences output, leaving significant gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has a single optional parameter 'kind' with no description, and the tool description does not mention it at all. There is no explanation of what values it accepts or how it affects the results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('List') and a specific resource ('theorem-like environments in the currently loaded paper'). It distinguishes from sibling tools like get_theorem or workspace_search_theorems by focusing on the current paper, though 'theorem-like' is somewhat vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as workspace_search_theorems or get_theorem. The description does not mention any conditions or preferred scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_arxiv_paperC
Download an arXiv source project and build its theorem graph.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| arxiv_id | Yes | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does state that the tool downloads a source project and builds a theorem graph, but it omits important behavioral details like whether it modifies the workspace, how refresh affects execution, whether network access is required, or what gets persisted. This is a meaningful but incomplete disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, and the core action is front-loaded. It is concise and easy to parse, though its brevity contributes to other shortcomings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters, no annotations, and no output schema, this one-sentence description is far from complete. It does not describe return values, optional parameter semantics, preconditions, side effects, or how this tool fits into a larger workflow, making it inadequate for reliable agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no parameter-level meaning. arxiv_id is implied by the tool name, but refresh and main_file are completely unexplained, so the agent cannot reason about their purpose or valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Download an arXiv source project and build its theorem graph.' This makes the core operation identifiable. However, it does not explicitly distinguish itself from sibling tools like load_arxiv_request or workspace_add_arxiv_paper, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as load_arxiv_request, load_paper, or the workspace_add_* variants. No prerequisites, exclusions, or alternative conditions are mentioned, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_arxiv_requestB
Validate a raw arXiv request, then load it only if unambiguous.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | ||
| refresh | No | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It does disclose a key trait: the load happens only when the request is unambiguous, implying validation failure or ambiguity blocks loading. But it does not describe error behavior, side effects of loading, or what happens on invalid input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It efficiently conveys the validation-then-load sequence and the ambiguity condition, though 'it' is slightly ambiguous about whether the request or paper is loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, no output schema, and three parameters, this description is too sparse. It omits parameter semantics, failure behavior, and clear routing relative to the many sibling tools, leaving an agent with insufficient information to invoke it correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at the 'input' parameter via 'raw arXiv request'. The 'refresh' and 'main_file' parameters are completely unexplained in both the schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action sequence: validate a raw arXiv request, then load it only if unambiguous. This distinguishes it from validation-only siblings and from loading already-validated papers, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you have a raw arXiv request that needs validation before loading, and only when the request is unambiguous. However, it gives no explicit when-not-to-use guidance or mention of alternatives like validate_arxiv_request or load_arxiv_paper.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_paperA
Load a local LaTeX paper and build its theorem graph.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions loading and building a theorem graph, which covers the main behavior, but it does not disclose whether the tool has side effects (e.g., saving state) or what it returns. Without annotations, some behavioral aspects remain unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no redundant or extraneous information. It is well-structured and immediately conveys the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description is sufficient for a simple tool, it omits details about return value (e.g., the built theorem graph), error handling, and specific conditions for use. Given the presence of many related tools, a bit more context on what the output or effect is would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for the 'path' parameter, and the description only indirectly implies it is the file path to a local LaTeX paper. It adds some meaning but does not explicitly define path format, required permissions, or relation to the graph building.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (load) and the specific resource (local LaTeX paper) and adds the purpose of building a theorem graph. This makes the tool's primary function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for local LaTeX files but does not explicitly contrast with sibling tools like load_arxiv_paper. It lacks guidance on when to choose this tool over alternatives, though the name and context provide some implicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_workspaceC
Open or initialize a persistent multi-paper workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only mentions persistence and the open/initialize action, but does not explain side effects, whether it creates a new workspace, whether it is idempotent, what it returns, or any required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. 'Persistent multi-paper workspace' is a meaningful qualifier, and the sentence is appropriately compact, though 'open or initialize' is slightly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and many workspace-related siblings, the description omits essential context such as expected path format, whether the workspace must already exist, return behavior, and when to call this tool relative to other workspace operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required 'path' parameter has no schema description and the description never explains what path means—whether it is an existing directory, a workspace identifier, or a path to be created. The tool name makes the inference plausible, but the description adds no explicit semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action on a resource: 'Open or initialize a persistent multi-paper workspace.' This distinguishes it from sibling tools focused on adding papers or loading individual documents, though the dual phrasing 'open or initialize' leaves some ambiguity about the exact operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus any of its many siblings, such as load_paper or workspace_add_paper. An agent has no explicit criteria to determine that this is the required first step before interacting with a workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_arxiv_inputC
Normalize arXiv ID and URL inputs and return the safe next action.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| text_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions normalization and returning a 'safe next action' but does not explain what 'safe next action' means, whether the tool performs network access, or whether it has side effects. The behavior is only superficially disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, front-loading the verb 'Normalize' and the resource. Every word contributes to the stated purpose. It is an example of concise, well-structured writing, even though additional detail is needed elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is insufficient for an agent to understand the return value 'safe next action' or how to interpret it. It also lacks context for when to call this tool in a workflow, especially with validate_arxiv_request so close in name and purpose. The agent cannot confidently select and invoke the tool correctly based solely on this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does map 'arXiv ID' to text_id and 'URL' to url, which is helpful, but it does not explain whether the parameters are mutually exclusive, which takes precedence, or what formats are accepted. Both parameters are optional, and the description gives no guidance on how to choose between url and text_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Normalize' and names the resource 'arXiv ID and URL inputs', clearly stating the tool's purpose. It also mentions the outcome 'return the safe next action,' which adds specificity. However, it does not differentiate from the similarly named sibling validate_arxiv_request, so it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It says nothing about preconditions, what input state is expected, or when validate_arxiv_request or load_arxiv_paper would be more appropriate. The presence of validate_arxiv_request as a sibling makes this omission particularly problematic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_arxiv_requestB
Validate a raw user arXiv request and return the safe next action.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does state the core behavior—validating a raw request and returning a safe next action—but it does not explain what 'safe next action' means, what the possible actions are, or how invalid requests are handled. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the primary action ('Validate a raw user arXiv request') and then states the output ('return the safe next action'). Every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one input, no output schema, and no annotations, the description is minimally sufficient: it identifies the input and the general nature of the return value. But 'safe next action' remains vague, and the relationship to the similarly named validate_arxiv_input tool is unexplained, leaving the agent without enough information to confidently select or invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter, 'input', with no description and 0% coverage. The tool description adds that the input is a 'raw user arXiv request', which gives meaningful context beyond the bare parameter name. However, it does not specify the expected format, structure, or example values, so it only partially compensates for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Validate' and the resource 'raw user arXiv request', and it indicates the outcome ('return the safe next action'). However, it does not distinguish itself from the sibling tool validate_arxiv_input, which appears to have a very similar purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as validate_arxiv_input or load_arxiv_request. The phrase 'safe next action' implies a decision-support role, but no explicit conditions or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
where_usedB
Return theorem-like nodes that reference the given theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, and the description only says that the tool returns referencing nodes. It does not disclose traversal depth, directness of references, or any side effects, though it is implied to be a read-only query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the verb and object front-loaded. There is no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple tool, especially since an output schema exists. However, it lacks contextual details about reference scope (direct vs. transitive) and how this relates to dependency/citation tools, which would help an agent select it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter theorem_id is named clearly and referred to as 'the given theorem', but the description adds little beyond the schema. It does not clarify whether the ID is a database key, external identifier, or how it should be formatted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the operation ('return') and target ('theorem-like nodes that reference the given theorem'), which distinguishes it from forward-dependency tools. However, 'theorem-like nodes' is somewhat vague and could be more explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus sibling tools such as get_dependencies or workspace_get_citations. It does not mention alternatives, limitations, or whether references are direct or transitive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_arxiv_paperB
Add or replace an arXiv LaTeX project in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| arxiv_id | Yes | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Add or replace' signals a mutating operation and hints at overwrite behavior, but it does not explain side effects, whether the project is downloaded from arXiv, what happens to existing files, or whether an active workspace is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence with no filler or repetition. It is well-formed and front-loaded, though very brief given the tool's parameter and behavioral complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter description coverage, this one-line description is not sufficient for an agent to call the tool correctly. It should at least explain what refresh and main_file do and clarify the replacement semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description mentions none of the three parameters. The tool name implies arxiv_id, but refresh and main_file are completely unexplained, leaving the agent unable to determine their meaning or how they affect the operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add or replace') and a specific resource ('arXiv LaTeX project') scoped to the active workspace. This clearly distinguishes it from siblings like workspace_add_local_paper and workspace_add_pdf_paper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in the active workspace' implies the tool operates on the currently open workspace, but the description gives no explicit guidance about when to choose this tool over alternatives such as load_arxiv_paper, workspace_add_pdf_paper, or workspace_add_local_paper. Usage context is implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_local_paperB
Add or replace a local LaTeX project in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose the add-or-replace behavior, but it does not explain side effects, whether replacement is keyed by paper_id, or any permission or validation requirements. For a mutating tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler; every word contributes to stating the operation and scope. It is appropriately succinct and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, critical invocation details are missing. The description does not define path or paper_id precisely, nor does it mention prerequisites or return behavior, so an agent may not be confident about how to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only names and types with 0% description coverage, so the description needed to explain path and paper_id. It only adds the general 'local LaTeX project' context and does not clarify whether path points to a file or directory, or what paper_id means. This is insufficient to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Add or replace'), a concrete resource ('local LaTeX project'), and the scope ('active workspace'). It clearly distinguishes the tool from siblings like workspace_add_arxiv_paper and workspace_add_pdf_paper, so an agent can identify the correct operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the intended scenario: adding a local LaTeX project to the active workspace, and it hints that this is not for arXiv or PDF inputs. However, it does not explicitly name alternatives or state when not to use the tool, leaving the agent to infer routing from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_pdf_paperC
Add or replace a born-digital PDF paper in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. 'Add or replace' usefully reveals mutating overwrite semantics, but it does not state what happens to the replaced paper's associated data (notes, theorem links, reading sessions), what prerequisites exist, or how failures (bad path, invalid PDF) manifest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 12-word sentence with the action verb front-loaded and zero filler words. It is well structured and efficient, but it is arguably too short relative to the information load required for a mutation tool with 0% schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a replace-capable operation with two opaque required parameters, no annotations, and no output schema, the description is incomplete. An agent cannot confidently determine what path and paper_id mean, whether the workspace must already be active, or what consequences replacement has on existing associated data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the two required parameters, and it explains neither. 'path' is weakly inferable as a file location, but 'paper_id' is genuinely ambiguous — it is unclear whether it identifies the paper being added, the paper to be replaced, or both — and the relationship between the two parameters is never clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add or replace') and resource ('born-digital PDF paper') with a location modifier ('active workspace'). The 'born-digital' qualifier and 'PDF' resource meaningfully differentiate it from workspace_add_arxiv_paper, though the boundary with workspace_add_local_paper is not clearly drawn.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, exclusions, or alternative-tool routing is present. The phrase 'active workspace' implies a prerequisite (a workspace must be open) but that is never made explicit, and the description gives no hint of when to choose this tool over workspace_add_arxiv_paper or workspace_add_local_paper.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_reading_noteC
Add a note or question to a reading session.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| note_type | No | note | |
| target_id | No | ||
| session_id | Yes | ||
| target_kind | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral details, but it only says 'Add a note or question.' It does not mention whether the operation mutates state, whether the reading session must already exist, whether the note is immediately persisted, or what the return behavior is. It is not misleading, but it is very thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler and the action is front-loaded. It could have added useful parameter context without becoming bloated, so it earns a 4 rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no annotations, and no output schema, this description is under-specified. The agent cannot determine what target_id/target_kind mean or how note_type should be used, and the required session_id is only implicitly tied to the phrase 'reading session.' The tool needs more context for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain session_id, note_type, target_id, or target_kind. 'Note or question' loosely hints at the text content, but it does not clarify how the optional fields control the note's type or target, so it fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Add') and a specific resource ('a note or question to a reading session'), so an agent knows what the tool accomplishes. It does not explicitly contrast this with sibling tools such as workspace_record_reading_checkpoint, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives like workspace_record_reading_checkpoint, workspace_create_reading_session, or workspace_apply_reading_queue_to_session. There are no exclusions, preferences, or conditions provided, leaving the agent to infer usage solely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_advance_reference_expansionB
Execute the saved policy: search, download arXiv sources, import unique strong candidates and record resolutions without per-paper prompts, within approved budgets. Resume with the same run ID.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| max_steps | No | ||
| time_budget_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool operates without per-paper prompts and within approved budgets, which is useful behavioral context. However, it does not explain what 'budgets' mean, whether it is idempotent, or what happens on partial failures. It doesn't contradict any annotations (since there are none), but it leaves significant behavioral unknowns for a complex workflow tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core purpose ('Execute the saved policy') and includes key operational constraints (no per-paper prompts, within approved budgets, resume with run ID). It is concise and structured, though it could be slightly more organized with separate sentences for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the workflow (search, download, import, record) and the lack of output schema, the description provides a broad overview but misses critical details: What constitutes 'strong candidates'? How are budgets defined and enforced? What does 'record resolutions' mean? The tool has three parameters (one required) and no output schema, so the description should clarify the workflow's lifecycle, but it leaves many operational details ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three parameters (run_id, max_steps, time_budget_seconds). The description mentions 'run ID' and 'budgets' but does not explain the meaning or format of any parameter in detail. For example, 'run_id' is mentioned as a resume key, but max_steps and time_budget_seconds are not elaborated. The description fails to fully clarify parameter semantics beyond what the schema provides (which is just names and types).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (execute the saved policy) and the resource (reference expansion workflow), and lists the main steps (search, download arXiv sources, import candidates, record resolutions). It distinguishes from siblings like workspace_pause_reference_expansion or workspace_get_reference_expansion by focusing on the execution/resume action. However, it doesn't explicitly name an alternative tool for when to use a different tool, so it doesn't fully differentiate from all siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'Resume with the same run ID,' which implies a prior run, but it doesn't explicitly state when to use this tool versus alternatives (e.g., when to use workspace_decide_reference_expansion or workspace_plan_external_imports_for_result). It provides context (resume with run ID) but lacks explicit exclusions or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_apply_reading_queue_to_sessionC
Apply reading queue items as checkpoints in a reading session.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | queued | |
| queue_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only implies mutation through 'apply' but does not state side effects such as whether checkpoints are created, whether the queue is modified, or whether existing session checkpoints are affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It conveys the essential operation efficiently, though it could add a bit more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without annotations or an output schema, this minimal description is insufficient for a mutation-like tool. An agent is left to guess side effects, required preconditions, the role of the optional status parameter, and what the operation returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning to queue_id, session_id, or the status parameter with default 'queued'. The parameter names are somewhat self-explanatory, but 'status' remains unexplained, and the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Apply reading queue items') and target context ('as checkpoints in a reading session'), which clarifies the tool's core purpose. It is distinguishable from siblings like workspace_record_reading_checkpoint, though it does not explicitly call out that distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidance is provided. The description does not explain when to choose this tool over workspace_record_reading_checkpoint or workspace_create_reading_queue, nor does it mention prerequisites like an existing session or queue.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_bootstrap_reading_projectC
Create starter artifacts for a first reading project.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | ||
| papers | Yes | ||
| artifact_dir | Yes | ||
| create_queue | No | ||
| project_title | No | ||
| create_session | No | ||
| max_candidates | No | ||
| workspace_path | Yes | ||
| target_result_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations provided (no readOnlyHint, destructiveHint). The description 'Create starter artifacts' indicates it performs creation (write) and likely modifies the workspace, but it doesn't disclose side effects, such as whether it overwrites existing files, requires specific directory permissions, or what artifacts are generated. The description carries the full burden but is very terse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded, which is good, but it is overly vague. It lacks detail that could be added without much length. It doesn't provide any concrete information about the artifacts or inputs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters, 0% schema description coverage, and no output schema or annotations, the description is highly inadequate. Without elaboration on parameters, side effects, or prerequisites, an agent cannot reliably know how to invoke this tool correctly. The description leaves critical gaps such as what 'papers' should look like and how the bootstrap relates to other workspace tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 9 parameters, but it doesn't. The description doesn't explain any parameters. It doesn't clarify what 'workspace_path', 'artifact_dir', or 'papers' mean in context, nor the purpose of optional flags like 'create_queue' and 'create_session'. The required parameters are only implied by their names, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Create starter artifacts for a first reading project' clearly identifies a verb (create) and a resource (starter artifacts for a first reading project). It implies a bootstrap action that initializes a project. It is distinct enough from siblings like 'workspace_plan_starter_project' (which plans) and 'workspace_create_reading_session' (which creates a session), but it doesn't explicitly name those alternatives in the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for initial setup of a reading project. The mention of 'first reading project' suggests this is for initializing new projects, which distinguishes it from siblings that manage existing papers or sessions. However, it doesn't explicitly state when not to use it or name alternatives like 'workspace_plan_starter_project' for planning beforehand.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_cancel_reference_expansionC
Permanently stop this run, preserving imported papers and history.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It discloses that the stop is permanent and that imported papers and history are preserved, which is useful. However, it does not explain what happens to other data (e.g., intermediate results, settings) or any side effects, leaving gaps in understanding the full impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action ('Permanently stop') and then adds a clarifying benefit. It is concise and structured logically, with no redundant filler, though it may be too sparse for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, irreversible action with no annotations and no output schema, the description is inadequate. It lacks information about prerequisites, side effects, reversibility (beyond 'permanently'), and how to distinguish this from pause/decide operations. The description should provide more context to allow safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not elaborate on the run_id parameter beyond the schema's bare 'Run Id'. It does not explain what a run_id is, how to obtain it, or its format. Since the description fails to compensate for the lack of schema detail, the meaning is under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (permanently stop) and resource (this run), and adds the preservation detail which distinguishes it from a plain stop. However, 'this run' is vague without explicitly referencing the reference-expansion run from the tool name, so it's clear but not maximally specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like workspace_pause_reference_expansion or workspace_decide_reference_expansion. The description does not mention exclusions, conditions, or alternative tools, leaving the agent to infer the usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_create_reading_queueC
Create a persistent reading queue for one stored result.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | ||
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must bear the full behavioral transparency burden. It discloses that the queue is 'persistent', but does not explain side effects, whether creating multiple queues for the same result is allowed, whether the operation is idempotent, or what happens to existing queues.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the core purpose, though the brevity crosses into under-specification for the parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description is the only source of behavioral and parameter context. It is too sparse to fully support correct invocation, especially for label and recursive, and it does not explain what the tool returns or how it fits into the reading-queue workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only hints at the meaning of result_id by saying 'one stored result', but it provides no explanation of label or recursive, both of which have defaults and likely affect queue behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create'), the resource ('persistent reading queue'), and the scope ('for one stored result'). It is specific enough to distinguish from similar concepts like creating a reading session, but it does not explicitly name or contrast sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as workspace_create_reading_session, workspace_list_reading_queues, or workspace_apply_reading_queue_to_session. There is no mention of prerequisites, expected workflow, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_create_reading_sessionC
Create a persistent reading session in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | ||
| paper_id | Yes | ||
| target_result_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It adds one useful trait ('persistent') and notes the active-workspace scope, but it does not explain side effects, whether creating another session is idempotent, how the session relates to a paper, or what the response will contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the action, then names the object and scope in a single efficient sentence. No words are wasted, though the brevity comes at the cost of the missing semantic detail captured in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter creation tool with no output schema and no annotations, one sentence is far from complete. It fails to define what a reading session is, how the parameters fit, or what the agent should expect after calling it, making the definition inadequate for reliable selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to explain the roles of paper_id, label, and target_result_id. It does not mention any of these parameters or their relationships, offering no meaning beyond the bare property names in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('create'), the resource ('persistent reading session'), and the scope ('active workspace'). It does not explicitly contrast with sibling tools like workspace_create_reading_queue or workspace_record_reading_checkpoint, but 'reading session' is a distinct enough resource to avoid major ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as workspace_create_reading_queue, workspace_add_reading_note, or workspace_record_reading_checkpoint. No preconditions, prerequisites, or exclusions are mentioned, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_create_reference_expansionC
Save an approved finite expansion policy. Creation does not search or import.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | No | ||
| root_paper_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It adds one useful negative ('does not search or import') and implies an 'approved' prerequisite, but omits side effects, validation, idempotency, persistence, or response behavior. For a state-changing create tool, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the core action is front-loaded and the clarifying negative behavior follows immediately. The description is appropriately sized and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must provide enough context to invoke the tool correctly, but it does not explain inputs, lifecycle relationship, or expected results. The minimal text leaves significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description names no parameters. root_paper_ids and policy remain unexplained, including requiredness and policy shape. The description fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Save') and resource ('approved finite expansion policy'), and clarifies that creation does not search or import. This gives a reasonably specific purpose, though it does not explicitly distinguish it from the sibling update tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are named. 'Creation does not search or import' implies a boundary against search/import tools, but the agent must infer when this tool should be chosen over related expansion lifecycle tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_decide_reference_expansionC
Record exactly one candidate_id, target, existing_paper_id, skip:true or retry:true. Advance separately to execute the approved choice.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| edge_id | Yes | ||
| decision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that execution is separate, which is helpful, but it does not explain side effects, persistence behavior, validation rules, idempotency, or what happens if a decision already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, and the core instruction is front-loaded. It is concise, though the terseness contributes to ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a nested decision object, no output schema, no annotations, and many related siblings, the description is incomplete. It omits the meaning of run_id and edge_id, the allowed shape of decision, and what the tool returns, leaving significant gaps for an agent trying to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only partially compensates by mentioning candidate_id, target, existing_paper_id, skip:true, and retry:true. It does not explain run_id, edge_id, or how the decision object should be structured, so an agent still lacks enough meaning to build a valid request.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'Record' and lists decision-related fields, suggesting this tool records a decision about a reference expansion. However, it never explicitly names the resource ('reference expansion decision') and the field list is cryptic, so an agent cannot confidently distinguish it from create/update siblings without extra inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Advance separately to execute the approved choice' gives a useful hint that this tool only records and does not execute, implying a separate follow-up step. But it does not explicitly name the alternative sibling or state when not to use this tool, leaving usage conditions largely implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_cross_paper_reading_planC
Export a deterministic Markdown reading plan for selected papers.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | ||
| paper_ids | Yes | ||
| max_candidates_per_paper | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses only that output is Markdown and 'deterministic', but does not say whether content is returned inline or written to a file, what the plan contains, ordering guarantees, or whether any permissions are needed for a mutation/export operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with no filler, and the artifact type is front-loaded. It is efficient, though the brevity contributes to the missing guidance rather than being a virtue of restraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, no annotations, and three parameters at 0% documentation leave the agent without enough to invoke this correctly. For an export tool nested among numerous near-named siblings, substantially more explanation is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
'selected papers' hints at paper_ids, but schema coverage is 0% and the description never explains the 'focus' filter or 'max_candidates_per_paper' (default 3), which are the parameters an agent most needs to reason about. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export), output artifact (deterministic Markdown reading plan), and scope (selected papers). However, it offers no differentiation from the many sibling export tools such as workspace_export_paper_reading_report or workspace_export_reading_bundle, which an agent must choose between.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this export versus the sibling exports (reading_bundle, result_reading_context, reading_session_summary, paper_reading_report), nor any prerequisites or exclusions. The agent is left to infer the routing entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_paper_reading_reportC
Export a deterministic Markdown reading report for one stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes | ||
| max_candidates | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. 'Deterministic' is a useful trait (stable output), but it says nothing about whether this is read-only, whether it writes a file, permission requirements, or how the report is delivered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. Efficient, though almost terse to the point of under-specifying for a two-parameter export tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, 0% param coverage, and a non-obvious 'max_candidates' parameter left unexplained. The description does not supply enough context for an agent to invoke this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter is explained. 'max_candidates' (default 5) is particularly opaque for a report on 'one stored paper' and the description does nothing to clarify it or the paper_id semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Export') and resource ('deterministic Markdown reading report') scoped to 'one stored paper'. It is distinguishable from export siblings like workspace_export_reading_bundle, but it never explicitly contrasts itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one stored paper' weakly implies the paper must already be in the workspace, but there is no explicit when-to-use, when-not-to-use, or alternative routing among the several other export tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_reading_bundleC
Export a paper-level evidence bundle for paper-reading consumers.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It only says 'Export a bundle,' without stating whether the tool is read-only, what format the bundle takes, whether it modifies workspace state, or any output/limits. The phrase is not contradictory but is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; the core action is front-loaded. It is concise, though some terms like 'evidence bundle' are undefined, which the conciseness itself does not resolve.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must explain what the returned bundle is and how it behaves; it does neither. For an export-style tool among many similar reading/export siblings, the current text is too incomplete for reliable selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions paper_id. The term 'paper-level' gives weak contextual confirmation that the ID refers to a paper, but no format, source, or usage example is provided, leaving the agent to infer from the parameter name alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb, 'Export', a resource ('paper-level evidence bundle'), and the intended audience ('paper-reading consumers'). It is more informative than a tautology, though it does not explicitly differentiate itself from sibling export tools like workspace_export_result_reading_context beyond the 'paper-level' qualifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over the many sibling export/reading tools. The description implies it operates at the paper level, but it never states when this bundle is preferable to workspace_export_result_reading_context or workspace_export_reading_session_summary, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_reading_session_summaryC
Export a deterministic recovery summary for a reading session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden for behavioral disclosure. It mentions the output is 'deterministic', which is a useful trait, but does not disclose side effects, format, export destination, or any operational behavior. This is insufficient for a tool with no structured annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly-worded sentence that front-loads the core verb and object. Every word adds meaning, and there is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and only a one-sentence description, the tool is under-specified. The exact content of a 'deterministic recovery summary', the output format, whether it creates files or returns text, and error behavior are all unknown, leaving an agent with insufficient context to invoke it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention session_id at all. While the parameter name is reasonably self-explanatory, the description provides no elaboration on what constitutes a valid session identifier or how it affects the summary, so it fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Export') and resource ('a reading session summary'), and adds the meaningful qualifier 'deterministic recovery summary'. It is not tautological and gives a reasonable sense of what the tool produces, though it does not explicitly distinguish itself from sibling export tools like workspace_export_reading_bundle or workspace_export_result_reading_context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus its siblings, no alternative names, and no exclusions. The phrase 'recovery summary' hints at a use case, but an agent has no explicit criteria for choosing this over the other export-related workspace tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_reference_expansionC
Return saved JSON or Markdown without writing client files or using network.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | json | |
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses two key behavioral traits: no client file writes and no network usage, which are useful. However, it doesn't explain other behaviors like the effect of the format parameter, error handling, or server-side side effects. The disclosure is minimal but not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence, which is concise and front-loads the core action. It is efficient in word count, but it omits essential details, making it borderline under-specified. The structure is clean, but the trade-off between brevity and completeness is not optimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 2 parameters (one required), no annotations, and no output schema, the description is insufficient. It doesn't clarify what run_id refers to, what the default format is, or what the return value looks like. An agent would have to rely heavily on inference or trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It hints at format ('JSON or Markdown') but does not explain the required run_id parameter at all. The agent cannot determine what run_id refers to or how to supply it correctly from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb (Return) and object (saved JSON or Markdown), but 'saved' is ambiguous and it doesn't explicitly tie to 'reference expansion' beyond the tool name. It doesn't differentiate from siblings like workspace_get_reference_expansion, which could also return a reference expansion. The action is clear but the scope is vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. There is no mention of context, exclusions, or how it differs from similar tools like workspace_get_reference_expansion. The description is purely functional with no usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_result_reading_contextB
Export focused evidence context for reading one result's proof.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, whether it has side effects, what the exported context contains, or what format is returned. 'Export' alone is ambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word contributes to identifying the operation and its target, making it highly concise while remaining readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description should explain what is returned or exported and when to use this instead of similar export/get tools. It leaves those gaps open, so an agent does not have enough context to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description only indirectly defines result_id as identifying 'one result's proof.' It does not explain what form the ID takes, how to obtain it, or how it relates to result contexts in sibling tools, so the schema gap is only minimally compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Export') and a clear target: 'focused evidence context for reading one result's proof.' This distinguishes it from session-level exports like workspace_export_reading_bundle or workspace_export_reading_session_summary, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for reading one result's proof' implies the usage context, but the description gives no explicit when-to-use guidance or exclusions relative to the many sibling tools (e.g., workspace_get_evidence, workspace_export_reading_bundle). An agent must infer which tool to choose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_citationsC
Return incoming or outgoing citation evidence for a stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes | ||
| direction | No | outgoing | |
| include_unresolved | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden of behavioral disclosure. It communicates that the operation returns data, but it does not explain the meaning of unresolved citations, whether incoming citations require a different lookup path, what happens for missing papers, or how resolved versus unresolved evidence is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence that front-loads the core behavior. There is no filler, repetition of the tool name, or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description is too thin for a tool with three parameters and zero schema coverage. Key behavioral semantics such as what 'citation evidence' includes, what 'unresolved' means, and how direction affects results are missing, leaving the agent to guess at important invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only indirectly clarifies the direction parameter via 'incoming or outgoing'. It provides no additional meaning for paper_id or include_unresolved, and it does not explain allowed direction values or the effect of include_unresolved on the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: it returns citation evidence for a stored paper, and it identifies the key incoming/outgoing axis. It does not explicitly distinguish itself from similar tools like where_used or workspace_get_evidence, but the basic purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus siblings such as where_used, get_dependencies, or workspace_get_external_result_mentions. There is no statement of prerequisites, exclusions, or alternative selection criteria, so the agent must infer usage solely from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_dependenciesC
Return dependencies of a globally identified stored theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| global_theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It states that dependencies are returned, implying a read operation, but it does not explain the effect of the 'recursive' flag, whether dependencies are direct or transitive by default, or any restrictions on which theorems qualify. This is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is appropriately sized for a simple tool. It is front-loaded with the action and resource, and contains no filler or redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although the tool has only two parameters and an output schema, the description omits important context needed for reliable use: what 'recursive' means, whether the tool is scoped to the current workspace, and how its results differ from related dependency/citation tools. The description alone would not let an agent confidently choose this tool among the many siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only hints that 'global_theorem_id' is a globally identifying string via 'globally identified,' and it says nothing about the meaning of 'recursive' or the default behavior. Most parameter semantics are left to the agent to infer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return dependencies') and identifies the target resource as a 'globally identified stored theorem.' This is specific enough to convey the core purpose, but it does not explicitly contrast with siblings like get_dependencies or workspace_get_proof_dependencies, leaving some ambiguity about the exact scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. The description does not mention workspace context, when recursive behavior might be needed, or how this differs from similar sibling tools such as workspace_get_proof_dependencies or get_dependencies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_dependency_diagnosticsC
Explain how workspace dependencies were extracted for one theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| global_theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. 'Explain how...were extracted' strongly implies a read-only diagnostic operation and clarifies the single-theorem scope. However, it does not disclose whether the call has side effects, how the recursive parameter affects behavior, or what form the explanation takes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. It is concise and readable, though its brevity leaves important parameter and usage details unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a two-parameter tool with no annotations and no output schema, so the description needs to provide substantial context. It omits the meaning/effect of recursive, the expected diagnostic output, and how it relates to the sibling get_dependency_diagnostics, leaving gaps that an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It only alludes to the required parameter via 'for one theorem' and says nothing about the optional recursive parameter, leaving half of the parameter surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Explain') and identifies the resource ('how workspace dependencies were extracted for one theorem'), making it clear this is a diagnostic tool rather than a dependency-computation tool. It does not explicitly distinguish itself from the similarly named sibling get_dependency_diagnostics, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool instead of alternatives such as get_dependency_diagnostics or workspace_get_dependencies. 'For one theorem' is a scoping hint, but no when-to-use or when-not-to-use context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_evidenceC
Return metadata and source spans for one evidence node or edge.
| Name | Required | Description | Default |
|---|---|---|---|
| node_or_edge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It only says 'Return', which implies a non-mutating read, but it does not disclose error behavior, whether node and edge IDs share a namespace, or any prerequisites. For a tool with zero annotation coverage, this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; the verb and object come first. It is appropriately short for the simple interface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description leaves the agent without enough context about valid IDs, return shape, and failure modes. The one-sentence description is too sparse to confidently invoke this tool among more than 40 siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the sole parameter node_or_edge_id is undocumented. The description mentions 'one evidence node or edge' but gives no ID format, examples, or guidance on how to distinguish node IDs from edge IDs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Return'), a specific target ('one evidence node or edge'), and the content type ('metadata and source spans'). It is clearly distinct from sibling getters like workspace_get_paper and workspace_get_result, though 'evidence node' is somewhat domain-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to choose this tool over siblings such as workspace_get_source_slice, workspace_get_citations, or workspace_get_dependencies. The phrasing only implies use when evidence metadata/source spans are needed, with no explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_external_result_mentionsB
Return external result mentions from a result's proof evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only says 'Return...'. It does not explicitly state that the operation is read-only, what happens if result_id does not exist, or whether results are paginated or ordered. The 'get' naming implies safety, but the description itself does not carry the behavioral transparency burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to identifying the operation and its source, making it appropriately concise for such a simple one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has low parameter complexity and an output schema, so the brief description is not catastrophic. However, it omits useful context about what 'external result mentions' are and how this tool relates to external-import planning or other result getters, leaving an agent to infer that from the sibling list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only labels result_id as 'Result Id' with no description. The description adds meaning by indicating that the ID belongs to a result whose proof evidence is searched. It does not specify ID format, how to obtain it, or what counts as an 'external result mention,' so it only partially compensates for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('external result mentions') and identifies the source ('from a result's proof evidence'), so an agent can tell what the tool returns. It does not explicitly contrast with sibling tools such as workspace_get_citations or workspace_get_evidence, so it stops just short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this tool over closely related siblings like workspace_get_evidence, workspace_get_citations, or workspace_plan_external_imports_for_result. The phrase 'from a result's proof evidence' provides a data source, but not selection context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_paperA
Return metadata and counts for one stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It states a read-only return behavior, which is useful, but it does not disclose error/not-found behavior, prerequisites such as an open workspace, or what 'counts' specifically refers to.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. Every word contributes to describing the tool's core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter getter, the description covers the basics, but with no annotations and no output schema, the vague phrase 'metadata and counts' leaves return structure and the meaning of 'counts' to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter, paper_id, with 0% description coverage. The description does not explain the format or source of paper_id, but the name plus 'one stored paper' makes the parameter's referent reasonably clear, adding limited semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and a specific resource ('one stored paper'), and indicates the result content ('metadata and counts'). It clearly differentiates this from list-type siblings, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one stored paper' implies this tool is for retrieving a single paper's metadata and counts, but the description gives no explicit when-to-use vs. when-not-to-use guidance and does not mention alternatives like workspace_list_papers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_paper_mapC
Return an evidence-first first-load map for one stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes | ||
| max_candidates | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must fully disclose behavior. It indicates a read-like operation ('Return') but does not state whether it is read-only, what happens on missing paper_id, whether max_candidates affects results or errors, or any side effects. The 'evidence-first first-load map' is left undefined, so the agent cannot anticipate the output shape or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence with no filler or repetition. It front-loads the core purpose and avoids redundancy. However, its brevity borders on under-specification, which is why it earns a 4 rather than 5 – it is concise but sacrifices essential detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has two parameters, no output schema, and no annotations, so the description is the sole source of context. It fails to explain what a 'map' is, what 'evidence-first' means, how max_candidates behaves, or what the return value looks like. For a tool with this complexity, the description is clearly insufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain both parameters. It mentions neither paper_id nor max_candidates. paper_id is self-explanatory, but max_candidates is cryptic; the agent has no idea what it limits (candidates for what?) or how it influences the result. The description adds zero value over the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns a specific artifact ('evidence-first first-load map') for a single stored paper, using a distinct verb ('Return'). It distinguishes itself from sibling tools like workspace_get_paper (which likely returns the paper object) and workspace_get_evidence (which returns evidence). However, the term 'evidence-first first-load map' is jargon and not explained, which slightly reduces clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of when it is appropriate, what scenarios it fits, or when to choose a sibling tool instead. The description is purely declarative with no usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_proof_dependenciesC
Return proof dependency evidence for one stored evidence result.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It signals a read operation via 'Return' but does not disclose behavior such as whether dependencies are direct or recursive, performance implications, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler, and the core action is front-loaded. It sacrifices helpful detail but is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and a very close sibling (workspace_get_dependencies), the description is too thin: it omits recursive behavior, result_id provenance, output shape, and when to use alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only implies that result_id identifies a stored evidence result. It does not explain the meaning or effect of the recursive parameter (default false), nor what an evidence result ID looks like.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Return') and resource ('proof dependency evidence for one stored evidence result'), which distinguishes it from generic dependency tools in the sibling list. However, it does not explain what 'proof dependency evidence' means or how this differs from workspace_get_dependencies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this tool over siblings such as workspace_get_dependencies or get_dependencies. The description only states what it returns, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_reading_queueB
Return one persistent reading queue with ordered items.
| Name | Required | Description | Default |
|---|---|---|---|
| queue_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does add useful traits: the queue is persistent and the items are ordered. However, it does not disclose behavior for invalid or missing queue_id, whether the operation is read-only, or what the returned object structure looks like beyond 'ordered items'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. 'Persistent' and 'ordered items' are meaningful qualifiers that add value without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description conveys the core returned item concept: one persistent reading queue with ordered items. However, given no output schema and no annotations, it leaves gaps around error behavior, queue existence, and how this getter fits with related queue operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain queue_id at all, and schema description coverage is 0%. The single parameter is inferable from its name and the tool's 'get' semantics, so this is not severely misleading, but the description fails to compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names the exact resource ('one persistent reading queue') while adding that it contains ordered items. This clearly distinguishes it from sibling tools like workspace_list_reading_queues, which lists many queues, and workspace_create_reading_queue, which creates a queue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as workspace_list_reading_queues or workspace_apply_reading_queue_to_session. There are no explicit conditions, prerequisites, or exclusions; the usage context is only implicitly suggested by the tool name and queue_id parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_reading_sessionB
Return one persistent reading session with checkpoints and notes.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral weight. 'Return' and 'persistent' suggest a read-only retrieval of stored data, but the description does not disclose error behavior, permissions, or what happens when the session_id does not exist. For a simple getter this is acceptable but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. Every word adds meaning: 'persistent' signals state, 'checkpoints and notes' signals returned content, and 'one' signals a singular lookup.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple single-parameter getter, and the description names the resource and returned content. However, with no output schema and no parameter semantics, an agent still lacks guidance on session_id sourcing and not-found behavior, leaving the description minimally viable rather than fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the only parameter, session_id, is documented only by its title. The description does not explain what session_id values look like, where to obtain them, or how they relate to the reading session being returned, so it fails to compensate for the sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('persistent reading session') and notes that it includes checkpoints and notes. This clearly distinguishes it from create/list/queue tools, though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a caller needs a single existing reading session with its checkpoints and notes, but it does not explicitly state when not to use it or name alternatives such as workspace_list_reading_sessions or workspace_get_reading_queue. The usage context is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_reference_expansionB
Read saved graph, counts, decisions and continuation actions without network.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose that the operation is a read and that it works without network access, which is meaningful. However, it does not describe side effects, error behavior, or what happens when the run_id is invalid or missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence starting with the verb 'Read'. It contains no filler and each phrase ('saved graph', 'counts', 'decisions', 'continuation actions', 'without network') adds distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool, this is reasonably complete: it names the output categories and the offline behavior. However, without an output schema or annotations, and with many related sibling tools, the agent would benefit from more detail on how 'saved graph, counts, decisions and continuation actions' are structured and how this tool relates to other reference-expansion tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention run_id at all. The parameter is simple and the schema title 'Run Id' provides some meaning, but the description adds no semantics about how run_id relates to the saved graph or how it identifies the expansion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('Read') and names a specific resource: 'saved graph, counts, decisions and continuation actions'. It distinguishes itself from the write/action-oriented siblings like workspace_create_reference_expansion and workspace_decide_reference_expansion by being read-only and offline. It does not explicitly mention 'reference expansion' in the description, but the tool name supplies that context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives such as workspace_list_reference_expansions, workspace_get_result, or workspace_resolve_external_reference. The phrase 'without network' hints at an offline use case, but no when-to-use or when-not-to-use conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_resultB
Return one stored evidence result with source spans.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'Return ... with source spans' conveys an idempotent read and the included content, which is adequate for a simple getter, but it does not disclose not-found behavior, error cases, or whether source spans are always present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, front-loading the action and the key output detail. Every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter getter, this is close to sufficient, but without an output schema or usage guidance the agent is left without a clear return shape beyond 'source spans' or context about when this tool is preferable to its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only loosely implies that result_id selects the stored evidence result. It does not explain the ID format, how to obtain it, or what 'source spans' contains, so the agent must infer parameter semantics from the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('return') and resource ('one stored evidence result') and adds the output characteristic 'with source spans'. This distinguishes it from list_results and result_proof tools, though it does not explicitly name a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus workspace_get_evidence, workspace_get_result_proof, or workspace_list_results. The use case is implied by the name but never stated, and no exclusions or alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_result_proofC
Return proof evidence for one stored evidence result.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. 'Return' indicates a retrieval operation with no mutation, and 'one stored evidence result' signals a single-record lookup, but the description does not disclose edge cases (e.g., missing result_id, not-found behavior) or the nature of the returned evidence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with no filler; the key action and target are front-loaded. It is concise without being a tautology, though the brevity comes at the cost of parameter and context detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and zero parameter coverage, the description leaves ambiguity about what 'proof evidence' consists of, what result_id references, and how this differs from sibling evidence/result getters. It is functional but under-specified for an agent that must select among many workspace_get_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explicitly define result_id. It only implies that result_id identifies the stored evidence result, leaving the agent to infer the type, format, and source of the ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('proof evidence') and scopes it to 'one stored evidence result,' which conveys the core function. It does not explicitly contrast with sibling tools like workspace_get_evidence or workspace_get_result, but the resource phrase is specific enough to avoid tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over workspace_get_evidence, workspace_get_proof_dependencies, or workspace_get_result. The single sentence implies use when proof evidence for a result is needed, but it provides no conditions, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_result_reading_pathB
Return deterministic local reading paths for one result.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It adds a genuinely useful behavioral trait by promising deterministic, local, read-oriented paths, suggesting a safe idempotent operation. However, it remains silent on path existence, error behavior for invalid result IDs, and the impact of the recursive flag, so transparency is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition, making it easy to scan. It is appropriately terse, though it could have used one more clause to convey the recursive parameter without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, and the description does not explain the return format, what a 'reading path' consists of, or how recursive changes the result. Given a required result_id and an optional recursive flag, an agent lacks enough context to know the exact effect of recursion or what to do with the returned paths. The tool is too opaque relative to its sibling-rich environment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only maps 'one result' to result_id. The recursive parameter—a boolean with default true—receives no explanation, leaving a core behavioral switch undocumented. This is partial compensation at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') with a concrete resource ('deterministic local reading paths') and scopes it to 'one result,' which differentiates it from sibling getters like workspace_get_result and workspace_get_reading_session. The modifiers 'deterministic' and 'local' add precision about what kind of path is returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states only what the tool returns; it gives no guidance about when to choose this tool over sibling tools such as workspace_get_result, workspace_get_result_proof, or workspace_export_reading_bundle. There are no prerequisites, exclusions, or alternative-selection conditions, so an agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_source_sliceC
Return bounded source text around one span, result, or proof.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| span_id | No | ||
| proof_id | No | ||
| result_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It indicates a read-only operation returning source text, which is helpful, but it does not explain what happens when multiple IDs are supplied, how 'context' affects the result, or whether errors occur for invalid IDs. Adequate for a simple read tool but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with the action and target front-loaded. There is no filler or redundancy, though the brevity sacrifices useful operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and four parameters lacking schema descriptions, the single sentence is not enough to invoke the tool confidently. The behavior of context, ID selection semantics, and expected result format are all missing, so an agent would likely need to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by tying span_id, result_id, and proof_id to their conceptual resources, but it fails to explain the 'context' parameter, the relationship among the three ID parameters, or the expectation that exactly one should be selected. A significant parameter is left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('bounded source text around one span, result, or proof'), making the core purpose clear. It is slightly ambiguous what 'bounded' means and it does not differentiate from sibling getter tools, but the intent is understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool versus alternatives like workspace_get_result, workspace_get_proof_dependencies, or workspace_get_evidence. No exclusions, prerequisites, or alternative routing are mentioned, leaving usage entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_external_reference_resolutionsC
List recorded external reference resolutions and selection provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. 'List recorded...' conveys that this is a read-style operation returning stored data, and 'selection provenance' hints at the output contents. However, it does not disclose behavior such as whether the result is scoped by paper_id, how results are ordered, or whether pagination applies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler words. Every phrase ('recorded', 'external reference resolutions', 'selection provenance') adds meaning, making it an appropriately concise definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema or annotations, the description is too minimal. It omits parameter semantics, any filtering behavior, and the shape/scope of the returned provenance. An agent would struggle to know whether and how to pass paper_id to get the intended result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions the single parameter, paper_id. An agent cannot tell whether paper_id filters the listed resolutions, scopes the provenance, or is simply ignored. The description does nothing to compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a specific resource ('recorded external reference resolutions') and adds 'selection provenance', which makes the tool's scope clearer than just 'list resolutions.' It distinguishes reasonably from related siblings like workspace_list_external_reference_searches, but it does not explicitly call out how it differs from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as workspace_resolve_external_reference, workspace_list_external_reference_searches, or workspace_resolve_external_reference_candidate. There is no when-to-use context, no exclusion criteria, and no mention of the optional paper_id relationship.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_external_reference_searchesC
List scholarly reference search runs and candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | No | ||
| blocked_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that the tool lists search runs and candidates, but it does not disclose whether this is a read-only operation, whether it returns both runs and candidates together, how results are ordered, or any side effects. The lack of annotation coverage makes this a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words. It is front-loaded with the verb and resource. However, it is under-specified for the tool's complexity, so the brevity is not fully a strength.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and two undocumented parameters, the description is incomplete. An agent cannot determine what the returned list contains, how the parameters affect the result, or how this tool relates to the many sibling list/search tools. More context is needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the two parameters (paper_id and blocked_id). It does not explain what these parameters mean, how they filter the list, or whether they are mutually exclusive. The description adds no parameter-level meaning beyond the schema's bare names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'List scholarly reference search runs and candidates' uses a clear verb ('List') and identifies the resource ('scholarly reference search runs and candidates'). However, it does not distinguish this from sibling tools like workspace_list_external_reference_resolutions or workspace_search_external_reference, so an agent may struggle to know which list tool to choose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as workspace_list_external_reference_resolutions or workspace_search_external_reference. The description implies a listing operation but does not state the context, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_papersB
List all papers stored in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only implies a read-only action via the verb 'list'. It does not disclose any potential side effects, default behaviors, or limitations beyond the basic listing operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It is concise and well-structured, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal but adequate for a simple list operation. However, it does not specify what information is returned (e.g., paper IDs, titles, metadata) or any inherent ordering or filtering, which could be useful for an agent to predict the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline score is 4. The description does not need to add parameter-specific meaning, and it correctly reflects that no arguments are required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (list) and the resource (papers) within the context of the active workspace. It is specific enough to distinguish from sibling tools that list other resources like theorems or results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternative list tools among the siblings. There is no mention of filters, sorting, or scenarios where other tools would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_reading_queuesC
List persistent reading queues in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It communicates that queues are persistent and scoped to the active workspace, and 'List' implies a read-only operation. It does not disclose filtering behavior, ordering, pagination, or side effects, though for a simple listing tool the lack of explicit side-effect disclosure is less critical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the action, resource, and scope; no filler. It is efficient and easy to parse, though it omits some useful supporting detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter list tool with no annotations, the description is underspecified: it ignores status and paper_id and gives no hint about how to filter or whether results are mutable. The output schema supplies return shape, which helps, but does not compensate for the missing parameter and behavioral context. Overall it is a viable one-line definition for the simplest 'list everything' call, but not complete for the full parameter set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has two optional parameters, status and paper_id, with zero description coverage, and the tool description does not mention either. The agent has no guidance on what values are valid or how filtering works. Thus the description adds no meaning beyond the parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation ('List'), a specific resource ('persistent reading queues'), and a scope ('active workspace'). It clearly reads as a list action and, by naming 'persistent', hints at the distinction from reading sessions. It does not explicitly contrast with workspace_create_reading_queue or workspace_get_reading_queue, so it stops short of fully differentiating among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in the active workspace' gives useful context for when the tool applies, and the verb 'List' implies it is the collection-view counterpart to workspace_get_reading_queue. However, it provides no explicit when-to-use/when-not-to-use guidance and does not mention alternatives or exclusions. Usage must be inferred from the name and the sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_reading_sessionsB
List persistent reading sessions in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. 'List' clearly signals a read-only operation, and 'persistent' plus 'active workspace' adds scoping context. There are no misleading or hidden behavioral claims.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence and contains no redundant or filler content. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so return values are likely documented elsewhere. However, the optional parameters are completely unexplained in both schema and description, leaving a meaningful gap for agents that need filtered results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the optional filters 'status' or 'paper_id' at all. An agent cannot determine valid values or filtering semantics from either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation ('List'), the resource ('persistent reading sessions'), and the scope ('in the active workspace'). This distinguishes it from singular getter tools like workspace_get_reading_session and from queue-listing siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as workspace_get_reading_session or workspace_list_reading_queues. The description only states what the tool does, not the conditions that should trigger its use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_reference_expansionsC
List saved expansion summaries, optionally filtered by state.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read-only listing operation but does not mention pagination, ordering, or what 'summaries' includes. There is no statement about side effects or return format, leaving significant ambiguity for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the verb and resource immediately. Every word earns its place, and there is no redundant elaboration. It is appropriately concise for a simple list operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and no output schema, the description is minimally adequate. It tells the agent the action and the optional filter, but omits details like allowed state values, pagination behavior, or what constitutes a 'summary'. Given the simplicity, a score of 3 reflects that it is acceptable but leaves room for essential information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the 'state' parameter, and the description only says 'optionally filtered by state' without explaining valid values or how the filter behaves. Since there are no enums and the type is anyOf string/null, the agent cannot infer what states are acceptable. The description adds minimal semantic value beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'saved expansion summaries', with an optional filter by state. It is distinct from singular get_reference_expansion, but does not explicitly name alternatives or differentiate from other list tools like workspace_list_external_reference_resolutions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It does not mention scenarios where one would prefer the singular get_reference_expansion, nor does it clarify the distinction between listing summaries and retrieving full details. No when-to-use or when-not-to-use context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_resultsC
List stored evidence results across the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It conveys that this is a read-only listing scoped to the active workspace, but it does not mention whether an active workspace must already be open, how results are ordered or paginated, or any side effects. This is minimal disclosure for a list operation but lacks meaningful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler; it front-loads the verb and resource and adds the workspace scope. It is appropriately tight, although that tightness comes at the cost of missing contextual detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three optional but effectively undefined parameters and no annotations, this description is too thin. It does not explain any of the filters, the relationship between 'results' and 'evidence', or when an agent should prefer this over related sibling tools. The output schema mitigates return-value questions but not invocation decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention kind, limit, or paper_id. An agent cannot infer what values kind accepts, what paper_id filters by, or how limit behaves from the description alone. The schema provides only names, types, and defaults, so the description adds no parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), a distinctive resource ('stored evidence results'), and a scope ('across the active workspace'). This distinguishes it from sibling list tools like workspace_list_papers and from single-result accessors like workspace_get_result, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over workspace_get_result, workspace_get_evidence, or the other list-related tools. No mention is made of intended workflows, prerequisites, or exclusions. The only implied usage is 'list results,' which is not enough for an agent to select it confidently among many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_pause_reference_expansionC
Pause scheduling after the active step. Advance explicitly to resume.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that it pauses scheduling and that resuming requires an advance, but it doesn't mention side effects, idempotency, what happens to the current step, or any permissions. For a state-changing tool, this is insufficient disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core instruction is front-loaded. It is efficient and easy to read, though it sacrifices information density for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and a single required parameter, the description leaves gaps: what is 'the active step', what does 'scheduling' refer to, and what is the tool's overall effect in the reference expansion workflow? An agent would need additional context to know if this is the correct action and what the consequences are.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter, run_id, with 0% coverage in the description. The description does not explain what run_id refers to (e.g., the run being paused), leaving the agent to infer its meaning from the name alone. The description adds no semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Pause') and resource ('scheduling after the active step'), making the core action clear. It doesn't explicitly mention 'reference expansion' in the text, but the tool name and sibling context imply it. It doesn't distinguish from siblings directly, but the purpose is discernible.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that resumption requires an explicit advance ('Advance explicitly to resume'), which hints at the alternative sibling workspace_advance_reference_expansion, but it doesn't name it or provide conditions for when to use this tool over other pause-like tools. The guidance is minimal and implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_plan_external_imports_for_paperC
Plan external arXiv imports visible in one stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of explaining behavior. It does not disclose whether this tool mutates the workspace, creates a queue, performs network fetches, or simply returns a read-only plan. The word 'Plan' hints at a non-executing action, but this is not explicit enough for an agent to safely predict side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler words. It front-loads the main action and object. The main issue is not length but ambiguity in the verb 'Plan'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and related sibling tools, the description is under-specified. An agent cannot determine the return value, whether executing the plan changes state, or how this paper-specific tool relates to the result- and queue-specific variants. More operational context is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter, paper_id, with 0% description coverage. The phrase 'one stored paper' at least implies that paper_id must refer to an existing paper in the workspace, which adds some meaning beyond the schema. However, it does not explain where to find valid paper IDs, what format they take, or how this paper_id relates to the imported arXiv papers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb and a resource: 'Plan external arXiv imports visible in one stored paper.' It conveys some scope, and the paper focus helps distinguish it from the result/queue sibling tools. However, 'plan' is ambiguous: it does not explain whether this returns a plan, schedules imports, or lists importable papers, so the purpose remains vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus the closely related workspace_plan_external_imports_for_result and workspace_plan_external_imports_for_queue. The only signal is the word 'paper' in the name and 'one stored paper' in the description. No exclusions, prerequisites, or alternative-selection criteria are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_plan_external_imports_for_queueC
Plan external arXiv imports referenced by one reading queue.
| Name | Required | Description | Default |
|---|---|---|---|
| queue_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral disclosure burden. 'Plan' is ambiguous: it could mean return a computed plan, persist something, or fetch external metadata. The description does not mention side effects, return values, or whether the operation is read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler or repetition. It is front-loaded with the core action and resource, though it is so brief that it misses important behavioral context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description is not complete enough for an agent to invoke the tool confidently. It does not explain what a 'plan' consists of, whether imports are modified, or what the agent should expect in response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter, queue_id, with 0% description coverage. The description adds only that the tool concerns 'one reading queue,' which mostly restates the parameter name. It does not explain queue_id format, where to find it, or any constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a specific verb ('Plan') and resource ('external arXiv imports referenced by one reading queue'), so an agent can tell this is about planning imports for a queue rather than executing them. It does not explicitly differentiate itself from the sibling tools for_result and for_paper, but the queue scope is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus sibling tools such as workspace_plan_external_imports_for_result or workspace_plan_external_imports_for_paper. It also does not state prerequisites like whether the reading queue must already exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_plan_external_imports_for_resultC
Plan external arXiv imports for one result's reading path.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must disclose side effects and behavior on its own. 'Plan external arXiv imports' suggests planning rather than executing, but it does not state whether the tool mutates state, accesses external services, requires permissions, or returns a plan. This leaves the agent guessing about observable effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. It loses a point only because the term 'plan' is left vague; otherwise it is appropriately compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and only one sentence of description, an agent lacks information about what the plan contains, how recursive changes behavior, and what side effects or prerequisites apply. This is insufficient for a non-trivial planning tool with two parameters and many planning siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the input schema's bare property names. It only implicitly explains result_id via 'one result' and says nothing about the recursive boolean, its default true behavior, or its effect on planning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (plan), resource (external arXiv imports), and scope (one result's reading path), so an agent can tell it apart from actual import tools like workspace_add_arxiv_paper. It does not explicitly distinguish from the queue/paper planning siblings, though the phrase 'for one result's' narrows the target.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one result's reading path' implies the tool is appropriate when planning imports for a single result, but the description gives no explicit when-to-use or when-not-to-use guidance, and no alternatives such as workspace_plan_external_imports_for_queue or workspace_plan_external_imports_for_paper are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_plan_starter_projectC
Plan a Workspace Starter run without writing files.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | ||
| papers | Yes | ||
| artifact_dir | Yes | ||
| create_queue | No | ||
| project_title | No | ||
| create_session | No | ||
| max_candidates | No | ||
| workspace_path | Yes | ||
| target_result_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses one important behavioral trait: the tool does not write files. However, with no annotations at all, it should also clarify effects, output, or side-effects; it does not. There is also potential ambiguity around the create_queue/create_session parameters, which sound like state-changing actions but are not reconciled with the no-write claim.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, and the key no-write constraint is front-loaded. However, for a tool with nine parameters and no schema descriptions, this is more under-specification than appropriately concise: it omits essential context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given high complexity (nine parameters, three required), zero schema description coverage, no annotations, and no output schema, the description leaves nearly everything undisclosed. It does not explain what a plan looks like, what the tool returns, how parameters interact, or what 'Starter run' means, so it is far from complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no meaning for any of the nine parameters. Nothing explains workspace_path, artifact_dir, papers, focus, project_title, max_candidates, or target_result_id, so the agent must guess at how to populate them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Plan') and identifies a distinct resource/action ('a Workspace Starter run'). The qualifier 'without writing files' adds useful scope and distinguishes this from file-writing operations, though the term 'Workspace Starter run' is domain jargon and not expanded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are provided. The phrase 'without writing files' implies a dry-run/planning context, but the description does not name sibling tools or state when this tool should be preferred over them, leaving usage selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_record_reading_checkpointC
Create or update a reading checkpoint in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| summary | No | ||
| evidence | No | ||
| target_id | Yes | ||
| session_id | Yes | ||
| target_kind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden of behavioral disclosure, but it only says 'Create or update.' It does not describe what a 'checkpoint' is, whether the operation is idempotent, what effects occur on the target, whether existing values are overwritten, or what happens on invalid inputs. This is minimal for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words and is front-loaded with the core action. It is appropriately concise, though it achieves that at the expense of needed detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters, required fields, no output schema, and no annotation support, this description is severely incomplete. It omits expected fields semantics, return behavior, and any differentiation from a large set of workspace reading siblings, leaving agents unable to confidently invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no additional meaning for the six parameters. It does not explain status values, target_kind options, how target_id maps to a resource, what evidence should contain, or how session_id relates to a reading session. The description entirely fails to compensate for the schema being undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair: 'Create or update a reading checkpoint' in the active workspace. It clearly identifies the operation and resource, though it does not differentiate it from sibling reading tools such as workspace_add_reading_note or workspace_create_reading_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, nor any conditions for create versus update. The phrase 'active workspace' hints at a prerequisite, but it is not explained, and no exclusions or routing cues are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_resolve_external_referenceC
Resolve a blocked external reference to a user-confirmed target.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | ||
| paper_id | Yes | ||
| blocked_id | Yes | ||
| artifact_dir | No | ||
| import_target | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It only says 'resolve', which implies an action, but doesn't disclose side effects, reversibility, permissions, or return value. Minimal behavioral info.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single sentence, very concise, but it lacks structure and any detail. It's not verbose, but it's under-specified. For conciseness alone, it's okay, but it doesn't earn its place because it doesn't provide value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, one nested object, no output schema, and no annotations, this description is grossly insufficient. It doesn't explain what a blocked external reference is, what the target object should contain, or what happens after resolution.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and description doesn't explain any parameters. Only 'target' is mentioned in prose but without meaning. No help for agent to fill paper_id, blocked_id, or import_target.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'resolve' and resource 'blocked external reference', and mentions 'user-confirmed target' which implies a confirmation step. However, it doesn't differentiate from sibling workspace_resolve_external_reference_candidate, which likely handles candidates. So the purpose is clear but lacks sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this vs alternatives. The description doesn't mention any context like prerequisites, or that it should be used after a search or candidate resolution. An agent has no way to know when to invoke this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_resolve_external_reference_candidateC
Apply a searched candidate as an explicit reference resolution.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes | ||
| overwrite | No | ||
| blocked_id | Yes | ||
| artifact_dir | No | ||
| candidate_id | Yes | ||
| import_target | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits, but it only states 'Apply' (implying a mutation) without details on side effects, reversibility, permissions, or impact on existing resolutions. It does not mention the overwrite flag or any consequences, leaving the agent without essential safety or impact information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it is under-specified rather than appropriately structured. It lacks front-loaded key details (e.g., the workflow, required parameters) and does not earn its place by conveying essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters, no annotations, and no output schema, the description is woefully incomplete. It does not explain the resolution workflow, what a successful application returns, error conditions, or how this tool fits into the external-reference pipeline. An agent cannot call this tool correctly with the given information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not describe any of the six parameters (paper_id, blocked_id, candidate_id, overwrite, import_target, artifact_dir) beyond the vague 'searched candidate.' This is a critical gap; the agent cannot understand how to supply the arguments correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Apply a searched candidate as an explicit reference resolution' clearly identifies the action (apply) and the resource (a searched candidate) and distinguishes it from sibling tools like workspace_search_external_reference (searching) and workspace_list_external_reference_resolutions (listing). It implies a workflow but doesn't explicitly state the relationship to search, so it's slightly below perfect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as workspace_resolve_external_reference, or on prerequisites like having performed a search first. The description does not mention any conditions, exclusions, or context that would help an agent decide between this and other external-reference tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_search_external_referenceC
Search scholarly metadata providers for a blocked external reference.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| paper_id | Yes | ||
| providers | No | ||
| blocked_id | Yes | ||
| max_candidates | No | ||
| resolver_version | No | deterministic_v2 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Search' and does not explain whether the tool queries external services, records results, has rate limits, or has side effects such as creating a search record.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundant wording. However, it is concise at the expense of substance, so it does not earn the highest rating.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters, no annotations, no output schema, and a crowded group of related reference-resolution tools, one clause is far too thin. An agent cannot infer return values, behavioral effects, preconditions, or how this tool fits into a broader workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description was expected to compensate for six undocumented parameters. It only weakly hints at the meaning of 'blocked_id' and 'providers'; refresh, max_candidates, resolver_version, and paper_id remain semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search'), a resource ('scholarly metadata providers'), and a target ('blocked external reference'). This makes the core purpose understandable, but it does not distinguish the tool from closely related siblings such as workspace_resolve_external_reference or workspace_list_external_reference_searches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no stated preconditions, and no definition of what makes a reference 'blocked.' The phrase implies a search context but leaves the exact triggering situation to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_search_theoremsB
Search theorem titles and bodies across the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| query | Yes | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the scope ('active workspace') but does not clarify search semantics such as exact vs. fuzzy matching, case sensitivity, whether filtering by kind or paper_id narrows scope, or whether the operation is read-only. The description is too thin for a search tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to the core purpose, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no annotations, and several closely related sibling tools, the description is under-specified. It does not explain how filters interact, what counts as a match, or when this tool should be chosen over list_theorems or get_theorem. The presence of an output schema reduces the need to describe return values, but significant operational context is still missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains the general search target ('theorem titles and bodies'). It does not clarify the meaning of 'kind', 'paper_id', or 'limit'. The query parameter's role is inferable from the description, but the filter and pagination parameters are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a clear verb ('Search'), a resource ('theorem titles and bodies'), and a scope ('across the active workspace'). This distinguishes it from sibling tools like list_theorems and get_theorem, which imply enumeration and single-item retrieval respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for query-based search across the active workspace, but it does not explicitly state when to prefer it over list_theorems, get_theorem, or where_used. No alternative tools or exclusions are mentioned, leaving usage conditions somewhat inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_update_reference_expansion_policyB
Explicitly revise numeric budgets; cumulative usage is retained.
| Name | Required | Description | Default |
|---|---|---|---|
| limits | Yes | ||
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It adds one useful non-obvious side effect: 'cumulative usage is retained' when budgets are revised. However, it does not disclose whether changes are reversible, how the update interacts with active reference expansions, or what response is returned, so transparency is incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single tight sentence with no filler. The action is front-loaded, and the behavioral note about cumulative usage is presented efficiently. Both clauses earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating policy-update tool with no annotations, no output schema, and an undocumented nested 'limits' object, this description is too sparse. It does not explain the required run_id, the expected shape of limits, or the tool's place among the many reference expansion lifecycle siblings. An agent would still need to infer or probe critical details before calling it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented parameters. It adds only that numeric budgets are revised, which loosely hints at the 'limits' object, but it does not explain the structure of 'limits' or the role of 'run_id' at all. The 'additionalProperties: true' schema leaves key names completely ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action, 'revise numeric budgets', which maps clearly to updating a reference expansion policy's limits. It is not a tautology and is distinct from sibling lifecycle tools like advance, pause, or cancel. However, it never explicitly names the policy resource, relying on the tool name to supply that context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as workspace_advance_reference_expansion or workspace_pause_reference_expansion. It also gives no prerequisites or exclusion conditions. The only implied context is that this tool is for revising budgets, but that is purpose, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.1.5- Changed
workspace_search_external_reference1 field changed- added
Input schema / properties / resolver_versionAdded value: +{ + "default": "deterministic_v2", + "title": "Resolver Version", + "type": "string" +}
9 tool updates
v1.1.4- Added
workspace_advance_reference_expansion - Added
workspace_cancel_reference_expansion - Added
workspace_create_reference_expansion - Added
workspace_decide_reference_expansion - Added
workspace_export_reference_expansion - Added
workspace_get_reference_expansion - Added
workspace_list_reference_expansions - Added
workspace_pause_reference_expansion - Added
workspace_update_reference_expansion_policy
5 tool updates
v1.1.3- Added
workspace_list_external_reference_resolutions - Added
workspace_list_external_reference_searches - Added
workspace_resolve_external_reference - Added
workspace_resolve_external_reference_candidate - Added
workspace_search_external_reference
2 tool updates
v1.1.0- Added
workspace_bootstrap_reading_project - Added
workspace_plan_starter_project
1 tool update
v0.11.1- Added
workspace_export_cross_paper_reading_plan
1 tool update
v0.11.0- Added
workspace_export_paper_reading_report
1 tool update
v0.10.0- Added
workspace_get_paper_map
30 tool updates
v0.9.3- Added
get_dependency_diagnostics - Added
get_environment_diagnostics - Added
load_arxiv_request - Added
validate_arxiv_input - Added
validate_arxiv_request - Added
workspace_add_pdf_paper - Added
workspace_add_reading_note - Added
workspace_apply_reading_queue_to_session - Added
workspace_create_reading_queue - Added
workspace_create_reading_session - Added
workspace_export_reading_bundle - Added
workspace_export_reading_session_summary - Added
workspace_export_result_reading_context - Added
workspace_get_dependency_diagnostics - Added
workspace_get_evidence - Added
workspace_get_external_result_mentions - Added
workspace_get_proof_dependencies - Added
workspace_get_reading_queue - Added
workspace_get_reading_session - Added
workspace_get_result - Added
workspace_get_result_proof - Added
workspace_get_result_reading_path - Added
workspace_get_source_slice - Added
workspace_list_reading_queues - Added
workspace_list_reading_sessions - Added
workspace_list_results - Added
workspace_plan_external_imports_for_paper - Added
workspace_plan_external_imports_for_queue - Added
workspace_plan_external_imports_for_result - Added
workspace_record_reading_checkpoint
8 tool updates
v0.4.1- Added
open_workspace - Added
workspace_add_arxiv_paper - Added
workspace_add_local_paper - Added
workspace_get_citations - Added
workspace_get_dependencies - Added
workspace_get_paper - Added
workspace_list_papers - Added
workspace_search_theorems
6 tool updates
v0.3.0- First observed
get_dependencies - First observed
get_theorem - First observed
list_theorems - First observed
load_arxiv_paper - First observed
load_paper - First observed
where_used
TDQS
Scored across 63 tools
Many tools have clear roles, but there are several near-duplicate pairs (validate_arxiv_input vs validate_arxiv_request, load_paper/load_arxiv_paper vs workspace_add_*_paper, get_dependencies vs workspace_get_dependencies) that an agent could easily confuse. The external-reference and expansion workflows also use many similar names, making boundaries unclear.
The dominant workspace_verb_noun pattern is readable and consistent, but a significant segment of unprefixed legacy-style tools (load_paper, list_theorems, get_theorem, get_dependencies, where_used, validate_arxiv_*) breaks the convention. It is mixed rather than chaotic, so it still earns a middle score.
With 63 tools, this server is far beyond a typical agent-friendly surface and fits the extreme-count end of the scale. Many tools are workflow-stage variants or duplicates that could be consolidated into fewer, more general operations.
The surface covers the core domain thoroughly: paper ingestion, theorem/dependency extraction, evidence and results, reading sessions/queues, exports, and external reference handling. Minor lifecycle gaps exist (no explicit delete/remove tools for papers, sessions, or queues), but agents can reasonably work around them.
Maintenance
Related MCP Connectors
Academic research MCP server for paper search, citation checks, graphs, and deep research.
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
arXiv MCP — preprint server search (free, no auth)
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Related MCP Servers
- FlicenseBqualityNot gradedmaintenanceEnables extraction of mathematical content from TeX papers and conversion to Lean code through a structured intermediate representation. Supports project scaffolding, entity management, and task tracking for mathematical formalization workflows.14-
- AlicenseAqualityCmaintenanceMCP server for searching and retrieving arXiv papers with full-text PDF extraction.52MIT
- AlicenseNot gradedqualityDmaintenanceEnables searching arXiv, fetching metadata, reading papers as section-aware Markdown, listing recent papers, and downloading PDFs via five MCP tools.9 npm2MIT
- AlicenseNot gradedqualityDmaintenanceRemotely-callable MCP server for academic paper search, full-text retrieval and image to LaTeX conversion across arXiv, Semantic Scholar, and OpenAlex.1MIT