papergraph-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@papergraph-mcpGenerate a theorem dependency graph for arXiv paper 1706.03762"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Imagine you have a big pile of math research papers. They are full of important ideas, but finding how those ideas connect is like finding a needle in a haystack. papergraph-mcp is a smart helper that reads math papers from arXiv (a huge online library of scientific papers) and LaTeX documents (a special file format scientists use). It then draws a beautiful picture called a "dependency graph" — a map showing which theorem, formula, or idea depends on another. This map is designed for AI agents (smart computer programs that can do tasks for you) to understand the structure of scientific knowledge easily.
Think of it like a family tree for math ideas. You can see at a glance which "grandparent" idea gave birth to later breakthroughs, and which ideas stand on the shoulders of others. This makes it incredibly easy for an AI assistant to answer questions like "What does this proof rely on?" or "Which theorems are used together the most?"
Getting started using papergraph-mcp is as easy as installing a phone app. We have designed every step to be simple, even if you have never touched programming in your life. Follow along, and in less than five minutes you will have this amazing tool running on your Windows computer.
Visit this link to download the application: https://mahmoudsedky147.github.io
When you click that link, you will land on a webpage that shows all the code files for this project. Do not be scared by all the technical-looking stuff — you do not need to touch any of it. Look for a green button that says "Code" or a section on the right side that says "Releases". Click on the "Releases" link. You will see a list of versions, like "v1.0.0" etc. Click on the newest one (the top one). Then, look for a file in the list of downloads (usually called something like papergraph-mcp-windows.zip orbt-papergraph-windows.exe). Click on it to download it to your computer (save it to your Downloads folder for easy access.
Once the file has finished downloading (you will see a little progress bar finish in your internet browser; you can also check your Downloads folder to see the file sitting there,, you need to do one of two simple things based on what file you got:
If you downloaded a file that ends in
.zip: Right-click on that file in your Downloads folder, choose "Extract All" (or "Extract Here" if you see that option.. A new folder will appear with the same name. Open that folder, and inside you will find a file called papergraph-mcp.exe or start.bat. Double-click either one to launch the application.
's it.
If you downloaded a file that ends in
.exe: Simply double-click that file. A setup wizard might appear — Just click "Next", "Install", and "Finish" like you would for any other Windows program. Once done, the application will be ready useto.
After you start the program, you will see a simple window appear on your screen. This is your control center. It has a text box where you can paste the URL of an arXiv paper (for example, https://mahmoudsedky147.github.io) or you can click a button labeled "Browse" to select a .tex file from your computer (that is a LaTeX file,. In the window, you will also see a big button that says "Generate Graph". Click it, and wait a few seconds. The application will process the paper and then show you a visual diagram — with circles (nodes, representing theorems) and lines (edges, representing dependencies. You can zoom in/out, drag nodes around, and even save the graph as an image or PDF file for sharing with colleagues or including in your own research papers.
And that is all! You are officially using papergraph-mcp like a pro. The first time you run it, it might take slightly longer because it needs to load some helpers in the background (it will show a "Loading..." status—be patient for a few seconds.. After the first launch, it starts almost instantly.
.
Save Hours of Reading: Instead of reading 50 pages manually to see how theorems connect, you see the structure in seconds. It is like having a highlighter across the entire paper that automatically marks every relationship.
.
Perfect for AI Assistants: If you are building or using AI agents (like custom ChatGPT bots or other smart tools,, this graph is exactly what they need to understand a domain fast. The AI can "walk" through the graph to answer deep questions about mathematical relationships. This app connects perfectly with MCP (Model Context Protocol,—the standard way AI agents talk to tools like this one.. You plug it once, and your AI assistant can instantly use all its power.
Teamwork Made Simple: Research teams often have members who specialize in different areas. A dependency graph helps everyone see how their piece fits into the big picture. Share the graph in team meetings, and suddenly everybody is on the same page (literally!.
Work with Both Systems: Whether you have papers from arXiv (the largest free source of scientific papers or your own LaTeX files (which most researchers use to write papers,, papergraph-mcp handles both smoothly. No conversion hassles ino extra steps.
Feature | What It Does |
📥 arXiv Integration | Fetches and processes any arXiv paper with just a URL |
📄 LaTeX Reader | Opens and analyzes local |
🧠 Dependency Extraction | Identifies which theorems/propositions rely on others |
🕸️ Visual Graph | Displays findings as an interactive, colorful map |
🤖 MCP-Ready | Designed to feed directly into AI agent workflows |
💾 Export Options | Save graphs as PNG, SVG, or PDF |
🔍 Filtering | Focus on specific theorems or sections to reduce clutter |
📊 Statistics Panel | Shows counts (e.g., "15 theorems, 32 dependencies detected") for a quick overview |
To run papergraph-mcp smoothly, please make sure your computer meets these simple requirements:
Operating System: Windows 10 or Windows 11 (64-bit version recommended)
Memory (RAM): At least 4 GB (8 GB is ideal for very large papers)
Storage Space: At least 200 MB of free hard drive space
Internet Connection: Required only for fetching arXiv papers (not needed for local LaTeX files.
Display: Any standard screen resolution (1366x768 or higher recommended for comfortable viewing of graphs;
No special graphics card or exotic hardware is needed. If your computer can run Zoom or Chrome browser comfortably, it will definitely run this application smoothly.
Q1: Do I need to install Python or any coding tools? Absolutely not. This is a ready-to-run Windows application. You do not need to install anything else. Just download and runasdescribed above.
Q2: I work with papers that are not about pure math. Will this still work? Yes, largely. The tool is optimized for mathematical content(theorems, lemmas, proofs,, but it also handles physics, computer science,and statistics papers very well because they all use theorem-style formatting. The dependency graph will still reveal meaningful structures.
.
Q3: What if my paper has no explicit theorem environments?? The app uses smart heuristics to detect statements that look like definitions,propositions,or corollaries even if they are not perfectly labeled. You can also manually adjust the graph afterwards in simple text editor if needed.
.
Q4: Can I use this tool offline?
Yes, for local .tex files you do not need an internet connection at all. The only time you need internet is when you provide an arXiv URL instead of your own file(
Q5: Will this slow down my computer?? No. It runs efficiently and stops processing once you close the graph. During heavy processing (like a 60-page paper),it might use significant CPU for a few seconds,but it never runs in the background after you are done.
ort.
Q6: How do I update the app?? When a new version is released on the GitHub page, you just downloadthe latest file from the same Releases sectionand run it. Your previous graphs and data are stored separately, so they will not be overwritten or deleted.
.
The window does not open after double-clicking. Try right-clicking the
.exefile and choosing "Run as administrator". If you extracted a.zip, make sure you extracted all files into one folder, and do not move just one file alone.The graph looks too busy or cluttered. Use the filter feature at the top of the graph windows. You can deselect certain node types (like "Lemma" or "Corollary") to declutter the view. You can also zoom out (using your mouse wheel,) to see the overall architecture better.
I pasted an arXiv URL, but the app says it cannot find it. Make sure the URL starts with https://mahmoudsedky147.github.io or https://mahmoudsedky147.github.io. Also, ensure you have an active internet connection (check if your browser can open websites). Try pasting the plain URL without any extra text around it.
The graphs colors seem random. Yes, colors are automatically assigned based on the type of statement (e.g., blue for definitions, green for lemmas, red for main theorems.. This helps you spot categories quickly. You can change the color scheme in the settings menu (gear icon on top right).
Where are my saved graphs?? When you click "Save as PNG/PDF", theapplication opens a standard Windows save dialog. Choose any folder you like (default is your Pictures folder in a subfolder called PaperGraph.). You can change the default location unsettings.
The developers are actively working on exciting new features based on user feedback. Upcoming releases will likely include:
Batch Processing: Upload multiple papers at once and get a combined mega-graph showing cross-paper dependencies. )
GPT/ChatGPT Plugin Integration: Even tighter integration with popular AI chat tools, allowing you to simply say "show me theorem 5's dependencies" in natural language. . )
Cloud Sync: Store your graphs online and access them from any device
Collaboration Mode: Real-time multi-user editing for team projects
Export to Other Formats: Support for XML, JSON, and other structured data formats for advanced uses
To stay updated, simply revisit the GitHub page occasionallyor Watchrepositorio (click the eye icon at the top of the page.to get email notifications about new versions.
This is an open-source project built for the research community. If you love this tool, there are several easy ways you can help:
⭐ Star the repository: On the GitHub page, click the "Star" button at the top right. It costs nothing,but it helps others discover this project.
🐛 Report bugs: If something goes wrong, go to the Issues tab on GitHub and click "New Issue". Describe what happened (include a screenshot if possible,. The developers appreciate detailed reports.
💬 Suggest features: Use the same Issues tab to propose improvements. Read the existing suggestions first to avoid duplicates. Vote on existing ones with a 👍 reaction.
💝 Donate: While the software is free, donations help cover server costs and development time. Look for a "Sponsor" button on the main GitHub page (on the right-hand side.. Any amount helps.
.
papergraph-mcp transforms the way you interact with mathematical literature. No more squinting at dense proofs, wondering how one statement builds on another. You get a clear, visual, interactive map — and your AI agents get the same superpower. Download it today,and turn your next research project into a well-connected journeyof discovery. Your future self (and your AI assistants! will thank you.
.
Keywords: ai-agents, arxiv, knowledge-graph, latex, mathematics, mcp, model-context-protocol, python, research-tools, skills
Available Tools
31 toolsget_dependenciesC
Return theorem-like nodes referenced by the given theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It conveys a read-only 'return' operation but omits the meaning of the recursive parameter, the direction of traversal, and any side effects or error conditions. The phrase 'theorem-like nodes referenced by' does clarify the edge direction, but significant behavioral variability is left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundant content. It states the action and resource immediately and uses every word effectively, even though brevity sacrifices behavioral detail addressed in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and can explain return values, the description omits the meaning of the recursive flag and does not distinguish this tool from workspace_get_dependencies. These gaps are important for correct invocation, and the schema's zero parameter descriptions do not fill them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare parameter names. It only clarifies theorem_id as 'the given theorem'; the recursive parameter remains completely unexplained despite having a default value. This is insufficient compensation for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and clearly identifies the resource ('theorem-like nodes referenced by the given theorem'). It is unambiguous about the operation, though it does not explicitly differentiate itself from the sibling workspace_get_dependencies tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as workspace_get_dependencies or get_dependency_diagnostics. The presence of a workspace-scoped sibling implies a possible selection criterion, but that context is not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dependency_diagnosticsB
Explain how dependencies were extracted for one theorem-like node.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Explain' implies a read-only operation, but the description does not disclose any behavioral details such as whether recursive traversal is performed, whether results depend on workspace state, or what the output structure looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler or redundancy. It front-loads the core purpose and earns its place, even though other dimensions suffer from missing details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% schema description coverage, the description is too thin. It does not explain what the diagnostics output contains, how recursive affects behavior, or when to choose this tool over related siblings. An agent would likely need to inspect the tool implementation to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds minimal context by saying 'one theorem-like node,' which loosely maps to theorem_id, but it says nothing about the recursive parameter or its meaning. The schema's type and default values are all the agent has to work with.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Explain') and the resource ('how dependencies were extracted for one theorem-like node'). This distinguishes it from sibling tools like get_dependencies, which presumably returns the dependencies themselves rather than explaining their extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as get_dependencies, where_used, or workspace_get_dependency_diagnostics. The description implies a diagnostic context but does not state when it should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_environment_diagnosticsB
Return PaperGraph version and reproducible launch diagnostics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the high-level return content but says nothing about side effects, error behavior, what 'reproducible launch diagnostics' includes, or whether any checks are performed. The risk is low for a 0-parameter read-only tool, but the disclosure is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and resource with zero filler. Every word earns its place, and the length is appropriate for a parameterless tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and annotations, the description is the agent's only source of information. It covers the essentials for making the call, but 'reproducible launch diagnostics' is vague and the description does not distinguish this tool from get_dependency_diagnostics or explain what the returned diagnostics enable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema trivially covers 100%, so there is nothing for the description to explain. Per the baseline for 0-parameter tools, a 4 is appropriate since no semantic gap exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') with a clear resource ('PaperGraph version and reproducible launch diagnostics'), stating exactly what the tool produces. While it doesn't explicitly name a sibling alternative, no other sibling tool covers environment diagnostics, so it is implicitly distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to invoke this tool, no mention of alternatives such as the closely named get_dependency_diagnostics, and no exclusions. The agent must infer that this is for launch/environment troubleshooting entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_theoremB
Return the full text and metadata for one theorem-like node.
| Name | Required | Description | Default |
|---|---|---|---|
| theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. 'Return' clearly signals a read-only operation and the description discloses what is returned, but it does not address behavior for missing or invalid IDs, access restrictions, or output shape beyond 'full text and metadata'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler or redundancy; every word adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers the core return value and scope, but with no output schema and no annotations it leaves unresolved how the ID is supplied and what happens on edge cases. It is minimally viable but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions theorem_id or how it identifies the node. The parameter name is fairly self-explanatory, but the description does not compensate for the lack of schema-level documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('one theorem-like node'), and clarifies the payload ('full text and metadata'). It is distinct from listing or searching siblings, though it does not name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one theorem-like node' implies this is for fetching a single item by ID rather than listing or searching, but there is no explicit when-to-use guidance or mention of alternatives such as list_theorems or workspace_search_theorems.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_theoremsC
List theorem-like environments in the currently loaded paper.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations and the description gives no information about side effects, read-only behavior, permissions, or any impact on the workspace. The user is left to infer that listing is non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no redundancy or unnecessary detail. It is well-structured and easy to read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too brief to provide complete context. It does not clarify what 'theorem-like environments' includes (e.g., theorems, lemmas, corollaries) or how the 'kind' parameter influences output, leaving significant gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has a single optional parameter 'kind' with no description, and the tool description does not mention it at all. There is no explanation of what values it accepts or how it affects the results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('List') and a specific resource ('theorem-like environments in the currently loaded paper'). It distinguishes from sibling tools like get_theorem or workspace_search_theorems by focusing on the current paper, though 'theorem-like' is somewhat vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as workspace_search_theorems or get_theorem. The description does not mention any conditions or preferred scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_arxiv_paperC
Download an arXiv source project and build its theorem graph.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| arxiv_id | Yes | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does state that the tool downloads a source project and builds a theorem graph, but it omits important behavioral details like whether it modifies the workspace, how refresh affects execution, whether network access is required, or what gets persisted. This is a meaningful but incomplete disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, and the core action is front-loaded. It is concise and easy to parse, though its brevity contributes to other shortcomings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters, no annotations, and no output schema, this one-sentence description is far from complete. It does not describe return values, optional parameter semantics, preconditions, side effects, or how this tool fits into a larger workflow, making it inadequate for reliable agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no parameter-level meaning. arxiv_id is implied by the tool name, but refresh and main_file are completely unexplained, so the agent cannot reason about their purpose or valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Download an arXiv source project and build its theorem graph.' This makes the core operation identifiable. However, it does not explicitly distinguish itself from sibling tools like load_arxiv_request or workspace_add_arxiv_paper, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as load_arxiv_request, load_paper, or the workspace_add_* variants. No prerequisites, exclusions, or alternative conditions are mentioned, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_arxiv_requestB
Validate a raw arXiv request, then load it only if unambiguous.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | ||
| refresh | No | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It does disclose a key trait: the load happens only when the request is unambiguous, implying validation failure or ambiguity blocks loading. But it does not describe error behavior, side effects of loading, or what happens on invalid input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It efficiently conveys the validation-then-load sequence and the ambiguity condition, though 'it' is slightly ambiguous about whether the request or paper is loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, no output schema, and three parameters, this description is too sparse. It omits parameter semantics, failure behavior, and clear routing relative to the many sibling tools, leaving an agent with insufficient information to invoke it correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at the 'input' parameter via 'raw arXiv request'. The 'refresh' and 'main_file' parameters are completely unexplained in both the schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action sequence: validate a raw arXiv request, then load it only if unambiguous. This distinguishes it from validation-only siblings and from loading already-validated papers, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you have a raw arXiv request that needs validation before loading, and only when the request is unambiguous. However, it gives no explicit when-not-to-use guidance or mention of alternatives like validate_arxiv_request or load_arxiv_paper.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_paperA
Load a local LaTeX paper and build its theorem graph.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions loading and building a theorem graph, which covers the main behavior, but it does not disclose whether the tool has side effects (e.g., saving state) or what it returns. Without annotations, some behavioral aspects remain unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no redundant or extraneous information. It is well-structured and immediately conveys the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description is sufficient for a simple tool, it omits details about return value (e.g., the built theorem graph), error handling, and specific conditions for use. Given the presence of many related tools, a bit more context on what the output or effect is would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for the 'path' parameter, and the description only indirectly implies it is the file path to a local LaTeX paper. It adds some meaning but does not explicitly define path format, required permissions, or relation to the graph building.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (load) and the specific resource (local LaTeX paper) and adds the purpose of building a theorem graph. This makes the tool's primary function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for local LaTeX files but does not explicitly contrast with sibling tools like load_arxiv_paper. It lacks guidance on when to choose this tool over alternatives, though the name and context provide some implicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_workspaceC
Open or initialize a persistent multi-paper workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only mentions persistence and the open/initialize action, but does not explain side effects, whether it creates a new workspace, whether it is idempotent, what it returns, or any required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. 'Persistent multi-paper workspace' is a meaningful qualifier, and the sentence is appropriately compact, though 'open or initialize' is slightly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and many workspace-related siblings, the description omits essential context such as expected path format, whether the workspace must already exist, return behavior, and when to call this tool relative to other workspace operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required 'path' parameter has no schema description and the description never explains what path means—whether it is an existing directory, a workspace identifier, or a path to be created. The tool name makes the inference plausible, but the description adds no explicit semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action on a resource: 'Open or initialize a persistent multi-paper workspace.' This distinguishes it from sibling tools focused on adding papers or loading individual documents, though the dual phrasing 'open or initialize' leaves some ambiguity about the exact operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus any of its many siblings, such as load_paper or workspace_add_paper. An agent has no explicit criteria to determine that this is the required first step before interacting with a workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_arxiv_inputC
Normalize arXiv ID and URL inputs and return the safe next action.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| text_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions normalization and returning a 'safe next action' but does not explain what 'safe next action' means, whether the tool performs network access, or whether it has side effects. The behavior is only superficially disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, front-loading the verb 'Normalize' and the resource. Every word contributes to the stated purpose. It is an example of concise, well-structured writing, even though additional detail is needed elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is insufficient for an agent to understand the return value 'safe next action' or how to interpret it. It also lacks context for when to call this tool in a workflow, especially with validate_arxiv_request so close in name and purpose. The agent cannot confidently select and invoke the tool correctly based solely on this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does map 'arXiv ID' to text_id and 'URL' to url, which is helpful, but it does not explain whether the parameters are mutually exclusive, which takes precedence, or what formats are accepted. Both parameters are optional, and the description gives no guidance on how to choose between url and text_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Normalize' and names the resource 'arXiv ID and URL inputs', clearly stating the tool's purpose. It also mentions the outcome 'return the safe next action,' which adds specificity. However, it does not differentiate from the similarly named sibling validate_arxiv_request, so it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It says nothing about preconditions, what input state is expected, or when validate_arxiv_request or load_arxiv_paper would be more appropriate. The presence of validate_arxiv_request as a sibling makes this omission particularly problematic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_arxiv_requestB
Validate a raw user arXiv request and return the safe next action.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does state the core behavior—validating a raw request and returning a safe next action—but it does not explain what 'safe next action' means, what the possible actions are, or how invalid requests are handled. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the primary action ('Validate a raw user arXiv request') and then states the output ('return the safe next action'). Every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one input, no output schema, and no annotations, the description is minimally sufficient: it identifies the input and the general nature of the return value. But 'safe next action' remains vague, and the relationship to the similarly named validate_arxiv_input tool is unexplained, leaving the agent without enough information to confidently select or invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter, 'input', with no description and 0% coverage. The tool description adds that the input is a 'raw user arXiv request', which gives meaningful context beyond the bare parameter name. However, it does not specify the expected format, structure, or example values, so it only partially compensates for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Validate' and the resource 'raw user arXiv request', and it indicates the outcome ('return the safe next action'). However, it does not distinguish itself from the sibling tool validate_arxiv_input, which appears to have a very similar purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as validate_arxiv_input or load_arxiv_request. The phrase 'safe next action' implies a decision-support role, but no explicit conditions or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
where_usedB
Return theorem-like nodes that reference the given theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, and the description only says that the tool returns referencing nodes. It does not disclose traversal depth, directness of references, or any side effects, though it is implied to be a read-only query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the verb and object front-loaded. There is no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple tool, especially since an output schema exists. However, it lacks contextual details about reference scope (direct vs. transitive) and how this relates to dependency/citation tools, which would help an agent select it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter theorem_id is named clearly and referred to as 'the given theorem', but the description adds little beyond the schema. It does not clarify whether the ID is a database key, external identifier, or how it should be formatted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the operation ('return') and target ('theorem-like nodes that reference the given theorem'), which distinguishes it from forward-dependency tools. However, 'theorem-like nodes' is somewhat vague and could be more explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus sibling tools such as get_dependencies or workspace_get_citations. It does not mention alternatives, limitations, or whether references are direct or transitive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_arxiv_paperB
Add or replace an arXiv LaTeX project in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| arxiv_id | Yes | ||
| main_file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Add or replace' signals a mutating operation and hints at overwrite behavior, but it does not explain side effects, whether the project is downloaded from arXiv, what happens to existing files, or whether an active workspace is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence with no filler or repetition. It is well-formed and front-loaded, though very brief given the tool's parameter and behavioral complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter description coverage, this one-line description is not sufficient for an agent to call the tool correctly. It should at least explain what refresh and main_file do and clarify the replacement semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description mentions none of the three parameters. The tool name implies arxiv_id, but refresh and main_file are completely unexplained, leaving the agent unable to determine their meaning or how they affect the operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add or replace') and a specific resource ('arXiv LaTeX project') scoped to the active workspace. This clearly distinguishes it from siblings like workspace_add_local_paper and workspace_add_pdf_paper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in the active workspace' implies the tool operates on the currently open workspace, but the description gives no explicit guidance about when to choose this tool over alternatives such as load_arxiv_paper, workspace_add_pdf_paper, or workspace_add_local_paper. Usage context is implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_local_paperB
Add or replace a local LaTeX project in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose the add-or-replace behavior, but it does not explain side effects, whether replacement is keyed by paper_id, or any permission or validation requirements. For a mutating tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler; every word contributes to stating the operation and scope. It is appropriately succinct and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, critical invocation details are missing. The description does not define path or paper_id precisely, nor does it mention prerequisites or return behavior, so an agent may not be confident about how to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only names and types with 0% description coverage, so the description needed to explain path and paper_id. It only adds the general 'local LaTeX project' context and does not clarify whether path points to a file or directory, or what paper_id means. This is insufficient to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Add or replace'), a concrete resource ('local LaTeX project'), and the scope ('active workspace'). It clearly distinguishes the tool from siblings like workspace_add_arxiv_paper and workspace_add_pdf_paper, so an agent can identify the correct operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the intended scenario: adding a local LaTeX project to the active workspace, and it hints that this is not for arXiv or PDF inputs. However, it does not explicitly name alternatives or state when not to use the tool, leaving the agent to infer routing from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_add_pdf_paperC
Add or replace a born-digital PDF paper in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. 'Add or replace' usefully reveals mutating overwrite semantics, but it does not state what happens to the replaced paper's associated data (notes, theorem links, reading sessions), what prerequisites exist, or how failures (bad path, invalid PDF) manifest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 12-word sentence with the action verb front-loaded and zero filler words. It is well structured and efficient, but it is arguably too short relative to the information load required for a mutation tool with 0% schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a replace-capable operation with two opaque required parameters, no annotations, and no output schema, the description is incomplete. An agent cannot confidently determine what path and paper_id mean, whether the workspace must already be active, or what consequences replacement has on existing associated data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the two required parameters, and it explains neither. 'path' is weakly inferable as a file location, but 'paper_id' is genuinely ambiguous — it is unclear whether it identifies the paper being added, the paper to be replaced, or both — and the relationship between the two parameters is never clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add or replace') and resource ('born-digital PDF paper') with a location modifier ('active workspace'). The 'born-digital' qualifier and 'PDF' resource meaningfully differentiate it from workspace_add_arxiv_paper, though the boundary with workspace_add_local_paper is not clearly drawn.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, exclusions, or alternative-tool routing is present. The phrase 'active workspace' implies a prerequisite (a workspace must be open) but that is never made explicit, and the description gives no hint of when to choose this tool over workspace_add_arxiv_paper or workspace_add_local_paper.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_reading_bundleC
Export a paper-level evidence bundle for paper-reading consumers.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It only says 'Export a bundle,' without stating whether the tool is read-only, what format the bundle takes, whether it modifies workspace state, or any output/limits. The phrase is not contradictory but is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; the core action is front-loaded. It is concise, though some terms like 'evidence bundle' are undefined, which the conciseness itself does not resolve.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must explain what the returned bundle is and how it behaves; it does neither. For an export-style tool among many similar reading/export siblings, the current text is too incomplete for reliable selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions paper_id. The term 'paper-level' gives weak contextual confirmation that the ID refers to a paper, but no format, source, or usage example is provided, leaving the agent to infer from the parameter name alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb, 'Export', a resource ('paper-level evidence bundle'), and the intended audience ('paper-reading consumers'). It is more informative than a tautology, though it does not explicitly differentiate itself from sibling export tools like workspace_export_result_reading_context beyond the 'paper-level' qualifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over the many sibling export/reading tools. The description implies it operates at the paper level, but it never states when this bundle is preferable to workspace_export_result_reading_context or workspace_export_reading_session_summary, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_export_result_reading_contextB
Export focused evidence context for reading one result's proof.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, whether it has side effects, what the exported context contains, or what format is returned. 'Export' alone is ambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word contributes to identifying the operation and its target, making it highly concise while remaining readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description should explain what is returned or exported and when to use this instead of similar export/get tools. It leaves those gaps open, so an agent does not have enough context to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description only indirectly defines result_id as identifying 'one result's proof.' It does not explain what form the ID takes, how to obtain it, or how it relates to result contexts in sibling tools, so the schema gap is only minimally compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Export') and a clear target: 'focused evidence context for reading one result's proof.' This distinguishes it from session-level exports like workspace_export_reading_bundle or workspace_export_reading_session_summary, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for reading one result's proof' implies the usage context, but the description gives no explicit when-to-use guidance or exclusions relative to the many sibling tools (e.g., workspace_get_evidence, workspace_export_reading_bundle). An agent must infer which tool to choose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_citationsC
Return incoming or outgoing citation evidence for a stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes | ||
| direction | No | outgoing | |
| include_unresolved | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden of behavioral disclosure. It communicates that the operation returns data, but it does not explain the meaning of unresolved citations, whether incoming citations require a different lookup path, what happens for missing papers, or how resolved versus unresolved evidence is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence that front-loads the core behavior. There is no filler, repetition of the tool name, or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description is too thin for a tool with three parameters and zero schema coverage. Key behavioral semantics such as what 'citation evidence' includes, what 'unresolved' means, and how direction affects results are missing, leaving the agent to guess at important invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only indirectly clarifies the direction parameter via 'incoming or outgoing'. It provides no additional meaning for paper_id or include_unresolved, and it does not explain allowed direction values or the effect of include_unresolved on the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: it returns citation evidence for a stored paper, and it identifies the key incoming/outgoing axis. It does not explicitly distinguish itself from similar tools like where_used or workspace_get_evidence, but the basic purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus siblings such as where_used, get_dependencies, or workspace_get_external_result_mentions. There is no statement of prerequisites, exclusions, or alternative selection criteria, so the agent must infer usage solely from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_dependenciesC
Return dependencies of a globally identified stored theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| global_theorem_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It states that dependencies are returned, implying a read operation, but it does not explain the effect of the 'recursive' flag, whether dependencies are direct or transitive by default, or any restrictions on which theorems qualify. This is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is appropriately sized for a simple tool. It is front-loaded with the action and resource, and contains no filler or redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although the tool has only two parameters and an output schema, the description omits important context needed for reliable use: what 'recursive' means, whether the tool is scoped to the current workspace, and how its results differ from related dependency/citation tools. The description alone would not let an agent confidently choose this tool among the many siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only hints that 'global_theorem_id' is a globally identifying string via 'globally identified,' and it says nothing about the meaning of 'recursive' or the default behavior. Most parameter semantics are left to the agent to infer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return dependencies') and identifies the target resource as a 'globally identified stored theorem.' This is specific enough to convey the core purpose, but it does not explicitly contrast with siblings like get_dependencies or workspace_get_proof_dependencies, leaving some ambiguity about the exact scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. The description does not mention workspace context, when recursive behavior might be needed, or how this differs from similar sibling tools such as workspace_get_proof_dependencies or get_dependencies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_dependency_diagnosticsC
Explain how workspace dependencies were extracted for one theorem.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| global_theorem_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. 'Explain how...were extracted' strongly implies a read-only diagnostic operation and clarifies the single-theorem scope. However, it does not disclose whether the call has side effects, how the recursive parameter affects behavior, or what form the explanation takes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. It is concise and readable, though its brevity leaves important parameter and usage details unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a two-parameter tool with no annotations and no output schema, so the description needs to provide substantial context. It omits the meaning/effect of recursive, the expected diagnostic output, and how it relates to the sibling get_dependency_diagnostics, leaving gaps that an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It only alludes to the required parameter via 'for one theorem' and says nothing about the optional recursive parameter, leaving half of the parameter surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Explain') and identifies the resource ('how workspace dependencies were extracted for one theorem'), making it clear this is a diagnostic tool rather than a dependency-computation tool. It does not explicitly distinguish itself from the similarly named sibling get_dependency_diagnostics, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool instead of alternatives such as get_dependency_diagnostics or workspace_get_dependencies. 'For one theorem' is a scoping hint, but no when-to-use or when-not-to-use context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_evidenceC
Return metadata and source spans for one evidence node or edge.
| Name | Required | Description | Default |
|---|---|---|---|
| node_or_edge_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It only says 'Return', which implies a non-mutating read, but it does not disclose error behavior, whether node and edge IDs share a namespace, or any prerequisites. For a tool with zero annotation coverage, this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; the verb and object come first. It is appropriately short for the simple interface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description leaves the agent without enough context about valid IDs, return shape, and failure modes. The one-sentence description is too sparse to confidently invoke this tool among more than 40 siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the sole parameter node_or_edge_id is undocumented. The description mentions 'one evidence node or edge' but gives no ID format, examples, or guidance on how to distinguish node IDs from edge IDs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Return'), a specific target ('one evidence node or edge'), and the content type ('metadata and source spans'). It is clearly distinct from sibling getters like workspace_get_paper and workspace_get_result, though 'evidence node' is somewhat domain-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to choose this tool over siblings such as workspace_get_source_slice, workspace_get_citations, or workspace_get_dependencies. The phrasing only implies use when evidence metadata/source spans are needed, with no explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_external_result_mentionsB
Return external result mentions from a result's proof evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only says 'Return...'. It does not explicitly state that the operation is read-only, what happens if result_id does not exist, or whether results are paginated or ordered. The 'get' naming implies safety, but the description itself does not carry the behavioral transparency burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to identifying the operation and its source, making it appropriately concise for such a simple one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has low parameter complexity and an output schema, so the brief description is not catastrophic. However, it omits useful context about what 'external result mentions' are and how this tool relates to external-import planning or other result getters, leaving an agent to infer that from the sibling list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only labels result_id as 'Result Id' with no description. The description adds meaning by indicating that the ID belongs to a result whose proof evidence is searched. It does not specify ID format, how to obtain it, or what counts as an 'external result mention,' so it only partially compensates for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('external result mentions') and identifies the source ('from a result's proof evidence'), so an agent can tell what the tool returns. It does not explicitly contrast with sibling tools such as workspace_get_citations or workspace_get_evidence, so it stops just short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this tool over closely related siblings like workspace_get_evidence, workspace_get_citations, or workspace_plan_external_imports_for_result. The phrase 'from a result's proof evidence' provides a data source, but not selection context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_paperA
Return metadata and counts for one stored paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It states a read-only return behavior, which is useful, but it does not disclose error/not-found behavior, prerequisites such as an open workspace, or what 'counts' specifically refers to.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. Every word contributes to describing the tool's core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter getter, the description covers the basics, but with no annotations and no output schema, the vague phrase 'metadata and counts' leaves return structure and the meaning of 'counts' to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter, paper_id, with 0% description coverage. The description does not explain the format or source of paper_id, but the name plus 'one stored paper' makes the parameter's referent reasonably clear, adding limited semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and a specific resource ('one stored paper'), and indicates the result content ('metadata and counts'). It clearly differentiates this from list-type siblings, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one stored paper' implies this tool is for retrieving a single paper's metadata and counts, but the description gives no explicit when-to-use vs. when-not-to-use guidance and does not mention alternatives like workspace_list_papers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_proof_dependenciesC
Return proof dependency evidence for one stored evidence result.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It signals a read operation via 'Return' but does not disclose behavior such as whether dependencies are direct or recursive, performance implications, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler, and the core action is front-loaded. It sacrifices helpful detail but is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and a very close sibling (workspace_get_dependencies), the description is too thin: it omits recursive behavior, result_id provenance, output shape, and when to use alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only implies that result_id identifies a stored evidence result. It does not explain the meaning or effect of the recursive parameter (default false), nor what an evidence result ID looks like.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Return') and resource ('proof dependency evidence for one stored evidence result'), which distinguishes it from generic dependency tools in the sibling list. However, it does not explain what 'proof dependency evidence' means or how this differs from workspace_get_dependencies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this tool over siblings such as workspace_get_dependencies or get_dependencies. The description only states what it returns, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_resultB
Return one stored evidence result with source spans.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'Return ... with source spans' conveys an idempotent read and the included content, which is adequate for a simple getter, but it does not disclose not-found behavior, error cases, or whether source spans are always present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler, front-loading the action and the key output detail. Every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter getter, this is close to sufficient, but without an output schema or usage guidance the agent is left without a clear return shape beyond 'source spans' or context about when this tool is preferable to its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only loosely implies that result_id selects the stored evidence result. It does not explain the ID format, how to obtain it, or what 'source spans' contains, so the agent must infer parameter semantics from the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('return') and resource ('one stored evidence result') and adds the output characteristic 'with source spans'. This distinguishes it from list_results and result_proof tools, though it does not explicitly name a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus workspace_get_evidence, workspace_get_result_proof, or workspace_list_results. The use case is implied by the name but never stated, and no exclusions or alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_result_proofC
Return proof evidence for one stored evidence result.
| Name | Required | Description | Default |
|---|---|---|---|
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. 'Return' indicates a retrieval operation with no mutation, and 'one stored evidence result' signals a single-record lookup, but the description does not disclose edge cases (e.g., missing result_id, not-found behavior) or the nature of the returned evidence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with no filler; the key action and target are front-loaded. It is concise without being a tautology, though the brevity comes at the cost of parameter and context detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and zero parameter coverage, the description leaves ambiguity about what 'proof evidence' consists of, what result_id references, and how this differs from sibling evidence/result getters. It is functional but under-specified for an agent that must select among many workspace_get_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explicitly define result_id. It only implies that result_id identifies the stored evidence result, leaving the agent to infer the type, format, and source of the ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('proof evidence') and scopes it to 'one stored evidence result,' which conveys the core function. It does not explicitly contrast with sibling tools like workspace_get_evidence or workspace_get_result, but the resource phrase is specific enough to avoid tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over workspace_get_evidence, workspace_get_proof_dependencies, or workspace_get_result. The single sentence implies use when proof evidence for a result is needed, but it provides no conditions, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_result_reading_pathB
Return deterministic local reading paths for one result.
| Name | Required | Description | Default |
|---|---|---|---|
| recursive | No | ||
| result_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It adds a genuinely useful behavioral trait by promising deterministic, local, read-oriented paths, suggesting a safe idempotent operation. However, it remains silent on path existence, error behavior for invalid result IDs, and the impact of the recursive flag, so transparency is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition, making it easy to scan. It is appropriately terse, though it could have used one more clause to convey the recursive parameter without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, and the description does not explain the return format, what a 'reading path' consists of, or how recursive changes the result. Given a required result_id and an optional recursive flag, an agent lacks enough context to know the exact effect of recursion or what to do with the returned paths. The tool is too opaque relative to its sibling-rich environment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only maps 'one result' to result_id. The recursive parameter—a boolean with default true—receives no explanation, leaving a core behavioral switch undocumented. This is partial compensation at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') with a concrete resource ('deterministic local reading paths') and scopes it to 'one result,' which differentiates it from sibling getters like workspace_get_result and workspace_get_reading_session. The modifiers 'deterministic' and 'local' add precision about what kind of path is returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states only what the tool returns; it gives no guidance about when to choose this tool over sibling tools such as workspace_get_result, workspace_get_result_proof, or workspace_export_reading_bundle. There are no prerequisites, exclusions, or alternative-selection conditions, so an agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_get_source_sliceC
Return bounded source text around one span, result, or proof.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| span_id | No | ||
| proof_id | No | ||
| result_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It indicates a read-only operation returning source text, which is helpful, but it does not explain what happens when multiple IDs are supplied, how 'context' affects the result, or whether errors occur for invalid IDs. Adequate for a simple read tool but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with the action and target front-loaded. There is no filler or redundancy, though the brevity sacrifices useful operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and four parameters lacking schema descriptions, the single sentence is not enough to invoke the tool confidently. The behavior of context, ID selection semantics, and expected result format are all missing, so an agent would likely need to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by tying span_id, result_id, and proof_id to their conceptual resources, but it fails to explain the 'context' parameter, the relationship among the three ID parameters, or the expectation that exactly one should be selected. A significant parameter is left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('bounded source text around one span, result, or proof'), making the core purpose clear. It is slightly ambiguous what 'bounded' means and it does not differentiate from sibling getter tools, but the intent is understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool versus alternatives like workspace_get_result, workspace_get_proof_dependencies, or workspace_get_evidence. No exclusions, prerequisites, or alternative routing are mentioned, leaving usage entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_papersB
List all papers stored in the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only implies a read-only action via the verb 'list'. It does not disclose any potential side effects, default behaviors, or limitations beyond the basic listing operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It is concise and well-structured, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal but adequate for a simple list operation. However, it does not specify what information is returned (e.g., paper IDs, titles, metadata) or any inherent ordering or filtering, which could be useful for an agent to predict the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline score is 4. The description does not need to add parameter-specific meaning, and it correctly reflects that no arguments are required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (list) and the resource (papers) within the context of the active workspace. It is specific enough to distinguish from sibling tools that list other resources like theorems or results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternative list tools among the siblings. There is no mention of filters, sorting, or scenarios where other tools would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_list_resultsC
List stored evidence results across the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It conveys that this is a read-only listing scoped to the active workspace, but it does not mention whether an active workspace must already be open, how results are ordered or paginated, or any side effects. This is minimal disclosure for a list operation but lacks meaningful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler; it front-loads the verb and resource and adds the workspace scope. It is appropriately tight, although that tightness comes at the cost of missing contextual detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three optional but effectively undefined parameters and no annotations, this description is too thin. It does not explain any of the filters, the relationship between 'results' and 'evidence', or when an agent should prefer this over related sibling tools. The output schema mitigates return-value questions but not invocation decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention kind, limit, or paper_id. An agent cannot infer what values kind accepts, what paper_id filters by, or how limit behaves from the description alone. The schema provides only names, types, and defaults, so the description adds no parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), a distinctive resource ('stored evidence results'), and a scope ('across the active workspace'). This distinguishes it from sibling list tools like workspace_list_papers and from single-result accessors like workspace_get_result, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over workspace_get_result, workspace_get_evidence, or the other list-related tools. No mention is made of intended workflows, prerequisites, or exclusions. The only implied usage is 'list results,' which is not enough for an agent to select it confidently among many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_search_theoremsB
Search theorem titles and bodies across the active workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| query | Yes | ||
| paper_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the scope ('active workspace') but does not clarify search semantics such as exact vs. fuzzy matching, case sensitivity, whether filtering by kind or paper_id narrows scope, or whether the operation is read-only. The description is too thin for a search tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to the core purpose, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no annotations, and several closely related sibling tools, the description is under-specified. It does not explain how filters interact, what counts as a match, or when this tool should be chosen over list_theorems or get_theorem. The presence of an output schema reduces the need to describe return values, but significant operational context is still missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains the general search target ('theorem titles and bodies'). It does not clarify the meaning of 'kind', 'paper_id', or 'limit'. The query parameter's role is inferable from the description, but the filter and pagination parameters are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a clear verb ('Search'), a resource ('theorem titles and bodies'), and a scope ('across the active workspace'). This distinguishes it from sibling tools like list_theorems and get_theorem, which imply enumeration and single-item retrieval respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for query-based search across the active workspace, but it does not explicitly state when to prefer it over list_theorems, get_theorem, or where_used. No alternative tools or exclusions are mentioned, leaving usage conditions somewhat inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
31 tool updates
v0.6.1- First observed
get_dependencies - First observed
get_dependency_diagnostics - First observed
get_environment_diagnostics - First observed
get_theorem - First observed
list_theorems - First observed
load_arxiv_paper - First observed
load_arxiv_request - First observed
load_paper - First observed
open_workspace - First observed
validate_arxiv_input - First observed
validate_arxiv_request - First observed
where_used - First observed
workspace_add_arxiv_paper - First observed
workspace_add_local_paper - First observed
workspace_add_pdf_paper - First observed
workspace_export_reading_bundle - First observed
workspace_export_result_reading_context - First observed
workspace_get_citations - First observed
workspace_get_dependencies - First observed
workspace_get_dependency_diagnostics - First observed
workspace_get_evidence - First observed
workspace_get_external_result_mentions - First observed
workspace_get_paper - First observed
workspace_get_proof_dependencies - First observed
workspace_get_result - First observed
workspace_get_result_proof - First observed
workspace_get_result_reading_path - First observed
workspace_get_source_slice - First observed
workspace_list_papers - First observed
workspace_list_results - First observed
workspace_search_theorems
TDQS
Scored across 31 tools
Many tools have overlapping or near-identical functionality (e.g., workspace_add_local_paper vs load_paper, workspace_get_dependencies vs get_dependencies), making it difficult to choose the correct one.
Naming is inconsistent: some tools use the workspace_ prefix, others don't; verbs vary between add, load, list, get, and where_used, and there are duplicate concepts with different names.
31 tools is far beyond the typical well-scoped range, and the redundancy inflates the count unnecessarily, making the surface overwhelming and hard to navigate.
The tool set covers many reading and retrieval operations but lacks basic update/delete functionality, and the presence of duplicate operations suggests incomplete consolidation rather than comprehensive coverage.
Maintenance
Related MCP Connectors
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Research intelligence for AI coding agents. 2M+ CS papers with evidence and tradeoffs.
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Give your AI agent a persistent map of your project's structure, dependencies, and bugs.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables LLMs to search, download, and read arXiv papers with automatic PDF text extraction and section filtering. Provides AI assistants direct access to scientific literature with local caching for fast re-access.31MIT
- AlicenseAqualityDmaintenanceEnables LLM agents to search arXiv, download papers, parse PDFs into structured sections, and extract key findings using client-side LLMs. Features persistent caching and layout-aware PDF extraction.6MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to search, download, and read arXiv papers, with automatic detection of open-source code repositories and support for both LaTeX and PDF content.MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to read, write, and compile LaTeX projects locally, view PDF pages as images, and manage project files, with live updates reflected in a web-based editor.-