mathematica-wstp
A headless MCP server for auditable, stateful Wolfram Language (Mathematica) computation: it runs code in a persistent kernel, drives .nb notebooks with per-cell identity, and can render or inspect them without a GUI.
Evaluate code in a persistent kernel where definitions survive between calls; on timeout the evaluation is aborted but kernel state stays intact, so you can retry a smaller piece.
Interrupt work with
abort()(orkernel(action='abort')) instead of destroying the kernel, and administer it:state,restart,subkernels,reap.Check health via
status: server, kernel, installation, and orphaned kernels.Manage notebook sessions over on-disk
.nbfiles:open,create,list,info,save,close.List or read cells, filtered by style (e.g.
Input) with pagination and optional content.Evaluate cell spans in document order with state carrying between them, using
indexorfrom_+to, per-run timeout,stop_on_error, and auto/full/summary detail for large ranges.Edit notebooks by inserting content at a position/anchor or deleting a cell.
Render headlessly (offscreen, no X server) — typeset an expression, rasterise a cell, export a notebook, or check renderer availability. Rendering never evaluates.
Inspect kernel state via
vars: list, get, set, clear, or clear all `Global`` symbols, optionally by pattern.Batch tools in one round trip for fixed setup sequences, stopping at the first failure unless disabled.
Read
.nbfiles from disk without a kernel session in outline, markdown, wolfram, plain, or json mode.Verify a derivation step by step, returning the first step that does not follow from the previous one, with optional assumptions.
Read usage guidance by topic: workflow, abort, errors, notebooks, performance.
Note: the README describes two further tools — replay (per-cell manifests, reconciliation, resumable audited runs) and supervisor (a kernel in its own process that outlives the client) — that do not appear in this schema, so those capabilities are not actually exposed here.
Allows AI agents to execute Wolfram Language code in a persistent Mathematica kernel, manage and replay notebooks, render typeset mathematics and graphics headlessly, abort runaway evaluations without losing state, and verify symbolic derivations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mathematica-wstpIntegrate that, and stop if it takes more than ten seconds."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DISCLAIMER: The code is LLM generated, so there is always room for improvement. So treat this as a functional implementation of an architecture that enables one to use LLM agents in a fully headless containerised environment, that work with long-duration stateful Wolfram kernels for complex research calculations that may even span days.
P.S. I know the README is a bit verbose, but please remember that this is meant more for the agent than for the user.
Mathematica MCP over WSTP
Auditable execution for agent-driven symbolic computation in Mathematica.
An agent can always ask a computer algebra system for an answer. The harder problem is to figure out what actually happened: which cell ran, in which kernel, whether an interrupt came from the user or the code itself, whether an output followed from this run or another one, and whether work survived when the client disappeared. Afterall, an agent's own report of what it did is not hard evidence for what actually transpired. Which has to do with the fact that all generative models can, and do lie.
This server is built around that problem.
The motivating workflow
Many research workflows are not centered on one notebook or one long interactive session. A more common pattern is hub-and-spoke work:
┌── computational session A
│
main research ─┼── computational session B
session │
└── computational session CThe hub is the session where the larger research problem is being developed and integrated.
The spokes are short-term computational investigations. They do not need to communicate with one another directly. Each may have its own notebook, kernel state, assumptions, intermediate results, and failure history. Each spoke can be used for investigating specific questions that emerged naturally in the course of research in the hub, and would require a Mathematica-driven workflow to adequately address them.
A harness like Claude Code enables agentic message passing across sessions, and the user can also prompt in any session to read up the transcript of another one, and also look up the artifacts generated therein. One of these sessions could be the hub, and the others can play the role of spokes. The artifacts must be generated in a zero trust manner - they should not change depending on the choice of the harness, model, model-effort, etc.
A typical workflow:
Work in the main research session.
Reach a question that benefits from a separate symbolic calculation.
Open a separate session with its own stateful kernel and notebook.
Perform the calculation there, including exploratory attempts and corrections.
Finalize the result into a durable notebook or artifact.
Return to the main session.
Read the auxiliary session's transcript and inspect its finalized computational record.
Use the result in the larger line of reasoning.
Several such investigations may happen sequentially or concurrently.
Role of the MCP
In this model, the computational system is not primarily a graphical notebook frontend. The primary working surface may be an agent harness, terminal, editor, or research conversation. The notebook is still important, but as a durable scientific record, not necessarily as the place where all interaction happens.
That makes several properties more important than GUI integration:
a long-lived stateful kernel;
explicit ownership of that kernel;
the ability to interrupt runaway work without discarding useful state;
a durable identity for each scientific execution;
clear recovery after a client or session disappears;
notebook records that can be inspected later from another session;
reproducible finalization into a self-contained artifact;
provenance strong enough that one session can rely on results produced in another.
The goal is therefore not merely to let an agent execute code in a computer algebra system. It is closer to let independent computational sessions behave like small scientific laboratories whose results can later be inspected, independently verified, and incorporated into a larger research process.
A computational session can be thought of as a temporary laboratory:
temporary laboratory
= agent session
+ stateful kernel
+ notebook
+ execution record
+ generated artifactsRoles of the transcript and the notebook
The transcript and the notebook serve different purposes.
Transcript
explains what the auxiliary session concluded
and why it considered the result relevant.
Notebook / artifact
records what was actually computed.The transcript is the reasoning-level handoff. The notebook or artifact is the computational evidence. A robust workflow needs both.
This becomes especially important when one session consumes a result produced by another. The receiving session should not have to trust a sentence such as "the other session found that this identity holds." It should be possible to inspect the durable calculation behind that statement.
Related MCP server: jupyter-kernel-mcp
Advantages of a headless design
A GUI-centric workflow is excellent when the human and agent are collaborating inside one visible notebook. A headless MCP becomes attractive when the unit of work is instead:
research session
↕
computational session
↕
stateful kernel
↕
durable notebook / artifactThis makes it natural to run several independent investigations without forcing them into one interface or one global kernel state. Each computational session can remain isolated, while the main research session decides which conclusions and artifacts to bring back.
That separation also reduces accidental state coupling. Rather than sharing invisible kernel state between investigations, results move between sessions explicitly through notebooks, artifacts, transcripts, and recorded execution evidence. Even if sessions stop, the results computed while they were running must be preserved as demanded by the user.
The same architecture could also be useful for a single long-running session. The hub-and-spoke workflow explains why features like persistent kernels, recovery, audit trails, and durable notebook finalization matter, but it is not a requirement for using the server.
The execution model in one minute
There are three levels to keep apart.
A persistent Wolfram kernel. Definitions and package state survive from one tool call to the next.
A notebook replay record.
replayruns executable cells one at a time, giving each one an identity and recording the plan in a durable manifest before execution begins.Optional supervisor ownership. The default server owns its kernel. With the supervisor selected, a separate process owns it, so a client can disappear while the computation continues.
A few terms appear throughout the documentation:
ordinal: the nth
InputorCodecell, counting only executable cells. It is stable when output cells are inserted or deleted; raw notebook indices are not.manifest: the replay-side record written to disk before the first child runs. It records the intended run and the progress of its cells.
execution token / request id: identities attached to one cell execution.
idempotency key: a stable key for an intended child execution, so a reconnecting client can refer to the existing work rather than accidentally creating a duplicate.
reconcile: inspect a previous replay after interruption and distinguish completed, still-running, never-submitted, changed-source, and output-state cases.
backend: the component that owns execution: either this process's direct kernel or the separate supervisor-owned kernel.
The distinction between ordinal and index matters immediately. Evaluating an input can insert an output cell, shifting every later index. "The seventh input cell" remains the seventh input cell.
What a replay buys you
For ordinary one-off evaluation, use evaluate. For a span where cell-by-cell identity does not matter, evaluate_cells is the simpler baseline. For long-running or auditable notebook work, use replay.
A replay provides:
one execution identity per executable cell;
a manifest persisted before the first cell is submitted;
per-child timeout policy and progress;
output-to-execution provenance;
detection when the source cell has changed underneath an existing run;
reconciliation after interruption instead of guessing where to resume.
With the supervisor selected, reconciliation also survives loss of the client that originally submitted the work. The original MCP call does not survive a client/session exit; the scientific computation can. A new client asks the ledger what happened by durable key and continues from there.
Why use WSTP
WSTP provides an evaluation channel and a separate out-of-band message channel. That lets the server interrupt a kernel that is currently busy without destroying the kernel merely to regain control.
That gives four practical properties:
Interrupt a running evaluation and preserve state. A timeout or
abort()targets the evaluation rather than replacing the whole kernel.Distinguish a dead link from a slow computation. Link failure becomes a typed failure rather than an indefinite wait.
Keep process ownership explicit. Subkernels are tracked and can be closed without throwing away the master kernel's definitions.
Observe control separately from scientific output. The execution layer can record that a user requested an abort even when Wolfram code catches that interrupt and returns a normal value.
The architecture and its two channels are described indocs/architecture.md.
Quick start
You need Mathematica 14 or newer (15 recommended) and uv. There is no compiler step and no Wolfram SDK to build; the transport binds to the WSTP library that ships with Mathematica.
git clone https://github.com/sudeepan/mathematica-mcp-wstp.git
cd mathematica-mcp-wstp
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e .
claude mcp add --scope user mathematica-wstp -- "$PWD/.venv/bin/mathematica-wstp"Restart your client and try a small evaluation.
--scope usermatters.claude mcp adddefaults to--scope local, which registers the server for one directory only. Check from somewhere else (cd /tmp && claude mcp list) if a server seems to vanish outside the project.
The installation, kernel and WSTP library are discovered automatically, including relocated installs reachable only through a symlink on PATH. Override them with MATHEMATICA_WSTP_KERNEL, MATHEMATICA_WSTP_INSTALL, or MATHEMATICA_WSTP_LIB.
What it looks like in use
"Integrate that, and stop if it takes more than ten seconds."
evaluate("Integrate[Sqrt[1 + x^4], x]", timeout=10)
=> timed_out: true
kernel_state: "intact, the evaluation was aborted rather than the kernel""Replay this notebook so I can resume if something interrupts us."
notebooks(action="open", path="analysis.nb")
replay(action="run", timeout=300)
=> run_id: ...
executed: 276
failed: 0
execution_timeout_seconds: 300
manifest: .../.mcp-replays/....jsonThe timeout is per executable cell. If one cell is expected to run for hours, set that value explicitly rather than inheriting the default.
"The client died while a long cell was running. What happened?"
With the supervisor selected:
replay(action="reconcile")
=> c1 COMPLETE
c2 STILL_RUNNING
c3 NEVER_SUBMITTEDSTILL_RUNNING is not an error. It means the scientific evaluation is alive without the original client attached.
"Show me what cell 39 looks like."
render(action="cell", index=39)
=> [typeset PNG from a headless front end]Choosing the right notebook path
Use replay when the work is long, interruptible, or needs an audit trail. Use evaluate_cells when you deliberately want one span-level operation and do not need durable per-cell identity.
replay addresses executable cells by ordinal, not raw notebook index. Writing outputs mutates the document, so indices move during the run.
Before replaying an unfamiliar notebook, run notebooks(action="dependencies") to discover which files it reads and writes
and classify them automatically. The fresh-agent procedure in docs/replaying-a-notebook.md is designed for exactly this case.
Work that lasts hours or days
By default, this server owns its kernel. If the owning client/server process is killed, the kernel goes with it.
For work that must outlive the client, deliberately opt into the supervisor:
supervisor(action="start")
supervisor(action="use")Then open the notebook and start the replay. A notebook belongs to the kernel that opened it, so choose the backend before opening the document.
The supervisor protects against client loss, not against a kernel or machine failure. Domain-level checkpoints are still necessary for work whose recomputation cost is measured in hours or days.
See docs/long-running-work.md.
Tools
Sixteen consolidated tools rather than a wide flat surface:
Tool | Purpose |
| Run Wolfram Language in the persistent kernel |
| Interrupt the running evaluation, keeping all state |
|
|
| Kernel, installation and tracked-process health |
|
|
| List or read cells of an open notebook |
| Replay a span of cells, state carrying between them |
| Per-cell replay with identity and a resumable manifest |
| Insert, replace or delete a cell |
| Typeset an expression, rasterise a cell, export a notebook |
| Inspect, set or clear the kernel's `Global`` symbols |
| A kernel in its own process, outliving this one |
| Run several tools in one round trip |
| Check a chain of expressions step by step |
| Read a |
| Usage notes by topic |
render drives the Wolfram front end with -platform offscreen: no display and no X server are required. Rendering never owns scientific evaluation; that stays on the kernel link where interrupt and liveness semantics are defined.
For Wolfram Language documentation, use Wolfram's MCP
This server answers questions about your kernel, your state, and your documents. It is not a replacement for Wolfram's reference documentation.
Run Wolfram's own MCP server alongside it and use WolframLanguageContext for questions about built-ins, options, and language behavior. It uses a different kernel and cannot see definitions created here.
What the language means → Wolfram's MCP. What your scientific session contains → this server.
Documentation
Document | Start here when... |
You want to drive the server correctly and efficiently | |
You are replaying a notebook or onboarding a fresh agent | |
A cell may run for hours/days or must survive a dropped client | |
You want the execution, replay and supervisor model | |
You want the observed ways a plausible result can still be wrong | |
You want measured costs and trade-offs |
guide(topic=...) carries the short form inside the server:
workflow · abort · errors · notebooks · state · parallel · performance.
Requirements
Mathematica 14 or newer, with the WSTP library that ships with it
Python 3.10+
Linux or macOS. Windows is untested.
One third-party Python dependency:
mcp
Not for untrusted input or multi-tenant hosts. An evaluation is arbitrary code execution.
Tests
python3 tests/test_kernel.py
.venv/bin/python tests/test_server_mcp.py
.venv/bin/python tests/test_supervisor.py
python3 tests/test_recorder_foundation.py
python3 tests/test_recorder_core.py
python3 tests/test_recorder_annotations.py
python3 tests/test_recorder_finalize.py
python3 tests/test_recorder_adversarial.py
python3 tests/test_recorder_hardening.py
.venv/bin/python tests/test_recorder_server.pyThe recorder tests use their own __main__ runner and do not require pytest. test_recorder_server.py calls the server's tool functions against a live kernel, so it needs the mcp package. Point MATHEMATICA_WSTP_TEST_NOTEBOOK at any .nb to exercise notebook tools against a real document; those checks are skipped when it is unset.
Available Tools
14 toolsabortA
Interrupt the evaluation the kernel is running right now. The kernel survives with all state intact. Use this instead of kernel(action='restart') for a runaway computation -- restart destroys every definition.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description takes on full behavioral disclosure. It clearly states the key effect—interrupting the running evaluation—and the crucial consequence that the kernel survives with all state intact. It does not cover what happens to the interrupted evaluation or whether abort is reversible, but the most important behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place: the first states the action and its immediate outcome, the second says when to use it and contrasts with the destructive alternative. Information is front-loaded and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with an output schema and a clear sibling (kernel), the description covers all an agent needs to call it correctly: the action, the effect on kernel state, and when to prefer it over a destructive alternative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters)Skip; no parameter semantics are needed because the input schema is empty and schema description coverage is 100%. The description adds no parameter information, but this is not a gap for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Interrupt') and a clear resource ('the evaluation the kernel is running right now'), precisely identifying what the tool does. It also contrasts itself with kernel(action='restart'), making its role unambiguous among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool: for a runaway computation, instead of kernel(action='restart'). It also explains why this choice is better, because restart destroys every definition while abort preserves state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batchA
Run several of this server's tools in one round trip. ops: [{"tool": "evaluate", "args": {"code": "1+1"}}, ...]. Stops at the first failure unless stop_on_error is false. Useful for a fixed setup sequence; not a substitute for one compound Wolfram expression.
| Name | Required | Description | Default |
|---|---|---|---|
| ops | Yes | ||
| stop_on_error | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It explains round-trip batching, the ops format, and the key failure behavior: "Stops at the first failure unless stop_on_error is false." This is meaningful behavioral context beyond what the schema alone provides, though it does not discuss side effects or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose, then the parameter format, then the usage guidance. Every sentence adds value and no space is wasted on restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low schema coverage and missing annotations, the description is surprisingly complete: purpose, usage context, parameter format, and failure behavior are all covered. The output schema exists, so explaining the return value is unnecessary. Minor gaps like edge cases around nested tool failures are not critical for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It explains ops with a concrete JSON example showing the tool and args structure, and it explains stop_on_error's semantics explicitly. An agent can construct a valid call using only the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: "Run several of this server's tools in one round trip." It specifies a concrete verb (run), a resource (this server's tools), and the batching behavior that distinguishes it from the sibling evaluation tools. It also adds a clarifying exclusion: "not a substitute for one compound Wolfram expression."
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context for when to use the tool: "Useful for a fixed setup sequence." It also provides a when-not-to-use signal by stating it is not a substitute for a compound Wolfram expression. It does not name a specific sibling alternative, but the guidance is enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cellsB
List or read cells of an open notebook. Use style to filter (e.g. 'Input'). Cell indices are positions in the document and are what evaluate_cells takes.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| style | No | ||
| offset | No | ||
| notebook | No | ||
| include_content | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are attached, so the description carries the full burden of disclosure. It implies a read operation ('list or read') and discloses that indices map to evaluate_cells input, which is useful behavioral context. But it does not reveal pagination semantics implied by limit/offset defaults, what happens with notebook=null, or the output shape. Partial disclosure with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the purpose front-loaded and the sibling-routing hint placed at the end. No filler or redundant phrasing, though the guidance is thin relative to the tool's parameter count.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema covers return values, but the tool has five parameters with no schema docstrings and zero annotation coverage. The description omits notebook-selection semantics, pagination, and content flags, so a correct call is under-specified for an agent. Incomplete given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the five parameters, yet it only addresses style (with the 'Input' example) and introduces the index concept. limit, offset, include_content, and notebook remain unexplained, leaving an agent to guess at pagination, content-inclusion, and notebook-selection behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-resource pair ('List or read cells of an open notebook') and adds the clarifying fact that cell indices are positions in the document, linking to the sibling evaluate_cells. This distinguishes it from notebook-management siblings like notebooks and edit_cells, though the dual 'list or read' phrasing leaves scope slightly ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one routing signal — that the returned indices are what evaluate_cells consumes — which indirectly tells an agent this tool is the source for cell index data. However, it never explicitly states when to prefer this tool over siblings (e.g., edit_cells, render) or when not to use it, and there are no exclusions or preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_cellsC
Insert or delete a cell. actions: write(content,style,position) | delete(index).
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | ||
| style | No | Input | |
| action | Yes | ||
| anchor | No | ||
| content | No | ||
| notebook | No | ||
| position | No | end |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the full burden of behavioral disclosure. It conveys that delete(index) is destructive and write takes content/style/position, but it does not explain consequences, reversibility, notebook context, or how position/anchor affect insertion. For a mutation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with the operation front-loaded and no filler. Every token contributes to understanding the available actions and their parameter groupings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no annotations, this is incomplete. It does not explain the role of notebook or anchor, how write uses position, or when index is required. The output schema reduces the need to describe return values, but the input semantics remain under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully maps write to content/style/position and delete to index, but it omits notebook and anchor and does not clarify what position or anchor mean. The mapping helps but remains incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource: 'Insert or delete a cell', and enumerates the two action variants. It is obvious that this tool mutates cells, though it does not explicitly contrast itself with sibling tools like cells or evaluate_cells.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus alternatives such as cells, evaluate_cells, or batch. There are no exclusions, preconditions, or examples to help an agent decide correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluateA
Evaluate Wolfram Language code in the persistent kernel. State carries between calls. On timeout the evaluation is ABORTED but the kernel and all its definitions survive, so you can retry a smaller piece.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| timeout | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It transparently discloses that state persists across calls and that a timeout aborts the evaluation but preserves kernel definitions. This is significant behavioral context that goes beyond a generic 'evaluate' tool. It does not cover error handling or result format, but the core stateful and timeout behaviors are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core purpose is front-loaded ('Evaluate Wolfram Language code in the persistent kernel'), and the timeout behavior is added concisely. Every sentence contributes necessary operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (which presumably describes return values), the description covers the key operational aspects: persistent state and timeout behavior. It does not discuss error handling or how to interact with the 'abort' sibling, but for a basic evaluation tool this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds meaning to 'code' by implying it is Wolfram Language code and that state from previous calls is available. For 'timeout', it explains the consequence of a timeout but does not specify units or range. This adds some value but does not fully cover the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool evaluates Wolfram Language code in a persistent kernel, giving a specific verb and resource. It does not explicitly differentiate from the sibling 'evaluate_cells', but the emphasis on 'persistent kernel' and the absence of cell context makes it distinct enough for an agent to infer the intended use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful guidance on statefulness and timeout behavior: state carries between calls, and on timeout the evaluation is aborted but the kernel survives, advising a retry with a smaller piece. However, it does not explicitly compare against alternatives like evaluate_cells or batch, leaving the agent to infer when to prefer this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_cellsA
Evaluate notebook cells in document order, in the persistent kernel, with state carrying between them. Give either index, or from_+to for a range. A long replay can be interrupted with abort(). Large ranges come back summarised (counts, failures, messages, slowest cells); set detail='full' to force per-cell output, or 'summary' to force the compact form.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from_ | No | ||
| index | No | ||
| detail | No | auto | |
| timeout | No | ||
| notebook | No | ||
| stop_on_error | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: persistent kernel state, state carrying between cells, document-order execution, summarization of large ranges, and forced full/summary output modes. This gives an agent a reliable mental model of side effects and performance behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences cover core mechanics, range syntax, interruption, and output modes without filler. Key behavioral facts are front-loaded, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the most important behavioral and usage aspects, and an output schema exists so return-format details need not be repeated. Minor gaps remain for `timeout`, `stop_on_error`, and `notebook`, but defaults and names make these reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies `index` vs. `from_`+`to` and explains `detail` values well, but it leaves `timeout`, `stop_on_error`, and `notebook` semantics undocumented, forcing inference from names and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool does: evaluate notebook cells in document order in the persistent kernel, with state carrying between cells. It also differentiates itself from simpler siblings by explaining index vs. range evaluation, making the scope unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit calling patterns: use either `index` or `from_`+`to` for a range, and use `detail='full'` or `'summary'` to control output verbosity. It also mentions `abort()` for interrupting long replays, but does not explicitly contrast with the `evaluate` or `batch` sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
guideB
Usage notes for this server. topics: workflow | abort | errors | notebooks | performance.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | workflow |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Usage notes for this server' indicates an informational, non-destructive tool, which is useful, but it does not explicitly state that no state is modified, what the returned notes contain, or any limiting behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely compact: one clear statement of purpose plus a brief topic list. Every element contributes, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a low-complexity tool with a single optional enum parameter and an output schema, so the description does not need to explain return values. The description plus schema are sufficient for an agent to request a valid topic and understand it is receiving usage guidance, though a bit more detail about each topic would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only repeats the enum values without explaining what each topic returns or how to decide which topic to request. The parameter's meaning remains thin: an agent knows topic names but not their content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Usage notes for this server' clearly identifies the tool as an instructions/help resource, and the topic list conveys its scope. It distinguishes itself from sibling tools by signaling that it provides guidance about topics like abort and notebooks rather than performing those operations, though it lacks an explicit verb such as 'retrieve' or 'return'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that this tool should be consulted when the agent needs usage guidance on the listed topics, while the sibling tools are the actual operations. However, it does not explicitly state when to use this tool instead of a sibling or provide guidance on choosing among the topics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kernelA
Kernel administration. actions: state | restart | abort | subkernels | reap. 'restart' clears ALL definitions and closes subkernels properly; prefer abort() for a merely slow evaluation.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | state |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burdenches and meaningfully warns that restart 'clears ALL definitions and closes subkernels properly'. It also signals that abort() is a lighter-weight alternative for slow evaluations. It could disclose more about subkernels and reap, but the most destructive behavior is exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first establishes the tool's domain and action set, and the second delivers the critical caution about restart and the preferred alternative. No filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema may cover return values, but for an admin tool with five distinct actions, only restart is given meaningful semantic detail. 'state', 'subkernels', and especially 'reap' are jargon-heavy and remain unexplained, leaving an agent to guess at their effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides the enum values but no per-value descriptions, and schema description coverage is 0%. The description repeats the action list and explains restart, which adds some meaning beyond the bare enum, but state, subkernels, and reap are left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource as kernel administration and lists the five supported actions, so an agent can see the tool's scope. It does not fully distinguish the kernel 'abort' action from the sibling 'abort' tool, though the preference hint points toward that relationship.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit selection guidance: 'restart' is a heavy operation that clears all definitions, while abort() is preferred for a merely slow evaluation. It does not provide when-to-use guidance for state, subkernels, or reap, but the action enum and kernel-administration framing carry some of that weight.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notebooksC
Notebook sessions over .nb files on disk. actions: open(path) | create(title,path) | list | info | save(path) | close. Cells are evaluated from their original stored boxes, so nothing is lost in translation.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| title | No | Untitled | |
| action | No | list | |
| notebook | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the behavioral burden but only discloses one useful trait: cells are evaluated from original stored boxes. It doesn't disclose side effects of open/create/save, session statefulness, persistence, or failure behavior, which matter for a mutation-capable tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The three sentences are compact and front-loaded, first establishing the resource and then listing actions. The final sentence about evaluation fidelity earns its place by explaining a non-obvious behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with an output schema present, this is a multiplexing tool with four parameters and no required fields, so the agent needs action-level parameter semantics and usage context. The description gives only a skim overview and leaves notebook, defaults, and state behavior unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The action signatures map path and title to specific actions, adding meaning the schema doesn't convey, but schema coverage is 0% and the description omits the notebook parameter entirely. This partial guidance helps but leaves critical parameter relationships (especially around info/save and notebook) unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (.nb files on disk) and enumerates the supported actions (open/create/list/info/save/close), making the general purpose clear. It doesn't explicitly contrast with siblings like read_notebook_file or evaluate, so some inference is still needed to tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when this session manager should be used instead of sibling tools, nor when one action should be preferred over another. The action list implies usage contexts but leaves the routing decision entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_notebook_fileA
Read a .nb file from disk without opening a kernel session for it. modes: outline (headings only) | markdown | wolfram (code cells only) | plain | json. Use notebooks(action='open') instead when you intend to evaluate anything.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | outline | |
| path | Yes | ||
| limit | No | ||
| offset | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It usefully reveals that this works without opening a kernel session and defines what each mode returns (headings only, code cells only, etc.). It does not describe failure behavior or provide much detail beyond the mode list, but the core behavioral differentiation is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: a one-sentence core behavior, a mode list with clarifying parentheticals, and a routing note. Every element earns its place and the most important behavioral trait is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return format does not need to be described. The description covers the tool's central non-evaluation behavior.Paginaking semantics are the only notable gap, and path is inferable from the tool name and description. This is nearly complete for a read-only file inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does map the mode enum to meaningful behavior, which is valuable. However, path, limit, and offset are left semantically unexplained; limit and offset strongly suggest pagination but the description never says so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Read a .nb file from disk" and immediately adds the key scoping trait "without opening a kernel session for it." The mode list further clarifies what the tool produceschery. This clearly distinguishes it from the evaluate and notebooks siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when not to use this tool: "Use notebooks(action='open') instead when you intend to evaluate anything." This is an unambiguous routing signal that mentions the relevant alternative tool and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
renderA
Render with the Wolfram front end, headlessly: typeset an expression, rasterise a notebook cell, or export a notebook. actions: expression(code) | cell(index) | export(path) | available. This RENDERS only -- it never evaluates through the front end; use evaluate() for that. Export renders what is visible, so collapsed cell groups export collapsed; pass open_groups=True for the whole document.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| code | No | ||
| path | No | ||
| index | No | ||
| action | No | expression | |
| notebook | No | ||
| open_groups | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the headless behavior, the non-evaluation guarantee, and the export visibility behavior (collapsed groups export collapsed, open_groups=True changes that). It does not mention side effects like file creation or whether rendering is synchronous, but the core behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the main purpose and action list come first, the critical non-evaluation warning second, and the export nuance last. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema and no annotations, the description covers the action semantics, the key behavioral boundary (render vs. evaluate), and the open_groups nuance. It does not explain dpi or notebook, and there is no mention of return values, but the core invocation logic is sufficiently complete for an agent to select and call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the action enum values ('expression(code) | cell(index) | export(path) | available') and the open_groups parameter's effect on export. It does not explain dpi or notebook, but the action mapping covers the most important parameters and the default action is implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Render') with the Wolfram front end, headlessly, and enumerates the concrete actions: typeset an expression, rasterise a notebook cell, or export a notebook. It also explicitly contrasts itself with evaluate(), a sibling, so an agent can distinguish it from the most likely confusable tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'This RENDERS only -- it never evaluates through the front end; use evaluate() for that.' This is a clear when-to-use vs. when-not-to-use statement naming the alternative. It also gives a conditional usage hint: 'Export renders what is visible, so collapsed cell groups export collapsed; pass open_groups=True for the whole document.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusB
Server, kernel and installation status, plus any orphaned kernels.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral disclosure. It lists what status covers but never explicitly states that the operation is read-only, has no side effects, or requires any prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence front-loads the core scope and adds a specific extra detail (orphaned kernels) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with an output schema present, the description covers the report scope succinctly. It would be slightly stronger with an explicit read-only note or usage guidance, but nothing essential is missing given the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the description has no parameter burden to carry. Schema coverage for the empty parameter set is effectively complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource and scope: server, kernel, installation status, and orphaned kernels. It lacks an explicit verb, but 'status' plus these objects makes the purpose recognizable and distinct from sibling tools like kernel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to call status versus alternatives such as kernel or evaluate. The sibling list suggests it is a read-only monitoring tool, but the description never states when it should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
varsB
Inspect or change the kernel's Global` symbols. actions: list | get(name) | set(name,value) | clear(name) | clear_all. Use this to see what a notebook replay actually defined, or to clear one symbol without restarting.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| value | No | ||
| action | No | list | |
| pattern | No | ||
| include_system | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It mentions that actions include clear_all, implying it can clear all globals. However, it doesn't disclose side effects like persistence, whether changes are permanent, or if there are any side effects on the running kernel. The description says 'change' but doesn't elaborate on mutability beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main purpose. The action list is useful and not verbose. It is structured as a single sentence, which is efficient, but could be slightly clearer with the parameter behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 parameters, 5 actions) and the lack of schema parameter descriptions, the description is incomplete. It doesn't explain the 'pattern' parameter or 'include_system' flag, which an agent might need to call correctly. However, an output schema exists, which may provide some guidance on return values, but not on parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no descriptions for parameters. The description lists actions but doesn't explain what each parameter does. For example, 'pattern' is a parameter but not mentioned in the description, and its purpose is unclear. The description only hints at 'name' and 'value' but leaves out 'pattern' and 'include_system'. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to inspect or change kernel globals, listing specific actions. It is distinct from siblings like 'kernel' or 'status' by focusing on variable manipulation. However, it could be more specific about the resource (kernel's global symbols) being unique, as 'inspect or change' is a bit broad.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit use cases: 'use this to see what a notebook replay actually defined, or to clear one symbol without restarting.' This explains when to use it, but doesn't explicitly mention when not to use it or compare with alternatives. However, given the simple nature, it's adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_derivationB
Check a chain of expressions step by step: each step must equal the one before it. Returns the first step that does not follow. steps are Wolfram expressions as strings, in order.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | ||
| timeout | No | ||
| assumptions | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses core behavior: it checks step-by-step equality and returns the first step that does not follow. It also states that steps are Wolfram expressions as strings. With no annotations available, it does not cover what happens when all steps pass, how equality is determined (e.g., evaluation vs structural comparison), or how 'assumptions' and 'timeout' affect behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, two sentences long, and front-loads the main purpose before adding the key parameter detail. There is slight redundancy between 'step by step' and 'in order', but no meaningful waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value details are already covered. The description adequately explains the required 'steps' parameter and the core behavior, but because there are no annotations, the complete silence on 'timeout' and 'assumptions' leaves gaps for an agent trying to invoke the tool with non-default options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the missing parameter documentation. It usefully explains 'steps' as Wolfram expression strings in order, but it says nothing about the 'timeout' or 'assumptions' parameters, leaving two of three parameters semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation (verify a chain of expressions) and defines the key property (each step must equal the previous one). It returns the first non-following step, which gives a clear behavioral signature. However, it does not explicitly contrast with sibling tools such as 'evaluate', so it falls slightly short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Check a chain of expressions...' implies the intended use case: verifying a derivation. But there is no explicit when-to-use vs when-not-to-use guidance and no mention of alternative tools like 'evaluate' or 'batch'. Usage is inferable but not directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.1.0- First observed
abort - First observed
batch - First observed
cells - First observed
edit_cells - First observed
evaluate - First observed
evaluate_cells - First observed
guide - First observed
kernel - First observed
notebooks - First observed
read_notebook_file - First observed
render - First observed
status - First observed
vars - First observed
verify_derivation
TDQS
Scored across 14 tools
Most tools map cleanly to different resources or actions: kernel evaluation, notebook/cell access, rendering, variable inspection, and batch orchestration. There are a couple of overlapping boundaries—abort vs kernel(action='abort') and status vs kernel(state)—but the descriptions generally clarify which one is intended.
Names are readable and mostly action-oriented, but the convention is mixed: bare verbs (evaluate, abort, render), bare nouns (kernel, status, notebooks, cells, vars), and underscore compounds (evaluate_cells, read_notebook_file, verify_derivation). It is not chaotic, but there is no single predictable verb_noun pattern.
Fourteen tools is within the well-scoped range and each tool addresses a genuinely distinct part of the server's purpose: persistent evaluation, kernel admin, notebook sessions, cell operations, rendering, symbol inspection, batching, and guidance. The count is substantial but not padded.
The surface covers the main lifecycle for both the persistent kernel (evaluate, abort, restart, vars) and notebooks (open/create/list/save/close, read cells, write/delete cells, evaluate cells), plus rendering and derivation verification. There are no obvious dead ends or missing core operations for this domain.
Related MCP Connectors
- mcp-serverOAuthai.cdbx
Build Apps and run code in 30 languages — sandboxed, with persistent sessions for agent loops.
Hosted runtime for persistent agent teams, durable workflows, memory, schedules, and goals.
Build, validate, and deploy multi-agent AI solutions from any AI environment.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Related MCP Servers
- FlicenseAqualityDmaintenanceAllows LLMs to execute Wolfram Language code in a secure, session-based environment by providing an interface to interact with a Wolfram Mathematica kernel.33-
- AlicenseAqualityDmaintenanceEnables AI agents to execute Python, TypeScript, and JavaScript code in persistent Jupyter kernels with stateful variables and imports across interactions.74MIT
- AlicenseCqualityBmaintenanceEnables AI agents to run Mathematica code, control live notebooks, and verify results through natural language.48238 PyPI50MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to execute Jupyter notebook cells with persistent kernel state, output persistence, and structured JSON control surface.2-