Skip to main content
Glama
sudeepan

mathematica-wstp

by sudeepan

DISCLAIMER: The code is LLM generated, so there is always room for improvement. So treat this as a functional implementation of an architecture that enables one to use LLM agents in a fully headless containerised environment, that work with long-duration stateful Wolfram kernels for complex research calculations that may even span days.

P.S. I know the README is a bit verbose, but please remember that this is meant more for the agent than for the user.

Mathematica MCP over WSTP

Auditable execution for agent-driven symbolic computation in Mathematica.

An agent can always ask a computer algebra system for an answer. The harder problem is to figure out what actually happened: which cell ran, in which kernel, whether an interrupt came from the user or the code itself, whether an output followed from this run or another one, and whether work survived when the client disappeared. Afterall, an agent's own report of what it did is not hard evidence for what actually transpired. Which has to do with the fact that all generative models can, and do lie.

This server is built around that problem.

Python 3.10+ Mathematica 14+ Transport: WSTP


The motivating workflow

Many research workflows are not centered on one notebook or one long interactive session. A more common pattern is hub-and-spoke work:

               ┌── computational session A
               │
main research ─┼── computational session B
session        │
               └── computational session C

The hub is the session where the larger research problem is being developed and integrated.

The spokes are short-term computational investigations. They do not need to communicate with one another directly. Each may have its own notebook, kernel state, assumptions, intermediate results, and failure history. Each spoke can be used for investigating specific questions that emerged naturally in the course of research in the hub, and would require a Mathematica-driven workflow to adequately address them.

A harness like Claude Code enables agentic message passing across sessions, and the user can also prompt in any session to read up the transcript of another one, and also look up the artifacts generated therein. One of these sessions could be the hub, and the others can play the role of spokes. The artifacts must be generated in a zero trust manner - they should not change depending on the choice of the harness, model, model-effort, etc.

A typical workflow:

  1. Work in the main research session.

  2. Reach a question that benefits from a separate symbolic calculation.

  3. Open a separate session with its own stateful kernel and notebook.

  4. Perform the calculation there, including exploratory attempts and corrections.

  5. Finalize the result into a durable notebook or artifact.

  6. Return to the main session.

  7. Read the auxiliary session's transcript and inspect its finalized computational record.

  8. Use the result in the larger line of reasoning.

Several such investigations may happen sequentially or concurrently.

Role of the MCP

In this model, the computational system is not primarily a graphical notebook frontend. The primary working surface may be an agent harness, terminal, editor, or research conversation. The notebook is still important, but as a durable scientific record, not necessarily as the place where all interaction happens.

That makes several properties more important than GUI integration:

  • a long-lived stateful kernel;

  • explicit ownership of that kernel;

  • the ability to interrupt runaway work without discarding useful state;

  • a durable identity for each scientific execution;

  • clear recovery after a client or session disappears;

  • notebook records that can be inspected later from another session;

  • reproducible finalization into a self-contained artifact;

  • provenance strong enough that one session can rely on results produced in another.

The goal is therefore not merely to let an agent execute code in a computer algebra system. It is closer to let independent computational sessions behave like small scientific laboratories whose results can later be inspected, independently verified, and incorporated into a larger research process.

A computational session can be thought of as a temporary laboratory:

temporary laboratory
    = agent session
    + stateful kernel
    + notebook
    + execution record
    + generated artifacts

Roles of the transcript and the notebook

The transcript and the notebook serve different purposes.

Transcript
    explains what the auxiliary session concluded
    and why it considered the result relevant.

Notebook / artifact
    records what was actually computed.

The transcript is the reasoning-level handoff. The notebook or artifact is the computational evidence. A robust workflow needs both.

This becomes especially important when one session consumes a result produced by another. The receiving session should not have to trust a sentence such as "the other session found that this identity holds." It should be possible to inspect the durable calculation behind that statement.

Related MCP server: jupyter-kernel-mcp

Advantages of a headless design

A GUI-centric workflow is excellent when the human and agent are collaborating inside one visible notebook. A headless MCP becomes attractive when the unit of work is instead:

research session
    ↕
computational session
    ↕
stateful kernel
    ↕
durable notebook / artifact

This makes it natural to run several independent investigations without forcing them into one interface or one global kernel state. Each computational session can remain isolated, while the main research session decides which conclusions and artifacts to bring back.

That separation also reduces accidental state coupling. Rather than sharing invisible kernel state between investigations, results move between sessions explicitly through notebooks, artifacts, transcripts, and recorded execution evidence. Even if sessions stop, the results computed while they were running must be preserved as demanded by the user.

The same architecture could also be useful for a single long-running session. The hub-and-spoke workflow explains why features like persistent kernels, recovery, audit trails, and durable notebook finalization matter, but it is not a requirement for using the server.

The execution model in one minute

There are three levels to keep apart.

  1. A persistent Wolfram kernel. Definitions and package state survive from one tool call to the next.

  2. A notebook replay record. replay runs executable cells one at a time, giving each one an identity and recording the plan in a durable manifest before execution begins.

  3. Optional supervisor ownership. The default server owns its kernel. With the supervisor selected, a separate process owns it, so a client can disappear while the computation continues.

A few terms appear throughout the documentation:

  • ordinal: the nth Input or Code cell, counting only executable cells. It is stable when output cells are inserted or deleted; raw notebook indices are not.

  • manifest: the replay-side record written to disk before the first child runs. It records the intended run and the progress of its cells.

  • execution token / request id: identities attached to one cell execution.

  • idempotency key: a stable key for an intended child execution, so a reconnecting client can refer to the existing work rather than accidentally creating a duplicate.

  • reconcile: inspect a previous replay after interruption and distinguish completed, still-running, never-submitted, changed-source, and output-state cases.

  • backend: the component that owns execution: either this process's direct kernel or the separate supervisor-owned kernel.

The distinction between ordinal and index matters immediately. Evaluating an input can insert an output cell, shifting every later index. "The seventh input cell" remains the seventh input cell.

What a replay buys you

For ordinary one-off evaluation, use evaluate. For a span where cell-by-cell identity does not matter, evaluate_cells is the simpler baseline. For long-running or auditable notebook work, use replay.

A replay provides:

  • one execution identity per executable cell;

  • a manifest persisted before the first cell is submitted;

  • per-child timeout policy and progress;

  • output-to-execution provenance;

  • detection when the source cell has changed underneath an existing run;

  • reconciliation after interruption instead of guessing where to resume.

With the supervisor selected, reconciliation also survives loss of the client that originally submitted the work. The original MCP call does not survive a client/session exit; the scientific computation can. A new client asks the ledger what happened by durable key and continues from there.

Why use WSTP

WSTP provides an evaluation channel and a separate out-of-band message channel. That lets the server interrupt a kernel that is currently busy without destroying the kernel merely to regain control.

That gives four practical properties:

  • Interrupt a running evaluation and preserve state. A timeout or abort() targets the evaluation rather than replacing the whole kernel.

  • Distinguish a dead link from a slow computation. Link failure becomes a typed failure rather than an indefinite wait.

  • Keep process ownership explicit. Subkernels are tracked and can be closed without throwing away the master kernel's definitions.

  • Observe control separately from scientific output. The execution layer can record that a user requested an abort even when Wolfram code catches that interrupt and returns a normal value.

The architecture and its two channels are described indocs/architecture.md.

Quick start

You need Mathematica 14 or newer (15 recommended) and uv. There is no compiler step and no Wolfram SDK to build; the transport binds to the WSTP library that ships with Mathematica.

git clone https://github.com/sudeepan/mathematica-mcp-wstp.git
cd mathematica-mcp-wstp
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e .
claude mcp add --scope user mathematica-wstp -- "$PWD/.venv/bin/mathematica-wstp"

Restart your client and try a small evaluation.

--scope user matters. claude mcp add defaults to --scope local, which registers the server for one directory only. Check from somewhere else (cd /tmp && claude mcp list) if a server seems to vanish outside the project.

The installation, kernel and WSTP library are discovered automatically, including relocated installs reachable only through a symlink on PATH. Override them with MATHEMATICA_WSTP_KERNEL, MATHEMATICA_WSTP_INSTALL, or MATHEMATICA_WSTP_LIB.

What it looks like in use

"Integrate that, and stop if it takes more than ten seconds."

evaluate("Integrate[Sqrt[1 + x^4], x]", timeout=10)
=> timed_out: true
   kernel_state: "intact, the evaluation was aborted rather than the kernel"

"Replay this notebook so I can resume if something interrupts us."

notebooks(action="open", path="analysis.nb")
replay(action="run", timeout=300)
=> run_id: ...
   executed: 276
   failed: 0
   execution_timeout_seconds: 300
   manifest: .../.mcp-replays/....json

The timeout is per executable cell. If one cell is expected to run for hours, set that value explicitly rather than inheriting the default.

"The client died while a long cell was running. What happened?"

With the supervisor selected:

replay(action="reconcile")
=> c1  COMPLETE
   c2  STILL_RUNNING
   c3  NEVER_SUBMITTED

STILL_RUNNING is not an error. It means the scientific evaluation is alive without the original client attached.

"Show me what cell 39 looks like."

render(action="cell", index=39)
=> [typeset PNG from a headless front end]

Choosing the right notebook path

Use replay when the work is long, interruptible, or needs an audit trail. Use evaluate_cells when you deliberately want one span-level operation and do not need durable per-cell identity.

replay addresses executable cells by ordinal, not raw notebook index. Writing outputs mutates the document, so indices move during the run.

Before replaying an unfamiliar notebook, run notebooks(action="dependencies") to discover which files it reads and writes and classify them automatically. The fresh-agent procedure in docs/replaying-a-notebook.md is designed for exactly this case.

Work that lasts hours or days

By default, this server owns its kernel. If the owning client/server process is killed, the kernel goes with it.

For work that must outlive the client, deliberately opt into the supervisor:

supervisor(action="start")
supervisor(action="use")

Then open the notebook and start the replay. A notebook belongs to the kernel that opened it, so choose the backend before opening the document.

The supervisor protects against client loss, not against a kernel or machine failure. Domain-level checkpoints are still necessary for work whose recomputation cost is measured in hours or days.

See docs/long-running-work.md.

Tools

Sixteen consolidated tools rather than a wide flat surface:

Tool

Purpose

evaluate

Run Wolfram Language in the persistent kernel

abort

Interrupt the running evaluation, keeping all state

kernel

state, restart, stop, abort, subkernels, close_subkernels, reap

status

Kernel, installation and tracked-process health

notebooks

open, create, list, info, save, close, dependencies, verify

cells

List or read cells of an open notebook

evaluate_cells

Replay a span of cells, state carrying between them

replay

Per-cell replay with identity and a resumable manifest

edit_cells

Insert, replace or delete a cell

render

Typeset an expression, rasterise a cell, export a notebook

vars

Inspect, set or clear the kernel's `Global`` symbols

supervisor

A kernel in its own process, outliving this one

batch

Run several tools in one round trip

verify_derivation

Check a chain of expressions step by step

read_notebook_file

Read a .nb without opening a session

guide

Usage notes by topic

render drives the Wolfram front end with -platform offscreen: no display and no X server are required. Rendering never owns scientific evaluation; that stays on the kernel link where interrupt and liveness semantics are defined.

For Wolfram Language documentation, use Wolfram's MCP

This server answers questions about your kernel, your state, and your documents. It is not a replacement for Wolfram's reference documentation.

Run Wolfram's own MCP server alongside it and use WolframLanguageContext for questions about built-ins, options, and language behavior. It uses a different kernel and cannot see definitions created here.

What the language means → Wolfram's MCP. What your scientific session contains → this server.

Documentation

Document

Start here when...

docs/agent-guide.md

You want to drive the server correctly and efficiently

docs/replaying-a-notebook.md

You are replaying a notebook or onboarding a fresh agent

docs/long-running-work.md

A cell may run for hours/days or must survive a dropped client

docs/architecture.md

You want the execution, replay and supervisor model

docs/pitfalls.md

You want the observed ways a plausible result can still be wrong

docs/benchmarks.md

You want measured costs and trade-offs

guide(topic=...) carries the short form inside the server: workflow · abort · errors · notebooks · state · parallel · performance.

Requirements

  • Mathematica 14 or newer, with the WSTP library that ships with it

  • Python 3.10+

  • Linux or macOS. Windows is untested.

  • One third-party Python dependency: mcp

Not for untrusted input or multi-tenant hosts. An evaluation is arbitrary code execution.

Tests

python3 tests/test_kernel.py
.venv/bin/python tests/test_server_mcp.py
.venv/bin/python tests/test_supervisor.py
python3 tests/test_recorder_foundation.py
python3 tests/test_recorder_core.py
python3 tests/test_recorder_annotations.py
python3 tests/test_recorder_finalize.py
python3 tests/test_recorder_adversarial.py
python3 tests/test_recorder_hardening.py
.venv/bin/python tests/test_recorder_server.py

The recorder tests use their own __main__ runner and do not require pytest. test_recorder_server.py calls the server's tool functions against a live kernel, so it needs the mcp package. Point MATHEMATICA_WSTP_TEST_NOTEBOOK at any .nb to exercise notebook tools against a real document; those checks are skipped when it is unset.

Available Tools

14 tools
abortA

Interrupt the evaluation the kernel is running right now. The kernel survives with all state intact. Use this instead of kernel(action='restart') for a runaway computation -- restart destroys every definition.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description takes on full behavioral disclosure. It clearly states the key effect—interrupting the running evaluation—and the crucial consequence that the kernel survives with all state intact. It does not cover what happens to the interrupted evaluation or whether abort is reversible, but the most important behavioral traits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each earning its place: the first states the action and its immediate outcome, the second says when to use it and contrasts with the destructive alternative. Information is front-loaded and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with an output schema and a clear sibling (kernel), the description covers all an agent needs to call it correctly: the action, the effect on kernel state, and when to prefer it over a destructive alternative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters)Skip; no parameter semantics are needed because the input schema is empty and schema description coverage is 100%. The description adds no parameter information, but this is not a gap for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Interrupt') and a clear resource ('the evaluation the kernel is running right now'), precisely identifying what the tool does. It also contrasts itself with kernel(action='restart'), making its role unambiguous among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to use this tool: for a runaway computation, instead of kernel(action='restart'). It also explains why this choice is better, because restart destroys every definition while abort preserves state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batchA

Run several of this server's tools in one round trip. ops: [{"tool": "evaluate", "args": {"code": "1+1"}}, ...]. Stops at the first failure unless stop_on_error is false. Useful for a fixed setup sequence; not a substitute for one compound Wolfram expression.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYes
stop_on_errorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It explains round-trip batching, the ops format, and the key failure behavior: "Stops at the first failure unless stop_on_error is false." This is meaningful behavioral context beyond what the schema alone provides, though it does not discuss side effects or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary purpose, then the parameter format, then the usage guidance. Every sentence adds value and no space is wasted on restating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low schema coverage and missing annotations, the description is surprisingly complete: purpose, usage context, parameter format, and failure behavior are all covered. The output schema exists, so explaining the return value is unnecessary. Minor gaps like edge cases around nested tool failures are not critical for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does. It explains ops with a concrete JSON example showing the tool and args structure, and it explains stop_on_error's semantics explicitly. An agent can construct a valid call using only the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: "Run several of this server's tools in one round trip." It specifies a concrete verb (run), a resource (this server's tools), and the batching behavior that distinguishes it from the sibling evaluation tools. It also adds a clarifying exclusion: "not a substitute for one compound Wolfram expression."

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context for when to use the tool: "Useful for a fixed setup sequence." It also provides a when-not-to-use signal by stating it is not a substitute for a compound Wolfram expression. It does not name a specific sibling alternative, but the guidance is enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cellsB

List or read cells of an open notebook. Use style to filter (e.g. 'Input'). Cell indices are positions in the document and are what evaluate_cells takes.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
styleNo
offsetNo
notebookNo
include_contentNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are attached, so the description carries the full burden of disclosure. It implies a read operation ('list or read') and discloses that indices map to evaluate_cells input, which is useful behavioral context. But it does not reveal pagination semantics implied by limit/offset defaults, what happens with notebook=null, or the output shape. Partial disclosure with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the purpose front-loaded and the sibling-routing hint placed at the end. No filler or redundant phrasing, though the guidance is thin relative to the tool's parameter count.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema covers return values, but the tool has five parameters with no schema docstrings and zero annotation coverage. The description omits notebook-selection semantics, pagination, and content flags, so a correct call is under-specified for an agent. Incomplete given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the five parameters, yet it only addresses style (with the 'Input' example) and introduces the index concept. limit, offset, include_content, and notebook remain unexplained, leaving an agent to guess at pagination, content-inclusion, and notebook-selection behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource pair ('List or read cells of an open notebook') and adds the clarifying fact that cell indices are positions in the document, linking to the sibling evaluate_cells. This distinguishes it from notebook-management siblings like notebooks and edit_cells, though the dual 'list or read' phrasing leaves scope slightly ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one routing signal — that the returned indices are what evaluate_cells consumes — which indirectly tells an agent this tool is the source for cell index data. However, it never explicitly states when to prefer this tool over siblings (e.g., edit_cells, render) or when not to use it, and there are no exclusions or preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_cellsC

Insert or delete a cell. actions: write(content,style,position) | delete(index).

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNo
styleNoInput
actionYes
anchorNo
contentNo
notebookNo
positionNoend

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the full burden of behavioral disclosure. It conveys that delete(index) is destructive and write takes content/style/position, but it does not explain consequences, reversibility, notebook context, or how position/anchor affect insertion. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with the operation front-loaded and no filler. Every token contributes to understanding the available actions and their parameter groupings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no annotations, this is incomplete. It does not explain the role of notebook or anchor, how write uses position, or when index is required. The output schema reduces the need to describe return values, but the input semantics remain under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully maps write to content/style/position and delete to index, but it omits notebook and anchor and does not clarify what position or anchor mean. The mapping helps but remains incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: 'Insert or delete a cell', and enumerates the two action variants. It is obvious that this tool mutates cells, though it does not explicitly contrast itself with sibling tools like cells or evaluate_cells.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool versus alternatives such as cells, evaluate_cells, or batch. There are no exclusions, preconditions, or examples to help an agent decide correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluateA

Evaluate Wolfram Language code in the persistent kernel. State carries between calls. On timeout the evaluation is ABORTED but the kernel and all its definitions survive, so you can retry a smaller piece.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It transparently discloses that state persists across calls and that a timeout aborts the evaluation but preserves kernel definitions. This is significant behavioral context that goes beyond a generic 'evaluate' tool. It does not cover error handling or result format, but the core stateful and timeout behaviors are clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core purpose is front-loaded ('Evaluate Wolfram Language code in the persistent kernel'), and the timeout behavior is added concisely. Every sentence contributes necessary operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (which presumably describes return values), the description covers the key operational aspects: persistent state and timeout behavior. It does not discuss error handling or how to interact with the 'abort' sibling, but for a basic evaluation tool this is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds meaning to 'code' by implying it is Wolfram Language code and that state from previous calls is available. For 'timeout', it explains the consequence of a timeout but does not specify units or range. This adds some value but does not fully cover the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates Wolfram Language code in a persistent kernel, giving a specific verb and resource. It does not explicitly differentiate from the sibling 'evaluate_cells', but the emphasis on 'persistent kernel' and the absence of cell context makes it distinct enough for an agent to infer the intended use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful guidance on statefulness and timeout behavior: state carries between calls, and on timeout the evaluation is aborted but the kernel survives, advising a retry with a smaller piece. However, it does not explicitly compare against alternatives like evaluate_cells or batch, leaving the agent to infer when to prefer this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_cellsA

Evaluate notebook cells in document order, in the persistent kernel, with state carrying between them. Give either index, or from_+to for a range. A long replay can be interrupted with abort(). Large ranges come back summarised (counts, failures, messages, slowest cells); set detail='full' to force per-cell output, or 'summary' to force the compact form.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNo
from_No
indexNo
detailNoauto
timeoutNo
notebookNo
stop_on_errorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: persistent kernel state, state carrying between cells, document-order execution, summarization of large ranges, and forced full/summary output modes. This gives an agent a reliable mental model of side effects and performance behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences cover core mechanics, range syntax, interruption, and output modes without filler. Key behavioral facts are front-loaded, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the most important behavioral and usage aspects, and an output schema exists so return-format details need not be repeated. Minor gaps remain for `timeout`, `stop_on_error`, and `notebook`, but defaults and names make these reasonably inferable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies `index` vs. `from_`+`to` and explains `detail` values well, but it leaves `timeout`, `stop_on_error`, and `notebook` semantics undocumented, forcing inference from names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool does: evaluate notebook cells in document order in the persistent kernel, with state carrying between cells. It also differentiates itself from simpler siblings by explaining index vs. range evaluation, making the scope unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit calling patterns: use either `index` or `from_`+`to` for a range, and use `detail='full'` or `'summary'` to control output verbosity. It also mentions `abort()` for interrupting long replays, but does not explicitly contrast with the `evaluate` or `batch` sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guideB

Usage notes for this server. topics: workflow | abort | errors | notebooks | performance.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoworkflow

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Usage notes for this server' indicates an informational, non-destructive tool, which is useful, but it does not explicitly state that no state is modified, what the returned notes contain, or any limiting behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely compact: one clear statement of purpose plus a brief topic list. Every element contributes, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a low-complexity tool with a single optional enum parameter and an output schema, so the description does not need to explain return values. The description plus schema are sufficient for an agent to request a valid topic and understand it is receiving usage guidance, though a bit more detail about each topic would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only repeats the enum values without explaining what each topic returns or how to decide which topic to request. The parameter's meaning remains thin: an agent knows topic names but not their content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Usage notes for this server' clearly identifies the tool as an instructions/help resource, and the topic list conveys its scope. It distinguishes itself from sibling tools by signaling that it provides guidance about topics like abort and notebooks rather than performing those operations, though it lacks an explicit verb such as 'retrieve' or 'return'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that this tool should be consulted when the agent needs usage guidance on the listed topics, while the sibling tools are the actual operations. However, it does not explicitly state when to use this tool instead of a sibling or provide guidance on choosing among the topics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kernelA

Kernel administration. actions: state | restart | abort | subkernels | reap. 'restart' clears ALL definitions and closes subkernels properly; prefer abort() for a merely slow evaluation.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNostate

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burdenches and meaningfully warns that restart 'clears ALL definitions and closes subkernels properly'. It also signals that abort() is a lighter-weight alternative for slow evaluations. It could disclose more about subkernels and reap, but the most destructive behavior is exposed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first establishes the tool's domain and action set, and the second delivers the critical caution about restart and the preferred alternative. No filler or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema may cover return values, but for an admin tool with five distinct actions, only restart is given meaningful semantic detail. 'state', 'subkernels', and especially 'reap' are jargon-heavy and remain unexplained, leaving an agent to guess at their effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides the enum values but no per-value descriptions, and schema description coverage is 0%. The description repeats the action list and explains restart, which adds some meaning beyond the bare enum, but state, subkernels, and reap are left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as kernel administration and lists the five supported actions, so an agent can see the tool's scope. It does not fully distinguish the kernel 'abort' action from the sibling 'abort' tool, though the preference hint points toward that relationship.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit selection guidance: 'restart' is a heavy operation that clears all definitions, while abort() is preferred for a merely slow evaluation. It does not provide when-to-use guidance for state, subkernels, or reap, but the action enum and kernel-administration framing carry some of that weight.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notebooksC

Notebook sessions over .nb files on disk. actions: open(path) | create(title,path) | list | info | save(path) | close. Cells are evaluated from their original stored boxes, so nothing is lost in translation.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
titleNoUntitled
actionNolist
notebookNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears the behavioral burden but only discloses one useful trait: cells are evaluated from original stored boxes. It doesn't disclose side effects of open/create/save, session statefulness, persistence, or failure behavior, which matter for a mutation-capable tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The three sentences are compact and front-loaded, first establishing the resource and then listing actions. The final sentence about evaluation fidelity earns its place by explaining a non-obvious behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even with an output schema present, this is a multiplexing tool with four parameters and no required fields, so the agent needs action-level parameter semantics and usage context. The description gives only a skim overview and leaves notebook, defaults, and state behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The action signatures map path and title to specific actions, adding meaning the schema doesn't convey, but schema coverage is 0% and the description omits the notebook parameter entirely. This partial guidance helps but leaves critical parameter relationships (especially around info/save and notebook) unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (.nb files on disk) and enumerates the supported actions (open/create/list/info/save/close), making the general purpose clear. It doesn't explicitly contrast with siblings like read_notebook_file or evaluate, so some inference is still needed to tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when this session manager should be used instead of sibling tools, nor when one action should be preferred over another. The action list implies usage contexts but leaves the routing decision entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_notebook_fileA

Read a .nb file from disk without opening a kernel session for it. modes: outline (headings only) | markdown | wolfram (code cells only) | plain | json. Use notebooks(action='open') instead when you intend to evaluate anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNooutline
pathYes
limitNo
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It usefully reveals that this works without opening a kernel session and defines what each mode returns (headings only, code cells only, etc.). It does not describe failure behavior or provide much detail beyond the mode list, but the core behavioral differentiation is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: a one-sentence core behavior, a mode list with clarifying parentheticals, and a routing note. Every element earns its place and the most important behavioral trait is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return format does not need to be described. The description covers the tool's central non-evaluation behavior.Paginaking semantics are the only notable gap, and path is inferable from the tool name and description. This is nearly complete for a read-only file inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does map the mode enum to meaningful behavior, which is valuable. However, path, limit, and offset are left semantically unexplained; limit and offset strongly suggest pagination but the description never says so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: "Read a .nb file from disk" and immediately adds the key scoping trait "without opening a kernel session for it." The mode list further clarifies what the tool produceschery. This clearly distinguishes it from the evaluate and notebooks siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when not to use this tool: "Use notebooks(action='open') instead when you intend to evaluate anything." This is an unambiguous routing signal that mentions the relevant alternative tool and the condition that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renderA

Render with the Wolfram front end, headlessly: typeset an expression, rasterise a notebook cell, or export a notebook. actions: expression(code) | cell(index) | export(path) | available. This RENDERS only -- it never evaluates through the front end; use evaluate() for that. Export renders what is visible, so collapsed cell groups export collapsed; pass open_groups=True for the whole document.

ParametersJSON Schema
NameRequiredDescriptionDefault
dpiNo
codeNo
pathNo
indexNo
actionNoexpression
notebookNo
open_groupsNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the headless behavior, the non-evaluation guarantee, and the export visibility behavior (collapsed groups export collapsed, open_groups=True changes that). It does not mention side effects like file creation or whether rendering is synchronous, but the core behavioral traits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the main purpose and action list come first, the critical non-evaluation warning second, and the export nuance last. Every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no output schema and no annotations, the description covers the action semantics, the key behavioral boundary (render vs. evaluate), and the open_groups nuance. It does not explain dpi or notebook, and there is no mention of return values, but the core invocation logic is sufficiently complete for an agent to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the action enum values ('expression(code) | cell(index) | export(path) | available') and the open_groups parameter's effect on export. It does not explain dpi or notebook, but the action mapping covers the most important parameters and the default action is implied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Render') with the Wolfram front end, headlessly, and enumerates the concrete actions: typeset an expression, rasterise a notebook cell, or export a notebook. It also explicitly contrasts itself with evaluate(), a sibling, so an agent can distinguish it from the most likely confusable tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'This RENDERS only -- it never evaluates through the front end; use evaluate() for that.' This is a clear when-to-use vs. when-not-to-use statement naming the alternative. It also gives a conditional usage hint: 'Export renders what is visible, so collapsed cell groups export collapsed; pass open_groups=True for the whole document.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusB

Server, kernel and installation status, plus any orphaned kernels.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. It lists what status covers but never explicitly states that the operation is read-only, has no side effects, or requires any prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence front-loads the core scope and adds a specific extra detail (orphaned kernels) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool with an output schema present, the description covers the report scope succinctly. It would be slightly stronger with an explicit read-only note or usage guidance, but nothing essential is missing given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the description has no parameter burden to carry. Schema coverage for the empty parameter set is effectively complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and scope: server, kernel, installation status, and orphaned kernels. It lacks an explicit verb, but 'status' plus these objects makes the purpose recognizable and distinct from sibling tools like kernel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to call status versus alternatives such as kernel or evaluate. The sibling list suggests it is a read-only monitoring tool, but the description never states when it should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

varsB

Inspect or change the kernel's Global` symbols. actions: list | get(name) | set(name,value) | clear(name) | clear_all. Use this to see what a notebook replay actually defined, or to clear one symbol without restarting.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
valueNo
actionNolist
patternNo
include_systemNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It mentions that actions include clear_all, implying it can clear all globals. However, it doesn't disclose side effects like persistence, whether changes are permanent, or if there are any side effects on the running kernel. The description says 'change' but doesn't elaborate on mutability beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the main purpose. The action list is useful and not verbose. It is structured as a single sentence, which is efficient, but could be slightly clearer with the parameter behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, 5 actions) and the lack of schema parameter descriptions, the description is incomplete. It doesn't explain the 'pattern' parameter or 'include_system' flag, which an agent might need to call correctly. However, an output schema exists, which may provide some guidance on return values, but not on parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the schema provides no descriptions for parameters. The description lists actions but doesn't explain what each parameter does. For example, 'pattern' is a parameter but not mentioned in the description, and its purpose is unclear. The description only hints at 'name' and 'value' but leaves out 'pattern' and 'include_system'. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to inspect or change kernel globals, listing specific actions. It is distinct from siblings like 'kernel' or 'status' by focusing on variable manipulation. However, it could be more specific about the resource (kernel's global symbols) being unique, as 'inspect or change' is a bit broad.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases: 'use this to see what a notebook replay actually defined, or to clear one symbol without restarting.' This explains when to use it, but doesn't explicitly mention when not to use it or compare with alternatives. However, given the simple nature, it's adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_derivationB

Check a chain of expressions step by step: each step must equal the one before it. Returns the first step that does not follow. steps are Wolfram expressions as strings, in order.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYes
timeoutNo
assumptionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses core behavior: it checks step-by-step equality and returns the first step that does not follow. It also states that steps are Wolfram expressions as strings. With no annotations available, it does not cover what happens when all steps pass, how equality is determined (e.g., evaluation vs structural comparison), or how 'assumptions' and 'timeout' affect behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences long, and front-loads the main purpose before adding the key parameter detail. There is slight redundancy between 'step by step' and 'in order', but no meaningful waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details are already covered. The description adequately explains the required 'steps' parameter and the core behavior, but because there are no annotations, the complete silence on 'timeout' and 'assumptions' leaves gaps for an agent trying to invoke the tool with non-default options.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the missing parameter documentation. It usefully explains 'steps' as Wolfram expression strings in order, but it says nothing about the 'timeout' or 'assumptions' parameters, leaving two of three parameters semantically unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation (verify a chain of expressions) and defines the key property (each step must equal the previous one). It returns the first non-following step, which gives a clear behavioral signature. However, it does not explicitly contrast with sibling tools such as 'evaluate', so it falls slightly short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Check a chain of expressions...' implies the intended use case: verifying a derivation. But there is no explicit when-to-use vs when-not-to-use guidance and no mention of alternative tools like 'evaluate' or 'batch'. Usage is inferable but not directly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 14 tool updatesv0.1.0
    • First observedabort
    • First observedbatch
    • First observedcells
    • First observededit_cells
    • First observedevaluate
    • First observedevaluate_cells
    • First observedguide
    • First observedkernel
    • First observednotebooks
    • First observedread_notebook_file
    • First observedrender
    • First observedstatus
    • First observedvars
    • First observedverify_derivation

TDQS

A3.6/5.0

Scored across 14 tools

Disambiguation4/5

Most tools map cleanly to different resources or actions: kernel evaluation, notebook/cell access, rendering, variable inspection, and batch orchestration. There are a couple of overlapping boundaries—abort vs kernel(action='abort') and status vs kernel(state)—but the descriptions generally clarify which one is intended.

Naming Consistency3/5

Names are readable and mostly action-oriented, but the convention is mixed: bare verbs (evaluate, abort, render), bare nouns (kernel, status, notebooks, cells, vars), and underscore compounds (evaluate_cells, read_notebook_file, verify_derivation). It is not chaotic, but there is no single predictable verb_noun pattern.

Tool Count5/5

Fourteen tools is within the well-scoped range and each tool addresses a genuinely distinct part of the server's purpose: persistent evaluation, kernel admin, notebook sessions, cell operations, rendering, symbol inspection, batching, and guidance. The count is substantial but not padded.

Completeness5/5

The surface covers the main lifecycle for both the persistent kernel (evaluate, abort, restart, vars) and notebooks (open/create/list/save/close, read cells, write/delete cells, evaluate cells), plus rendering and derivation verification. There are no obvious dead ends or missing core operations for this domain.

Related MCP Connectors

Related MCP Servers