Skip to main content
Glama
johnhenry

mallory-grapher

by johnhenry

math-grapher

npm version CI license

Full documentation: opensource.johnhenry.me/math/math-grapher

A headless, DOM-less session runtime for the @johnhenry/math family's reactive compute graph (CellGraph), agent-drivable over MCP.

Usage

# stdio (the transport MCP hosts speak natively -- e.g. `claude mcp add`)
npx @johnhenry/math-grapher

# Streamable HTTP on http://localhost:3920/mcp (or a custom port)
npx @johnhenry/math-grapher --http
npx @johnhenry/math-grapher --http 8123

Tools: session_open (kind generic or graph-theory), session_close, session_list, session_set_cell, session_get_cell, session_list_cells, session_explain_cell (a cell's own op/args/ immediate dependencies with their current values, one level -- issue #5), session_snapshot/session_resume (serialize a session's free-cell values + define-specs; reconstruct an equivalent session, possibly on a different process -- issue #6), session_define. Computed cells are declared as JSON define-specs over a server-side op catalog (math_eval, graph_parse_edge_list, graph_analyze, graph_bfs/dfs/dijkstra) with {"$cell": "name"} live references — see docs/design.md §5. An op MAY declare a requiresCapability (issue #7); session_define rejects it unless session_open/session_resume's own optional capabilities arg granted it for that session (default none — matching the existing write-path gating precedent below, just per-op instead of one global switch). No op in the catalog above declares one yet.

Resource guards default modest and are overridable: MATH_GRAPHER_MAX_SESSIONS (16), MATH_GRAPHER_MAX_CELLS (512), MATH_GRAPHER_EVAL_BUDGET_MS (250), MATH_GRAPHER_MAX_PAYLOAD_BYTES (262144).

Related MCP server: chuk-mcp-runtime

Platform

The . entry (SessionTable, buildServer, CellGraph, the op catalog) is Node- and browser-safe — no node:* import in its graph, verified by bundling it with esbuild under platform: 'browser' (see test/browser-bundle.test.ts). randomUUID uses globalThis.crypto (Node ≥19, every modern browser) rather than node:crypto. The bin (npx @johnhenry/math-grapher, src/cli.ts) is Node-only — it owns the --http transport (node:http) — and is never imported by .; embedding this runtime in a browser page or bundler build means importing . directly (e.g. via an in-page MCP Client/InMemoryTransport pair, ORRERY's own pattern), not the CLI.

Status

v1 implemented. docs/design.md is the settled design; the session runtime, op catalog, MCP tool surface, and both transports are built and tested.

Why this exists

mallory's (the graphing-calculator app, formerly mallory-graph) in-page WebMCP tools (useCellGraphTools: ${prefix}_list_cells/get_cell/set_cell) already let an agent drive a live, reactive CellGraph — but only from inside a rendered browser tab. The server-side MCP endpoint mallory ships today (@johnhenry/math-plus-mcp, math-plus's packages/mcp) only covers stateless math tools (Symbolic eval, guarded tensor/linalg) plus read-only, serialized gallery access (gallery_list/gallery_get read NotebookState.blocks[] JSON — no computed/derived-cell evaluation, no reactivity).

Real session parity — "an agent could run an entire modeling session headlessly" — means running the reactive compute graph itself server-side, with no DOM/React tree at all. That's a materially different, bigger project than either of the above, so it lives here instead of inside mallory.

Split out of mallory#163 after that issue's own audit trail: a feasibility spike (cell-graph-headless-spike.test.ts) already confirmed CellGraph itself has zero window/document references — set/define/get/ subscribe/subscribeAll are plain data-structure + closure code. What doesn't exist yet is the actual session API around it.

Relationship to mallory

Optional, not coupled. math-grapher does not depend on mallory (the app), and mallory does not need to depend on math-grapher to function. mallory may choose to mount a math-grapher-backed MCP route the same way it mounts @johnhenry/math-plus-mcp today (src/routes/api.mcp.ts) — a separate integration issue, not a prerequisite for this repo to exist or ship v1.

Concretely, this repo owns:

  • The headless session runtime (open a session, drive its cells, read results) — no rendering, no React, no DOM.

  • An MCP tool surface over that runtime.

It deliberately does NOT own:

  • Canvas/WebGL rendering — session parity is about the same get/set/list contract WebMCP already gives an in-page agent, not pixels.

  • mallory's specific panel components, gallery storage, or UI.

What's known so far (carried over from #163's audit)

  • CellGraph's core (cell-graph.ts in mallory, ~500 lines) has no structural blocker to running headless — proven empirically, not just asserted.

  • Most panels' useXGraph() seed step reads window.location.hash / getComputedStyle for URL-state hydration and theming. Both are already guarded with typeof window !== "undefined" checks (existing SSR-safety code), so they degrade gracefully rather than crash — but a real agent-drivable session needs a way to seed state from something other than a browser's URL bar (an MCP tool argument, presumably).

  • Write-path auth precedent: mallory's gallery_save tool (#163 item 1, shipped) is gated OFF by default behind an explicit env var (MALLORY_GRAPH_ENABLE_MCP_WRITE=1), mirroring llmtm's LLMTM_HUB_ENABLE_* convention. A session-runtime write surface should follow the same default-off, explicit-opt-in posture.

Available Tools

10 tools
session_closeB

Close a session and free its cells.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does communicate a main effect—freeing cells—but it does not disclose whether the session must be active, whether this is destructive/reversible, or what happens to unsaved state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one short, front-loaded sentence. It includes the core action and a meaningful side effect with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple one-parameter, no-output-schema tool, so a brief description can be acceptable. However, the lack of annotations and lack of state/error/return context leave only a minimal viable level of completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, sessionId, has no description coverage, and the description does not explain it or mention how the ID should be obtained. The parameter name is self-explanatory enough for basic inference, but the description adds no value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and object: 'Close a session', and adds meaningful scope with 'free its cells.' This clearly distinguishes it from siblings like session_open, session_list, and session_set_cell.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool, when not to use it, or what alternatives exist. It does not mention whether it should be used only after a session is no longer needed, nor does it reference sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_defineA

Define a computed cell from a catalog op. args values may be literal JSON or live cell references ({"$cell": "name"}) -- referenced cells become reactive dependencies, so the cell recomputes when they change. Available ops:

  • math_eval: Evaluate an @johnhenry/math Symbolic expression string over named numeric variables. args: { expr: string, vars?: { name: number | {"$cell": ...} } }. value: number.

  • graph_parse_edge_list: Parse a from to [weight]-per-line edge list into a graph value. args: { text: string, directed?: boolean (default true) }. value: graph (opaque; project with session_get_cell).

  • graph_analyze: Structural analysis of a graph cell. args: { graph: graph }. value: { hasCycle, connectedComponents, stronglyConnectedComponents, topologicalOrder, adjacencyMatrix: { matrix, order } }.

  • graph_bfs: Breadth-first traversal order. args: { graph: graph, start: string }. value: string[].

  • graph_dfs: Depth-first traversal order. args: { graph: graph, start: string }. value: string[].

  • graph_dijkstra: Dijkstra shortest-path distances from a start vertex. args: { graph: graph, start: string }. value: [{ vertex, distance }].

ParametersJSON Schema
NameRequiredDescriptionDefault
opYes
argsYes
cellYes
sessionIdYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does disclose a genuine behavioral trait: referenced cells become reactive dependencies so the cell recomputes on change. It also lists per-op return types. It omits error behavior for invalid ops and whether redefining overwrites an existing cell.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core reactive-args rule, then a tight per-op bullet list pairing args with produced value types. Dense but every line earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-op dispatcher with no output schema and no annotations, the description supplies args and return shapes per op plus reactivity semantics, which is substantial. Remaining gaps are sessionId semantics, failure modes for bad ops, and overwrite behavior on redefine.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it enumerates every legal 'op' value with its 'args' shape, and explains the {'$cell': ...} reference syntax for args. However 'sessionId' and the 'cell' naming semantics are never explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line gives a specific verb+resource ('Define a computed cell from a catalog op') and the op catalog makes the scope concrete. It does not explicitly distinguish itself from close siblings like session_set_cell or session_get_cell, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the op catalog rather than stated: no when-to-use/when-not guidance and no comparison against session_set_cell. The one real routing hint is 'project with session_get_cell' for opaque graph values, which points to a sibling but only for that narrow case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_explain_cellA

Explain a cell's own derivation (issue #5): its role (free input vs computed), the op and raw args that defined it (if computed), its immediate upstream cells with their current values, and its own current value. One level only -- call again on a listed dependency's "cell" name to go deeper.

ParametersJSON Schema
NameRequiredDescriptionDefault
cellYes
sessionIdYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return structure and the one-level recursion limit well, implying a safe read operation. It doesn't mention error behavior for invalid cell names or missing sessions, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the purpose and followed by the recursion rule. Every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe what is returned, and it does so precisely (role, op, args, upstream cells, values). Combined with the one-level recursion note, an agent has enough to call and iterate correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It clarifies that 'cell' is a cell name that appears in dependency listings, which adds meaning, but sessionId is never explained and no format hints are given for either parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Explain a cell's own derivation') and enumerates exactly what the explanation contains: role (free input vs computed), op and raw args, immediate upstream cells with values, and current value. This distinguishes it from session_get_cell, which would only return the value.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly bounds the scope ('One level only') and tells the agent how to go deeper: re-invoke on a listed dependency's cell name. It doesn't explicitly compare against sibling tools like session_get_cell or session_list_cells, but the recursion guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_get_cellB

Read a cell's current value (recomputing it if stale). Rich values project to typed JSON (a graph cell returns { "$type": "graph", vertices, edges, directed }).

ParametersJSON Schema
NameRequiredDescriptionDefault
cellYes
sessionIdYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full transparency burden. It reveals a key behavioral trait—recomputing stale values—rather than presenting a plain read, and explains rich-value projection with a concrete graph example. It does omit error/edge behaviors, but the core side-effectful recomputation is explicitly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with a compact illustrative parenthetical. Every clause earns its place: the core read behavior, the stale recomputation side effect, and a concrete example of the rich JSON output. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter reader, it covers the critical behaviors (stale recompute and typed JSON projection). However, it doesn't explain the basic return format for non-rich cells, the session-open prerequisite, or failure behaviors. Since there is no output schema and no annotations, the description carries a heavier burden that it only partially satisfies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not compensate: it never explains the semantics of sessionId or cell, how cells are referenced, or what format is expected. The parameter names are somewhat self-explanatory, but the description adds no operational meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description provides a concrete action and target: 'Read a cell's current value'. It is distinct from the sibling set_cell/write operations in effect, though it does not explicitly name alternatives. The additional detail about recomputing stale values further clarifies intended semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus siblings like session_set_cell or session_list_cells. It implies a read operation but does not state prerequisites, exclusions, or alternative conditions. This leaves the agent to infer usage purely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_listA

List open sessions with kind, cell count, and creation time.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full transparency burden. It makes the read-only list behavior clear and discloses the returned attributes, but it does not mention ordering, pagination, or other behavioral details such as whether closed sessions are excluded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that starts with the action and immediately delivers the key semantic information. Every word is meaningful, with no redundancy or boilerplate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

As a straightforward zero-parameter list tool with no output schema, the description is largely complete: it identifies the resource and the returned fields. A brief note on ordering or cell-level access would improve completeness, but the current information is sufficient for basic selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, which matches the baseline of 4 for tools without parameters. The description needs to explain parameter-level semantics, and it does so implicitly by defining the scope of the returned list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') plus a specific resource ('open sessions') and even names the returned fields: kind, cell count, and creation time. This clearly communicates what the tool does and distinguishes it from sibling tools like session_list_cells, which operate on cells rather than sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool—when you need a list of open sessions and their basic metadata. However, it does not explicitly state when not to use it or point to alternatives like session_list_cells for cell-level information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_list_cellsA

List every cell in a session with its role ("free" input vs "computed") and, for computed cells, the op that defines it.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output structure (role and op) but does not explicitly state that the operation is read-only, whether session must be open, or error behavior. The 'List' verb implies safety, but more could be added.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded with the verb 'List' and immediately conveys the key details. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description covers the main function and return values, but lacks prerequisite context (e.g., session must be open) and parameter clarification. It's adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the sessionId parameter. While the name is self-explanatory, the description fails to specify that it references an existing session or how to obtain it, leaving a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists every cell in a session, specifying the information returned (role and op for computed cells). This distinguishes it from siblings like session_get_cell (single cell) and session_list (sessions).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for viewing all cells in a session, but it doesn't explicitly mention when to use this over alternatives like session_get_cell. Sibling names provide context, but no direct guidance or exclusion is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_openA

Open a reactive cell session. kind "generic" starts empty; "graph-theory" pre-wires an edge-list -> analysis -> BFS pipeline (input cells: edgeListText, directed, startVertex; computed cells: parsed, analysis, bfsOrder). Optional seed sets input cells in the same call (overriding preset defaults). Optional capabilities (issue #7) grants this session use of ops that declare a requiresCapability -- see session_define's own description for which ops (if any) need one; default none. Sessions are in-memory and die with the server.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
seedNo
capabilitiesNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses that sessions are in-memory and die with the server, that seed overrides preset defaults, and that capabilities gate ops declaring requiresCapability. It omits idempotency, error behavior, and any session-count/resource limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core purpose is in the first sentence and the kind semantics immediately follow. Parenthetical asides and the cross-reference to session_define add length without being strictly wasteful, but the text is less scannable than it could be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should say what opening a session returns (e.g. a session identifier needed by sibling calls) but never does. Lifecycle ('in-memory, die with the server') is covered and partially compensates, but the return contract is a real gap for a session-creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it defines both enum values of kind, explains seed as setting input cells (and names the graph-theory input cells), and explains capabilities as granting access to requiresCapability ops. The capabilities explanation is deferred to session_define's description rather than fully self-contained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ('Open a reactive cell session') and then details what each kind produces, which lets an agent distinguish this from session_resume. It does not explicitly contrast itself with the sibling that resumes an existing session, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what 'generic' vs 'graph-theory' means and how seed/capabilities behave, which implicitly guides parameter choice. It never states when to call session_open instead of session_resume or session_list, so the when-to-use guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_resumeA

Reconstruct a session from a session_snapshot document. A resumed session is freshly opened, not re-authorized from wherever it paused -- every existing resource guard (session/cell/payload limits) applies exactly as it would to session_open, and capabilities (issue #7) are NOT carried forward from the original session: pass this call's own optional capabilities to grant any, default none.

ParametersJSON Schema
NameRequiredDescriptionDefault
snapshotYes
capabilitiesNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that resource guards (session/cell/payload limits) apply exactly as for session_open, and the critical gotcha that capabilities are NOT carried forward from the original session, defaulting to none unless passed explicitly. This is exactly the non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the behavioral caveats. Dense but every sentence earns its place except the minor '(issue #7)' aside. Slightly long for a two-parameter tool but no real waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does not state what a resumed session returns, but it thoroughly covers the behavioral risks. The main gap is the undocumented nested snapshot structure, which is not explained in either schema or description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the optional 'capabilities' parameter semantics well ('pass this call's own optional capabilities to grant any, default none'), but the required nested 'snapshot' object (v, kind, free, defines) is described only as a 'session_snapshot document' with no field-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Reconstruct a session from a session_snapshot document') and immediately differentiates it from session_open by clarifying the session is 'freshly opened, not re-authorized from wherever it paused.' An agent can distinguish it from session_snapshot (creates the doc) and session_open (starts fresh) without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly establishes the context of use (given a session_snapshot document) and implicitly contrasts with session_open via the 'freshly opened' framing. It does not explicitly state when NOT to use it or name session_open as the alternative, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_set_cellA

Set an input cell's value (any JSON). Setting a previously-computed cell demotes it to a plain input. Dependent computed cells recompute lazily on their next get.

ParametersJSON Schema
NameRequiredDescriptionDefault
cellYes
valueNo
sessionIdYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral transparency. It discloses two important side effects: demotion of computed cells to plain inputs and lazy recomputation of dependent cells. This is solid, though it omits potential edge cases like whether the session must be open.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the main action and immediately followed by important behavioral details. Every sentence adds value with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core action and key side effects, which is good for a simple setter. However, it omits sessionId semantics, the optional nature of value, and error/edge-case behavior, making it not fully complete for an unannotated tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain all parameters, but it does not mention sessionId at all. It also says 'Set an input cell's value,' implying value is required, while the schema marks value as optional, creating ambiguity about whether a value must be provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the primary action: 'Set an input cell's value (any JSON).' It also distinguishes the tool from siblings like session_get_cell and session_define by explicitly mentioning the behavior for computed cells, making the scope and effect unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful usage context by explaining that setting a previously-computed cell demotes it to a plain input and that dependents recompute lazily. It implies when the tool is appropriate, though it does not explicitly name alternative tools or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_snapshotA

Serialize a session's portable state (issue #6): current free-cell values plus define-specs. Does NOT include computed-cell cache -- those re-derive from define() on resume. Hand the returned document to session_resume, possibly on a different math-grapher process, to reconstruct an equivalent session.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the computed-cell cache is deliberately omitted because cells re-derive from define() on resume, and that the output is a portable document usable in another process. It does not discuss side effects or failure modes, but for a serialize operation the key behavioral fact (what state is and isn't captured) is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences that each add substance: what is serialized, what is not, and how to consume it. The parenthetical '(issue #6)' is the one piece of noise that does not help an agent decide or invoke, slightly undercutting otherwise tight prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema or annotation set, yet the description explains what the returned document contains and what the caller should do with it, which covers the important return-value concern. The only real gap is parameter origin/format for sessionId, which leaves the definition slightly short of fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single sessionId parameter, so the description would need to compensate and largely does not — it never says where sessionId comes from (presumably session_open) or what form it takes. The session-centric framing implies the parameter's meaning, which is the minimum viable level given the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (serialize) and resource (a session's portable state), then enumerates exactly what is included (free-cell values, define-specs) and excluded (computed-cell cache). It also names the counterpart sibling session_resume, so an agent can distinguish it from session_get_cell or session_list without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent what to do with the result: hand the document to session_resume, possibly across a different math-grapher process. That is a clear use context (session transfer/portability), though it does not state when NOT to use this tool or contrast with a sibling alternative beyond session_resume.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.0.1
    • Addedsession_explain_cell
    • Changedsession_open1 field changed
      • addedInput schema / properties / capabilities
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Addedsession_resume
    • Addedsession_snapshot
  2. 7 tool updatesv0.0.0
    • First observedsession_close
    • First observedsession_define
    • First observedsession_get_cell
    • First observedsession_list
    • First observedsession_list_cells
    • First observedsession_open
    • First observedsession_set_cell

TDQS

A3.7/5.0

Scored across 10 tools

Disambiguation4/5

Session lifecycle tools and graph/math operations are mostly distinct, with clear boundaries between get, set, list, and explain cell operations. However, session_define lists catalog ops that are also exposed as standalone tools (math_eval, graph_*), which could confuse whether to invoke an op directly or via define.

Naming Consistency4/5

All tool names use snake_case with predictable prefixes: session_* for session lifecycle and graph_* for graph operations. Minor deviations like math_eval lacking a prefix and graph_bfs using an acronym rather than verb_noun keep it from perfect consistency.

Tool Count4/5

The actual tool list contains 16 tools (despite the stated count of 10), which is slightly above the 3-15 sweet spot. The set is still reasonably scoped for a reactive-cell session server with graph and math operations, and each tool covers a distinct lifecycle or op.

Completeness4/5

The surface covers session open/resume/snapshot/close/list, cell get/set/list/explain/define, and a useful graph/math op catalog. A few gaps exist, such as no explicit cell deletion or session update/capability modification after opening, but core workflows are supported.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers