Skip to main content
Glama

Axiom — Advanced Math MCP Server

npm License: GPL v3+ Node.js >=20 MCP CI codecov

Exact symbolic and numerical mathematics for LLMs — a real computer algebra system (Giac/Xcas) behind the Model Context Protocol, and behind a shell command. Published as axiom-math.

Axiom catching a wrong derivative, then computing an exact integral

Quick start

As a CLI, straight away:

npx -y axiom-math compute 'integrate(sin(x)^3,x)'   # -cos(x)+cos(x)^3/3
npx -y axiom-math verify 'diff(x^3,x) = 3*x^2'      # exit 0 — it holds

As an MCP server, in any client's config:

{ "command": "npx", "args": ["-y", "axiom-math"] }

As an agent skill — drop in skills/axiom-math/SKILL.md, which teaches an agent the three commands and their exit codes.

Related MCP server: math-logic-mcp

Why Axiom?

LLMs often make calculation errors, especially with symbolic math, exact fractions, and multi-step problems. Axiom provides verified, exact results through two layers:

  • math.js — Fast numerical evaluation (arithmetic, trigonometry, matrices)

  • Giac/Xcas WASM — Symbolic computation (calculus, algebra, equation solving)

Benchmark Results (GLM-5.1, May 2026)

Dataset

Baseline

+MCP

Delta

GSM8K (100)

96.0%

98.0%

+2.0%

MATH L3 (50)

70.0%

80.0%

+10.0%

MATH L4 (50)

50.0%

62.0%

+12.0%

MATH L5 (50)

38.0%

52.0%

+14.0%

CAS-quick (60)

55.0%

70.0%

+15.0%

Omni-MATH ≥7 (50)

0.0%

0–4%

(ceiling)

Key insights:

  • Phase 0 grader (LaTeX/Unicode normalization + symbolic equivalence) is the dominant value driver across all datasets

  • CAS-quick lifted from 26.7% (April pre-grader) to 70% (post-grader) — the biggest single jump

  • Omni-MATH ≥7 is at ceiling for current LLM+CAS setups; needs fundamentally different approaches (Lean/Coq, fine-tuning, RAG)

Full results: benchmark/results/ and docs/superpowers/specs/ (per-phase analysis)


Features

Axiom exposes 3 MCP tools. Almost everything flows through compute, a single gateway that parses a CAS-style problem string and routes it to the right internal engine — so callers learn one tool, not dozens.

Tool

Purpose

compute

Solve any math problem. Pass a CAS-style string (solve(...), diff(...), det([[...]]), C(10,3), 2+3*sin(pi/4)) or any Giac/Xcas expression.

verify

Independently check a mathematical claim (identity, solution, or computation) via symbolic and/or numeric methods.

plot

Render a 2D function graph as an SVG image.

What compute covers

compute recognizes CAS-style verbs and dispatches across these domains. Anything it doesn't recognize falls through to raw Giac/Xcas evaluation.

Domain

Verbs / examples

Arithmetic & units

2+3*sin(pi/4), 100 km/h to m/s

Equation solving

solve(x^2-4=0, x), csolve(...) (complex), solve_system([x+y=5, x-y=1], [x,y])

Calculus

diff, int, limit, taylor, desolve (ODEs of any order, and linear constant-coefficient systems)

Multivariable calculus

gradient, hessian, jacobian, divergence, curl, partial, iint/iiint (multiple integrals), critical_points, lagrange, tangent_plane, directional_derivative

Algebra

factor, simplify, expand, partfrac

Linear algebra

det, inv, eigenvals, eigenvects, rref, rank, tran, ker, qr, lu, cholesky, svd, norm, cond

Number theory

ifactor, isprime, euler, analyze

Combinatorics

C(n,k), P(n,k), stirling, bell, catalan, derangements, multinomial

Probability

binomial, normal, poisson, geometric, hypergeometric, chi_square, student_t, f_distribution, beta, exponential

Hypothesis testing

t_test (one/two/paired), anova, chi_square_test

Numerical methods

newton, bisection, secant, romberg, simpson

2D geometry

distance, midpoint, slope, area_*, perimeter, circumference, line_intersection, point_line_distance, angle_between_lines

3D geometry

distance3d, midpoint3d, dot, cross, vector_norm, angle_vectors, plane_from_points, point_plane_distance, line_plane_intersection, plane_plane_angle, line_line_distance, volume_tetrahedron, volume_sphere, volume_parallelepiped

Transforms & series

laplace, ilaplace, fourier/fft/ifft, sum, product

Exact values

to_exact, to_decimal, simplify_fraction

Regression & sequences

linear_regression/fit, polynomial_regression, sequence (pattern identification)


Installation

The package is axiom-math on npm. Nothing to install for normal use — npx fetches and caches it:

npx -y axiom-math compute '2+2'

Or install it so the axiom-math command is on your PATH:

npm install -g axiom-math

Node.js >= 20 required. The first run downloads about 3.8 MB (the CAS engine compiled to WebAssembly) and takes a few seconds; later runs come from the npx cache.

From source

For contributors, or to run a modified build:

git clone https://github.com/tufantunc/axiom-advanced-math-mcp.git
cd axiom-advanced-math-mcp
npm install
npm run build

Docker

# Build and run
docker-compose -f docker/docker-compose.yml up -d

# Check logs
docker-compose -f docker/docker-compose.yml logs -f

# Stop
docker-compose -f docker/docker-compose.yml down

Usage

CLI (STDIO Transport)

# Run with stdio transport (default)
npm start

# Development mode
npm run dev

Claude Desktop integration:

// ~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": [
    {
      "name": "axiom-math",
      "command": "npx",
      "args": ["-y", "axiom-math"]
    }
  ]
}

Running from a local checkout instead of npm — point args at the built entry point:

"args": ["/path/to/axiom-advanced-math-mcp/dist/cli.js"]

Command line

The same binary works as a one-shot CLI, so agents can use it as a skill with no MCP configuration. With no arguments it is the MCP server; with a subcommand it runs one computation and exits.

npx -y axiom-math compute 'integrate(sin(x)^3,x)'
npx -y axiom-math compute -q 'solve(x^2-4=0,x)'     # {-2, 2}
npx -y axiom-math verify 'sin(x)^2+cos(x)^2 = 1'    # exit 0 if true
npx -y axiom-math plot 'sin(x)' -o wave.svg
echo 'diff(x^3,x)' | npx -y axiom-math compute -q   # 3*x^2

Flag

Meaning

-q

print one value only, for scripting

--json

structured output

--latex

LaTeX-focused text (compute only)

-h, --help

usage, or usage for a subcommand

Exit codes: 0 success · 1 tool or usage error · 2 verify checked the claim and it is false.

2 is a mathematical verdict, so a claim that never got checked does not use it: one that fails to parse, or that the CAS cannot evaluate, exits 1 with nothing on stdout. axiom-math verify '...' && ... therefore never reads a syntax error as a disproof.

A ready-to-use agent skill is in skills/axiom-math/SKILL.md.

HTTP Transport

# Start HTTP server (default: http://127.0.0.1:3000)
npm run start:http

# Development HTTP
npm run dev:http

The HTTP transport is stateless: every POST /mcp is handled independently, no Mcp-Session-Id is issued, and no session state is kept between requests. This server sends no server-initiated notifications, so nothing is lost — and it scales horizontally with no shared state.

Method

Path

Behaviour

POST

/mcp

Handles a JSON-RPC message

GET

/mcp

405 — no SSE stream is offered

DELETE

/mcp

405 — there are no sessions to terminate

GET

/health

200 when ready, 503 when the CAS engine is not

Security: there is no authentication and no rate limiting. The default bind address is 127.0.0.1, but docker/docker-compose.yml sets MCP_HOST=0.0.0.0. If you expose the port, put it behind a reverse proxy that authenticates and rate-limits — docker/reverse-proxy/ is a working, tested one (nginx + basic auth + per-client concurrency cap, with the app publishing no port of its own). SECURITY.md documents the full posture — what is protected, what is not, and how to report a vulnerability.

POST /mcp also validates the Host header against an allowlist (localhost, 127.0.0.1, [::1] by default) to block DNS rebinding — a malicious page can make a victim's browser resolve an attacker domain to 127.0.0.1 and reach this server through it. If you reach the server by a LAN address, hostname, or reverse-proxy domain other than loopback, set MCP_ALLOWED_HOSTS or every POST /mcp request will get a 403. This check is not authentication — it only constrains which host names may reach the endpoint, nothing about who is asking.

Environment variables:

Variable

Default

Description

MCP_PORT

3000

HTTP server port

MCP_HOST

127.0.0.1

HTTP server host

MCP_ALLOWED_HOSTS

loopback only (localhost, 127.0.0.1, [::1])

Comma-separated Host header allowlist for POST /mcp (DNS-rebinding protection). An explicit value replaces the default rather than extending it.

AXIOM_EVAL_TIMEOUT_MS

10000

Per-evaluation timeout, in milliseconds. Bounds one CAS call and one js-compute call (arbitrary-precision integer work, arithmetic, plot sampling), so lowering it tightens both. Accepts a plain number or a ms/s suffix (10000, 250ms, 10s), clamped into what a timer can hold (12147483647ms); a value with no number in it, or a non-positive one, falls back to 10000 with a warning. An unset or blank value takes the default silently.

AXIOM_INTEGRATION_BUDGET_MS

max(3 × AXIOM_EVAL_TIMEOUT_MS, 30000)

Wall-clock budget for one multi-call numerical routine (integration, root finding). Bounds the SUM of CAS calls, where AXIOM_EVAL_TIMEOUT_MS bounds one. Validated like AXIOM_EVAL_TIMEOUT_MS, and a rejected value is named on stderr.

AXIOM_JS_COMPUTE_HEAP_MB

512

Heap ceiling for the child process that runs arbitrary-precision integer work and mathjs evaluation. Exceeding it fails the computation that caused it — calls queued behind it are re-sent to the replacement worker — and leaves the server up. Accepts whole MB, optionally suffixed (512, 512MB, 2GB). A value is floored to whole MB and clamped into [16, 1048576], never to a looser ceiling than the one written: below 16 the child cannot boot, and above 1048576 this server caps it (V8 honours more, but past ~1.76e13 its size_t arithmetic wraps to a smaller heap than requested). Only a value with no number in it falls back to 512. Every correction is named on stderr. The mathjs-backed tools (quick_calc, plot) need at least 48 — measured, their startup import needs ~46MB and 44MB dies — and below that they are refused by name with the cause, rather than failing as a worker fault.

AXIOM_COMPUTE_HYGIENE

unset

Set to 1 to enable compute output post-processing

One bound is not configurable: a result over 100,000 characters is refused rather than returned, so an expression like 1:2000000 reports its element count instead of shipping 24 million characters into the caller's context.

Some inputs are refused rather than answered, because any answer would be meaningless. Arithmetic that evaluates to NaN (such as 0/0) is an error; an infinite result is returned with a warning, because a true infinity and a value that overflowed the range of a double are indistinguishable once computed. A t-test needs variation in whatever it actually tests — paired_t compares the differences, so it is those that must vary, while Welch's two_sample_t needs only one of the two samples to vary. A contingency table needs non-negative counts, no all-zero row or column, rows of equal length, and more than one row and column. A one-way ANOVA needs some within-group variation and more observations than groups. And any of these is refused when the values are large enough that the statistic itself overflows to infinity, because an overflowed statistic is no longer the statistic. A numerical method is refused when its expression does not depend on the variable it is solved or integrated over, or when the CAS answers symbolically rather than with a number — previously the leading term of that symbolic answer was reported as the result.

A system of differential equations written as a list — desolve([y'=z, z'=-y], x) — is rewritten into the matrix form the CAS solves and returns a solution for every function. The components come back in the order the equations were written, and the JSON envelope names them in a components field, because [[cos(x),-sin(x)]] is not interpretable without it.

Initial conditions must be given for every function, at the same point, or not at all — a partial set is refused rather than ignored. Also refused, each with its own reason: a system that is not linear in the unknown functions; coefficients that depend on the independent variable; a derivative of order above one (rewrite y''=z as y'=w, w'=z); more than nine equations; and a system the CAS cannot finish.

The infinite-result rule covers arithmetic evaluation. A symbolic +infinity from the CAS routes — a limit, a divergent integral — is a normal answer and carries no warning.

MCP Inspector

npm run inspect

Tool Reference

compute

The single gateway for all math. Pass a CAS-style problem string; the router parses it and dispatches to the right engine.

Parameter

Type

Description

problem

string (required)

CAS-style problem, e.g. solve(x^2-4=0, x), diff(x^3, x), det([[1,2],[3,4]]), gradient(x^2+y^2, [x,y]).

domain

real | complex | numeric | exact

Domain hint (default real). complex → complex solutions; numeric → force numerical methods; exact → exact symbolic form.

precision

integer 1–50

Decimal places (default 10).

format

text | latex | json

Output format (default text). json returns a structured envelope.

Examples:

{ "problem": "solve(x^2 - 5*x + 6 = 0, x)" }
{ "problem": "int(x^2*sin(x), x)", "format": "latex" }
{ "problem": "lagrange(x*y, x+y, 1, [x, y])" }
{ "problem": "volume_tetrahedron([0,0,0],[1,0,0],[0,1,0],[0,0,1])" }
{ "problem": "binomial cdf n=10 k=3 p=0.5", "format": "json" }

verify

Independently check a mathematical claim. Useful as a second, tool-grounded opinion on a result the model produced.

Parameter

Type

Description

claim

string (required)

The claim, e.g. "sin(x)^2 + cos(x)^2 = 1" (identity), "x=2 satisfies x^2-4=0" (solution), "diff(x^3, x) = 3*x^2" (computation).

method

numeric | symbolic | both

Verification method (default both).

Returns four fields: verified, evaluated, confidence, and checks_performed.

evaluated is the one to read first. It is false when no check produced a usable answer — the claim did not parse, or the CAS could not evaluate it — in which case verified: false means "unknown", not "refuted". Treating the two as the same turns a syntax error into a disproof.

plot

Render a 2D function as an SVG image.

Parameter

Type

Description

expression

string (required)

Function to plot, e.g. "sin(x)", "x^2 - 3*x + 1".

variable

string

Variable name (default x).

x_min, x_max

number

X range (default −10 … 10).

y_min, y_max

number

Y range (auto-detected if omitted).

width, height

number

Image size in px (default 600 × 400).

title

string

Optional chart title.

Returns a base64-encoded SVG image (axes, grid, labels, asymptote detection) plus a text caption.

Prompts

The server also registers guided MCP prompts that chain compute/verify for multi-step workflows: solve-step-by-step, analyze-function, verify-identity, convert-units, analyze-dataset, solve-ode-system, and regression-workflow.


Run Benchmarks

Default production recipe (grader-v2 included automatically):

cd benchmark
npm install

# Set provider API key (one of):
export ZAI_API_KEY=...
export ANTHROPIC_API_KEY=...
export OPENROUTER_API_KEY=...

# Run benchmarks (provider defaults from --zai/--anthropic/--openrouter flags)
npm run cas:quick:zai      # CAS-quick (60 problems, ~30 min)
npm run gsm8k:quick:zai    # GSM8K-quick (100 problems, ~30 min)
npm run math:quick:zai     # MATH L3-L5 quick (150 problems, ~75 min)

Optional ablation features (off by default)

  • --features=output-hygiene — tool output post-processing (Unicode normalize, optional simplify, silent-failure warning). Marginal +1pp on CAS in live measurement.

  • --features=grader-v3 — equation-RHS extraction + bare-comma-list set match. Marginal +1pp on CAS.

  • --features=self-consistency — N=3 majority voting (variance reduction; 3× cost; no accuracy gain on CAS).

Example:

npm run cas:quick:zai -- --features=output-hygiene,grader-v3

See docs/superpowers/specs/2026-05-*-results.md for live ablation analysis of every flag.

What we tried that didn't work

This project went through extensive ablation across five phases (Phase 0–4). The following experimental approaches were tested live and rejected:

  • Phase 1: Structured JSON output with \boxed{} trailers — model paraphrased boxed content into LaTeX style, breaking answer extraction. Net regression on CAS.

  • Phase 2: 8K token budget (tokens-8k) — gave the model more room to wander rather than recovering from truncation. Net regression −6.7pp on CAS.

  • Phase 3: Self-consistency for accuracy — N=3 voting did not lift accuracy (Wang et al. literature gain not reproducible on CAS); kept as a methodology tool for variance reduction only.

  • Phase 4: Olympiad-specific scaffolding prompt — engagement improved (no-tool-call rate 84% → 74%) but accuracy stayed at 0%. Olympiad-tier problems are out of scope for prompt-engineering interventions.

Each phase's per-problem analysis is in docs/superpowers/specs/2026-05-*-results.md. The honest documentation of failures is preserved as a project archive.


Architecture

Compute gateway → router → domain handlers

┌─────────────────────────────────────────────────────────────┐
│              MCP Protocol Layer (stdio / HTTP)               │
└─────────────────────────────────────────────────────────────┘
                              │
        ┌─────────────────────┼─────────────────────┐
        ▼                     ▼                     ▼
   ┌─────────┐          ┌──────────┐          ┌─────────┐
   │ compute │          │  verify  │          │  plot   │
   └────┬────┘          └──────────┘          └─────────┘
        │  route() → extract args → dispatch
        ▼
┌─────────────────────────────────────────────────────────────┐
│  Domain handlers: calculus, algebra, matrix, multivariable,  │
│  geometry / geometry3d, combinatorics, probability,          │
│  hypothesis testing, number theory, numerical methods, …     │
└─────────────────────────────────────────────────────────────┘
        │                     │                     │
        ▼                     ▼                     ▼
┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│   math.js    │     │  Giac/Xcas   │     │ Exact engine │
│ (numerical)  │     │  (symbolic)  │     │ (fractions)  │
└──────────────┘     └──────────────┘     └──────────────┘

compute never asks the caller to pick a handler. The router matches the problem string against ordered rules, the matching extractor parses arguments, and the dispatcher calls the corresponding domain handler. Unmatched input falls through to raw Giac/Xcas.

Response Format

Text-format responses are line-structured so LLMs (and the benchmark grader) can extract answers reliably:

{
  "content": [
    { "type": "text", "text": "Result: 400/11" },
    { "type": "text", "text": "Decimal: 36.3636363636" },
    { "type": "text", "text": "LaTeX: \\frac{400}{11}" },
    { "type": "text", "text": "" },
    { "type": "text", "text": "The answer is 400/11 (≈ 36.36)" }
  ],
  "isError": false
}

Benchmark Results

Datasets

Dataset

Problems

Difficulty

GSM8K

100

Grade school math (arithmetic)

MATH L3

50

High school math

MATH L4

50

Advanced high school math

MATH L5

50

Olympiad-level math

Omni-MATH ≥7

50

Expert-level math

How to Run

See Run Benchmarks above for the commands. In short, from the repository root:

npm run benchmark:zai         # quick sample, GLM-5.1
npm run benchmark:full:zai    # all datasets
npm run benchmark:l5:zai      # one difficulty tier

Swap :zai for :openrouter to change provider. The benchmark/ directory is a separate npm project with finer-grained scripts (cas:quick:zai, gsm8k:quick:zai, …); npm run benchmark:* from the root delegates to them.

Environment variables:

Variable

Required for

Description

ZAI_API_KEY

zai provider

Your z.ai API key

OPENROUTER_API_KEY

openrouter provider

Your OpenRouter API key


Development

Scripts

Command

Description

npm run build

Compile TypeScript to dist/ and copy the WASM asset

npm start

Run STDIO server

npm run dev

Run in development mode (tsx)

npm run start:http

Run HTTP server

npm run dev:http

Run HTTP server in dev mode

npm test

Unit tests — no build required

npm run test:integration

Integration tests — builds first, exercises dist/

npm run test:watch

Unit tests in watch mode

npm run test:coverage

Unit tests with coverage report

npm run typecheck

Type-check without emitting

npm run lint

Lint with oxlint

npm run lint:fix

Auto-fix linting issues

npm run format

Format with Prettier

npm run format:check

Check formatting without writing

npm run inspect

Open the MCP Inspector against the stdio server

Testing

The suites are split. npm test runs the unit tests and needs no build; npm run test:integration builds first and exercises the packaged dist/ output, so it catches things the unit suite cannot — the shipped binary's argument dispatch, the MCP handshake, exit codes.

npm test                  # unit
npm run test:integration  # integration (runs npm run build first)
npm run test:watch        # unit, watch mode
npm run test:coverage     # unit, with coverage

Test coverage: unit + integration suite, 100% pass rate. Run npm test for the current count — it changes too often to keep a number here in sync.

WASM Build (Giac)

npm run build:giac:wasm

# Build a specific upstream ref instead of master
GIAC_REF=v1.9.x npm run build:giac:wasm

This runs scripts/build-giac-wasm.sh, which builds docker/build-giac-wasm/Dockerfile with docker build (no Compose file involved) and writes giac.wasm.js straight into src/server/giac/ — no manual copy step needed. Requires Docker Desktop (or another Docker daemon) running locally. Per-task build logs land under logs/giac-build/.


Contributing

Bug reports and pull requests are welcome — see CONTRIBUTING.md for the setup, the checks CI runs, and the few things about this codebase that are not obvious from reading it.


License

GNU General Public License v3.0 or later — see LICENSE.

Axiom embeds Giac/Xcas, which is GPL-3.0-or-later, so the combined work carries the same license. Details and attribution: THIRD-PARTY-NOTICES.md.

Does the GPL affect my agent?

No. Your agent talks to Axiom over the Model Context Protocol — a separate process, over stdio or HTTP. Separate programs communicating at arm's length are not a combined work, so running Axiom alongside your own agent puts no license obligation on your code, whatever license it uses. Running the software is unrestricted under the GPL, including running it as a service.

The copyleft terms apply when you redistribute Axiom itself — shipping it (modified or not) inside a product you hand to someone else. In that case, pass along the source under GPL-3.0 and keep the notices intact.

Available Tools

3 tools
computeA

Solve any math problem: equations, calculus, algebra, matrices, combinatorics, probability, statistics, geometry, number theory, and more. Pass a CAS-style problem string (e.g., "solve(x^2-4=0, x)", "diff(x^3, x)", "det([[1,2],[3,4]])", "C(10,3)", "2+3*sin(pi/4)") or any Giac/Xcas expression.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNoDomain hint: real (default) — real solutions complex — complex solutions (csolve, cfactor) numeric — force numerical methods exact — exact symbolic form
formatNoOutput format: text (default) — human-readable result latex — LaTeX-focused output json — structured ComputeEnvelope
problemYesMathematical problem to solve. Use CAS-style function calls for clarity: solve(x^2-4=0, x) — solve equation diff(x^3, x) — differentiate int(x^2, x, 0, 1) — definite integral limit(sin(x)/x, x, 0) — limit taylor(exp(x), x=0, 5) — Taylor series factor(x^2-4) — factorize simplify((x^2-1)/(x-1)) — simplify expand((x+1)^3) — expand det([[1,2],[3,4]]) — matrix determinant C(10,3) — combinations ifactor(2310) — prime factorization 2+3*sin(pi/4) — arithmetic Or any valid Giac/Xcas expression as fallback.
precisionNoDecimal precision (default: 10)

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It accurately indicates a computational tool via examples, but does not describe potential limitations, output formatting behaviors, or error handling. It does not contradict any annotations, but lacks richer behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused paragraph that front-loads the core purpose and then provides illustrative examples. While somewhat long due to the many examples, each example adds practical value for understanding supported syntax, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description covers the purpose and input syntax well but omits details about return values, output structure (unless using 'format' parameter), and situational guidance relative to siblings. It is adequate for invoking the tool but leaves some gaps in full context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description enriches the 'problem' parameter with detailed CAS-style examples (e.g., 'solve(x^2-4=0, x)', 'det([[1,2],[3,4]])'), which adds meaningful guidance beyond the schema. Other parameters are well-documented in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool solves math problems across many domains and provides explicit CAS-style examples. It distinguishes itself from siblings primarily through its focus on computation, but does not explicitly contrast with 'plot' or 'verify'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for math problem-solving through examples, but it does not provide explicit when-to-use or when-not-to-use guidance. It does not mention alternatives or contexts where 'verify' or 'plot' might be more appropriate, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plotA

Plot a mathematical function as an SVG graph. Returns an image showing the function curve with axes, grid, and labels.

Examples:

  • plot sin(x) from -2pi to 2pi

  • plot x^2 - 3*x + 1 from -5 to 5

  • plot exp(-x^2) (Gaussian curve)

  • plot 1/x with asymptote detection

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoChart title (optional)
widthNoImage width in pixels (default: 600)
x_maxNoMaximum x value (default: 10)
x_minNoMinimum x value (default: -10)
y_maxNoMaximum y value (auto-detected if omitted)
y_minNoMinimum y value (auto-detected if omitted)
heightNoImage height in pixels (default: 400)
variableNoVariable name (default: "x")
expressionYesMathematical expression to plot (e.g., "sin(x)", "x^2 - 3*x + 1")

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses the output type (SVG image), the presence of axes/grid/labels, and asymptote detection in an example. It does not mention error handling or limitations, but covers the key behavior of returning an image.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core statement, then a brief output description, followed by concise, well-chosen examples. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 9 parameters and no output schema, the description is relatively complete: it states the return type, mentions asymptote detection, and shows usage patterns. It lacks explicit details on expression syntax limitations, but the examples plus full schema coverage make it sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value with examples that illustrate usage of x_min/x_max and expression syntax, going beyond the schema's individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool plots a mathematical function as an SVG graph, with a specific verb and resource. It distinguishes itself from sibling tools (verify, compute) by focusing on graphical output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context through examples, showing typical usage like 'plot sin(x) from -2*pi to 2*pi'. It does not explicitly exclude alternatives or name when not to use, but the examples and focus on graphing imply its intended use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verifyA

Verify a mathematical claim using symbolic and/or numeric checks. Supports identity verification (e.g., "sin(x)^2+cos(x)^2 = 1"), solution checking (e.g., "x=2 satisfies x^2-4=0"), and computation assertions.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimYesMathematical claim to verify. Examples: "sin(x)^2 + cos(x)^2 = 1" — identity check "x=2 satisfies x^2-4=0" — solution check "diff(x^3, x) = 3*x^2" — computation check
formatNoOutput format: text (default) — human-readable verdict json — structured VerifyResult
methodNoVerification method (default: "both")

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It mentions that verification uses symbolic and/or numeric checks, adding some context, but does not describe limitations, error behavior, or what happens when a claim cannot be verified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the purpose statement, followed by clear, relevant examples. Every sentence adds value without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with strong schema coverage, the description adequately covers purpose and examples. However, without an output schema or annotations, a brief note about the verdict structure (beyond the format parameter) would improve completeness, though it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% parameter coverage with detailed descriptions for claim, format, and method, including enums and examples. The description adds no significant parameter semantics beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies mathematical claims using symbolic and/or numeric checks. It specifies three distinct use cases (identity verification, solution checking, computation assertions) and distinguishes it from sibling tools compute and plot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through its examples but does not explicitly state when to use verify over alternatives like compute or plot. No exclusions or alternative recommendations are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.2
    • First observedcompute
    • First observedplot
    • First observedverify

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation4/5

Verify and compute have some overlap in that compute can evaluate assertions, but their primary purposes are distinct (checking claims vs. solving problems). Plot is clearly separate. Descriptions help differentiate them, though an agent might occasionally misselect.

Naming Consistency5/5

All three tool names are single lowercase verbs (verify, compute, plot), forming a simple and consistent pattern. No naming ambiguities or mixed conventions.

Tool Count4/5

Three tools is slightly on the lower end but reasonable for a focused math MCP. Each tool covers a broad category (verification, computation, plotting), so the count feels appropriate rather than sparse.

Completeness5/5

The tool surface covers the core mathematical workflows: solving/computing, verifying claims, and visualizing functions. 'Compute' is comprehensive enough to handle simplification, integration, and other operations, leaving no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    A Model Context Protocol server that exposes 8 mathematical tools (arithmetic, algebra, calculus, matrix operations, statistics, probability, unit conversions) to any MCP-compatible AI agent, enabling mathematical computations without code.
    8
    9 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides a token-efficient exact math engine for AI agents, enabling computation of derivatives, integrals, equations, and optimized Python/NumPy code via a single MCP tool.
    4
    MIT