Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
completions
{}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
list_languagesA

List every language codecalc can execute, with extension, compile flag, and what this machine resolved. status is installed (its command was found on the sandbox PATH) or supported (nothing for it here); status_basis is resolved, meaning nothing was executed to check. Run codecalc doctor --deep to promote a runtime to available by actually running it.

tier is a DIFFERENT axis from status: status says whether THIS machine resolved the command (resolution); tier says whether codecalc's own CI has actually executed this language and asserted on its output (reliability) — tested, best_effort (declared, plausibly works, never CI-checked), or plan_only (never validated anywhere). A language can be installed here and still be best_effort or worse — that combination is exactly "the toolchain resolved and may still be broken".

list_execution_providersA

List execution providers (execution BACKENDS — local subprocess, gVisor-strict, remote) and their machine-readable capabilities.

This is about which BACKEND runs your code, not which LANGUAGE it runs — a provider's ready/strict fields are resolution facts about the backend itself. Per-language reliability (how much codecalc's CI has actually verified a given language's toolchain, vs merely resolved it) is a separate axis reported by list_languages/runtimes_status/codecalc doctor as tier; a ready provider says nothing about whether a specific language running through it has ever been execution-tested.

execute_codeA

Execute code in language in a sandbox. Use this, not execute_code_stream/run_submit/session_run, for one program whose result you can wait for within a 120s cap.

Returns stdout, stderr, exit_code, duration_ms, cpu_ms, peak_memory_kb, verdict (OK/TLE/MLE/OLE/RTE).

  • session_id: run inside a session workspace (see session_start); with a stateful session (python3/node) interpreter state persists across calls. Also reports artifacts_created (files just created/modified) since that workspace outlives the call; a sessionless run has none. See session_run for the same field plus inline content blocks.

  • max_output_kb: the 240 KiB hard ceiling is what the anthropic/maxResultSizeChars this tool advertises in its _meta already assumes — the cap leaves no headroom to raise past it without the real result exceeding that hint. A run whose real output needs more than 240 KiB belongs in a session instead: leave max_output_kb at its default (0) with session_id set (below), and oversized output SPILLS to a full-fidelity file readable via session_read_file rather than truncating — see the spill paragraph further down. An EXPLICIT max_output_kb, even under the 240 KiB ceiling, is honoured as a literal cap with no spill.

  • no_net: Linux enforces in-kernel via a seccomp-bpf filter. macOS / no-seccomp kernel: best-effort symbol shim, disclosed in unenforced when that's the only guarantee that held. See SECURITY.md.

  • compact: never drops unenforced, output_error, artifacts_created, or dependencies — if a guarantee you asked for was not applied, or a declared install failed, a compact result still says so.

  • dependencies: merged with a PEP 723 # /// script block for python3, deduped by normalized name with this argument winning; the only source for node. A block ALONE, with no dependencies argument, is enough to trigger an install — see SECURITY.md. Installed BEFORE the sandboxed step, through the same confined install_package path — never inside the sandbox — and refused (capability_not_requested) without installing anything when no_net=True or the capability policy denies or strictly limits network. Bounded by a fixed install-time budget separate from timeout (120s aggregate across every dependency; codecalc.dependencies.DEFAULT_DEPENDENCY_INSTALL_BUDGET_SECONDS) and, session-less, by CODECALC_SESSION_DISK_QUOTA_MB on the run's own workdir — either one exceeded refuses the run with a stamped, coded error naming the ceiling. See the dependencies field on the result.

With session_id set and max_output_kb left at its default, output that would otherwise be truncated is instead SPILLED: the inline stdout/stderr still carry the same truncated prefix as before, and stdout_spill/stderr_spill name a codecalc://session/{sid}/files/... resource carrying the fuller stream (session_read_file or the resource route reads it back) — capped at 4 MiB, ..._spill_capped: true if even that was not enough to hold everything. Passing an EXPLICIT max_output_kb is honoured as a literal ceiling with no spill, same as before. Session-LESS runs (no session_id) have no workspace to spill into and keep the old truncate-and-drop behaviour.

session_startA

Start a persistent session. python3/node get a stateful REPL worker (variables/imports persist across execute_code calls); other languages get a persistent workspace directory. Returns session_id.

session_stopA

Stop a session: kill its REPL worker (if any) and delete its workspace. Also deletes every session_snapshot saved for it, unless keep_snapshots=True.

session_listA

List active sessions and their languages/state.

session_filesA

List workspace files, optionally using a bounded cursor page. Use session_artifacts, not this, for only the files executed code produced; use session_read_file for one file's contents.

session_write_fileA

Write a file into a session workspace (relative path, no escapes). Use this to seed input data for executed code.

session_artifactsA

List files created by executed code in a session (excluding runner internals like main.py/run.out).

session_snapshotA

Archive or restore a session's workspace files. action:

  • "save": tar.gz the session's current files (same rules as session_artifacts: .codecalc-run/ excluded, symlinks/hardlinks refused) into a snapshot stored OUTSIDE the workspace, so sandboxed code can never read or tamper with it. Returns snapshot_id.

  • "restore": extract snapshot_id's files into a brand-new session (default) or, with replace=True, wipe and recreate session_id's OWN workspace first. Only files are restored — a python3/node session's REPL variables/imports are never part of a snapshot.

  • "list": snapshots saved for session_id, oldest first.

  • "delete": remove one snapshot (snapshot_id required).

Snapshots are deleted when their session is stopped (session_stop(keep_snapshots=True) to keep them).

install_packageA

Install a package for a language (uv pip / npm / gem / go get / cargo add...).

Asks the caller to confirm before installing (a protocol-level gate, not just the anthropic/requiresUserInteraction _meta hint — see codecalc/confirmation.py); a declined or malformed confirmation refuses with no install attempted.

With session_id, executed code in that session can import the result.

NETWORK: yes, always. The package manager fetches from its registry (PyPI, npm, rubygems, crates.io). codecalc opens no socket itself; the child process does.

NOT SANDBOXED: the installer runs as a direct subprocess of the server, so install-time hooks (npm postinstall, Python build backends, Cargo build scripts) execute with the server user's filesystem access. The environment is still restricted to the allowlist, so secrets do not leak, but the filesystem is not confined. Do not point this at untrusted input. See SECURITY.md.

execute_code_streamA

Execute code and STREAM progress + partial output as it runs. Use this, not execute_code/run_submit/session_run, for the same run when you want output while it runs, up to a 300s cap.

Reports progress notifications to the client while the program runs, so agents can see output before the process finishes. Returns the same result shape and applies the SAME ceilings as execute_code: max_memory_mb, max_output_kb and max_cpu are forwarded to the executor exactly as execute_code forwards them, including the same 240 KiB per-stream clamp.

dependencies: same as execute_code's (PEP 723 merge, no_net/policy refusal, 120s budget, workdir quota) — installed before streaming starts; a refusal/failed install is the stream's only event.

trace_executionA

Debug WHY, line by line, for the ONE input you actually ran it on: which statements fired, in what order, with what variable values at each step, and which if/elif/while/for/try branch was taken versus never taken. Want just the printed output instead? Use execute_code.

Returns events: ordered {step, line, event, func, locals}, one entry per traced line/call/return/exception in YOUR code only (library internals excluded). locals on each entry is only the names that changed since the previous step in that same call — not a full dump every line. A return entry also carries return_value; an exception entry carries exception_type/exception_message.

Also returns branches (hit count per if/elif/while/for/try line), lines_executed / lines_never_executed (coverage from a static parse), and truncated/truncated_reason when max_events or an internal size ceiling stopped RECORDING early (the underlying stdout/exit code are unaffected either way).

TRUST: the trace is produced BY the traced program at its OWN privilege — a debugging aid, not an attestation of behaviour, exactly as trustworthy as that program's own stdout. discarded_events / events_consistent are a best-effort tamper/corruption signal (never a guarantee) computed independently of the file's own content. unenforced may additionally note "only the main thread is traced" (sys.settrace is per-thread) or, fallback backend only, an OLE exit_code race.

For a structural Big-O guess with nothing executed, use analyze_complexity.

branch_reachabilityA

Which if/elif/else arms and while/for loops of this python3 function can ever run, which are dead code, and what inputs reach each — decided with z3, without running the program.

Use trace_execution instead to see what happened on one run. Use z3_check, not this, when you already have an SMT-LIB2 script to solve directly rather than Python source to translate.

Unannotated parameters default to int. Each branch reports verdict (reachable/dead/unknown), a witness when reachable, and boundary_inputs (min/max/equality-edge for each comparison in its own guard) — every input dict is shaped to drop straight into compare_edge_cases's test_inputs. Refuses, naming the construct and line, prior to any z3 call: floats, attribute access, comprehensions, try/except, imports, data-dependent loop bounds, and anything else outside + - * // %, and/or/not, == != < <= > >=, and abs/min/max/len on int/bool/str. A for loop of at most 32 iterations is unrolled exactly; a longer for, or a while, is checked one iteration at a time — a branch can still come back reachable there, but never dead, and anything past the loop that depends on what it computed comes back unknown rather than a guess.

run_submitA

Submit code for BACKGROUND execution; returns a run_id immediately. Use this, not execute_code/execute_code_stream/session_run, when you do not want to hold the call open — poll run_inspect(run_id), and run_cancel(run_id) to stop it early.

Same request shape as execute_code minus session_id (a run is a standalone process, not a session workspace). timeout bounds the WORK itself, not how long you wait to collect it.

This call's own reply carries no output — a small run_id handle — so the anthropic/maxResultSizeChars hint lives on run_inspect instead, which returns the full envelope, same shape execute_code returns, once the run lands.

Admission is capped (CODECALC_MAX_ACTIVE_RUNS, default 64): past that many runs still running/cancelling at once, this returns a resource_exhausted error rather than growing without bound — call run_inspect/run_cancel to make room, or wait for one to finish.

Retention: see run_inspect.

dependencies: same semantics as execute_code's own. A refusal is returned directly with no run created. Otherwise this call still returns immediately: the install itself runs on the background worker, ahead of the code, and a failed install becomes the run's own terminal error — readable via run_inspect(run_id) like any other outcome.

run_inspectA

Poll a background run started with run_submit.

While running: {"ok": True, "state": "running"|"cancelling", "run_id", "provider_id", "started_at", "deadline"}.

Once terminal (state "finished"/"cleaned"/"recovered"), this returns the SAME result shape execute_code returns — stdout/stderr/exit_code/ verdict/unenforced/provider (the interface_version/provider_id/limits receipt)/... — merged with a small set of run_* extras (run_id, provider_id, started_at, deadline, state, cleaned; see server.py's _RUN_EXTRA_KEYS). This terminal reply carries the same anthropic/maxResultSizeChars _meta execute_code advertises (see server.py's _LARGE_RESULT_TOOLS) — it is the same envelope, once the run started with run_submit has finished, and run_submit's own max_output_kb is clamped the same way execute_code's is so that value stays true here too. Read ok and verdict on a terminal result to tell a clean finish from a failure; a run stopped by run_cancel is only reflected there for a provider that actually supports cancellation (see run_cancel's own docstring) — check the result the same way you would any other run.

Retention: a finished run's result stays inspectable for the life of this server process — call this as many times as you like; nothing is consumed by reading it. What IS released on the first terminal read is the PROVIDER's own resources for that run (RunSupervisor.cleanup(), idempotent on repeat calls) — the in-memory record of the run itself is not evicted; there is no cap or TTL on it here, deliberately: the durable state machine, leases and TTL-based eviction are out of this residual's scope (see run_supervisor.py's own docstring). A long-lived server that calls run_submit very many times will grow this table; the on-disk crash-recovery journal underneath it is already bounded (RunSupervisor.max_completed), independent of this.

run_cancelA

Cancel a background run started with run_submit.

Idempotent: calling this on a run that is already finished/cleaned reports cancelled: false, state: <its actual terminal state> rather than erroring — matching execute_code's own "no partial result" rule, there is nothing partial to hand back either way.

Propagation depends on the SELECTED PROVIDER (see list_execution_providers' cancel capability). The built-in local provider does not support stopping a run once it has started; that is reported honestly here rather than silently pretended to have worked — the computation keeps running to completion and its result stays available via run_inspect, so bound it in advance with run_submit's own timeout instead. A provider that DOES advertise cancel: true reaches the full spawned process tree the same way execute_code's own cancellation does — RunSupervisor already owns that; this tool only calls it.

evaluate_expressionA

Use evaluate_expression, not calc_exact, for something other than plain arithmetic on literal values. Symbolically evaluate to a value or closed form via sympify: 'integrate(x2, x)', 'sqrt(144) + 210'. Not simplification — for simplified/factored/expanded forms, use symbolic(op="simplify"). Returns value (if the result is a number) or the evaluated expression, plus type.

truth_tableA

Build the truth table for a boolean expression over and/or/not/xor/ implies/iff (plus true/false constants and variables): 'a and b or not c', 'p xor q', 'a implies b'. Use z3_check, not this, for satisfiability over inequalities or non-boolean variables; use evaluate_expression for symbolic (non-boolean) math.

Returns variables (sorted names) and rows (one dict per assignment, each variable name -> bool plus result), plus row_count, satisfiable (any row true), and tautology (every row true).

z3_checkA

Use z3_check, not symbolic(op="solve"), for satisfiability over inequalities, boolean combinations, or several variables at once: sat/ unsat/unknown plus a model. Example: '(declare-const x Int)(assert (> x 5))(check-sat)'.

unsat is graded solver_proven — see grade_basis for the engine version and timeout bound it was decided within. sat is graded ungraded: it's a real decided answer, just not a proof — reserving solver_proven for unsat means a counterexample can never wear a proof grade. unknown carries no proof either way and is also graded ungraded.

matrixA

Structured matrix operations: det, inverse, eigenvalues, transpose, rank, trace.

evaluate_expression refuses Matrix([[1,2],[3,4]]) on purpose — [/] are denied there to block subscript-based RCE escapes, and a matrix literal is collateral from that (correctly aimed) screen. This tool is the structured replacement: rows is a JSON array of arrays (row-major), never a string to parse. Each entry is either a JSON number, used directly, or a scalar expression string ('1/2', 'sqrt(2)', 'x+1'), screened per-entry the same way evaluate_expression screens its input before anything reaches SymPy. Example: rows=[[1,2],[3,4]], op='det' -> -2.

analyze_complexityA

Estimate the asymptotic (Big-O) time complexity of a code snippet via structural analysis.

benchmarkA

Empirically measure time complexity by running code at increasing input sizes.

Contract: the code must read an integer N from stdin (first line) and do work sized by N. codecalc runs it at each size in sizes and fits the growth curve to estimate Big-O (O(1), O(log n), O(n), O(n log n), O(n^2)...). Example python: 'import sys\nn=int(sys.stdin.readline()); s=0\nfor i in range(n): s+=i\nprint(s)'

compare_executionA

Run the same code in multiple languages side by side.

Returns per-language stdout/stderr/exit/duration plus which was fastest. Example: {"python3": "print(67)", "node": "console.log(67)"}

This tool fans out across every language with no per-language install plumbing behind it; use install_package/execute_code(dependencies=...) beforehand instead. A # /// script block in a snippet is likewise never installed, but is DISCLOSED, not dropped: a python3 row that carries one gets dependencies: {"status": "unsupported", "reason": ...}.

runtimes_statusA

Check every language runtime for available updates (NON-MUTATING).

Reports current vs latest version per language, which package manager owns it (mise/rustup/swiftly/apt/npm/uv), and the exact command that would run.

Each entry also carries tier (registry.RELIABILITY_TIERS) — see list_languages for what tested/best_effort/plan_only mean. A current, up-to-date toolchain can still be best_effort: tier is orthogonal to whether the update check below found a newer version.

NETWORK: yes. Non-mutating refers to this machine's runtimes, not to traffic — each package manager is asked what the latest version is, and they answer by contacting their own remote index.

update_runtimesA

Update language runtimes. SAFE BY DEFAULT: with apply=False this is a dry run — it returns the update commands that WOULD run without changing anything. Pass apply=True to actually execute them (mise up, rustup update, swiftly update, apt-get upgrade of language packages, npm -g update, uv tool upgrade).

apply=True asks the caller to confirm first (a protocol-level gate, not just the anthropic/requiresUserInteraction _meta hint — see codecalc/confirmation.py); apply=False is never gated, since nothing runs.

PRIVILEGE: the apt manager updates system packages and its command begins with sudo. Those commands do NOT run unless the HOST has set CODECALC_ALLOW_RUNTIME_APPLY=1; without it they are reported as skipped with ok: false and the variable named, and the rest still run. Every entry carries an elevated flag either way. mise/rustup/swiftly/npm/uv touch user-owned toolchains and are never gated.

NETWORK: yes, on both paths. apply=False still asks each manager what the latest version is, which is a remote lookup; apply=True additionally downloads and installs. "Dry run" bounds what changes on disk, not what is sent.

session_read_fileA

Read a file from a session workspace.

Text files return content. With as_image=True (or for image files), the file is returned as an inline image the model can see. Use session_files to discover paths; session_artifacts lists what executed code produced.

session_runA

Run a multi-file program already written into a session workspace (via session_write_file). Use this, not execute_code/execute_code_stream/run_submit, when entry_file may import other files already in that workspace (helper.py, data/...).

Runs as a fresh process in the session workdir (not the REPL worker), so relative imports and data files resolve. Returns stdout/stderr/verdict plus the entry file's path. Oversized output spills into the session workspace the same way execute_code's does — see its docstring for stdout_spill/stderr_spill.

Reports artifacts_created and inlines small ones as extra content blocks (image/text/link), capped at 8 blocks / 4 MiB encoded; truncated_inline: true past either cap. dependencies installs packages before running, same rule as execute_code's — see its docstring.

This tool takes no max_output_kb (its inline stdout/stderr stay at the 64 KiB default and spill past that, same as execute_code's session branch) but the anthropic/maxResultSizeChars _meta it advertises covers only that text envelope — the JSON result serialized as the reply's content text block. The inlined artifact blocks above (image/text/link, up to 8 of them within the 4 MiB encoded budget) are SEPARATE MCP content blocks, outside the text block this hint bounds.

Every run copies entry_file's own source into the runner's private scratch subdirectory before executing it — never into a root-level main.<ext> file a session's own files could collide with. A session's own main.py (or the equivalent for another language) at the session root is never touched by running a different entry file.

convert_unitsA

Convert a value between units (dimensional analysis via sympy).

Supports metric/imperial length, mass, time, speed, energy, power, force, pressure, temperature (°C/°F/K), volume, area, data sizes, frequency. Examples: ('60','mph','km/h'), ('100','celsius','fahrenheit'), ('1','gb','mib'). Use list_units for the full alias table.

physical_constantsA

Look up a physical constant (speed_of_light, planck, avogadro, gravity, electron_mass, gas_constant, ...) or list all 22 with values.

list_unitsA

List every supported unit alias — all aliases and spellings — for convert_units.

calc_exactA

Use calc_exact, not evaluate_expression, for a literal arithmetic expression with no symbols in it. EXACT arithmetic: 0.1 + 0.2 == 0.3 is True here (False in plain Python).

Everything is an exact rational, integers are arbitrary precision. Supports

      • / // % ** comparisons, bitwise ops (& | ^ << >> ~) on integers, and whitelisted math functions (sqrt, log, sin, ...) plus pi/e/tau. Use BEFORE asserting any computed number: thresholds, ratios, overflows, 'X is N% of Y'. Examples: '2**64 - 1', 'comb(52,5)', '0.1+0.2 == 0.3', '0xff & 0x0f'.

compare_thresholdA

Exact threshold check with a verdict and the shortfall when it fails. Use calc_exact, not this, when you want the computed VALUE rather than a threshold comparison.

a OP b. Both sides are evaluated exactly and printed as fractions — a threshold comparison written out cannot be gotten backwards. Example: ('1/25', '>', '0.05').

percentageA

Exact share and percentage of PART / TOTAL. Use calc_exact for a single arithmetic expression, or compare_threshold to check the result against a threshold rather than just compute it.

calc_statsA

Mean, median, sample stdev, and coefficient of variation (CV) for a sample of numbers. Pairs with percentiles for distribution shape (p50/p90/p95/p99) on the same sample, and with benchmark or verify_optimization, which are common sources of the timing samples this tool summarizes. CV > 0.2 flags run-to-run noise that swamps the effect. Returns n/mean/median/stdev/cv plus a cv_note.

percentilesA

p50/p90/p95/p99 (the 50th/90th/95th/99th percentile cutoffs) by nearest-rank AND linear interpolation. Pairs with calc_stats, which gives mean/median/stdev/CV on the same sample instead of these distribution points.

Warns when n < 100 that p99 is just the maximum wearing a label.

collision_probabilityA

Birthday-bound hash collision probability: 1 - exp(-n^2 / (2*2^b)).

Sizes hashes: 1e6 items into 64 bits is ~2.7e-8; 1e5 into 32 bits is ~0.69 — the answer to 'can I truncate this to 8 hex chars?' (no).

data_sizesA

Byte counts for a plain integer, both binary and decimal — the gap between them is where '291 MB' and '277 MiB' silently disagree by 5%. For units other than bytes, use convert_units. For a duration, not a byte count, use human_duration. Returns bytes plus binary and decimal dicts of unit -> value.

human_durationA

Convert a SPAN of elapsed seconds into a humanised duration (e.g. '2d 3h 4m 5s') plus per-day and per-30d rates. For an epoch timestamp to a calendar date, use epoch_time instead. For byte counts, not seconds, use data_sizes. Returns human, per_day, per_30d, and the echoed seconds.

epoch_timeA

Epoch seconds/millis/micros/nanos to ISO 8601 UTC (implausible readings suppressed).

bitsA

Programmer-mode integer facts and operations, selected by mode — replaces the four former standalone tools bit_analysis, bitop, int_widths and base_repr, retired in 0.12.0 (CHANGELOG.md). Every mode returns exactly its former tool's own result, plus mode (additive).

mode="analysis" (was bit_analysis) — facts about a single N: popcount, bit length, trailing zeros, power-of-two check, next power of two. Used by this mode: n (required), align (optional).

mode="op" (was bitop) — combine two integers a/b with and/or/xor/nand/nor/xnor/not/shl/shr/sar/rol/ror at a fixed width (8/16/32/64). Every result shows unsigned, signed (two's complement), hex, octal and binary. shr is logical (zero-fill); sar is arithmetic (sign-propagating) — 0x80 shr 1 = 0x40 (+64) but 0x80 sar 1 = 0xC0 (-64); rol/ror rotate bits around the width instead of shifting them out. A left shift that drops bits says OVERFLOW and shows the unbounded answer. Used by this mode: a, op (required), b (required unless op="not"), width (optional).

mode="widths" (was int_widths) — which widths (i8..i64/u8..u64) hold n, and the wrapped value where they do not; flags anything past 2^53 as unable to round-trip through a JS number or JSON float. Used by this mode: n (required).

mode="repr" (was base_repr) — hex/oct/bin of n; with width, two's complement and signed-overflow detection. Used by this mode: n (required), width (optional).

radix_convertA

Convert a value between ANY bases 2..36, fractions included; bases that cannot represent the fraction (e.g. 0.1 in base 2) are flagged non-terminating. radix_convert('zz', 36, 7) is one call.

float_reprA

What binary64 actually stores for X: exact value, raw bits, ULP, both neighbours, and whether the literal is representable. float_repr(0.1) shows 0.1000000000000000055511151231257827...; float_repr(0.25) says EXACT. Above 2^53 warns consecutive integers are indistinguishable.

algebraic_equivA

Are two expressions algebraically identical? 'is (ab)/c the same as a(b/c)?' answered exactly. Use symbolic(op="simplify"), not this, to see one expression's own simplified/factored/expanded forms rather than compare two; use verify_translation to compare running PROGRAMS, not expressions. Caveat: symbolic identity says nothing about float rounding, integer truncation or modular overflow.

symbolicA

Symbolic algebra, selected by op — replaces the four former standalone tools solve_expression, solve_linear, simplify_expression and limit_expression, retired in 0.12.0 (CHANGELOG.md). Every op returns exactly its former tool's own result, plus op (additive).

op="solve" (was solve_expression) — the roots of one equation: 'x**2 - 4 = 0', '2*x + 1 = 7'. For a system of several equations, use op="solve_linear". For general constraint satisfiability (inequalities, boolean constraints, multiple solvers), use z3_check. Returns solutions as a list of strings alongside the parsed equation and variable. Used by this op: expr (required), var (optional).

op="solve_linear" (was solve_linear) — a system of equations sharing variables. Example: system='x + y = 10; x - y = 2', variables='x, y'. Used by this op: system, variables (both required).

op="simplify" (was simplify_expression) — simplify, factor, and expand an expression — algebraic forms, not solving (use op="solve") and not a numeric value (use calc_exact). Returns simplified, factored, and expanded as strings alongside the parsed original. Used by this op: expr (required).

op="limit" (was limit_expression) — asymptotic behaviour: limit of expr as var -> point. 'symbolic("limit", "n*log(n)/n**2", "n")' returns 0 — settles complexity arguments faster than arguing. Used by this op: expr (required), var (optional), point (optional).

verify_translationA

PROVE that a port is equivalent: run both programs, compare their output.

You write the translation — you are the language model. This runs your source and your port on the same inputs and reports, per input, whether they matched, diverged, or could not be compared (a runtime that is missing or a program that failed on both sides is INCONCLUSIVE, never a pass).

Use it after porting anything: python3 -> go, node -> rust, a rewritten function against the original. Pair with compare_edge_cases to find the inputs worth testing.

Matching tolerates only line-ending/trailing-whitespace noise; stdout_raw carries what actually ran.

A pass is graded cross_checked (two independent implementations, run and agreeing — see grade_basis for which runtimes). A non-pass is graded ungraded: never a softer positive grade.

compare_edge_casesA

Run the same logic in N languages on edge-case inputs and flag divergence.

Default inputs cover empty, zero, negative, and float-precision cases: ['', '0', '1', '-1', '10', '100', '0.1\n0.2']. Returns a per-input matrix plus a divergences list where languages disagree on identical input.

verify_optimizationA

PROVE an optimisation: same outputs, measurably AND SIGNIFICANTLY faster.

Two gates, in order. Correctness: runs candidate against original on shared inputs — a faster-but-wrong candidate fails here and is never timed. Speed: times both at increasing sizes; accepts only when the median ratio clears min_speedup AND a one-sided Mann-Whitney U test rejects "not faster" at every counted size (2-3), or a Bonferroni-corrected majority above that — one size never accepts alone. See inference for the per-size U statistic, p-value, effect size.

A rejection names which gate failed and by how much, e.g. "correct, 1.3x median, but only 1/4 sizes significant."

Accepted grades cross_checked; any rejection — wrong, not faster enough, not significant — grades ungraded: correctness alone earns no grade for the speed claim this tool answers.

extract_functionA

Extract a named function (with its imports + referenced helpers) into a standalone program and run it in the sandbox.

python3 gets exact ast extraction; other languages best-effort block extraction (pass call to execute non-python). Returns the extracted program and per-input runs.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
verify_translation viewInteractive per-case comparison table for a verify_translation result, with first-differing-line highlighting. Ignored by hosts without MCP Apps support.
verify_optimization viewInteractive per-size timing chart and significance table for a verify_optimization result. Ignored by hosts without MCP Apps support.

TDQS

A3.9/5.0

Scored across 49 tools

Disambiguation3/5

Many tools are clearly distinct (e.g., session_*, run_*, symbolic vs calc_exact), but there is notable overlap among execution tools: execute_code, execute_code_stream, run_submit, session_run, and compare_execution all run code with different modes, and several math tools (calc_exact, evaluate_expression, symbolic, calc_stats, percentiles) have adjacent purposes. The descriptions are detailed enough to disambiguate with careful reading, but the sheer number of similar execution/math tools creates real selection risk.

Naming Consistency4/5

Most tools follow a clear verb_noun pattern (execute_code, session_start, run_inspect, convert_units, verify_optimization). There are minor deviations: 'bits' and 'symbolic' are noun/adjective-only names, and 'matrix' is a single noun, but these are documented as mode-based consolidated tools. Overall the pattern is consistent and predictable.

Tool Count2/5

49 tools is a very large surface for a code calculation/execution server. While the server covers many domains (execution, sessions, math, units, verification, runtimes), the count is heavy and includes several consolidated mode-based tools that could reduce the count further. It exceeds the typical well-scoped range and will burden agent tool selection.

Completeness4/5

The server covers its apparent domains thoroughly: code execution (sync, stream, background, session), file/session management, math/units/constants, verification (optimization, translation, edge cases), and runtime management. Minor gaps exist (e.g., no explicit session_run cancellation, no direct file deletion tool), but the core workflows are well covered and there are no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessResponsive