Skip to main content
Glama
Axiomatic-AI

axiomatic-mcp

Official
by Axiomatic-AI

AxArgmin_execute_code

Execute Python code in a sandboxed environment with numerical libraries, return exported results, and include verification certificates and diagnostics for solver outputs.

Instructions

Execute Python code in a sandboxed environment with numpy, math, and the ax_core.argmin numerical library available. Code must call export(name, value) at least once to return results. Typically used to run code produced by the generate_code tool, but also accepts hand-written or modified code.

success reports only that the code ran. The Verification: line leading the response is the verdict on whether the answers are solved; read it first. The response also carries the exports and, from a backend that supports it, a verification payload holding the certificate and diagnosis of every exported solver result — both as structured content and as a JSON text block.

Reading verification for the detail behind the verdict line:

  • _summary.all_passed is an input to that line, not a substitute for it. The line reports a pass only when the summary's four counts are all present, readable and adding up, at least one solve was counted, none was counted as failed or unverified, AND the per-export entries show a successful solver_success behind every solve counted. So a payload claiming a pass over nothing checked, beside a non-zero n_failed/n_unknown, with counts that cannot be read, or on certificates alone is reported as unverified. Where the two disagree, the line wins.

  • a certificate is not a solver verdict. An export carrying a bare certificate — the multistart idiom export('best_certificate', best['certificate']) — gets Verification: certificate only: the certificate passed at the point returned, but a failed solve's certificate can pass there too, so nothing says the solve converged. Export the result object, or the whole record {'success': ..., 'status': ..., 'certificate': ...}, to get the verdict as well.

  • _warnings names what did not check out, and by how much, per export. It may also carry an advisory that does not bear on the verdict — an export name colliding with a reserved key, say — so a warning is not by itself a failure.

  • each per-export entry carries the certificate (the KKT / residual / integration-accuracy check re-evaluated at the point actually returned) and, on failure, a diagnosis whose kind names the failure class and whose suggestion says what to change. Pass those on rather than only that it failed.

  • to get a certificate back at all, the code only has to export the result object itself (export('result', result)); the certificate and diagnosis travel with it.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
codeYesPython code to execute. Must call export(name, value) to return results.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.17

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility and does an excellent job. It explains that success only means the code ran, that the Verification line is the verdict, details the verification payload structure, warnings, certificate semantics, and how to get certificates. It even clarifies that a certificate alone is not a solver verdict, making behavior highly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but lengthy, spanning multiple paragraphs with extensive detail about verification payloads and edge cases. While well-structured and logically ordered, it is not concise; many sentences could be condensed without losing essential guidance, though the complexity of the tool may justify some verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, a simple one-parameter schema, no annotations, and no output schema, the description is exceptionally complete. It covers execution requirements, output interpretation (success vs verification), the structure of the verification payload, warnings, certificates, and how to obtain them. Nothing an agent needs to correctly use the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single 'code' parameter with 100% coverage, including the requirement to call export(). The description adds context about typical usage (with generate_code) and output interpretation, but does not meaningfully enhance the parameter semantics beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Execute Python code') and the specific resource (sandboxed environment with numpy, math, and ax_core.argmin). It distinguishes itself from other execute_code siblings by naming the argmin library and mentions its typical pairing with generate_code, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it is typically used to run code from the generate_code tool but also accepts hand-written code, which gives clear context for when to use it. It doesn't explicitly exclude other execute_code tools, but the library-specific scope strongly implies it's for argmin-related code, so the guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools