Python Code Validator
Server Details
Proves AI-generated Python does what you asked: lint, types, security, sandbox run, exact fixes.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- jkanselaar/python-code-validator
- GitHub Stars
- 0
- Server Listing
- Python Code Validator
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: validate only diagnoses, repair diagnoses and fixes, execute diagnoses, fixes, and runs. The descriptions explicitly state the differences and alternatives, leaving no ambiguity about which to choose.
All tools follow the same verb_noun pattern: validate_python, repair_python, execute_python. The naming is perfectly consistent and predictable.
Three tools is a well-scoped set for a Python code validator. Each tool adds a distinct level of functionality (diagnose, fix, run), and there are no redundant or unnecessary tools.
The toolset covers the full lifecycle of Python code validation: diagnose (validate), fix (repair), and verify (execute). The options within the tools (e.g., transpile, optimize, examples) further round out the surface, leaving no critical gaps.
Available Tools
3 toolsexecute_pythonExecute PythonAIdempotentInspect
Everything repair does, and then RUNS the code in a throwaway container — no network, read-only filesystem, killed at options.timeout_s — reporting exit code, stdout and stderr. Any '>>>' examples in the code are run too, and one that does not print what it says is an error the other tools cannot see. This is a side effect: do not submit code you do not want executed. Use it when you need proof that the code runs, or that it does what it says. Alternatives: validate_python for the diagnosis and repair_python for the fix, neither of which runs anything. Auth: a key is required. This call needs a paid key and answers HTTP 402 without one. Credits are bought without an account, 10 per call: GET /v1/pricing says where to send the xDAI. Or pay for this one call with no key at all: call it without one and the result carries x402 payment requirements ($0.1 in USD Coin on eip155:8453); sign them and repeat the call with the payment in _meta['x402/payment']. Arguments: code: the whole file, 1..200000 bytes of UTF-8 measured after encoding (empty is refused with 400, larger with 413); a fragment is fine, but line and column numbers in the answer count from 1 in what you sent. language: must be 'python'; anything else is 400, and the field may be omitted. options.max_iterations (1..10, default 3) caps the fix/verify rounds: raise it for a file with several independent faults, leave it for a snippet. options.optimize (default false) additionally folds constants and drops dead code, and is only worth setting when you asked for a rewrite anyway. options.transpile_to (e.g. 'javascript') returns a translation of the repaired source in transpiled, not of what you sent. fixed_code is null when nothing could be proven safe to change, so treat null as 'no fix', not as an error. options.timeout_s (seconds, default 5) is the wall clock for the run; the schema allows up to 60 but this deployment caps it at 30 and refuses a larger value with 400. options.expected_output compares stdout byte for byte and adds an 'expected-output' diagnostic (valid=false) when it differs, which is how you ask for 'it did the right thing' rather than 'it ran'. options.examples is the same question for code with no output: pass what you asked for as doctest lines ('>>> total([1, 2])' then '3') or assertions ('assert total([1, 2]) == 3'), and each is run against the code -- one that does not hold is a 'python:example-mismatch' error, and repair looks for a single-token change that makes them all pass. Send it whenever you know what you asked for: without it, code that runs but returns the wrong answer looks perfect from here. The program that runs is the repaired one, so read fixed_code before you trust runtime.stdout, and it runs exactly once however many rounds the repair took. Returns valid, score 0..1, diagnostics (rule, message, line, column), security findings, fixes, fixed_code and runtime; see outputSchema. The code and its verdict are retained to improve the service.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The source to check, as a whole file where possible: diagnostics carry the line and column of the text you send, and a fragment hides the imports and definitions the type check needs. A deployment may accept fewer bytes than the 200000 here. | |
| options | No | Tuning knobs. Most of them only take effect in the mode that does the corresponding work; see each field. | |
| language | No | The language of the code. A service that does not handle it refuses the request rather than guessing; the enum is shared across services, so it lists more than any one of them accepts. | python |
Output Schema
| Name | Required | Description |
|---|---|---|
| meta | Yes | |
| fixes | No | |
| score | Yes | |
| valid | Yes | |
| runtime | No | |
| security | No | |
| fixed_code | No | |
| transpiled | No | |
| diagnostics | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It adds critical context beyond the annotations: the code is actually executed, so malicious or unwanted code is a side effect; the container is no-network, read-only, and killed at timeout; the repaired code is what runs; and data is retained to improve the service. It also discloses auth and x402 payment behavior. No explicit contradiction with the annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but it is structured and front-loaded: it starts with what the tool does, then covers side effects, alternatives, auth, parameters, and returned values. The length is justified by the complexity of a side-effecting, execution tool with payment requirements and repair semantics, though some parameter details are summarized the schema already contains.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description, together with the output schema, provides everything needed to choose and call this tool correctly: behavioral constraints, alternatives, auth flow, argument semantics, edge cases like null fixed_code, and return shape. The agent is not left to guess or infer critical behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema description coverage is 100%, the description explains practical operational semantics: code is measured in encoded UTF-8 bytes, language is restricted to Python despite the schema enum, the deployment caps timeout_s at 30, and expected_output/examples are the mechanism for proving correctness rather than just execution. This guidance is not inferable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource: run the submitted code in a throwaway container and report exit code, stdout, and stderr. It also orients the agent by explaining that the tool does everything repair_python does and then executes the code, which clearly separates it from validate_python and repair_python.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent exactly when to choose this tool: when proof of execution is needed, such as 'that it runs or that it does what it says'. It explicitly names the alternative tools, validate_python and repair_python, and notes neither runs anything, which prevents selection mistakes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repair_pythonRepair PythonARead-onlyIdempotentInspect
Everything validation does, plus deterministic fixes: the corrected source comes back in fixed_code, and the original is kept whenever the fix cannot be proven safe. The code is still never run. Use it when validation failed and you want the fix rather than the diagnosis. Alternatives: validate_python when the diagnosis is enough; execute_python when the fix has to be proven to run. Auth: a key is required. This call needs a paid key and answers HTTP 402 without one. Credits are bought without an account, 3 per call: GET /v1/pricing says where to send the xDAI. Or pay for this one call with no key at all: call it without one and the result carries x402 payment requirements ($0.03 in USD Coin on eip155:8453); sign them and repeat the call with the payment in _meta['x402/payment']. Arguments: code: the whole file, 1..200000 bytes of UTF-8 measured after encoding (empty is refused with 400, larger with 413); a fragment is fine, but line and column numbers in the answer count from 1 in what you sent. language: must be 'python'; anything else is 400, and the field may be omitted. options.max_iterations (1..10, default 3) caps the fix/verify rounds: raise it for a file with several independent faults, leave it for a snippet. options.optimize (default false) additionally folds constants and drops dead code, and is only worth setting when you asked for a rewrite anyway. options.transpile_to (e.g. 'javascript') returns a translation of the repaired source in transpiled, not of what you sent. fixed_code is null when nothing could be proven safe to change, so treat null as 'no fix', not as an error. options.timeout_s, options.examples and options.expected_output do nothing here: nothing is run, so there is no clock, no stdout, and no way to check an example. Returns valid, score 0..1, diagnostics (rule, message, line, column), security findings, fixes, fixed_code and runtime; see outputSchema. The code and its verdict are retained to improve the service.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The source to check, as a whole file where possible: diagnostics carry the line and column of the text you send, and a fragment hides the imports and definitions the type check needs. A deployment may accept fewer bytes than the 200000 here. | |
| options | No | Tuning knobs. Most of them only take effect in the mode that does the corresponding work; see each field. | |
| language | No | The language of the code. A service that does not handle it refuses the request rather than guessing; the enum is shared across services, so it lists more than any one of them accepts. | python |
Output Schema
| Name | Required | Description |
|---|---|---|
| meta | Yes | |
| fixes | No | |
| score | Yes | |
| valid | Yes | |
| runtime | No | |
| security | No | |
| fixed_code | No | |
| transpiled | No | |
| diagnostics | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds critical behavior beyond annotations: code is never run, fixes are only returned when provably safe, fixed_code is null when no fix is possible, and authentication/payment requirements are disclosed. It also states which options are ignored because nothing executes. Annotations already indicate read-only and non-destructive, and the description does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but logically organized by sections (behavior, use cases, auth, arguments, returns). It front-loads the core distinction from siblings and then systematically covers constraints. It is somewhat verbose, especially around payment details, but every section contributes decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, and the description still explains the null fixed_code semantics, return fields, auth failure behavior, and constraint enforcement. Nothing an agent needs to call this tool correctly appears to be missing, including what happens with unsupported options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description enriches every parameter: it details byte limits for code, explains that fragments shift line/column numbers, states language must be 'python' and other values elicit 400, and clarifies that options like timeout_s, examples, and expected_output have no effect in this mode. This goes far beyond raw schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: repair Python code, returning deterministic fixes in fixed_code and preserving the original when safety cannot be proven. It also explicitly contrasts itself with the sibling tools validate_python and execute_python, so an agent can distinguish it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit selection criteria: use when validation failed and you want a fix, not just a diagnosis. It also names the alternatives ('validate_python when the diagnosis is enough; execute_python when the fix has to be proven to run').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_pythonValidate PythonARead-onlyIdempotentInspect
Check Python source without running it: parse, lint (ruff), type-check (mypy), AST security policy, credential scan. Safe on code you do not trust. Use it on every Python file you generated or edited, before writing it to disk. Alternatives: repair_python to get the corrected source instead of the diagnosis; execute_python to prove the code runs. Auth: a key is required. A free key covers this call, 25 per day, then HTTP 429; get one with POST /v1/keys. Credits are bought without an account, 1 per call: GET /v1/pricing says where to send the xDAI. Or pay for this one call with no key at all: call it without one and the result carries x402 payment requirements ($0.01 in USD Coin on eip155:8453); sign them and repeat the call with the payment in _meta['x402/payment']. Arguments: code: the whole file, 1..200000 bytes of UTF-8 measured after encoding (empty is refused with 400, larger with 413); a fragment is fine, but line and column numbers in the answer count from 1 in what you sent. language: must be 'python'; anything else is 400, and the field may be omitted. Of options only transpile_to (e.g. 'javascript', which returns a translated copy in transpiled) acts here; timeout_s, max_iterations, optimize, examples and expected_output need a pass that rewrites or runs the code, so send code alone. Ignored options are not refused, so a call that sets them looks like it worked; and code that does not parse is answered rather than refused: valid=false with the syntax error located, which is the point. Returns valid, score 0..1, diagnostics (rule, message, line, column), security findings, fixes, fixed_code and runtime; see outputSchema. The code and its verdict are retained to improve the service.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The source to check, as a whole file where possible: diagnostics carry the line and column of the text you send, and a fragment hides the imports and definitions the type check needs. A deployment may accept fewer bytes than the 200000 here. | |
| options | No | Tuning knobs. Most of them only take effect in the mode that does the corresponding work; see each field. | |
| language | No | The language of the code. A service that does not handle it refuses the request rather than guessing; the enum is shared across services, so it lists more than any one of them accepts. | python |
Output Schema
| Name | Required | Description |
|---|---|---|
| meta | Yes | |
| fixes | No | |
| score | Yes | |
| valid | Yes | |
| runtime | No | |
| security | No | |
| fixed_code | No | |
| transpiled | No | |
| diagnostics | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already mark readOnlyHint=true and destructiveHint=false, the description goes far beyond them: it promises safety on untrusted code ('Safe on code you do not trust'), discloses that it never runs code, and highlights silent no-op options ('Ignored options are not refused, so a call that sets them looks like it worked'). It also reveals rate limits and auth behavior (HTTP 429, key requirement, x402 payment flow) and the fact that submitted code is retained to improve the service. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place. It is clearly structured with labeled sections (Auth, Arguments, Returns) and front-loaded with the core purpose before diving into details. No filler or redundant repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description does not need to enumerate return fields in depth; it still gives a high-level list. It covers input constraints, authentication requirements, call-limits, error behaviors, sibling tool relations, and retention. An agent has everything necessary to call the tool safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema already covers parameters 100%, the description adds critical meaning: it notes empty/large code fails with HTTP 400/413, line/column numbers count from the submitted fragment, the language must be exactly 'python' despite the broad schema enum, and that only transpile_to among options actual effects in this static tool. It explains that other options are ignored rather than rejected, which is not communicated by the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Check Python source without running it' followed by an explicit checklist (parse, lint/ruff, type-check/mypy, AST security policy, credential scan). It explicitly distinguishes the tool from its siblings: 'repair_python to get the corrected source instead of the diagnosis; execute_python to prove the code runs.' No ambiguity remains about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states exactly when to use this tool ('Use it on every Python file you generated or edited, before writing it to disk') and clearly names alternatives and their purpose. The when-not-to-use is implicit but clear: if you need a corrected file, use repair_python; if you need runtime proof, use execute_python. This is explicit enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Changed
execute_python1 field changed- added
Input schema / $defs / Options / properties / examplesAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "What the code is supposed to do, as doctest lines ('>>> f(2)' on one line, '4' on the next) or as plain assertions ('assert f(2) == 4'). In execute mode they are run in the sandbox: an example that does not hold is a 'python:example-mismatch' error and makes the response invalid, and repair searches for a single-token change that makes every one of them pass. This is the only way the service can tell code that runs from code that is right, so send it whenever you know what you asked for. Examples already written in the code ('>>> ' in any string) are used the same way without this option. Ignored in the other modes, which run nothing.", + "title": "Examples" +}
- Changed
repair_python1 field changed- added
Input schema / $defs / Options / properties / examplesAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "What the code is supposed to do, as doctest lines ('>>> f(2)' on one line, '4' on the next) or as plain assertions ('assert f(2) == 4'). In execute mode they are run in the sandbox: an example that does not hold is a 'python:example-mismatch' error and makes the response invalid, and repair searches for a single-token change that makes every one of them pass. This is the only way the service can tell code that runs from code that is right, so send it whenever you know what you asked for. Examples already written in the code ('>>> ' in any string) are used the same way without this option. Ignored in the other modes, which run nothing.", + "title": "Examples" +}
- Changed
validate_python1 field changed- added
Input schema / $defs / Options / properties / examplesAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "What the code is supposed to do, as doctest lines ('>>> f(2)' on one line, '4' on the next) or as plain assertions ('assert f(2) == 4'). In execute mode they are run in the sandbox: an example that does not hold is a 'python:example-mismatch' error and makes the response invalid, and repair searches for a single-token change that makes every one of them pass. This is the only way the service can tell code that runs from code that is right, so send it whenever you know what you asked for. Examples already written in the code ('>>> ' in any string) are used the same way without this option. Ignored in the other modes, which run nothing.", + "title": "Examples" +}
3 tool updates
- First observed
execute_python - First observed
repair_python - First observed
validate_python
Related MCP Connectors
Production-safety audits for AI-generated code, with a fix for every finding.
Deterministic validation for AI-generated artifacts: JSON Schema, OpenAPI response, SQL syntax.
Validates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes.
Lints + auto-fixes how AI coding agents discover any new product. 24 rules, 6 tools, score 0-100.
Related MCP Servers
- AlicenseAqualityDmaintenanceValidates AI-generated code against actual codebases to catch hallucinations, dead code, and API mismatches before runtime.123 npm1MIT
- AlicenseNot gradedqualityBmaintenanceRuns, tests, and finds issues in your Python services with zero code changes, then helps your AI agent fix what breaks and proves it with acceptance tests.46Apache 2.0
- AlicenseAqualityCmaintenanceAI-powered characterization test generator that reads Python functions or class methods, synthesizes inputs, captures behavior in a sandbox, and emits pytest files to lock legacy code behavior for safe refactoring.4Apache 2.0
- FlicenseNot gradedqualityDmaintenanceProvides a secure, containerized Python sandbox for executing LLM-generated code with multi-layer isolation, along with JSON/CSV validation and workspace state snapshots.-
Glama MCP Gateway
Add one secure layer between your agents and this server.