Skip to main content
Glama

audit_coding

Use this before committing to a coding decision you've drafted — a proposed diff before merging, an architectural choice before adopting, a security assessment before signing off, a migration plan before scheduling. Four frontier models stress-test it for blind spots, then revise their critiques in light of specific counter-positions. Returns a structured critique with severity tags and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. Pass the REAL diff text, never a summary or an abridged version — elision manufactures findings about what was elided. LARGE DIFFS: past roughly 50 KB, split into 2-3 calls chunked by file or section instead of one oversized call — a single huge diff can cut off the panel's structured verdict (you'll see degraded: true with degraded_reason=assessment_parse_failure and an empty findings list). Input fields over 1 MB are clipped with a notice in the response. For a fast broad take use synthesize_coding; for an open-ended decision use deliberate_coding.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
diffNo
testsNo
gate_diffNo
gate_repoNo
constraintsNo
eval_case_idNo
relevant_codeNo
gate_context_idNo
proposed_actionNo
continuation_tokenNo
target_hunk_hashesNo
architectural_contextNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and discharges it thoroughly: it discloses the ~2-5 min runtime with no mid-call progress, the degraded:true failure signature (degraded_reason=assessment_parse_failure with an empty findings list) triggered by oversized diffs, the 1 MB input clipping notice, and the return shape (structured critique with severity tags plus a recommended action class). This is effectively a self-contained annotation layer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description runs long but every sentence is operational, and it is tightly sequenced: triggering scope, mechanism, return value, timing with a user-facing instruction, input fidelity requirement, size-limit handling, degradation signature, then alternative routing. Information is front-loaded (the when-to-use scoping comes first) and the length is justified by the tool's complexity — 12 parameters, a long runtime, and multiple failure modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with no output schema and 0% schema coverage, the description covers the core workflow well — what it returns (critique with severity tags, action class), when it degrades, how long it runs. The critical gap is parameter breadth: 11 of 12 parameters remain semantically unexplained, and no output schema exists to compensate. The main agent-facing path (passing a diff) is complete, but the broader parameter surface is not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 12 parameters, so the description must compensate — but it only documents one: `diff`. That parameter is handled excellently (pass the REAL text, never a summary; split past 50 KB; clipping behavior), yet the other 11 parameters (tests, gate_diff, gate_repo, constraints, eval_case_id, relevant_code, gate_context_id, proposed_action, continuation_token, target_hunk_hashes, architectural_context) have zero semantic explanation in either the schema or the description. Several are technical (continuation_token, target_hunk_hashes) and entirely opaque to an agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise triggering condition — 'before committing to a coding decision you've drafted' — and enumerates concrete scenarios (pre-merge diff, architectural choice, security assessment, migration plan). It names the mechanism (four frontier models stress-testing for blind spots), and differentiates itself from siblings by naming the fast/open-ended alternatives it is not. An agent can select this tool over synthesize_coding and deliberate_coding without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and operational: it states exactly when to call (before committing a drafted decision), how to behave pre-call (tell the user it's working, since no progress shows mid-call), and how to handle large inputs (split diffs past 50 KB into 2-3 chunked calls). It ends by naming the two alternatives with their selecting conditions — 'fast broad take' → synthesize_coding, 'open-ended decision' → deliberate_coding. No exclusion condition is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.