Panel Review
Server Details
Guardian agent for AI coding: four frontier models review risky diffs and commits before they ship.
- Status
- Healthy
- Uptime
- 46.5% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- TruVerifAI/init
- GitHub Stars
- 2
- Server Listing
- Panel Review
TDQS
Scored across 10 tools
Each tool occupies a clearly distinct role: the coding/financial split is reinforced by mode (synthesize vs. audit vs. deliberate), and the gate utilities (confirm_floor, record_gate_skip), outcome reporting, and ping are unambiguous. Cross-references between tools explicitly disambiguate when to use one mode over another.
Core tools follow a consistent verb_noun pattern (_coding/_financial suffixes), with synthesize/audit/deliberate as the verb prefix. Utility tools (confirm_floor, record_gate_skip, record_outcome) follow the same verb_noun convention; ping is the only deviation and is a standard health-check name.
Ten tools is well-scoped for this server: three review modes across two domains plus gate-handling, outcome reporting, and health check. Every tool earns its place and there is no redundancy or bloat.
The tool surface covers the full decision workflow for both coding and financial domains: fast synthesis, draft stress-testing, open-ended deliberation, gate unblocking, outcome feedback, and health verification. No obvious dead ends or missing lifecycle stages for the stated purpose.
Available Tools
10 toolsaudit_codingAInspect
Use this before committing to a coding decision you've drafted — a proposed diff before merging, an architectural choice before adopting, a security assessment before signing off, a migration plan before scheduling. Four frontier models stress-test it for blind spots, then revise their critiques in light of specific counter-positions. Returns a structured critique with severity tags and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. Pass the REAL diff text, never a summary or an abridged version — elision manufactures findings about what was elided. LARGE DIFFS: past roughly 50 KB, split into 2-3 calls chunked by file or section instead of one oversized call — a single huge diff can cut off the panel's structured verdict (you'll see degraded: true with degraded_reason=assessment_parse_failure and an empty findings list). Input fields over 1 MB are clipped with a notice in the response. For a fast broad take use synthesize_coding; for an open-ended decision use deliberate_coding.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | ||
| tests | No | ||
| gate_diff | No | ||
| gate_repo | No | ||
| constraints | No | ||
| eval_case_id | No | ||
| relevant_code | No | ||
| gate_context_id | No | ||
| proposed_action | No | ||
| continuation_token | No | ||
| target_hunk_hashes | No | ||
| architectural_context | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and discharges it thoroughly: it discloses the ~2-5 min runtime with no mid-call progress, the degraded:true failure signature (degraded_reason=assessment_parse_failure with an empty findings list) triggered by oversized diffs, the 1 MB input clipping notice, and the return shape (structured critique with severity tags plus a recommended action class). This is effectively a self-contained annotation layer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description runs long but every sentence is operational, and it is tightly sequenced: triggering scope, mechanism, return value, timing with a user-facing instruction, input fidelity requirement, size-limit handling, degradation signature, then alternative routing. Information is front-loaded (the when-to-use scoping comes first) and the length is justified by the tool's complexity — 12 parameters, a long runtime, and multiple failure modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with no output schema and 0% schema coverage, the description covers the core workflow well — what it returns (critique with severity tags, action class), when it degrades, how long it runs. The critical gap is parameter breadth: 11 of 12 parameters remain semantically unexplained, and no output schema exists to compensate. The main agent-facing path (passing a diff) is complete, but the broader parameter surface is not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters, so the description must compensate — but it only documents one: `diff`. That parameter is handled excellently (pass the REAL text, never a summary; split past 50 KB; clipping behavior), yet the other 11 parameters (tests, gate_diff, gate_repo, constraints, eval_case_id, relevant_code, gate_context_id, proposed_action, continuation_token, target_hunk_hashes, architectural_context) have zero semantic explanation in either the schema or the description. Several are technical (continuation_token, target_hunk_hashes) and entirely opaque to an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise triggering condition — 'before committing to a coding decision you've drafted' — and enumerates concrete scenarios (pre-merge diff, architectural choice, security assessment, migration plan). It names the mechanism (four frontier models stress-testing for blind spots), and differentiates itself from siblings by naming the fast/open-ended alternatives it is not. An agent can select this tool over synthesize_coding and deliberate_coding without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is explicit and operational: it states exactly when to call (before committing a drafted decision), how to behave pre-call (tell the user it's working, since no progress shows mid-call), and how to handle large inputs (split diffs past 50 KB into 2-3 chunked calls). It ends by naming the two alternatives with their selecting conditions — 'fast broad take' → synthesize_coding, 'open-ended decision' → deliberate_coding. No exclusion condition is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_financialAInspect
Use this before acting on a financial decision with real money at stake — a drafted trade thesis before placing the trade, a risk assessment before approving, an approve/decline call before committing, a backtest interpretation before deploying. Four frontier models critique it for tail risks, unstated assumptions, and asymmetric downside; specific weaknesses are routed back as targeted challenges; each model revises in light of specific counter-positions. Returns a structured critique with severity tags and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. For a fast broad take use synthesize_financial; for an open-ended decision use deliberate_financial.
| Name | Required | Description | Default |
|---|---|---|---|
| assumptions | No | ||
| constraints | No | ||
| eval_case_id | No | ||
| relevant_data | No | ||
| proposed_action | No | ||
| tests_backtests | No | ||
| continuation_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the process (four frontier models critique, weaknesses routed back, models revise), the runtime ('~2-5 min'), and the lack of progress ('no progress shown mid-call'), including an actionable instruction to tell the user it's working. However, it does not explicitly state whether the tool has side effects (e.g., writes to storage), though its read-only nature is implied. A small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose and usage, then details the process and runtime. It is somewhat verbose but every sentence contributes value. The structure is logical, moving from use case to behavior to return format to alternatives. It could be tightened but is not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the return format at a high level ('structured critique with severity tags and a recommended action class') but lacks detail on the exact structure. More importantly, it does not explain any of the parameters, which are all optional but essential for the agent to know what to supply. The absence of parameter semantics makes the tool under-specified for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description provides no explanation of any of the 7 parameters (assumptions, constraints, eval_case_id, relevant_data, proposed_action, tests_backtests, continuation_token). The description does not mention parameter names or their intended content, leaving the agent to guess what to pass. This is a critical omission given the zero coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to audit financial decisions with real money at stake before acting. It provides concrete examples (trade thesis, risk assessment, approve/decline calls, backtest interpretation) and specifies the output (structured critique with severity tags and recommended action class). It also distinguishes from siblings by naming synthesize_financial and deliberate_financial as alternatives for different needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly defines when to use the tool ('before acting on a financial decision with real money at stake') and gives specific scenarios. It also provides clear guidance on alternatives: 'For a fast broad take use synthesize_financial; for an open-ended decision use deliberate_financial.' This gives the agent explicit routing conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_floorAInspect
Cheap, FREE second opinion to clear a FALSE floor-gate block. When a floor-class gate (auth / secret / money / migration / guard, or a repo-defined custom floor) fired on what you believe is token-shape noise, this runs ONE budget model on your gate_diff: if it agrees the change is noise it mints a receipt that releases the change's FLOOR hunks. If it finds the change material — or is uncertain, errors, or the change touches the gate's own files — it releases NOTHING (run audit_coding instead). It covers FLOOR hunks only: if the gate's 'Still uncovered' line also lists NON-floor hunks, the gate keeps blocking on those — release them with record_gate_skip (a judgment reason) or an audit_coding PASS, which covers both kinds in one call. Forward BOTH the gate_context_id AND the target_hunk_hashes line the block printed (write and commit gates both print them) — the hashes are the deterministic binding tier, so the release binds to the gate's own recorded hunks even when your gate_diff is cosmetically drifted (drift is flagged, not dropped). The id alone does NOT bind unless your gate_diff genuinely overlaps the fire's hunks, so never omit the hashes. Free (no credits).
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | Alias for gate_diff (same spelling accepted across audit_coding / synthesize_coding / confirm_floor). If both are supplied, gate_diff wins. | |
| gate_diff | No | The change being gated (your staged diff, or the Edit's content). Judged by the model then discarded — never stored server-side. `diff` is accepted as an alias. | |
| gate_repo | Yes | The gate_repo fingerprint the block printed. | |
| gate_context_id | No | The `gate_context_id = "gc_…"` the block printed — copy verbatim. Identifies the fire; forward target_hunk_hashes ALONGSIDE it — the id alone binds only when your gate_diff genuinely overlaps the fire's hunks. | |
| gate_session_id | No | The gate_session_id from the block, if it provided one. | |
| target_hunk_hashes | No | The `target_hunk_hashes = [...]` line the WRITE gate printed — copy verbatim. Deterministic binding: releases even if the write gate's diff shape differs from your gate_diff. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses that it runs one budget model, is free, covers only FLOOR hunks, releases nothing on uncertainty/error, requires both gate_context_id and target_hunk_hashes, explains the deterministic binding of hashes, flags drift without dropping, and notes the diff is never stored. No annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with purpose and each sentence adds functional detail (binding, conditions, alternatives, parameter usage). It is structured and not repetitive, though it could be trimmed slightly without losing key facts. Still, the density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and six parameters, the description covers the decision logic, required inputs, binding mechanics, failure modes, and sibling routes. An agent can correctly decide when to call it and how to construct the call without missing critical constraints. The only minor omission is what the receipt looks like, but that's not needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaningful semantics: it explains the diff alias precedence, the binding role of target_hunk_hashes, the requirement to forward both id and hashes, and the fact that gate_diff is judged then discarded. These go beyond the schema descriptions, though some redundancy exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (confirm/clear) and resource (floor-gate block) and explicitly frames it as a cheap second opinion to clear false positives. It distinguishes itself from siblings by naming audit_coding as the alternative for material changes and record_gate_skip for non-floor hunks, leaving no ambiguity about its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use (floor-class gate fired on token-shape noise), when not to (if material, uncertain, errors, or touches gate's own files → run audit_coding), and explicitly names alternatives for non-floor hunks (record_gate_skip or audit_coding). The conditions are concrete and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliberate_codingAInspect
Use this when a coding decision is hard to undo and there's more than one defensible answer — deployment safety, architectural choices, dependency upgrade strategy, migration timing, security review judgment, contested merge calls. Four frontier models reason independently; conflicts between their reasoning are identified and routed back as specific points each model must defend or revise; refined responses are synthesized. Unlike peer-ranking approaches, models engage with specific counter-positions on the points where they disagreed. Returns a reasoned conclusion, agreement signal, dimensions of disagreement, and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. For a fast broad take use synthesize_coding; for stress-testing a draft answer use audit_coding.
| Name | Required | Description | Default |
|---|---|---|---|
| question | No | ||
| gate_diff | No | ||
| gate_repo | No | ||
| constraints | No | ||
| eval_case_id | No | ||
| relevant_code | No | ||
| relevant_paths | No | ||
| gate_session_id | No | ||
| continuation_token | No | ||
| options_considered | No | ||
| architectural_context | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It thoroughly explains the process (four models reason independently, conflicts routed back, refined responses synthesized), the output (reasoned conclusion, agreement signal, disagreement dimensions, recommended action class), and operational details (runtime ~2-5 min, no progress shown, instructs to inform the user). This goes far beyond the schema's silent fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively long but well-structured: it leads with purpose and examples, then explains process, output, runtime, and ends with alternatives. Each sentence contributes meaningful information; no filler or repetition. It could be slightly tightened, but the density is justified given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers usage context, process, output, runtime, and alternative routing, which is strong for a complex tool. However, it omits any explanation of the 11 parameters, making it incomplete for an agent that needs to construct a valid call. The output schema is absent, but the description does list return components. Overall, significant gaps remain regarding parameter usage, so completeness is only partial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description offers zero explanation of any of the 11 parameters (question, gate_diff, gate_repo, constraints, eval_case_id, relevant_code, relevant_paths, gate_session_id, continuation_token, options_considered, architectural_context). An agent has no guidance on what values to provide or how these fields relate to the tool's operation. The description entirely fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with an explicit condition for use ('coding decision is hard to undo and there's more than one defensible answer') and enumerates concrete examples (deployment safety, architectural choices, dependency upgrades, migration timing, security review, contested merges). It clearly names the tool's purpose as deliberate multi-model reasoning, distinguishing it from the listed siblings through later contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use criteria (hard-to-undo, multiple defensible answers) and directly names alternatives with their appropriate use cases: 'For a fast broad take use synthesize_coding; for stress-testing a draft answer use audit_coding.' This leaves no ambiguity about when to select this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliberate_financialAInspect
Use this when a financial decision has asymmetric downside and there's more than one defensible answer — trade-idea robustness, risk-model adequacy, backtest interpretation, lending judgment, fraud signal interpretation, portfolio reasoning. Four frontier models reason independently against tail-risk and asymmetric-downside framing; conflicts are routed back as specific points each model must defend or revise; refined responses are synthesized. Returns a reasoned conclusion, agreement signal, dimensions of disagreement, and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. For a fast broad take use synthesize_financial; for stress-testing a draft thesis use audit_financial.
| Name | Required | Description | Default |
|---|---|---|---|
| question | No | ||
| assumptions | No | ||
| constraints | No | ||
| eval_case_id | No | ||
| relevant_data | No | ||
| continuation_token | No | ||
| options_considered | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It states that four frontier models reason independently, conflicts are routed back for defense/revision, and refined responses are synthesized. It also discloses runtime (~2-5 min), lack of mid-call progress, and instructs to tell the user it's working. This is transparent about the process and output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence adds value: trigger conditions, process, output, runtime, and alternatives. It is a bit long but not wasteful; the structure front-loads usage context and then provides operational details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex with 7 parameters and no output schema, yet the description provides no guidance on how to populate parameters, what is required, or how to provide data. The output is described, but input handling is underspecified, leaving the agent without enough information to make a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the 7 parameters (question, assumptions, constraints, eval_case_id, relevant_data, continuation_token, options_considered). The agent has no guidance on what to supply for each field, making it impossible to correctly construct a call without external knowledge.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it is for financial decisions with asymmetric downside and multiple defensible answers, and it describes the output (reasoned conclusion, agreement signal, dimensions of disagreement, recommended action class). It also distinguishes itself from siblings by naming synthesize_financial and audit_financial and their different use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use conditions (asymmetric downside, more than one defensible answer) and provides alternatives: 'For a fast broad take use synthesize_financial; for stress-testing a draft thesis use audit_financial.' It also mentions runtime and the need to inform the user, which is practical usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingAInspect
Health check — confirms the TruVerifAI MCP server is reachable and your API key is valid. Returns immediately, runs no model orchestration, and consumes no credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It explicitly states that the tool returns immediately, runs no model orchestration, and consumes no credits, which gives the agent meaningful expectations about cost and latency. It does not specify the exact response shape, but for a ping that is not a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence that front-loads the core purpose ('Health check'), then adds the most important behavioral facts. Every phrase earns its place and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-side-effect health check, the description covers all necessary invocation context: purpose, reachability, authentication validity, latency, and lack of orchestration/credit cost. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters handschema coverage is trivially 100%. The description adds no parameter details because none exist, and the no-parameters case is fully handled; this meets the baseline for parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a health-check operation with a specific purpose: confirming server reachability and API key validity. It is unambiguous and distinct from all sibling tools, which are clearly heavier audit/deliberation functions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: use this when you need to verify connectivity and credentials before any work that requires model orchestration. It does not explicitly state when not to use it or name alternatives, but it strongly implies a lightweight verification role distinct from the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_gate_skipAInspect
Release a TruVerifAI proactive-invocation gate WITHOUT running (another) review, by logging an explicit reason. Use it when a gate fired (before a commit or a Write/Edit) and EITHER the review is genuinely unnecessary — e.g. a false positive, a trivial/test/docs-only change, or generated/vendored code — OR you ran ONE review and want to proceed on it: apply its findings and pass recommendations_applied, or pass review_deferred_to_commit to defer a batch to the commit gate. NO judgment skip reason releases a FLOOR hunk (auth/secrets/money/migrations/removed-guard/gate-self — the gate's own code — or a repo-defined custom floor from .truverifai/risk.json) — not even test_or_docs_only / generated_or_vendored_code, because the path doesn't change what the hunk IS (a real credential in a test file is still a live credential, and .github/workflows/ classifies as test/docs). The ONE exception is recommendations_applied WITH your recent review on record: after you ran a review and applied its findings, it releases the re-fired change's floor hunks too (lineage-verified, minutes-TTL, logged distinctly as 'findings applied' — never as an audited PASS). Otherwise, while a floor hunk is UNREVIEWED every skip is denied: cover the floor first — audit_coding, confirm_floor (free), synthesize_coding, or, as a last resort, ship it un-reviewed with accept_risk_no_review (a logged override + substantive pre-mortem) — and the same skip then becomes admissible and releases the change's remaining NON-floor hunks. On a MERGE commit whose branch content was already reviewed, branch_already_reviewed releases the non-floor hunks in one call (merge fires only; name where it was reviewed in reason_text). (Note: 'already reviewed' is NOT a skip — a real prior PASS releases the gate automatically.) On a non-floor change the skip releases it on retry. FREE — no credits; the reason is logged. Pass the gate_repo from the gate's message, plus the gate_context_id it printed (REQUIRED — the server verifies a gate truly fired and releases only the hunks IT recorded, so you never supply hunks yourself). There is no way to skip without it: if the gate printed no id (rare), don't skip — run audit_coding with gate_repo + gate_diff.
| Name | Required | Description | Default |
|---|---|---|---|
| score | No | Optional. The score from the gate's gate_signal line. | |
| gate_repo | Yes | The repo fingerprint from the gate's message (the gate_repo value). | |
| reason_code | Yes | Why you're releasing the gate without running (another) review. Pick the closest fit. Single-call model: after ONE panel-review call, use 'recommendations_applied' (you got findings, applied them — server-verified against that review; releases floor at both gates, and a floor hunk released at the WRITE gate is still re-audited at commit) or 'review_deferred_to_commit' (defer ALL review to the commit gate; releases the write for this session/area, the batch is re-reviewed at commit). On a FLOOR block (auth/secrets/money/migrations/removed-guard/gate-self, or a repo-defined custom floor) match the tool to your situation: a genuine floor change you want reviewed → audit_coding (a PASS releases; recommended); you believe the gate mis-fired → confirm_floor (free) or synthesize_coding (each releases only if it agrees it's not risky). 'accept_risk_no_review' is the LAST RESORT — only after the real paths above genuinely don't fit (the gate mis-fired, you're deadlocked, or you're consciously shipping un-reviewed): it ships the floor hunk UN-reviewed, needs a substantive pre-mortem reason_text, is logged as a distinct override to the human, and expires in minutes. 'other', 'disagree_with_classification', and 'accept_risk_no_review' REQUIRE reason_text (and so do the judgment codes at the write and commit gates). Note: 'prior_pass_receipt_match' is NOT a skip — a real prior audit PASS releases automatically; if the gate fired, re-review the changed hunks. | |
| reason_text | No | 1 sentence on why this skip is justified — REQUIRED for 'other' / 'disagree_with_classification', and for the judgment codes (false_positive_not_risky, trivial_change, reviewed_outside_truverifai, time_critical_hotfix) at the write and commit gates. For 'accept_risk_no_review' it must be a SUBSTANTIVE pre-mortem (assume it IS a real issue: name the failure, who it affects, why it's acceptable). No secrets, file paths, or proprietary identifiers; general terms only. LIMITS (fix A10): ordinary reasons are clipped past 500 chars; an 'accept_risk_no_review' pre-mortem gets 4000 chars (its substance IS the audit trail). Clipping never fails the call and the response says when it happened (reason_text_truncated) — don't retry to shorten. | |
| gate_context_id | No | REQUIRED. The `gate_context_id = "gc_…"` the gate's block message printed — copy it verbatim. The server verifies a gate truly fired and releases only the hunks IT recorded, so you never supply hunks yourself. There is no way to skip without it: if the message has no id (rare), don't skip — run `audit_coding` with gate_repo + gate_diff, whose PASS releases floor and non-floor hunks alike. | |
| risk_categories | No | Optional. The risk_categories from the gate's gate_signal line. | |
| classifier_version | No | Optional. The classifier_version from the gate's gate_signal line. Forwarding it (and score/risk_categories) sharpens the data that improves the classifier — no source, just the labels the gate already showed you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it delivers richly: the tool is free, the reason is logged, the server verifies a gate truly fired and releases only hunks IT recorded, recommendations_applied is lineage-verified with minutes-TTL and logged distinctly (never as an audited PASS), accept_risk_no_review expires in minutes and is logged as a distinct override, and reason_text clipping never fails the call. These are exactly the behavioral traits an agent needs to predict side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense ~400-word wall of text with no paragraph breaks or bullet structure, which hurts scannability for an agent. The length is largely justified by the tool's genuine complexity (floor rules, exceptions, 14 reason codes), and the core purpose is front-loaded. However, there is meaningful redundancy with the schema's already-extensive reason_code and reason_text descriptions, so not every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with no output schema and no annotations, the description covers every critical dimension: what it does, when to use it, floor-hunk restrictions, the single exception (recommendations_applied), merge behavior, alternatives, required parameters, the no-gate_context_id fallback (run audit_coding), cost, and logging semantics. An agent has everything needed to decide whether and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds substantial decision-making value beyond the schema: it explains WHY gate_context_id is required (server verification, never supply hunks yourself), provides the routing logic for choosing among the 14 reason_code values, and clarifies the floor-hunk consequences of each choice. The description's decision tree is genuinely additive to the per-parameter schema text, especially for reason_code and gate_context_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Release a TruVerifAI proactive-invocation gate WITHOUT running (another) review, by logging an explicit reason.' This precisely distinguishes it from sibling review tools (audit_coding, synthesize_coding, confirm_floor) by making clear it is the skip/release path, not a review path. The scope is unambiguous and the tool's role in the gate workflow is immediately identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is explicitly conditional: 'Use it when a gate fired... and EITHER the review is genuinely unnecessary... OR you ran ONE review and want to proceed on it.' The description names concrete alternatives (audit_coding, confirm_floor, synthesize_coding, accept_risk_no_review) and states exclusions — no judgment skip releases a FLOOR hunk. It even covers the merge-commit case (branch_already_reviewed) and the 'already reviewed is NOT a skip' caveat. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_outcomeAInspect
Report whether a prior TruVerifAI MCP call (synthesize / deliberate / audit) was useful and whether it changed the decision you would have made without it. Call this AFTER you've acted on (or explicitly rejected) the response from the prior call. Free — no credits charged. The user sees the aggregate on their TruVerifAI dashboard; outcome reporting is how they evaluate whether the tool is worth keeping.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | 1-2 sentences on what specifically shifted (or didn't). REQUIRED when useful=false OR changed_decision=false OR category='other' — the no-op cases are the most informative and we want a concrete reason. Max 500 chars. DO NOT include confidential or code-specific details (no proprietary file paths, no function or class names from the user's codebase, no secret values, no internal system identifiers, no copy-pasted source). Describe the decision in general terms only. | |
| impact | Yes | Decision blast radius. HIGH = hard to reverse (multi-file refactor to undo) OR touches a security/safety boundary OR affects load-bearing logic many callers depend on. MEDIUM = recoverable with effort; bounded blast radius. LOW = trivially reversible. | |
| useful | Yes | True if the prior response informed your decision-making in any way (caught something, confirmed something, or surfaced a tradeoff you hadn't considered). False if it was noise or duplicated what you already knew. | |
| call_id | Yes | The mcp_<uuid> request_id from the prior MCP call — find it in the prior response body's top-level post_action.call_id (or usage.request_id); it is in the body, not _meta. | |
| category | Yes | The kind of decision this MCP call was about. Pick the closest single fit. Use 'other' only when nothing else applies (and explain in notes). Powers the per-category dashboard slicing the user relies on to evaluate which decision types TruVerifAI helps with. | |
| changed_decision | Yes | True if your action AFTER reading the response differs from what you were about to do BEFORE the call. False if you proceeded as originally planned (even if the call was still useful as confirmation). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It adds valuable context: the operation is free, it records to the user's dashboard, and the aggregate is used for tool evaluation. It could further disclose idempotency or duplicate-call behavior, but the main behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four tight sentences, each adding new information: purpose, timing, cost, and user-facing value. It front-loads the core action and avoids repeating schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich parameter schema and the absence of an output schema, the description provides enough operational context: when to call, what the call does, and why it matters to the user. A small gap is not describing the tool's return value or behavior on an invalid call_id, but that is not critical for a side-effect-oriented reporting tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in detail. The description itself does not add parameter-level meaning, but it also does not need to; the baseline of 3 applies because the structured descriptions carry the weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Report whether a prior TruVerifAI MCP call ... was useful and whether it changed the decision.' It names the prior call types (synthesize / deliberate / audit) and makes clear this is a feedback-recording tool, not one of the analysis tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit timing guidance: 'Call this AFTER you've acted on (or explicitly rejected) the response from the prior call.' It also notes the call is free, which is useful context, but it does not explicitly contrast this tool with the sibling record_gate_skip or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
synthesize_codingAInspect
Use this when you want a multi-model second opinion fast — coding questions where being right matters but you don't have time for full deliberation. Library/API choices in production code, idiomatic patterns for new domains, anywhere a single-model answer might miss a viewpoint. Four frontier models answer in parallel and the result is synthesized into one answer with an alignment signal. ~15-30s. For high-stakes decisions reach for deliberate_coding; for stress-testing a draft answer reach for audit_coding.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | ||
| context | No | ||
| question | No | ||
| gate_diff | No | ||
| gate_repo | No | ||
| eval_case_id | No | ||
| gate_context_id | No | ||
| gate_session_id | No | ||
| continuation_token | No | ||
| target_hunk_hashes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It reveals meaningful traits: parallel execution of four models, a ~15-30s latency budget, and a synthesized output with an alignment signal. The only gap is that it never states whether the call has side effects (e.g., logging, gating) or is purely a read/compute operation, which matters given the gate_* and record_* sibling tools suggest a pipeline context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: trigger condition, example use cases, execution mechanism plus latency, and routing to alternatives. The key decision signal ('multi-model second opinion fast') is front-loaded, and the sibling routing is compactly delivered at the end. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For the tool-selection decision (when to call this vs siblings), the description is essentially complete. But the tool has 10 undocumented parameters and no output schema, and the description only vaguely gestures at what inputs are expected and what the return looks like beyond 'one answer with an alignment signal.' Given the parameter complexity, this leaves material gaps in what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 string parameters, so the description must compensate and largely does not. It hints that the question/content concerns coding decisions, which loosely maps to the `question` param, but diff, context, gate_*, continuation_token, and target_hunk_hashes are never explained or even acknowledged. An agent receives almost no help understanding what to populate for this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a specific verb+resource: 'multi-model second opinion fast' for coding questions, and explicitly names the mechanism ('Four frontier models answer in parallel... synthesized into one answer with an alignment signal'). It distinguishes itself from siblings by naming deliberate_coding and audit_coding as the alternatives for different stakes, so an agent can tell them apart without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance ('when you want a multi-model second opinion fast', 'where being right matters but you don't have time for full deliberation'), concrete example scenarios (library/API choices, idiomatic patterns), and explicit exclusions with named alternatives ('For high-stakes decisions reach for deliberate_coding; for stress-testing a draft answer reach for audit_coding'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
synthesize_financialAInspect
Use this when you want a multi-model second opinion fast on a financial question where being right matters but you don't have time for full deliberation. Interpretations of market data or model output, factor explanations relevant to live decisions, comparisons of how different framings change an analysis — anywhere a single-model answer might miss a viewpoint. Four frontier models answer in parallel and the result is synthesized into one answer with an alignment signal. ~15-30s. For high-stakes financial decisions reach for deliberate_financial; for stress-testing a trade thesis or risk assessment reach for audit_financial.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| question | No | ||
| eval_case_id | No | ||
| continuation_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well: it discloses that four frontier models run in parallel, the answer is synthesized, an alignment signal is included, and latency is ~15-30s. This explains the tool's core behavior and constraints, though it omits minor details like failure modes or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the key usage decision and alternative routes, and every sentence adds information about mechanism or latency. It is slightly long and meandering in the middle, but still economical for the amount of routing guidance it provides.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, latency, mechanism, and alternatives, but it lacks return-value detail (beyond 'alignment signal') and leaves the parameters under-documented. Given the absence of an output schema and annotations, this is a noticeable but not crippling gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate beyond implying that 'question' holds a financial question. The roles of 'context', 'eval_case_id', and 'continuation_token' are entirely unexplained, so an agent may struggle to invoke the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pairing: synthesize a fast multi-model second opinion on financial questions. It explicitly distinguishes itself from deliberate_financial and audit_financial, so an agent can select it without inspecting other tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use context ('fast', 'right matters but you don't have time for full deliberation') and names the exact alternatives for other cases, including high-stakes decisions and trade-thesis stress-testing. This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
- First observed
audit_coding - First observed
audit_financial - First observed
confirm_floor - First observed
deliberate_coding - First observed
deliberate_financial - First observed
ping - First observed
record_gate_skip - First observed
record_outcome - First observed
synthesize_coding - First observed
synthesize_financial
Publisher details
- Operator
- TruVerifAI
- Operator website
- https://truverif.ai/panel-review
- Vendor relationship
- Not applicable
- Documentation
- https://truverif.ai/settings/mcp#setup · Publisher source
- Trust center
- Not available
- Restrictions
- Not applicable
Related MCP Connectors
Agentic code review, no signup to try: reality gates + frontier-model review, with veto.
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
Governance layer for AI coding agents: knowledge-graph grounding, session audit, policy controls.
Multi-LLM council: 25+ frontier models in parallel, consensus scoring, verdict-first code review.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA runtime gate for coding agents. Blocks the tool calls that wreck a repo (force-push main, rm -rf, secret exfiltration, CI wipe) and lets normal build and commit work through. Machine-checked git-branch core (z3); the rest is high-precision heuristics. Tested on 3,790 real CI commands, 0 false blocks.1MIT
- AlicenseAqualityDmaintenanceAdversarial review system that spawns three independent contrarian reviewers to catch issues before AI coding agents execute critical changes.38 npmMIT
- AlicenseNot gradedqualityBmaintenanceA three-stage guardrail agent for LLM-powered coding assistants that reviews proposed actions before execution, blocking destructive commands and maintaining an audit trail.3MIT

CodePeel MCP Serverofficial
FlicenseAqualityDmaintenanceEnables AI agents to review code diffs for bugs, security issues, and bad patterns, and generate fixes.4-
Glama MCP Gateway
Add one secure layer between your agents and this server.