Skip to main content
Glama

Verify Performance

verify_performance
Destructive

Prove or fail performance requirements against measured uncertainty bands, returning pass, fail, or indeterminate verdicts with provenance. Escalate indeterminate results instead of guessing.

Instructions

Prove (or fail to prove) every requirement declared with declare_performance — the step that turns "a solver printed 0.29" into a claim with a band and a provenance.

A verdict has THREE states. pass and fail each require the measurement's whole uncertainty band to sit on one side of the limit; a band that straddles it is indeterminate, meaning "escalate", not "probably fine". A correlation reading Cd = 0.28 ± 10 % against a limit of 0.30 spans 0.252–0.308 and has NOT shown the part passes — collapsing that to a pass is how a spec silently goes unmet.

tier picks the evidence:

  • 'screen' — each requirement's cheap estimator only. Milliseconds, no solver.

  • 'solver' — the real solve for every requirement.

  • 'auto' (default) — screen first, escalate only what the screen could not decide or what declares fidelity_floor: 'solver'. This is the ladder that keeps a design loop cheap: cheap measurements eliminate candidates, solves confirm survivors.

Trust is part of the measurement, not a footnote: trust: {converged: true} or a band_max_pct cap makes an unconverged (or insufficiently mesh-converged) solve come back indeterminate with the reason, never pass.

Solver-tier measurements are asynchronous, so this returns EITHER the finished verdict (screen-only, or everything already decided) or {job_id, status, pending, results} — poll job_result for the completed verdict. Never raises on a failing requirement; a failure is a row.

Every verdict is also RECORDED on the part, stamped with a geometry signature of the shape it measured (#261). That record is what merge_assembly, substitutability_check and component_contract_check consult, since a gate has to answer synchronously and this may not have: an in-flight solve records rows the gates read as unverified, and editing the part invalidates the signature so they read stale — never a pass on either path.

Returns {handle, tier, ok, n_requirements, passed, failed, indeterminate, escalate, results: [{name, tier, metric, state, measured, limit, band_pct, worst_case, best_case, margin, margin_pct, detail, trust_reasons?, screen?, job_id?}]}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tierNoauto
handleYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint:false, destructiveHint:true), the description details three verdict states, the trust conditions, asynchronous job returns, and the recording side effect with geometry signature invalidation. It also states it never raises on failure. This adds substantial behavioral context essential for correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured, with a clear opening statement and organized sections for tier and behavior. It front-loads the purpose and uses bullet-like formatting. Some redundancy exists (e.g., re-explaining verdict states), but every sentence adds value for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, asynchronous behavior, lack of output schema, and side effects, the description covers all necessary aspects: return format, recording behavior, invalidation, and how gates read the records. It leaves no critical gaps for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents only tier and handle with no descriptions (0% coverage). The description explains the tier values and their semantics ('screen', 'solver', 'auto') and provides context for handle as a part identifier, but it does not explicitly define what handle refers to or its format. The tier explanation compensates partially, but the required handle parameter is left vague.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'Prove or fail to prove' with a clear resource: every requirement declared with declare_performance. It distinguishes its role from declare_performance and from other verify tools like verify_intent or verify_contract, making its scope explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the tier ladder and when screen vs solver is used, and mentions the asynchronous path for solver-tier, but it does not explicitly state 'use this when you have declared performance requirements' or name alternatives to avoid. However, the context is clear that this is the verification step for performance declarations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools