Skip to main content
Glama

verdict_stats

Report debate health and panel behavior per chain and lab over a chosen window, including sign-off rates, rounds, objections, dropouts, cost, and independence skew. Descriptive metrics only; no verdict changes.

Instructions

How the debate mechanism itself is doing, per chain and per lab, across a window of days: sign-off rate, mean rounds to sign-off, objections raised, withdrawals vs accepted proposals, dropouts, unparseable replies, shape-only critique rounds (rounds spent entirely on document shape rather than substance), per-lab independence skew (novel-objection rate, solo-signoff rate, a low-independence flag), mean cost and wall time per run, and the largest prompt file per stage type. Derived from report.json and *.usage.json on disk - nothing is recorded anywhere else and nothing leaves this machine. Descriptive only: never reweights a panel or changes a verdict.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
daysNohow far back to look, in days. Defaults to 30.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.8.1

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it names the data source ('report.json and *.usage.json on disk'), asserts no side effects or hidden recording, and states nothing leaves the machine. It also pins the tool as descriptive-only, i.e. non-mutating. It omits any statement about cost of the scan or pagination/limits, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The scope statement is front-loaded ('How the debate mechanism itself is doing, per chain and per lab, across a window of days') before the metric list. The long enumerated list of metrics is dense but justified because there is no output schema, and the closing sentence cleanly states the non-mutating constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description appropriately enumerates the returned fields (sign-off rate, mean rounds, objections, dropouts, cost/wall time, etc.), plus the on-disk data provenance and the descriptive-only guarantee. It is complete enough to call correctly, with the only gap being sibling differentiation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (days) and schema description coverage is 100%, so the schema already documents the lookback window and its default. The phrase 'across a window of days' mirrors that meaning but adds no format, boundary, or aggregation detail beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete resource (debate-mechanism health metrics per chain and per lab over a day window) and enumerates exactly what is measured, so the agent knows precisely what this tool returns. It does not, however, distinguish itself from siblings with overlapping remits such as metrics_report or spend_report, which an agent would need to disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'Descriptive only: never reweights a panel or changes a verdict' tells the agent this is a read/observe tool rather than a control tool, which is useful context. But there is no explicit when-to-use, when-not-to-use, or routing to the sibling reporting tools (metrics_report, spend_report), so an agent has to infer the choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.