Skip to main content
Glama

Baseline Verify

baseline_verify
Read-only

Checks a rebuild against a pinned baseline by re-resolving every pinned item from the current registry and verifying its revision and content fingerprint, catching any drift.

Instructions

Verify a rebuild against a pinned baseline (issue #142, C3) — the reproducible- rebuild gate. Re-resolves every pinned item from the current registry + artifacts and checks it still matches the pinned rev AND content fingerprint; a drifted input (changed bytes, a bumped rev, a vanished item) is caught, never silently accepted.

baseline: path to the baseline sidecar. registry: path to the current items.json sidecar. base_dir: artifact root for fingerprinting (defaults to the registry's directory).

Returns {ok, label, drifted (item/field/expected/actual), missing} — ok is True iff the rebuild reproduces the baseline exactly.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
base_dirNo
baselineYes
registryYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states that the tool 're-resolves every pinned item' and checks 'rev AND content fingerprint', and that it catches drifted inputs 'never silently accepted'. This adds behavioral detail beyond the readOnlyHint annotation, which is appropriate. It does not explicitly say it is read-only, but the annotation already covers that. The description provides useful behavioral context about what comparison is done and the guarantee of not silently accepting drift, which is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized; it provides a clear, front-loaded explanation of the tool's purpose, then lists parameters, and finally describes the return value. Each sentence earns its place, including the behavioral guarantee and the default for base_dir. It is well-structured and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema and low schema coverage, the description explicitly states the return format ('{ok, label, drifted (item/field/expected/actual), missing}') and its meaning ('ok is True iff the rebuild reproduces the baseline exactly'). The annotations (readOnlyHint=true) cover the safety profile. For a verification tool with this complexity, the description is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, but the description explains each parameter: 'baseline: path to the baseline sidecar', 'registry: path to the current items.json sidecar', and 'base_dir: artifact root for fingerprinting (defaults to the registry's directory).' This compensates for the lack of schema descriptions, providing meaning for all three parameters including the optional one.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool verifies a rebuild against a pinned baseline, using a specific verb ('Verify') and resource ('a rebuild against a pinned baseline'). It explicitly distinguishes from other tools by naming it as the 'reproducible-rebuild gate' and describing its specific function of re-resolving pinned items, which is unique among siblings like items_validate or items_resolve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: as the final gate for reproducible rebuilds. It mentions the context of issue #142 and C3, which adds specificity. However, it does not explicitly state when not to use it or name alternative tools for other verification tasks, though the purpose is clear enough that an agent could infer the right usage scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools