Skip to main content
Glama

record_vehicle_validation

Records human-observed pass/fail/skipped results with evidence and measurements, rejecting changed exports or mismatched hashes and requiring pass evidence for all checks.

Instructions

Record human-observed pass/fail/skipped results with evidence and measurements. Refuses a changed export or mismatched hash. All seven checks need pass evidence for a passed run.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYes
testerYes
game_versionYes
observationsYes
export_sha256Yes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
checksYes
designYes
run_idYes
statusYes
testerYes
revisionYes
created_atYes
part_countYes
game_versionYes
vehicle_nameYes
vehicle_pathYes
export_sha256Yes
evidence_levelYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.1.1

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does add real value: it discloses a guard behavior ("Refuses a changed export or mismatched hash") and an integrity rule (all seven checks need pass evidence). What is missing is idempotency/re-call behavior (can a run be overwritten?), permission requirements, and partial-submission semantics, so it is not fully complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler, front-loaded with the core action and followed by the two hard constraints. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the key rejection and evidence rules. However, for a complex nested-observation mutation tool with 0% schema coverage and no annotations, the absence of any explanation for run_id/tester/game_version and the overwrite/idempotency behavior leaves gaps an agent may need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it only partially does: "evidence and measurements" maps to the Observation fields, "checks" maps to the enum check list, and "mismatched hash" alludes to export_sha256. It says nothing about run_id, tester, or game_version, leaving three of five required parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ("Record") and resource ("human-observed pass/fail/skipped results with evidence and measurements"), which clearly distinguishes it from siblings like prepare_vehicle_validation and get_vehicle_validation. It stops short of explicitly naming those siblings, but an agent can identify the tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by "human-observed" results and the rule that a passed run needs pass evidence, but there is no explicit guidance on when to use this versus prepare_vehicle_validation or get_vehicle_validation. No prerequisites about calling prepare first are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.