Skip to main content
Glama
Zythenth

Antigravity MCP Bridge

by Zythenth

Record review test evidence

antigravity_record_test

Records client-reported test results against a reviewed patch, binding command, exit code, and output for audit. Document evidence without re-running or exposing secrets.

Instructions

Record a test already executed by the client in the isolated copy. Bind command, exit code and output to the reviewed patch. These are client-reported results; the bridge does not execute or independently verify the command. Do not include secrets in output.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
outputNo
taskIdYes
commandYes
exitCodeYes
expectedSha256Yes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNo
outputNo
sha256No
sourceNo
attemptNo
commandNo
sandboxNo
exitCodeNo
truncatedNo
recordedAtNo
testTaskIdNo
treeSha256No
executionErrorNo
beforeTreeSha256No

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.4.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the agent knows this writes state non-destructively. The description adds valuable context beyond annotations: it clarifies the tool records client-reported results and that the bridge does not execute or independently verify the command. It also warns against including secrets. It does not describe failure modes or what happens on duplicate recordings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each front-loaded with essential information: action, binding scope, verification stance, security constraint. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. The description covers the key behavioral caveats (client-reported, no verification, no secrets) for a write tool with annotations. It is nearly complete, missing only usage alternatives and some parameter meaning (taskId, expectedSha256).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by naming what gets bound (command, exit code, output) and the security constraint on output. However, it does not explain taskId or expectedSha256 semantics, leaving two of five parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (record) and resource (client-executed test evidence) with the binding scope stated (command, exit code, output tied to the reviewed patch). It is distinguishable from siblings like antigravity_test (which presumably runs tests) by clarifying that the bridge does not execute the command. Sibling differentiation is implicit rather than naming an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: record a test the client already ran in the isolated copy. However, it does not state when to use this versus antigravity_test or antigravity_verify, nor any preconditions (e.g., patch must be reviewed first). Usage context is implied, not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.