Skip to main content
Glama

Affine Earth Math Court Remote

release_regression

Franklin's sign-off over an Affine IDE release regression record. The IDE's regression agent opens every study and writes a canonical tab-separated record of what it saw: one stage arm per study (route and the viewer actually showing), the rail, the chat, the en/es/ar locales, the served download page, the prover and one line per screenshot. This court grades that record with ReleaseRegressionLaw from LatticeRender - the same source the IDE's own prover compiles - and answers FRANKLIN_SIGNED_OFF only when every arm is ok, there is exactly one stage arm per study, the three locale arms, a served arm, the prover arm, at least one study with one screenshot per study and, when dmg_sha256 is present, an ok served arm named dmg_sha_matches. Otherwise the terminal is CURE (a ruling, not an error) with failing and missing each naming what is not proven, once; a stage arm whose detail shows the movie stage on a domain route is CURE even when flagged ok (the 0.2.6.3 shape). Returns terminal, digest (sha256 over the canonical bytes), arms_total, arms_ok, failing, missing, franklin (the narration lines), sealed_by and status. The court opens no app and runs no prover: the arms are the caller's measurements. It WRITES NOTHING on the cell and reads no clock; nine cells return the same bytes for the same record.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
dateNooptional: the caller's clock, echoed as graded_at on the receipt; the court reads no clock of its own. e.g. 2026-10-06T21:10:00Z
recordNothe canonical regression record: UTF-8 lines, tab-separated, LF endings, first line affine_ide_regression<TAB>1, then version, build, date, host, optional dmg_sha256, app_sha256, studies, one arm line per arm (arm<TAB>surface<TAB>name<TAB>ok|BAD<TAB>detail) and one shot line per screenshot (shot<TAB>name<TAB>bytes<TAB>sha256). e.g. affine_ide_regression 1 version 0.2.6.4 build 3282 date 2026-10-06T21:10:00Z host macOS 27.2 app_sha256 <64 lowercase hex> studies 1 arm stage study-47 ok route domain · showing DomainViewer ...
self_testNooptional: 1 asks the court to run its own self-test over the law instead of grading a record - every arm first shown a planted defect; answers RELEASE_REGRESSION_COURT_SELFTEST_OK with arms_total, arms_ok and the report lines. e.g. 1

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and discharges it: it opens no app, runs no prover, writes nothing on the cell, reads no clock, and is deterministic ('nine cells return the same bytes for the same record'). It also clarifies that CURE is a ruling rather than an error, which prevents a caller from misreading a non-ok terminal as a failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Roughly 340 words for a three-parameter tool, loaded with invented vocabulary ('the court', 'sealed_by', 'franklin', 'the 0.2.6.3 shape') that forces re-reading. The core behavior is not front-loaded; an agent must parse the sign-off predicate before learning what the tool actually consumes and returns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description enumerates the returned fields (terminal, digest, arms_total, arms_ok, failing, missing, franklin, sealed_by, status) and explains the digest's basis. Combined with the input contract and determinism guarantees, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents date, record, and self_test. The description restates the record's internal arm/shot line structure and the self_test purpose, adding orientation but no syntax beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action on a specific resource: it grades an Affine IDE release regression record and returns FRANKLIN_SIGNED_OFF when every arm is ok, otherwise CURE. That is distinguishable from the many verify_*/umc_* siblings, though the 'Franklin's sign-off' framing is metaphorical enough that an agent must read several sentences to extract the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It makes the precondition explicit ('the arms are the caller's measurements') and spells out the exact conditions that yield sign-off versus CURE, including the edge case where a stage arm shows a movie stage on a domain route. It does not name alternatives for grading, but the sibling list contains no competing grader, so the when-to-use picture is effectively complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources