Skip to main content
Glama

GnosisLab

A scientific-method runtime for AI agents

Turn hypotheses into experiments, experiments into evidence, and evidence into auditable conclusions.

DOI CI PyPI Python License piwheel downloads PyPI downloads PyPI Downloads Discord

Install · Documentation · How it works · Architecture · For researchers · For developers · Discord


What is GnosisLab?

Most AI agents are optimized to produce an answer.

GnosisLab is built to help an agent produce a defensible result.

It turns the scientific process into explicit software objects:

hypothesis → experiment → observation → evidence → claim → conclusion

Instead of leaving an agent's reasoning, experiments, and evidence scattered across prompts, shell history, notebooks, logs, and temporary files, GnosisLab gives the process a structured runtime.

Experiments are represented explicitly. Evidence is retained. Claims can be connected to their support. Execution is observable. State is separated by responsibility. The resulting research process can be inspected by a human rather than accepted as an opaque final answer.

GnosisLab is intended for machine-learning experimentation, computational research, autonomous research agents, evaluation systems, and other workflows where how a result was obtained matters as much as the result itself.


Related MCP server: Researcher AI

Why this exists

LLMs can generate code, search literature, run analyses, compare alternatives, and propose conclusions.

That creates a new problem:

How do we know what actually happened?

A useful research system needs more than tool access. It needs a disciplined process around tool use:

  • What hypothesis was being tested?

  • Which experiment addressed it?

  • What code actually ran?

  • What observations were produced?

  • Which evidence supports which claim?

  • Which conclusions remain uncertain?

  • What changed between one iteration and the next?

  • Can a human inspect the path from question to result?

GnosisLab makes those questions part of the system itself.


Quick start

Install from PyPI

uv tool install gnosislab

or:

pip install gnosislab

Full setup example

Register the five servers with your MCP client, then start the stack:

# see which clients are supported, preview without writing
gnosislab setup --list
gnosislab setup claude-code --dry-run    # or opencode, vscodium, vscode,
                                         # windsurf, cursor

# merge the server entries into the client's MCP config
gnosislab setup claude-code

# background all five servers (upstreams first)
gnosislab start all

# pid/ports/health per server
gnosislab status

# resolved ports + config file path
gnosislab config

# tail one server's log
gnosislab logs episteme

# open the lab status dashboard in your browser (agora by default)
gnosislab dashboard            # → ML_AGORA_GUI_URL (default http://localhost:38051)
gnosislab dashboard episteme   # any server's GUI works too

# stop everything (consumers first)
gnosislab stop all

Ports resolve as: built-in defaults < ~/.config/gnosislab/ports.env < environment variables — same KEY=VALUE format as the repo's ports.env. Runtime state always lives in ~/.ml-<name>/, regardless of install method. The dashboard URL comes from the resolved ML_<NAME>_GUI_URL, so remote/VM setups can point it at a reachable host without editing scripts.

Full documentation: https://loop.cloudcell.workers.dev/docs

Install from source

git clone https://github.com/cloudcell/gnosislab.git
cd gnosislab
uv sync
./gnosislab start all

Linux is currently required for sealed trial execution.
The executor uses bubblewrap mount namespaces and strace.

On Debian/Ubuntu:

sudo apt install bubblewrap strace

The idea in one picture

flowchart LR
    Q["Research question"] --> H["Hypothesis"]
    H --> Z["Zetesis<br/>search & evidence discovery"]
    Z --> E["Episteme<br/>experiment & observation"]
    E --> X["Sealed execution"]
    X --> O["Artifacts & observations"]
    O --> A["Anamnesis<br/>claims, evidence & memory"]
    A --> C["Conclusion"]
    C --> R["Arete<br/>adaptation & improvement"]
    R --> H

    G["Agora<br/>coordination & lab-level view"] -.-> Z
    G -.-> E
    G -.-> A
    G -.-> R

The point is not the Greek names. The point is separation of scientific responsibilities.

Each subsystem owns a distinct part of the research process rather than collapsing everything into one agent loop and one unstructured memory.


How it works

GnosisLab is composed of five MCP services.

Service

Role

Think of it as

Episteme

experiments, trials, observations and experimental state

the laboratory notebook

Anamnesis

claims, evidence relationships and persistent research memory

the evidence graph

Zetesis

search, investigation campaigns and evidence discovery

the investigator

Arete

improvement, comparison and meta-level adaptation

the critic / improver

Agora

coordination and lab-level observability

the control room

The services deliberately keep their state boundaries separate. This makes ownership clearer, failures easier to isolate, schemas easier to evolve, and each subsystem independently inspectable.


Architecture

flowchart TB
    U["Human / AI agent / MCP client"]

    subgraph GL["GnosisLab"]
        AG["Agora<br/>coordination"]
        ZE["Zetesis<br/>search"]
        EP["Episteme<br/>experiments"]
        AN["Anamnesis<br/>claims + evidence"]
        AR["Arete<br/>improvement"]

        AG --> ZE
        AG --> EP
        AG --> AN
        AG --> AR

        ZE --> EP
        EP --> AN
        AN --> AR
        AR --> ZE
    end

    subgraph EX["Execution boundary"]
        SE["Sealed experiment executor"]
        ART["Code, logs, outputs,<br/>observations, artifacts"]
        SE --> ART
    end

    U --> AG
    U --> ZE
    U --> EP
    U --> AN
    U --> AR

    EP --> SE
    ART --> EP
    ART --> AN

Default service ports

Service

GUI

MCP

Agora

38051

38050

Arete

38061

38060

Zetesis

38071

38070

Episteme

38081

38080

Anamnesis

38091

38090

Defaults can be overridden with:

~/.config/gnosislab/ports.env

or the corresponding ML_* environment variables.

Runtime state lives outside the repository under:

~/.ml-<name>/

For researchers

GnosisLab is designed around a simple principle:

A scientific result should remain inspectable after the agent that produced it has moved on.

That means treating research state as durable data rather than transient conversation context.

Explicit experimental structure

Research can be represented in terms of programmes, trials, observations, evidence and claims rather than buried inside a chat transcript.

Provenance

The system is designed to retain the relationship between experimental actions, artifacts, observations, evidence and downstream claims.

Reproducibility

Experiment execution is separated from conversational reasoning. The aim is to preserve the material needed to understand and reproduce what was actually done.

Human review

GnosisLab is not built around the assumption that an autonomous agent should be trusted by default. Its structure is intended to make intermediate state visible enough for review, correction and override.

Uncertainty belongs in the record

A research system should distinguish a claim from the evidence supporting it and should avoid turning an agent's confidence into an unexplained magic number.


For developers

GnosisLab is infrastructure, not a monolithic chatbot.

MCP-native

The distribution exposes five MCP server entry points:

ml-episteme-mcp
ml-anamnesis-mcp
ml-zetesis-mcp
ml-arete-mcp
ml-agora-mcp

This means an MCP-capable agent or client can work with the scientific process as a set of tools rather than requiring a proprietary frontend.

One command for lifecycle management

gnosislab start all
gnosislab status
gnosislab logs arete
gnosislab stop all

Sealed experiment execution

Trial execution uses Linux isolation primitives so experiments can run inside a constrained boundary rather than directly in the agent's host environment.

Observable by design

Every service exposes a local observability GUI. You can inspect the system while it is operating instead of treating agent activity as an invisible background process.

Separate state ownership

The major research loops keep their own persistent stores. This avoids turning one shared database into an ambiguous write boundary and lets each service evolve independently.

Small dependency surface

The published package targets Python 3.10+ and keeps the core runtime dependency set focused around MCP, Pydantic and HTTPX.


A research loop, expressed as software

A conventional agent loop often looks roughly like this:

prompt → tool calls → more prompting → answer

GnosisLab makes the scientific structure explicit:

question
  ↓
hypothesis
  ↓
search / prior evidence
  ↓
experiment design
  ↓
sealed execution
  ↓
observations + artifacts
  ↓
evidence-linked claims
  ↓
conclusion
  ↓
critique / adaptation
  ↓
next hypothesis

That difference matters when the work is long-running, iterative, expensive, safety-sensitive, or intended to survive peer review.


What GnosisLab is not

GnosisLab is not:

  • a replacement for a capable language model;

  • a claim that autonomous agents can replace scientific judgment;

  • a generic vector-memory wrapper;

  • a single prompt that tells an LLM to "act like a scientist";

  • an opaque autonomous system that asks you to trust the final answer.

It is an attempt to make the process around AI-assisted research explicit, inspectable and programmable.


Example use cases

Machine-learning research

An agent can propose an architecture change, execute a controlled trial, capture metrics and artifacts, attach evidence to the resulting claim, and use the result to choose the next experiment.

Reproducible benchmarking

Benchmark runs can be treated as experiments with explicit observations rather than as disconnected terminal commands and copied numbers.

Literature-to-experiment workflows

Search and evidence discovery can feed hypotheses and experimental design while keeping source evidence distinct from experimentally generated evidence.

Long-running autonomous investigation

The system can preserve research state across many agent turns without relying on a single ever-growing context window.

Human-supervised research agents

A researcher can inspect intermediate claims, evidence, trial outcomes and service state rather than reviewing only a final generated report.


Project layout

gnosislab/
├── src/                # gnosislab CLI + five ml_*_mcp packages
├── tests/              # test suite (pytest-xdist)
├── constants/          # grounded-constants registry + disclosure
├── diagnostics/        # diagnostics corpus
├── gnosislab           # lifecycle CLI entry point
├── labloop             # transition compatibility symlink
├── ml-*.toml           # per-server config
├── ports.env           # port assignments
├── pyproject.toml      # distribution metadata
├── LICENSE
├── NOTICE
└── README.md

The repository root is the runnable project:

uv sync

Configuration

Per-service configuration lives under:

ml-*.toml

Port configuration:

ports.env

User overrides:

~/.config/gnosislab/ports.env

Inspect the resolved configuration at any time:

gnosislab config

Observability

After startup, each service exposes a local GUI.

For example:

Agora      http://localhost:38051
Arete      http://localhost:38061
Zetesis    http://localhost:38071
Episteme   http://localhost:38081
Anamnesis  http://localhost:38091

Use:

gnosislab status

for process, port and health information.

Use:

gnosislab logs <service>

to inspect a service directly.


Testing

Run the full suite:

uv run pytest -q

The repository currently contains 1,289 tests.

For a serial run:

uv run pytest -n0

Development status

GnosisLab is alpha software.

The architecture is usable and the package is published, but interfaces, schemas and internal boundaries may continue to evolve.

If you are evaluating it for serious research, inspect the code, test the failure modes, and treat the current release as research infrastructure under active development.


Install and explore

uv tool install gnosislab
gnosislab start all
gnosislab status

Documentation: https://loop.cloudcell.workers.dev/docs

PyPI: https://pypi.org/project/gnosislab/

Repository: https://github.com/cloudcell/gnosislab


Contributing

GnosisLab is most useful when challenged by people who care about scientific rigor, agent engineering, reproducibility, provenance, evaluation and research infrastructure.

Useful contributions include:

  • adversarial testing of the scientific state machine;

  • new experiment and evidence workflows;

  • integrations with MCP clients and research agents;

  • reproducibility and provenance improvements;

  • better observability;

  • failure-mode documentation;

  • benchmark suites;

  • documentation and examples.

Open an issue with a concrete failure case, proposed experiment, or implementation idea.


Philosophy

AI makes it dramatically cheaper to generate hypotheses, code, analyses and prose.

It does not automatically make those outputs scientific.

The difficult part is preserving the chain between:

what was proposed → what was done → what was observed → what counts as evidence → what may reasonably be concluded

GnosisLab exists to make that chain a first-class object.


License

Apache License 2.0. See LICENSE.


Stop collecting runs. Start running programmes.

pip install gnosislab

GnosisLab — hypothesis in, auditable conclusion out.

Available Tools

46 tools
abandon_hypothesisA

Abandon a hypothesis — attributed, rationaled, never erased.

The exit for a hypothesis that can never conclude: falsified premises, superseded questions, trials that can only fail. Abandoning is NOT a verdict — no conclusion is recorded and no claim is minted; the hypothesis simply leaves the loop as 'abandoned', stamped with who decided and why. Trials already recorded stay on the record.

Applies to proposed and under_test hypotheses; terminal states (accepted|rejected|inconclusive|abandoned) refuse. To abandon every hypothesis in a programme at once, use close_programme(status="abandoned").

ParametersJSON Schema
NameRequiredDescriptionDefault
rationaleYesNon-empty justification — why this hypothesis leaves the loop without a verdict.
decided_byYesAttributable decider — 'human:<name>' or an agent identity.
programme_idYesID of the owning research programme.
hypothesis_idYesID of the hypothesis to abandon.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
decided_byNo
from_statusNo
hypothesis_idNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it clarifies the operation is not a verdict, mints no claim, records no conclusion, stamps attribute and rationale, preserves already-recorded trials, and enumerates the source states it applies to versus the terminal states it rejects. These are the exact preconditions and side-effect semantics an agent needs for a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core distinction (abandon vs. verdict) before the supporting conditions. It runs slightly long and repeats the 'not a verdict' idea across clauses, but every sentence still carries operational meaning, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. Combined with full param coverage, stated preconditions, side-effect transparency, and a named alternative for the bulk case, the definition gives an agent everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so the baseline is 3. The description adds semantic weight beyond the schema by framing the operation as needing attribution and rationale ('stamped with who decided and why'), which explains why decided_by and rationale are required rather than optional metadata.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

A specific verb+resource ('Abandon a hypothesis') with an immediate scope modifier ('attributed, rationaled, never erased'). It explicitly distinguishes itself from the conclude path ('Abandoning is NOT a verdict — no conclusion is recorded') and names close_programme for the bulk case, so an agent can route without opening sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States exactly when to use it ('The exit for a hypothesis that can never conclude: falsified premises, superseded questions, trials that can only fail'), when it is refused (terminal states accepted|rejected|inconclusive|abandoned), and which alternative covers the bulk operation (close_programme(status="abandoned")). When, when-not, and alternatives are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

acknowledge_violationA

Acknowledge an open integrity violation — insert-only.

The check log is append-only; this records the disposition (remediated | accepted-with-reason) against (check_name, object_ref) so the finding stops gating writes and leaves the digest's blockers. The underlying record is never touched — remediation itself is done by the corrective tools first.

ParametersJSON Schema
NameRequiredDescriptionDefault
check_nameYesName of the integrity check that flagged the violation (as shown in check_invariants or the status digest blockers).
decided_byYesWho acknowledges — 'agent:<name>' or 'human:<name>'. Attribution is required.
object_refYesThe flagged record's reference (e.g. trial-…), exactly as reported by the check.
dispositionYesWhat was done about it — e.g. 'remediated via correct_trial_status', 'accepted: trial ran before sealing was enforced'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
ack_idNo
statusNo
open_violationsNo
matched_open_violationNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses insert-only/append-only semantics, that the underlying record is never touched, and the concrete consequence (stops gating writes, clears digest blockers). It omits permission/auth requirements and reversibility of the acknowledgement itself, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action is front-loaded in the first sentence and the following sentences justify the append-only behavior without redundancy. Slightly more prose than strictly necessary, but each sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the mutation's behavioral impact thoroughly. The only gap is the absence of permission/attribution rules beyond the decided_by parameter, which is minor for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning by enumerating the canonical disposition values (remediated | accepted-with-reason) and framing the key as the (check_name, object_ref) pair. This enriches the schema's example-based disposition field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (acknowledge) and resource (an open integrity violation) and clarifies the effect — the finding stops gating writes and leaves the digest's blockers. It also implicitly distinguishes itself from the corrective tools by noting remediation is done elsewhere, so an agent can separate it from siblings like correct_trial_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context for when to call it — after remediation, which 'is done by the corrective tools first' — giving an ordering constraint against those alternatives. It stops short of an explicit when-not or a named alternative tool, so it lands at clear-context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_pending_programmesA

Archive terminal programmes that are still in the live DB.

This is a migration tool for programmes that were closed (completed or abandoned) before the automatic archiver was implemented. It scans the live DB for terminal programmes not yet archived and archives them.

dry_run=true reports the scan result ({would_archive, would_mark_archived, skipped}) without writing anything.

Under normal operation, this is a no-op — close_programme already archives automatically. This tool only catches programmes that slipped through (e.g. abandoned before Phase 3 was deployed).

Returns: {archived: [...], errors: [...], skipped: int}

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runYesWhen true, report what would be archived without mutating. Required — pass false to apply.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorsNo
dry_runNo
skippedNo
archivedNo
would_archiveNo
marked_archivedNo
would_mark_archivedNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does most of it: it discloses the mutation ('archives them'), the dry_run behavior ('reports the scan result without writing anything'), the no-op/idempotent nature, and the return shape. It does not mention required permissions, rate limits, or whether re-running is safe beyond the no-op claim, so it falls just short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then scope, then dry_run semantics, then the no-op caveat, then the return shape. Every sentence earns its place, though the trailing 'Returns:' line partially duplicates the existing output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter migration tool with an output schema provided, the description covers purpose, when-not-to-use, dry_run semantics, idempotency, and return shape. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is only one parameter, so the baseline would be 3. The description adds genuine value beyond the schema by naming the exact dry_run payload fields ({would_archive, would_mark_archived, skipped}), which tells the agent what to expect before committing to a write.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('archive terminal programmes'), and crisply distinguishes itself from the sibling close_programme by framing itself as a one-off migration catch-up tool. An agent can tell it apart from close_programme, list_archives, and get_archived_programme without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when this tool is and is not needed: 'Under normal operation, this is a no-op — close_programme already archives automatically. This tool only catches programmes that slipped through.' It names the alternative and the exact condition selecting this tool, so the agent knows to prefer close_programme in the normal path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_programmeB

Assess programme health: progressive vs degenerating (commitment 8).

Reads observations from state.db and best_trials from the optimizer.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
healthNo
programme_idNo
total_trialsNo
completed_trialsNo
total_hypothesesNo
total_conclusionsNo
total_observationsNo
best_trials_from_optimizerNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral burden. It does signal a read-oriented operation ('Reads observations... and best_trials'), but never explicitly states it is side-effect free, nor whether it is expensive or requires prior observations to exist. Output schema existence relieves it of return-value detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core purpose front-loaded and no filler. The trailing blank line and the unexplained '(commitment 8)' tag are the only minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema and a fully documented single parameter, the description supplies adequate purpose and data-source context for a diagnostic tool. It still omits whether the operation is purely read-only and what preconditions (e.g., existing observations/trials) must hold, leaving a gap given the absence of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (programme_id) with 100% schema description coverage, so the schema already documents it fully. The description adds no meaning about the ID's format or constraints, hitting the baseline for a well-covered single-parameter schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (assess) and resource (programme) plus the concrete judgment it produces (progressive vs degenerating), which distinguishes it from sibling mutation/query tools like get_trial_status or check_invariants. The parenthetical '(commitment 8)' is unexplained internal jargon but does not obscure the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to invoke this versus alternatives, no prerequisites, and no exclusions. The sentence about reading state.db and best_trials describes data sources, not usage conditions, so an agent gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_trialA

Cancel a running trial, or abandon a designed one.

running → failed: kills the subprocess, marks failed. The record's failed here reflects the cancellation act, not an observed execution failure — cancelled: true in the payload marks the distinction, and an already finished executor reply means the executor completed first. A cancel landing before the task registers tombstones the id so no child can spawn under the terminal row. designed → abandoned: no executor to kill, no evidence lost — the honest terminal for a design that will never run (e.g. a bundle locked to code that no longer exists). The FSM has always permitted designed → abandoned; close_programme uses the same transition for programme-scoped sweeps.

Enforcement: commitment 1 — the loop is the unit (orphan check).

ParametersJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
trial_idNo
cancelledNo
executor_outputNo

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses the kill-subprocess side effect, the fact that 'failed' reflects the cancellation act rather than an observed failure, the cancelled:true payload marker, the 'already finished' executor reply, the tombstoning of ids that land before registration, and that designed→abandoned loses no evidence. This is unusually rich side-effect and edge-case disclosure beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the body is long, jargon-dense (FSM, tombstones, orphan check), and includes historical/tangential notes like the FSM always permitting the transition and close_programme reusing it. These are informative but not proportionate to tool selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description still covers both terminal transitions and the key edge cases. Remaining gaps are minor (e.g., permission or retry semantics), so it is nearly complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both trial_id and programme_id are already documented in the schema. The description adds no syntax, format, or scoping detail about either parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Cancel a running trial, or abandon a designed one') and splits the resource into its two distinguishable states. This routes clearly against siblings like close_programme and abandon_hypothesis without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit conditions for each mode: running trials get cancelled/killed, designed trials get abandoned as 'the honest terminal for a design that will never run.' It also notes close_programme shares the designed→abandoned transition for programme-scoped sweeps, which helps disambiguate. No explicit 'do not use this when...' exclusion, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_bundleA

Seal the auxiliary bundle for a DESIGNED trial (commitment 6).

ORDERING: call AFTER design_experiment and BEFORE run_trial. The seal is pre-registration — the bundle fixes the auxiliary assumptions (code/env/seeds/splits) before any observation, so they cannot be retro-fitted to results. Trials that have left 'designed' are rejected. The post-run counterpart is executed_code.json — what actually ran, captured at finalization from the strace read-trace.

code_ref MUST be a path to a Python file that exposes: def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. The config is the same dict passed to design_experiment. Read the executor://contract resource for the full contract.

data_refs is an optional list of DataRef IDs (from prepare_data). When provided, the bundle records structured data provenance. When omitted, splits is used (backward-compatible).

extra_code_refs is an optional list of additional .py file paths that the trial depends on but cannot be discovered by AST import analysis — e.g. scripts invoked via subprocess.run(). These are captured into code_snippets and stored in code_hash_extra_json so the bundle is fully self-contained and rerunnable from archive (commitment 2 + 5 — the bundle must contain ALL code needed to reproduce, not just the executor).

Enforcement: commitment 1 — the loop is the unit (trial must exist). Enforcement: commitment 6 — the bundle must be controlled (code_ref validated). Concurrency: rejects if trial already has a bundle (no double capture).

Structured params (seeds, splits, data_refs, extra_code_refs) may be sent as JSON-encoded strings; seeds also accepts a bare int.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedsYesSeed set sealed into the bundle (pre-registration — fixes the assumption before any observation); list or JSON-encoded.
splitsYesSplit spec sealed into the bundle — which data each split used; object or JSON-encoded.
env_refYesEnvironment reference sealed into the bundle (pre-registration). Accepts a venv/conda directory (mounted read-only; its bin/python runs the trial) or a Python executable path (its venv root is mounted). Non-path values like 'python:3.12' are recorded as provenance but do not change the interpreter — the executor default runs.
code_refYesPath to a Python file exposing run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}} — see executor://contract.
trial_idYesID of the target trial.
data_refsNoDataRef IDs from prepare_data — structured data provenance (falls back to splits when omitted); list or JSON-encoded.
baseline_refNoReference to the baseline the trial compares against.
extra_code_refsNoAdditional code files to seal into the bundle; list or JSON-encoded.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
code_refNo
warningsNo
bundle_idNo
code_hashNo
code_hash_extraNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full load and delivers: rejection on non-designed trials, no double-capture on a trial that already has a bundle, enforcement commitments (loop is the unit; code_ref validated), and the required code_ref callable contract. Runtime safety profile and constraints are fully disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and ordering, then structured by headings (Enforcement, Concurrency). Slightly long with some restatement of 'pre-registration' and repeated commitment framing, but each block earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values and doesn't. It instead completes the operational picture: sequencing, error conditions, contract resource pointer, and JSON-encoding conventions for structured params. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds real meaning beyond it: code_ref must expose run_training(config) -> dict with specific return keys, extra_code_refs exists for subprocess-invoked files undetectable by AST analysis, and data_refs falls back to splits. This explanation of intent and edge cases exceeds the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Seal the auxiliary bundle') scoped to a DESIGNED trial, and immediately distinguishes the action from run_trial and design_experiment by naming the ordering. An agent knows exactly what this does and where it sits in the loop.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit ORDERING clause (AFTER design_experiment, BEFORE run_trial), explicit exclusions ('Trials that have left designed are rejected'), and names the post-run counterpart (executed_code.json). When/when-not/alternatives are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_bundle_from_code_hashA

Capture a bundle using a content address (code_hash) instead of a file path.

This is the rerun-from-archive path: the LLM reads an archived programme, gets the code_hash from the bundle, and captures a new bundle for a new trial using the same code content. No filesystem access is required — the code is loaded from code_snippets by hash, materialized as a real file next to the execution wrapper, and imported at run_trial time.

NOTE: code_hash is a "sha256:..." content address returned by a prior capture_bundle — it is NOT a bundle_id. To re-use an existing bundle's code, pass its code_hash field.

The code_hash must already exist in code_snippets (captured by a prior capture_bundle or prepare_data call). If it doesn't, the tool returns an error.

The bundle's code_ref is set to "code://{code_hash}" — a content address, not a file path. This is the carrier/content separation (Rule 5.4) made explicit: the bundle references the ICE directly.

code_hash_extra is an optional list of additional content addresses for locally-imported or subprocess-dispatched modules captured alongside the primary. Each must already exist in code_snippets. At run_trial time, these are materialized as real files under a _deps/ dir on sys.path, reconstructing each file's path suffix — so from src.mod import x resolves through the real import machinery (commitment 2 + 5).

All other parameters are identical to capture_bundle.

Structured params (seeds, splits, data_refs, code_hash_extra) may be sent as JSON-encoded strings; seeds also accepts a bare int.

Returns: {"bundle_id": "bundle-...", "status": "captured", "code_hash": ...}

ParametersJSON Schema
NameRequiredDescriptionDefault
seedsYesSeed set sealed into the bundle (pre-registration); list or JSON-encoded.
splitsYesSplit spec sealed into the bundle; object or JSON-encoded.
env_refYesEnvironment reference sealed into the bundle (pre-registration). Accepts a venv/conda directory (mounted read-only; its bin/python runs the trial) or a Python executable path (its venv root is mounted). Non-path values like 'python:3.12' are recorded as provenance but do not change the interpreter — the executor default runs.
trial_idYesID of the target trial.
code_hashYes'sha256:...' content address from a prior capture_bundle — NOT a bundle_id; must already exist in code_snippets.
data_refsNoDataRef IDs from prepare_data — structured data provenance; list or JSON-encoded.
baseline_refNoReference to the baseline the trial compares against.
code_hash_extraNoContent hashes of additional code files to seal; list or JSON-encoded.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
code_refNo
warningsNo
bundle_idNo
code_hashNo
code_hash_extraNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full behavioral burden, and it does: the prerequisite that code_hash must already exist in code_snippets or the tool errors, the side effect of materializing a real file next to the execution wrapper at run_trial time, the absence of filesystem access, and the code_ref format. Failure mode, side effects, and lifecycle timing are all disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and information-dense, with the routing distinction first and prerequisites later. A few sentences lean on internal jargon that adds little for tool selection ('the carrier/content separation (Rule 5.4) made explicit', 'commitment 2 + 5'), which is the only drag on an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out (though the description still gives the shape). Prerequisites, error behavior, parameter roles, and the run_trial-time consequences are all covered for an 8-parameter, 5-required tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds real meaning the schema lacks: code_hash is a content address and NOT a bundle_id, code_hash_extra entries are materialized under a `_deps/` dir on sys.path so `from src.mod import x` resolves, and structured params accept JSON-encoded strings. It also clarifies the resulting code_ref is 'code://{code_hash}' rather than a path.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb and resource and pins the distinguishing mechanism: 'Capture a bundle using a content address (code_hash) instead of a file path.' That contrast with the sibling capture_bundle is stated up front, so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames the use case ('This is the rerun-from-archive path: the LLM reads an archived programme, gets the code_hash from the bundle, and captures a new bundle for a new trial') and notes no filesystem access is required. The alternative (plain capture_bundle) is implied by 'instead of a file path' and 'all other parameters are identical to capture_bundle', but the when-not condition is never stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_pending_artifactsA

Capture artifact files for trials that predate artifact capture.

Scans all trials with an artifact_path but no rows in trial_artifacts. Reads files from disk and stores them in SQLite (gzip-compressed). Files over 50MB are skipped and marked oversized. Files that no longer exist on disk are marked lost.

dry_run=true reports the pending trial ids without reading or storing anything.

cleanup: when True, staging files are deleted after successful capture (the SQLite copy becomes authoritative). Default False keeps the originals — conservative for migration of old trials.

Under normal operation this is a no-op — _finalize_trial already captures artifacts automatically. This tool only catches trials that slipped through (finalized before artifact capture was implemented).

Returns: {captured: [...], oversized: [...], lost: [...], trials_scanned: int, trials_with_existing: int}

ParametersJSON Schema
NameRequiredDescriptionDefault
cleanupNoDelete staging files after successful capture (the SQLite copy becomes authoritative).
dry_runYesWhen true, report which trials would be captured without mutating. Required — pass false to apply.

Output Schema

ParametersJSON Schema
NameRequiredDescription
lostNo
capturedNo
oversizedNo
trials_scannedNo
trials_with_existingNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does: storage target (SQLite, gzip-compressed), >50MB files skipped and marked oversized, missing files marked lost, cleanup deletes staging files making SQLite authoritative, and dry_run performs no mutation. These are exactly the side effects an agent needs before invoking a mutating backfill.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in one line, then introduces the scan criteria, failure modes, dry_run, and cleanup. Slight redundancy with the schema descriptions for dry_run/cleanup and with the output schema's return list, but every section is scannable and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers scope, discovery criteria, failure handling, mutation semantics, and the no-op caveat for a mutating tool with no annotations. An output schema exists, so repetition of the return shape is harmless rather than a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented, capping the ceiling. The description still adds meaning beyond the schema by explaining the consequence of cleanup ('the SQLite copy becomes authoritative') and the conservative default rationale for migrating old trials.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('capture artifact files') and narrows scope precisely: trials with an artifact_path but no rows in trial_artifacts. The backfill framing distinguishes it from the normal capture path used by _finalize_trial, so an agent won't confuse it with capture_bundle or run_trial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it and when not to: 'Under normal operation this is a no-op — _finalize_trial already captures artifacts automatically. This tool only catches trials that slipped through.' It also documents the dry_run=false condition required to apply changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_invariantsA

Audit the experiment store against the loop's invariants.

Returns {status: ok|violations, checks: [{name, ok, violations, detail}]}. Detection complement to write-time enforcement: orphaned running trials, completed trials lacking observations, unsealed executions, strace divergence, undigested input data, budget overruns, stuck hypotheses, stalled trials. Report-only — nothing is repaired or mutated. The run is logged to /logs/ (retention: [integrity] log_max_files).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
checksNo
serverNo
statusNo
triggerNo
log_fileNo
checked_atNo
duration_msNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantially: it declares 'Report-only — nothing is repaired or mutated,' specifies the log destination and retention policy, and gives the exact return shape. Gaps remain around permissions/auth needs and runtime cost, but the mutation-safety claim and side-effect disclosure are exactly what an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and return shape in the first two sentences, then the detection list, then the report-only guarantee and log path. The enumeration of eight violation types is long but each item is a concrete detectable condition, so the length is earned rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the description still previews the return structure ({status, checks:[...]}), correctly avoiding redundant explanation of that schema's internals. It covers scope, safety, side effects, and logging, leaving only unstated items like authorization requirements and whether the check is read-only in the strict sense.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to clarify and it appropriately spends no words on argument syntax. It instead documents the return payload, which is the relevant detail for a no-arg tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Audit the experiment store against the loop's invariants.' No sibling tool performs invariant auditing, so the agent can pick it out immediately. It also enumerates the exact classes of problems detected (orphaned running trials, unsealed executions, budget overruns, etc.), which sharpens the purpose rather than restating the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Detection complement to write-time enforcement' frames the tool as the read-side counterpart to enforcement, which implies when to reach for it. However, no sibling is named as an alternative and there is no explicit when/when-not guidance or prerequisite (e.g. run after run_trial/list_trials, or health-check cadence). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_programmeA

Close a programme by setting its status to completed or abandoned.

For status="completed": rejects if any hypothesis is still under_test — the agent must call conclude_hypothesis first. Auto-marks proposed → abandoned (no effort was expended).

For status="abandoned": auto-marks all unconcluded hypotheses as abandoned (the whole programme is being abandoned).

Enforcement: state machine — rejects if programme is not active. Enforcement: commitment 1 — the loop is the unit (no closing with unconcluded hypotheses when completing).

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNocompleted | abandoned — 'completed' requires all hypotheses concluded; 'abandoned' sweeps the rest.completed
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
archivedNo
auto_markedNo
programme_idNo
archive_errorNo
trials_auto_markedNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: rejection conditions (under_test hypotheses for completed), side effects (proposed → abandoned; unconcluded → abandoned), and state-machine enforcement (rejects if programme is not active). This is rich behavioral disclosure an agent cannot get elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then cleanly sectioned by status value and enforcement rules. All content earns its place, though the trailing 'commitment 1' line is somewhat opaque framing rather than actionable instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described. The description covers purpose, per-mode behavior, prerequisites, side effects, and enforcement — everything an agent needs to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds meaning beyond the schema's enum note by explaining the automatic side effects each status value triggers (auto-marking proposed/unconcluded hypotheses). That is genuine added semantic value on top of structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (close a programme) and immediately specifies the two outcome modes via status. It is clearly distinguishable from siblings like conclude_hypothesis and abandon_hypothesis because it operates at programme scope rather than hypothesis scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit guidance for each status: 'completed' requires all hypotheses concluded and directs the agent to conclude_hypothesis first; 'abandoned' sweeps unconcluded hypotheses. It names the prerequisite sibling tool, though it does not explicitly contrast against abandon_hypothesis for the alternative path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

conclude_hypothesisA

Conclude a hypothesis by accepting/rejecting it.

Creates an immutable conclusion (commitment 8: programmes, not runs). When a claims role is wired, also mints a semantic claim with provenance edges — a byproduct, never a gate on the verdict. Enforcement: commitment 8 — rejects duplicate conclusions for the same hypothesis. Enforcement: commitment 10 — rejects if hypothesis not found or not in this programme. Enforcement: commitment 1 — rejects if no completed trials for the hypothesis. Enforcement: commitment 7 — rejects if any completed trial has no observations.

ParametersJSON Schema
NameRequiredDescriptionDefault
verdictYesaccepted | rejected | inconclusive — creates an immutable conclusion.
programme_idYesID of the target research programme.
hypothesis_idYesID of the target hypothesis.
evidence_summaryYesSummary of the evidence the verdict rests on.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
verdictNo
claim_errorNo
claim_statusNo
conclusion_idNo
hypothesis_idNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so richly: it declares the conclusion immutable, discloses the side effect of minting a semantic claim with provenance edges (explicitly framed as a byproduct, not a gate), and enumerates four distinct rejection conditions. This is exactly the mutation-side-effect disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first sentence and organizes preconditions as parallel 'Enforcement:' lines, which is scannable. Slightly verbose with repeated 'Enforcement: commitment N' framing, but every clause conveys a real rejection condition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with an output schema present, the description covers purpose, immutability, side effects, and all rejection preconditions. Nothing an agent needs to invoke it correctly is missing, and return values are handled by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the verdict enum. The description adds no format, syntax, or constraint detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Conclude a hypothesis by accepting/rejecting it,' and the verdict enum makes the operation concrete. It is distinguishable from abandon_hypothesis by the accept/reject/inconclusive framing, though the description never explicitly contrasts the two siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The four 'Enforcement' clauses effectively define the preconditions for calling it: completed trials must exist, each must have observations, no duplicate conclusion, and the hypothesis must belong to the programme. That is clear usage context, but it never names an alternative tool (e.g., abandon_hypothesis) or when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

correct_trial_statusA

Correct a terminal trial's status — the recorded repair act.

For mislabeled records: e.g. the executor's outer wrapper exited 0 while the recorded output documents a crash ("status": "error" / nonzero inner exit code in the payload). Source must be completed|failed|retryable (all terminal — a retryable source is a correction made in error or evidence re-read after the fact); target must be completed|failed|retryable. reason is mandatory. Correcting TO completed requires the record to evidence a completed run — a cancellation receipt, reaper note, or failure output does not qualify, so that direction is refused unless the executor record documents an actual completion.

The correction is appended to the trial's executor_output_json under 'corrections' — the record shows both what was claimed and what it was corrected to. Never hand-edit the database: a direct sqlite3 UPDATE bypasses this audit trail.

Enforcement: commitment 1 — the loop is the unit (orphan check).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy the correction is recorded — mandatory, appended to the trial's audit trail.
trial_idYesID of the target trial.
to_statusYesTarget status: completed | failed | retryable.
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toNo
fromNo
trial_idNo
corrected_atNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: the correction is appended to executor_output_json under 'corrections', both claimed and corrected states are retained, and the direction-toward-completed refusal rule is disclosed. The 'reason' mandate and the ban on direct DB edits add concrete behavioral context an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first line, and the subsequent sentences all carry weight (valid statuses, refusal rule, audit destination, DB warning). It is somewhat verbose and the closing 'commitment 1 — the loop is the unit (orphan check)' line is opaque, keeping it just short of maximal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description instead covers everything else an agent needs: valid status directions, mandatory reason, the refusal condition on correcting to completed, and the audit-trail side effect. Nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters and the baseline is 3. The description adds real meaning: it explicates that the source status is effectively terminal-only (completed|failed|retryable) and constrains to_status semantically rather than just enumerating it, plus reinforces that reason is mandatory and fed into the audit trail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Correct a terminal trial's status') and characterizes it as 'the recorded repair act', which distinguishes it from siblings like mark_retryable, run_trial, and cancel_trial. An agent immediately understands this is an audited correction of an existing terminal status, not a status transition driver.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit motivating scenario (executor wrapper exited 0 while output documents a crash) and states the valid source/target status pairs. It also states a negative condition — correcting TO completed is refused unless the record evidences an actual completion — and warns against direct sqlite3 UPDATE. When-to-use, when-refused, and the alternative-to-avoid are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_evaluation_contractA

Create an evaluation contract for a programme (RSI Phase 0).

Contracts are insert-only and versioned per programme: each new contract for the same programme gets version = max+1 and supersedes the previous. metrics and promotion_policy may be sent as JSON-encoded strings.

The powered-policy gate enforces presence of the design keys (sesoi_d, target_power, min_evidence_rung), not adequacy — target_power: 0.5 pre-registers fine; adequacy surfaces downstream as underpowered / power_acknowledged on the campaign, not as a mint refusal.

Enforcement: programme must exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
budgetNoBudget recorded on the contract; object or JSON-encoded string.
metricsYesDeclares 'primary_metric' (+ optional 'direction': max|min, secondaries); may be JSON-encoded.
holdoutsNoHeld-out evaluation slices the contract declares; object or JSON-encoded.
programme_idYesID of the target research programme.
promotion_policyYesPre-registered promotion rule (confidence bar, safety gates); object or JSON-encoded.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
versionNo
contract_idNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: insert-only semantics, version = max+1, supersession of the previous contract, JSON-encoded string tolerance, and the exact behavior of the powered-policy gate (enforces key presence, not adequacy; underpowered surfaces downstream rather than as a refusal). This is unusually rich behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core verb and resource, then the versioning and gate mechanics. Dense but each sentence (versioning, JSON encoding, gate semantics, enforcement) earns its place; the gate discussion is the only slightly elaborate portion.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema, annotations absent, and 100% schema coverage, the description supplies everything an agent needs to call correctly: versioning model, encoding flexibility, gate expectations, and the programme-existence precondition. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema by naming the design keys the gate checks (sesoi_d, target_power, min_evidence_rung) and clarifying that the JSON-string form for metrics/promotion_policy is accepted at call time.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create an evaluation contract for a programme') and situates it in the workflow ('RSI Phase 0'). The create-vs-get distinction from the sibling get_evaluation_contract is implied but never made explicit, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(RSI Phase 0)' tag implies where this sits in the flow, and 'programme must exist' is a precondition, but there is no explicit when-to-use guidance, no named alternative, and no statement of when creating a new version is warranted over reading the existing contract.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_programmeA

Create a research programme with a goal, constraints, and budget.

Constraints are typed fields, not strings (commitment 4). Wires the optimizer role for this programme.

Structured params (constraints, allowed_variables, budget) may be sent as JSON-encoded strings if your client cannot emit objects.

metric_direction: 'minimize' (e.g. perplexity, loss) or 'maximize' (e.g. accuracy). The optimizer uses this to steer search and rank best trials.

candidate_version_id: optional RSI Phase 0 correlation — which registered candidate created this programme. Must exist if given.

investigation_id: optional zetesis correlation — which open investigation declared this programme as its obligation (open_investigation requires_programme → link_programme). Recorded as claimed provenance: episteme has no read path into zetesis, so the id is stored, not verified.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYesResearch goal — what the programme is trying to learn.
budgetYesBudget object {max_trials (required), max_wall_time_hours}; extra keys recorded but not enforced. May be a JSON-encoded string.
constraintsYesTyped constraint fields {gpu_memory_gb, max_training_hours_per_trial, max_parameter_count} — extra keys allowed (commitment 4); may be JSON-encoded.
investigation_idNoZetesis investigation this programme answers to — claimed provenance, recorded but not verified against zetesis (plan-20261002-1929Z).
metric_directionNo'minimize' (loss-like: perplexity, error) | 'maximize' (quality: accuracy, F1) — steers search and best-trial ranking.maximize
allowed_variablesYesVariables the programme may search/vary; list or JSON-encoded list.
candidate_version_idNoRSI Phase-0 correlation — which registered candidate created this programme; must exist if given.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
programme_idNo
budget_extras_not_enforcedNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does meaningful work: it discloses that constraints are typed fields rather than strings, that structured params may arrive as JSON-encoded strings, that the call wires the optimizer role, that metric_direction steers search and best-trial ranking, and that investigation_id is stored but never verified against zetesis. It does not state idempotency or what happens on duplicate goal, but the disclosure goes well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and each subsequent paragraph covers a distinct concern (constraints typing, optimizer wiring, JSON encoding, metric_direction, two optional correlation ids). Several parenthetical asides ('commitment 4', 'plan-20261002-1929Z') are noise, so it falls short of a clean 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a 7-parameter create tool with no annotations, the description covers the two non-obvious optional ids, the metric_direction semantics, and the JSON-string fallback, leaving only minor gaps like side effects or duplicate-name behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema (goal, budget, constraints, investigation_id, metric_direction, allowed_variables, candidate_version_id). The description largely restates those semantics — metric_direction, typed constraints, unverified provenance — adding little not already present. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

First sentence gives a specific verb+resource ('Create a research programme') plus the defining fields (goal, constraints, budget), making it easy to separate from list_programmes or close_programme. It does not explicitly name sibling tools, but the create verb is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to create a programme versus, say, formulate_hypothesis or design_experiment, nor any prerequisites for use. Only 'Wires the optimizer role for this programme' hints at context, which is a consequence, not when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_blobA

Describe a content-addressed blob by digest (read-only).

The resolution check behind digest claims: returns {exists, resolved_in, size_bytes, content_type, captured_at} across the content stores — artifact_files (HTTP ingest), code_snippets, bundles.code_hash. resolved_in names every store holding the hash; a digest may live in more than one. Existence + metadata only — bytes are never returned. This is the read upstream servers use to verify a digest before accepting it on a registration call. For the bytes themselves, get_blob.

ParametersJSON Schema
NameRequiredDescriptionDefault
content_hashYesThe sha256:<64 hex> digest to look up.

Output Schema

ParametersJSON Schema
NameRequiredDescription
existsNo
size_bytesNo
captured_atNo
resolved_inNo
content_typeNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares read-only behavior, that bytes are never returned, and enumerates the stores consulted (artifact_files, code_snippets, bundles.code_hash). It lacks auth/permission or rate-limit context, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and resource, and every sentence is informative (scope, stores, output fields, cross-reference). It restates the return fields that the output schema already covers, which is slight redundancy but not bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description need not explain return values, yet it usefully clarifies resolved_in semantics (multiple stores may hold a hash). For a single-parameter lookup tool the definition is fully sufficient to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already specifies the "sha256:<64 hex>" format for content_hash. The description adds only "by digest" and the notion of a digest living in multiple stores, which is marginal over the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Describe a content-addressed blob by digest") plus scope (read-only, metadata only). It explicitly contrasts with the sibling get_blob, so an agent can distinguish the two without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the concrete use case ("the read upstream servers use to verify a digest before accepting it on a registration call") and routes to the alternative for bytes ("For the bytes themselves, get_blob"). When-to-use and the alternative are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_experimentB

Design an experiment: create a trial (data item) within the programme.

config may be sent as a JSON-encoded string if your client cannot emit objects.

Enforcement: commitment 5 — budget is an epistemic resource (trial count + wall time). Enforcement: commitment 10 — the agent is a scientist (hypothesis must exist).

ParametersJSON Schema
NameRequiredDescriptionDefault
configYesTrial configuration dict handed to run_training; may be JSON-encoded.
programme_idYesID of the target research programme.
hypothesis_idYesID of the target hypothesis.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
trial_idNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two real behavioral traits: budget is enforced on trial count and wall time, and a hypothesis must pre-exist. However, the 'commitment 5' / 'commitment 10' jargon is unexplained external-document shorthand that an agent cannot act on, and nothing is said about failure modes, idempotency, or side effects of creating the trial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the JSON-string note is brief. The two 'Enforcement: commitment N' lines consume real space but convey little without the referenced commitment definitions, so the sizing is acceptable but not fully earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, for a creation tool with no annotations and many adjacent siblings, the description omits its relationship to run_trial and the trial lifecycle, leaving the agent with a minimum-viable picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all three parameters, so the schema already documents programme_id, hypothesis_id, and config. The description's note that config may be a JSON-encoded string merely repeats the schema's own text, adding no new meaning; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource: 'Design an experiment: create a trial (data item) within the programme.' An agent can tell this creates a trial record rather than executing one, but the description never differentiates itself from close siblings like run_trial or get_next_experiment, so the boundary is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no named alternative. The 'hypothesis must exist' enforcement note implies a precondition, but nothing tells the agent when to call design_experiment versus run_trial or get_next_experiment, which is the critical routing decision in this family.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

formulate_hypothesisB

Formulate a falsifiable hypothesis (commitment 3).

Rejects a hypothesis with no failure criterion. variables_involved may be sent as a JSON-encoded string list.

ParametersJSON Schema
NameRequiredDescriptionDefault
statementYesThe falsifiable claim under test.
programme_idYesID of the target research programme.
failure_criterionYesWhat observation would falsify the hypothesis — required (commitment 3).
variables_involvedYesVariables the hypothesis ranges over; list or JSON-encoded list.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
hypothesis_idNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose one real behavior — that a hypothesis without a failure criterion is rejected — which is valuable validation context. However it says nothing about permissions, whether this persists a new hypothesis record, or side effects, leaving meaningful gaps for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded with the core action and the key constraint. The '(commitment 3)' reference appears twice (header and failure_criterion mention), which is minor redundancy but the text stays lean overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and all four params are documented in the schema. Still, for a no-annotation tool that apparently creates a hypothesis under a programme, the description omits state-changing context and the relationship to sibling lifecycle tools, leaving it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The one parameter note ('variables_involved may be sent as a JSON-encoded string list') largely restates the schema text 'list or JSON-encoded list', so it adds little beyond the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (formulate) and resource (hypothesis) with the qualifier 'falsifiable', which is meaningful. It does not name or contrast with siblings like conclude_hypothesis, abandon_hypothesis, or design_experiment, so an agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus conclude_hypothesis, abandon_hypothesis, or design_experiment, nor any prerequisites such as requiring an existing programme. The only contextual hint is the '(commitment 3)' tag, which is internal jargon rather than usable routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_archiveB

Get details about an archive, including its programmes.

Returns the archive metadata (path, sealed status, hash) and the list of programmes archived in it.

ParametersJSON Schema
NameRequiredDescriptionDefault
archive_idYesID of the target archive.

Output Schema

ParametersJSON Schema
NameRequiredDescription
archiveNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the shape of the result (metadata fields plus the programme list), which is useful, but says nothing about permissions, whether the archive may be sealed/locked, or pagination of the programme list — relevant for something that can return many programmes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose. The second sentence lightly duplicates what an output schema already conveys, but nothing is wasted and it is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one required param) and an output schema exists, so return values need not be explained. The description covers purpose and result content adequately; the only real gap is disambiguation from sibling archive tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, archive_id, and schema description coverage is 100%, so the baseline is 3. The description adds no meaning beyond the schema (e.g. how to obtain an archive_id or what form it takes).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Get details about an archive') and enumerates what is returned (metadata: path, sealed status, hash; plus archived programmes). It is clear what the tool does, but it never distinguishes itself from close siblings like list_archives, get_archived_programme, or verify_archive, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no named alternative. The only signal that this is per-archive retrieval is the required archive_id, leaving the agent to guess how it differs from list_archives or get_archived_programme.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_archived_programmeA

Get an archived programme's details from its archive file.

Returns the programme metadata, hypotheses, trials, beliefs, and conclusions from the archive SQLite file. The programme is read-only — it cannot be modified or restored to the live DB.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the archived programme to read (read-only — cannot be restored).

Output Schema

ParametersJSON Schema
NameRequiredDescription
idNo
goalNo
statusNo
trialsNo
beliefsNo
bundlesNo
data_refsNo
row_countNo
archive_idNo
created_atNo
hypothesesNo
archived_atNo
conclusionsNo
archive_pathNo
observationsNo
code_snippetsNo
artifact_filesNo
trial_artifactsNo
constraints_jsonNo
metric_directionNo
budget_max_trialsNo
candidate_versionsNo
candidate_version_idNo
evaluation_contractsNo
allowed_variables_jsonNo
budget_max_wall_time_hoursNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden, and it does disclose the key behavioral trait: the programme is read-only and cannot be modified or restored to the live DB. It also enumerates the returned content (metadata, hypotheses, trials, beliefs, conclusions), though it omits failure behavior for a missing ID or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, and no filler. There is mild redundancy between 'from its archive file' and 'from the archive SQLite file', which keeps it just below the top mark.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with an output schema that covers the return structure, the description supplies the essential immutability context and a summary of returned content. The remaining gap is minor: no explicit sibling routing or error/edge-case notes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, and the schema already documents it as the archived programme ID (read-only, cannot be restored). The description adds no format or sourcing detail beyond restating that it reads an archived programme, so the baseline of 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get an archived programme's details from its archive file') and makes clear it operates on archived data rather than the live store, implicitly distinguishing it from siblings like get_archive or list_programmes. It stops short of naming a sibling alternative explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'archived' framing and the read-only constraint, but there is no explicit guidance on when to choose this over get_archive, list_archives, or list_programmes, nor on how to obtain a valid programme_id. A capable agent can infer the context but nothing routes it decisively.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_blobA

Retrieve a content-addressed blob's bytes by digest (read-only).

The read-back half of describe_blob: returns the stored bytes (base64 in content_b64) plus {digest, size_bytes, content_type, captured_at, resolved_in, served_from}. The returned bytes are re-hashed and verified against the requested digest — a blob whose stored bytes don't match its key is refused with the computed digest reported, never served.

Errors: malformed_digest | not_found | no_bytes | digest_mismatch | too_large. Read-only: no state is mutated.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_bytesNoMaximum returned payload size in bytes (pre-base64). Blobs larger than this return too_large with metadata but no bytes.
content_hashYesThe sha256:<64 hex> digest to retrieve.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
detailNo
digestNo
size_bytesNo
captured_atNo
content_b64No
resolved_inNo
served_fromNo
content_typeNo
computed_digestNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that bytes are re-hashed and verified against the requested digest, that mismatched blobs are refused with the computed digest reported rather than served, the enumerated error set, and the read-only guarantee. This is behavior an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and read-only marker, then supporting detail. The inline list of returned fields is mildly redundant given an output schema exists, but the prose is tight and every other sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with an output schema, nothing an agent needs is missing: return contents, verification semantics, error taxonomy, and safety profile are all covered. The listing of returned fields is redundant with the output schema but not harmful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3; the description adds real value by explaining that max_bytes produces a too_large result with metadata but no bytes, and it frames content_hash as the digest being verified. It stops short of documenting digest syntax (left to the schema).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (content-addressed blob's bytes by digest) and explicitly positions itself as 'the read-back half of describe_blob', which distinguishes it from its closest sibling. No ambiguity about what is fetched or how it is addressed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'read-back half of describe_blob' framing clearly implies when to use this rather than the metadata-oriented sibling, but it never states an explicit rule (e.g. 'use describe_blob to inspect metadata without downloading bytes') or any exclusions. Clear context, no explicit when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_candidateB

Read one candidate version (read-only).

ParametersJSON Schema
NameRequiredDescriptionDefault
candidate_idYesCandidate version ID to read.

Output Schema

ParametersJSON Schema
NameRequiredDescription
candidateNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it does disclose the read-only, single-item nature of the call. However, it says nothing about what a 'candidate version' is relative to a candidate, whether errors occur for unknown IDs, or any access constraints. Adequate but thin for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with zero waste and the key qualifier front-loaded. It is arguably too terse given the dense sibling set, but by pure conciseness standards it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the single parameter is fully documented. What remains missing is conceptual grounding (what a candidate version is, how it differs from lineage or scorecard reads), leaving the definition minimally complete for this tool family.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single required parameter ('candidate_id') fully documented in the schema. The description adds no syntax, format, or sourcing detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one candidate version') and the parenthetical '(read-only)' clarifies the operation type. It implicitly distinguishes from list_candidates by saying 'one', but never mentions nearby siblings like get_candidate_lineage, get_candidate_scorecard, or get_incumbent, so differentiation requires opening the schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no named alternatives among the many candidate-related siblings (lineage, scorecard, list_candidates, get_incumbent). The agent must infer from the name alone which retrieval tool to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_candidate_lineageA

Walk a candidate's parent chain to genesis (self first).

Returns {"lineage": [{id, parent_id, model_ref, code_artifact_digest, ...}, ...]} — the full ancestry that any result attributed to this candidate inherits.

ParametersJSON Schema
NameRequiredDescriptionDefault
candidate_idYesCandidate whose parent chain to walk to genesis.

Output Schema

ParametersJSON Schema
NameRequiredDescription
lineageNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the traversal order ('self first') and the returned structure, implying a safe read. It says nothing about behavior for a nonexistent candidate, depth limits, or permissions, which for a graph-walk tool is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, no filler. The parenthetical '(self first)' is a precise, load-bearing detail about traversal order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema present, the description covers purpose and traversal order adequately; return values need not be explained. The only shortfall is the absence of any usage routing versus siblings like get_candidate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter's schema text ('Candidate whose parent chain to walk to genesis') mirrors the description's wording. The description adds no format, ID scheme, or validation semantics beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Walk a candidate's parent chain to genesis (self first)'. This clearly separates it from get_candidate or list_candidates, which deal with a single candidate or a collection rather than ancestry. It never names a sibling explicitly, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrase 'the full ancestry that any result attributed to this candidate inherits' — an agent can infer this is for tracing provenance/inheritance. However, there is no explicit when-to-use statement, no when-not, and no named alternative (e.g., get_candidate for the candidate itself).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_candidate_scorecardA

Descendant quality over attributed programmes (read-only).

For each programme created under this candidate (the Phase-0 candidate_version_id correlation), reports trial/observation counts and the best observed value per metric — 'best' read against the programme's metric_direction.

contract_id scopes the report to that contract's declared metrics (primary + secondary); without it all recorded metrics are reported per programme.

ParametersJSON Schema
NameRequiredDescriptionDefault
contract_idNoOptional — scopes the report to that contract's declared metrics (primary + secondary).
candidate_idYesCandidate whose attributed programmes are scored.

Output Schema

ParametersJSON Schema
NameRequiredDescription
totalsNo
programmesNo
contract_idNo
candidate_version_idNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares read-only, discloses the correlation key (Phase-0 candidate_version_id), and explains that 'best' is resolved against each programme's metric_direction. It omits permissions, emptiness behavior for candidates with no programmes, and cost/rate considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then the semantics of the report and the parameter scope. Slightly loose line wrapping but every sentence adds information and nothing is redundant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is not required, and the description still sketches the report contents (trial/observation counts, best value per metric). For a two-parameter read tool this leaves little an agent needs uncovered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented in the schema, and the description's contract_id wording largely restates the schema text. It adds only modest value by clarifying what 'declared metrics' means (primary + secondary).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the resource ('descendant quality over attributed programmes') and implies the reporting verb, plus marks it read-only. It is distinguishable from generic siblings like get_candidate or get_candidate_lineage, though it never names an alternative to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the contract_id scoping condition (declared primary+secondary metrics vs all recorded metrics), which is a real usage decision. However it gives no when-to-use vs when-not guidance relative to siblings such as get_candidate_lineage, assess_programme, or list_trials.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evaluation_contractA

Read one evaluation contract (read-only).

Contracts are insert-only and versioned per programme — the promotion pipeline reads them to learn which metrics a campaign is scored under.

ParametersJSON Schema
NameRequiredDescriptionDefault
contract_idYesID of the evaluation contract to read.

Output Schema

ParametersJSON Schema
NameRequiredDescription
contractNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that the resource is read-only and that contracts are insert-only and versioned per programme (so no mutation path exists), but says nothing about permissions, not-found behavior, or how versioning affects which version is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the read-only getter semantics in the first phrase and follows with compact domain context; every sentence earns its place. The fragmented line breaks are slightly awkward but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the immutability/versioning note covers the main behavioral question for a single-resource getter. Only caller-side permissions and error conditions remain unaddressed, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema already documents contract_id fully. The description adds no format, naming, or sourcing detail beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one evaluation contract') with explicit scope of a single record. It implicitly contrasts with sibling create_evaluation_contract via the read vs. create framing, though it does not name that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides domain context ('the promotion pipeline reads them to learn which metrics a campaign is scored under'), which implies when this data matters, but gives no explicit guidance on when an agent should call this versus list_promotion_decisions or create_evaluation_contract, and states no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_incumbentA

Derive the incumbent candidate from decisions (read-only).

Champion/challenger is derived from promotion_decisions, never stored: the latest 'promote' not superseded by a 'rollback' for the same candidate is the incumbent. Empty until the first promote — registered-but-undecided candidates are candidates, not the champion.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
candidateNo
candidate_idNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares read-only explicitly, discloses that the value is derived and never stored, gives the derivation rule ('latest promote not superseded by a rollback'), and defines the empty case. It stops short of mentioning pagination or error behavior, but the behavioral profile is largely covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the operation and read-only constraint front-loaded, then the derivation rule, then the empty-state edge case. No filler and every sentence adds distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return format need not be described. Combined with zero parameters and full coverage of the derivation and empty semantics, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already communicates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Derive the incumbent candidate') and immediately scopes it as read-only. It goes further by defining the domain concept itself (champion/challenger), so an agent knows exactly what it retrieves and how it differs from siblings like get_candidate or list_promotion_decisions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the semantic state this tool answers for and notes the empty-until-first-promote condition, which implicitly tells an agent when it applies. However, it never names an alternative or states when to prefer get_candidate / list_promotion_decisions instead, so routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_next_experimentB

Get the next experiment configuration from the optimizer role.

Reports remaining budget (commitment 5: budget is an epistemic resource). Enforcement: commitment 5 — rejects when budget is exhausted.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
next_configNo
programme_idNo
remaining_budgetNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a meaningful failure mode (rejection on exhausted budget) and budget reporting behavior. It stops short of covering permissions, idempotency, or the nature of the configuration returned, so it is partial rather than complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the core action. The repeated '(commitment 5)' parenthetical and 'Enforcement: commitment 5' are somewhat redundant, costing a little efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be described, and the single parameter is fully documented. Still missing is how this relates to sibling experiment tools and any explanation of the 'optimizer role'/'commitment 5' domain framing, which an agent must interpret unaided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single programme_id parameter, so the baseline is 3. The description adds no syntax, format, or constraint detail beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource ('Get the next experiment configuration') and adds the source ('from the optimizer role'). However, it does not distinguish itself from related siblings like design_experiment or get_incumbent, so an agent gets the purpose but not the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no alternative tool is named for comparison. The only condition given is a failure precondition ('rejects when budget is exhausted'), which is behavioral rather than routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_trial_statusA

Check the status of a running or completed trial.

Returns: {"trial_id": ..., "status": "running"|"completed"|"failed", "executor_output": ...} # raw output once finalized

executor_output is persisted on the trial row at finalize time — it survives server restarts (the executor's async cache does not).

artifact_path is a transient staging directory — its contents are captured content-addressed into SQLite at finalize and the files deleted. An empty artifact_path is by design, not lost data: recover contents via get_blob / the code:// resource.

Enforcement: commitment 1 — the loop is the unit (orphan check).

ParametersJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintNo
reasonNo
statusNo
progressNo
trial_idNo
cancelledNo
started_atNo
eta_secondsNo
finished_atNo
artifact_pathNo
elapsed_secondsNo
executor_outputNo
duration_secondsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses that executor_output is persisted on the trial row and survives restarts, that artifact_path is a transient staging directory whose contents are captured content-addressed into SQLite, and that an empty artifact_path is by design rather than data loss. It does not cover permission requirements, error behavior for a missing/invalid trial_id, or whether the call blocks — minor gaps against an otherwise rich disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in a single sentence, and the subsequent paragraphs each carry distinct payload (return shape, persistence semantics, artifact lifecycle, enforcement note) with no filler. The Returns block partially overlaps the output schema, but the surrounding lifecycle notes are not derivable from it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return-value explanation is partly redundant, yet the description adds persistence and artifact-recovery context that the schema cannot express. Combined with the enforcement note about the orphan check, it is close to complete; the missing pieces are error/permission behavior rather than core call semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (trial_id, programme_id) are already documented in the schema. The description adds no syntax, format, or relational detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Check the status of a running or completed trial.' An agent can tell it inspects a specific trial rather than enumerating (list_trials) or blocking (wait_trial). It never names those siblings explicitly, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Scope is implied by 'a running or completed trial', and the note routes artifact recovery to 'get_blob / the code:// resource', which is a genuine alternative-pointer. But there is no explicit when-to-use guidance relative to wait_trial, list_trials, or correct_trial_status, so the agent must infer the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_active_programmesA

List all active research programmes.

Call this BEFORE create_programme to check whether a programme already covers your goal. If one exists, add a hypothesis to it (formulate_hypothesis) instead of creating a duplicate programme.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
countNo
active_programmesNo

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It conveys the intended dedup workflow, but does not state that it is read-only, what 'active' excludes (e.g., closed/archived programmes), or any listing behavior. For a risk-free no-param list tool this is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose and followed by the actionable workflow. No filler or repetition; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and there are no parameters to cover. The description supplies purpose plus the dedup workflow, which is nearly complete for this tool; only the read-only nature and definition of 'active' are unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline for a no-param tool applies. Nothing in the schema needs compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with a scope qualifier: 'List all active research programmes.' The word 'active' distinguishes it from the sibling list_programmes, so an agent can tell what set is returned without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly prescribes when to call it ('BEFORE create_programme to check whether a programme already covers your goal') and what to do instead when a match is found, naming formulate_hypothesis as the alternative. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_archivesA

List all archive files (open and sealed).

Returns a list of archives with their IDs, paths, programme counts, and sealed status. Sealed archives are read-only; open archives are still accepting new programmes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
archivesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It helpfully discloses the read-only nature of sealed archives and that open archives still accept programmes — real domain semantics the schema cannot express. It does not, however, state that the call itself is a safe read, nor mention pagination, result size, or empty-list behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and scope, then two supporting sentences with no filler. The sentence enumerating returned fields overlaps with the output schema, which is mildly redundant, but the domain note about open vs. sealed archives earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description need not explain return values, yet it does so lightly. It supplies the essential domain model (sealed = read-only, open = accepting programmes) that makes the returned sealed status interpretable. Only the relationship to adjacent read tools like get_archive is left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is nothing for the description to clarify or compensate for at the argument level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource with an explicit scope qualifier: "List all archive files (open and sealed)." An agent can distinguish this from archive-mutating siblings like archive_pending_programmes. It does not, however, name the closely related read siblings (get_archive, get_archived_programme, verify_archive), so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the word "all" — the agent can infer this is the enumeration entry point that precedes get_archive. There is no explicit when-to-use, no exclusion such as "for a single archive use get_archive," and no mention of prerequisites or scope of freshness.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_candidatesA

Enumerate the candidate population, newest first (read-only).

The population listing the Loop-1 roster pulls — candidates are the researchers; champion/challenger status is derived from promotion_decisions (see get_incumbent), never stored.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax rows to return (pagination).
offsetNoRows to skip before returning (pagination).

Output Schema

ParametersJSON Schema
NameRequiredDescription
totalNo
candidatesNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden; it does disclose '(read-only)' and the newest-first ordering, which are genuinely useful traits. It does not address pagination semantics, permissions, or failure behavior, so coverage is only partial for a tool with zero annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The verb and read-only nature are front-loaded, and the two sentences stay short. The parenthetical 'The population listing the Loop-1 roster pulls' is slightly opaque jargon that costs a little clarity relative to its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be described, and both parameters are documented in the schema. The description adds the ordering guarantee and the key domain fact that status is derived rather than stored, which is what an agent most needs here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — limit and offset are already documented as pagination in the schema. The description adds no parameter-level detail (e.g., maximum limit or default ordering interaction), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Enumerate') and resource ('candidate population') plus an ordering guarantee ('newest first'). The clarification that candidates are the researchers and that champion/challenger status lives in promotion_decisions helps distinguish this from get_incumbent, though the relationship to get_candidate/get_candidate_lineage is not spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the whole-population listing and points to get_incumbent for derived status, which gives an agent a rough sense of when to reach for it. However, there is no explicit when/when-not guidance versus single-entity tools like get_candidate, so usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_hypothesesA

List all hypotheses in a programme with their status.

Read-only enumeration — the tool-level counterpart of the programme://{id}/hypotheses resource. Use it to recover a hypothesis id (e.g. in a new session) rather than querying the state database directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hypothesesNo
programme_idNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full behavioral burden. It usefully discloses the read-only nature of the call and points to the equivalent resource, but says nothing about ordering, pagination, or volume of results — relevant traits for a list tool despite the output schema existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and front-loaded: purpose first, then read-only framing, then the resource counterpart and usage hint. Every sentence contributes; the resource-path sentence is slightly incidental but still informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the single parameter is fully covered by the schema. The description adds purpose, read-only status, and a usage rationale, leaving only minor gaps like result ordering or limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single required parameter, so the schema already documents programme_id fully. The description only implies the programme scoping without adding format or syntax detail, which is the baseline 3 for well-covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List all hypotheses in a programme") plus scope ("with their status"). The purpose is immediately legible and distinguishable from formulate/conclude/abandon_hypothesis siblings, though the description does not explicitly name those siblings to contrast with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use ("recover a hypothesis id, e.g. in a new session") and names an alternative to avoid ("rather than querying the state database directly"). It stops short of explicit when-not-to-use guidance against the sibling mutation tools, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_programmesA

List programmes with optional attribution/status filters (read-only).

candidate_version_id filters to programmes created under that candidate (the RSI Phase-0 correlation) — this is how a candidate's descendants are enumerated. Each row includes candidate_version_id. Unlike list_active_programmes this is a paginated general listing across all statuses.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax rows to return (pagination).
offsetNoRows to skip before returning (pagination).
statusNoOptional status filter.
candidate_version_idNoRSI Phase-0 correlation — which registered candidate created this programme; must exist if given.

Output Schema

ParametersJSON Schema
NameRequiredDescription
totalNo
programmesNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the read-only nature and pagination, plus that each row includes candidate_version_id, which is useful. However, it omits auth requirements, rate limits, and other operational traits expected when no annotation coverage exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then elaborates on the correlation filter and sibling differentiation. Well-sized and mostly waste-free, though the multi-line phrasing is slightly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, paginated listing tool with full schema coverage and an output schema, the description covers purpose, filters, and sibling differentiation adequately, so return values need not be explained. Minor gaps in operational context remain but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds genuine meaning beyond the schema by explaining candidate_version_id as the 'RSI Phase-0 correlation' and the mechanism for enumerating a candidate's descendants, plus clarifying status filtering.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List programmes') with scope ('optional attribution/status filters') and explicitly distinguishes itself from the sibling list_active_programmes ('a paginated general listing across all statuses'). An agent can tell it apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (list_active_programmes) and the condition selecting it (all statuses vs. active only), and explains the candidate_version_id use case for enumerating a candidate's descendants. Context is clear, though it doesn't spell out explicit when-not-to-use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_promotion_decisionsA

List a candidate's verdict trail, oldest first (read-only).

Decisions are insert-only — the trail is the record; reversal is a new 'rollback' verdict, never a mutation.

ParametersJSON Schema
NameRequiredDescriptionDefault
candidate_idYesCandidate whose verdict trail to list.

Output Schema

ParametersJSON Schema
NameRequiredDescription
decisionsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so reasonably well: it declares read-only, states decisions are insert-only, gives ordering, and explains that reversal appears as a new 'rollback' verdict rather than a mutation. It does not cover pagination or return size, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the read-only and ordering constraints are front-loaded before the insert-only rationale. Every clause contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description supplies the semantics (immutable trail, rollback-as-new-verdict) an agent needs to interpret results. Minor gaps remain around pagination or trail length, but the definition is otherwise complete for a one-parameter listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single candidate_id parameter, so the schema already documents it. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — listing a candidate's verdict trail — and adds the ordering ('oldest first'), which an agent can act on directly. It does not name record_promotion_decision explicitly, but the 'insert-only' phrasing implicitly distinguishes this read tool from the write sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the trail is described as a record, so listing it is a read operation, which distinguishes it obliquely from record_promotion_decision. There is no explicit 'use this when...' or 'do not use for...' guidance, and no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_trialsA

List all trials in a programme with configs and status.

Read-only enumeration — the tool-level counterpart of the programme://{id}/trials resource. Use it to recover a trial id (e.g. in a new session) rather than querying the state database directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
trialsNo
programme_idNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key trait: 'Read-only enumeration', i.e. a non-mutating read. It also ties itself to the equivalent programme://{id}/trials resource. It omits scale/pagination behavior for listing all trials and error handling for an invalid programme id, so it is not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences: what it returns, its read-only nature and resource equivalence, then the routing guidance. No sentence is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description summarizes the payload anyway. Combined with the single fully-documented parameter, it is nearly complete, with only listing scale/pagination left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema description coverage, so the schema already defines programme_id. The description adds no format or constraint detail beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list), resource (trials), scope (all trials in a programme) and payload (configs and status). This implicitly separates it from single-trial siblings like get_trial_status, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use ('recover a trial id, e.g. in a new session') and an explicit alternative to avoid ('rather than querying the state database directly'). It stops short of naming the sibling tools an agent should prefer or avoid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_retryableA

Mark a running trial as retryable (infrastructure failure).

This is for trials that are stuck in 'running' due to infrastructure issues (server restart, executor state lost, timeout) — NOT for scientific failures. A retryable trial is terminal and does not count as evidence.

It does NOT unlock re-capture or re-run: a terminal trial's record is finished. To retry the work, design_experiment a new trial and capture_bundle on it (file-path or code://).

Enforcement: commitment 1 — the loop is the unit (orphan check).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy the trial is marked retryable — becomes part of the record.
trial_idYesID of the target trial.
programme_idYesID of the target research programme.

Output Schema

ParametersJSON Schema
NameRequiredDescription
reasonNo
statusNo
trial_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden — and it delivers: it states the trial becomes terminal, does not count as evidence, does not unlock re-capture/re-run, and references an enforcement rule (commitment 1, orphan check). This is exactly the behavioral context an agent needs before an irreversible mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the critical when/when-not distinction, then the do-not-unlock warning, then the enforcement note. Well organized, though slightly long and the trailing enforcement sentence is a bit cryptic without more context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. For a mutation tool with no annotations, the description supplies everything an agent needs: what qualifies, what it excludes, the terminal/evidence semantics, the non-unlock caveat, and the correct alternative route.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (programme_id, trial_id, reason). The description adds the useful detail that reason 'becomes part of the record', but no syntax or format guidance. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (mark) + resource (running trial as retryable) with explicit scoping: it names the exact condition (infrastructure failure: server restart, executor state lost, timeout) and distinguishes it from scientific failures. Among siblings like correct_trial_status and cancel_trial, the retryable-vs-scientific-failure distinction lets an agent pick it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (stuck-in-running infrastructure failures) and when-not (scientific failures), plus a named alternative path: design_experiment a new trial and capture_bundle on it. It even warns that this does NOT unlock re-capture/re-run, closing the main misuse route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_dataA

Prepare a dataset and return a DataRef ID.

Two regimes:

  • generated: runs the generator (via executor), stores output, computes hash. Requires generator_code_ref, generator_seed. Generator contract: the file must expose generate_data(config, output_path) — config carries 'seed' AND 'generator_seed' (same value — either spelling works) plus generator_params flattened at the top level; a seed/generator_seed key inside generator_params is superseded so the recorded seed always equals the seed delivered. The dataset is written to output_path. A generator may print a single-line {"error": "..."} JSON object to stdout to report a structured failure. The generator runs under the executor's default interpreter ([executor] python) — not a bundle env; only packages installed there are importable.

  • captured: registers an external URI, computes hash if accessible. Requires source_uri.

For stream capture (captured + temporal), provide capture_window_start and capture_window_end.

The DataRef is stored in the state DB and the data is stored in data storage (separate from code storage). Access is read-only.

Returns: {"data_ref_id": "data-ref-...", "split": ..., "regime": ...}

generator_params and capture_source_metadata may be sent as JSON-encoded strings.

ParametersJSON Schema
NameRequiredDescriptionDefault
splitYesDataset split: train | validation | test.
regimeYesgenerated (run generator via executor; needs generator_code_ref + generator_seed) | captured (register external URI; needs source_uri).
versionNoVersion tag for the captured data source.
source_uriNoExternal data URI — required for regime='captured'.
generator_seedNoGenerator seed — required for regime='generated'.
generator_paramsNoGenerator parameters; object or JSON-encoded.
capture_window_endNoTemporal bound for stream capture (captured + temporal).
generator_code_refNoPath/ref to generator code — required for regime='generated'. The file must expose generate_data(config, output_path) — two positional args; see the tool description for the full contract.
capture_window_startNoTemporal bound for stream capture (captured + temporal).
capture_source_metadataNoExtra provenance about the captured source; object or JSON-encoded.

Output Schema

ParametersJSON Schema
NameRequiredDescription
splitNo
regimeNo
data_ref_idNo
storage_uriNo
content_hashNo
reproducibility_riskNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations, the description carries the full burden and does so well: it discloses the generator contract, the executor default interpreter and package constraints, the stdout structured-failure protocol, the seed-aliasing/supersession rule, that data goes to data storage separate from code storage, and that access is read-only. This is unusually rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then structured into two clearly labelled regimes with a separate returns block. The seed explanation is dense and reads as a run-on, but every sentence conveys operational detail rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, two-mode tool with no annotations, the description covers prerequisites, execution environment, failure reporting, storage behavior, and cross-parameter encoding rules. The return block is slightly redundant given the output schema exists, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the generator_seed/generator_params 'seed' aliasing and supersession behavior, the top-level flattening of generator_params, and the JSON-encoded-string allowance for generator_params and capture_source_metadata. That extra semantics justifies exceeding baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Prepare a dataset and return a DataRef ID') and enumerates the two operating regimes with their required inputs, so the agent knows exactly what the tool produces. It does not, however, explicitly contrast itself with nearby siblings such as capture_bundle or verify_data, leaving that differentiation to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear when-to-use guidance is given per regime ('generated' runs the generator, 'captured' registers an external URI) plus the additional condition for stream capture (captured + temporal). There are no explicit exclusions or named alternatives telling the agent when to pick a sibling instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_resourceA

Read one of this server's resources by URI (read-only).

The MCP surface exposes resources only through resources/read — this tool gives tool-only clients the same read surface. Ownership is by construction: a URI outside this server's registered schemes resolves to nothing and errors rather than crossing the loop boundary. Never a write path.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYesResource URI to read — this server's own schemes only (protocol://, executor://, programme://, trial://, artifact://, dataref://, code://).

Output Schema

ParametersJSON Schema
NameRequiredDescription
uriNo
contentsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses read-only semantics ('Never a write path'), the error behavior for out-of-scope URIs, and the loop-boundary ownership model. Return/pagination details are left to the output schema, which exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then layered with rationale. Slightly redundant phrasing ('read-only' and 'Never a write path' restate the same idea) and some jargon ('crossing the loop boundary'), but generally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with an output schema and full schema coverage, the description supplies the safety and error context an agent needs. It is complete enough to call correctly, missing only explicit sibling routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already enumerates the accepted schemes, so the baseline is 3. The description adds meaning about how an out-of-scope URI behaves ('resolves to nothing and errors'), which is a modest value-add over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one of this server's resources by URI') and immediately scopes it as read-only. It distinguishes itself from the domain siblings (list_trials, get_blob, etc.) as a generic URI-addressed reader, though it never names an alternative tool directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the intended audience ('gives tool-only clients the same read surface'), which suggests when to reach for it, but there is no explicit when-to-use vs when-not guidance and no named alternative among the many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_observationA

Record an observation (data item about a quality).

metrics and variance may be sent as JSON-encoded strings.

Only completed trials can be observed — a failed trial produced no measurement; its failure lives in its status, executor_output, and artifacts, not here. All-zero variance is rejected (n identical outcomes are one effective measurement, not reproducibility). If nothing is concludable, close the programme 'abandoned' rather than fabricating variance.

Enforcement: commitment 7 — reproducibility is the price of admission. Rejects a single point estimate with no variance.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricsYesMeasured values {metric_name: float}; may be a JSON-encoded string.
trial_idYesID of the target trial.
varianceYesPer-seed variance {metric: float} — all-zero rejected (commitment 7); may be JSON-encoded.
spatiotemporal_regionYesWhere/when the observation was produced (run footprint tag).

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
observation_idNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose important behavioral rules: zero variance is rejected, single point estimates without variance are rejected, and it cites an enforcement policy ('commitment 7'). It does not cover permissions/authorization requirements or reversibility, which are meaningful gaps for a mutation tool with no annotation support, so it is a strong 4 rather than a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, followed by constraints, a fallback action, and the enforcement note. Each sentence carries information, though phrasing like 'All-zero variance is rejected (n identical outcomes are one effective measurement, not reproducibility)' is somewhat dense. Efficient overall with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, return values need not be described, and the description appropriately focuses on preconditions and rejection rules. It also contextualizes the failure case against trial status/executor_output/artifacts. The main missing piece is authorization or permission requirements for this write operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's note that metrics and variance 'may be sent as JSON-encoded strings' merely repeats what the schema already states in both parameter descriptions. It adds no format, range, or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The definition opens with a specific verb+resource: 'Record an observation', clarified as 'a data item about a quality'. This is clearly distinguishable from siblings like run_trial and get_trial_status by the fact that it records measurements against a completed trial. It stops short of explicitly contrasting itself with the nearest sibling tools, so it earns a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage conditions are explicit and cover both when and when-not: only completed trials may be observed, failed trials' information 'lives in its status, executor_output, and artifacts, not here', and if nothing is concludable the agent is routed to close_programme 'abandoned'. This is precisely the when/when-not/alternative structure that scores top marks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_promotion_decisionA

Record an attributed verdict on a candidate (RSI Phase 0).

verdict: promote | reject | hold | rollback. Decisions are insert-only — reversal is a new decision ('rollback'), never a mutation. evidence_refs may be sent as a JSON-encoded string list. decided_by is required: attribution is first-class.

Enforcement: verdict enum, non-empty rationale/decided_by, candidate (and contract, if given) must exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
bf_2lnNoThe measured 2 ln BF bound forwarded with computed_rung. Stored for audit.
verdictYespromote | reject | hold | rollback — insert-only; reversal is a new 'rollback' decision.
rationaleYesNon-empty justification — accountability is first-class.
decided_byYesAttributable decider (e.g. 'human:<name>' or a protocol/improver id).
contract_idNoOptional contract the decision was scored under; must exist if given.
candidate_idYesCandidate the verdict applies to; must exist.
claimed_rungNoEvidence rung this verdict claims (not_worth|positive|strong|very_strong — the Kass–Raftery ladder). Required on 'promote' when the cited contract declares min_evidence_rung.
computed_rungNoEvidence rung the promotion channel measured (forwarded by zetesis record_promotion_verdict — the rung close computed from the arm data). Stored beside claimed for audit; upstream cannot recompute it.
evidence_refsYesEvidence_ref IDs (from pull_evidence) the decision cites — ≥1 required; may be a JSON-encoded list.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
decision_idNo
claimed_rungNo
computed_rungNo
declared_rungNo
unverified_refsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses key traits: insert-only semantics, reversal via a new decision, JSON-encoded evidence_refs, required decided_by, and enforcement checks (verdict enum, non-empty rationale/decided_by, existence of candidate/contract). It does not cover permissions or rate limits, but the core mutation and validation behavior is well communicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and then uses short paragraphs to cover verdicts, insert-only semantics, and enforcement. It avoids fluff and every sentence contributes relevant constraints or behavior. Minor fragmentation exists, but overall it is efficient and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool with no annotations and an output schema, the description provides a complete enough picture: what the tool does, the allowed verdicts, insert-only behavior, required attribution, and validation rules. It does not explain permissions or return values, but the output schema covers returns and no annotations exist to contradict. The remaining gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all nine parameters in detail. The description repeats the verdict options, evidence_refs JSON-string format, and decided_by requirement, adding little semantic meaning beyond the structured schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Record an attributed verdict on a candidate (RSI Phase 0).' It clearly distinguishes the act of recording a verdict from list/read siblings like list_promotion_decisions, though it does not explicitly name any alternative tool. The scope is clear enough for an agent to understand the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides an important usage rule: decisions are insert-only and reversal is a new 'rollback' decision, never a mutation. This tells the agent how to handle reversals, but it does not state when to use this tool versus alternatives or what preconditions are needed beyond enforcement rules. Usage context is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_candidateA

Register a researcher version with parent lineage (RSI Phase 0).

The thing doing the research becomes a durable object. parent_id links this candidate to its parent — omit for a genesis candidate. Lineage is never "latest wins": every candidate keeps its ancestry.

capability_profile describes the tools/roles available; it may be sent as a JSON-encoded string.

Enforcement: parent_id must name an existing candidate; code/harness digests must resolve to an ingested blob or be 'none' — a hash of nothing is a claim with no referent.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_refYesModel identifier/reference backing this version.
parent_idNoParent version ID — omit for genesis; must name an existing row if given.
search_policy_refNoReference to the search policy this version uses, if any.
capability_profileYesTools/roles the version can use; object or JSON-encoded.
code_artifact_digestYessha256:<64 hex> digest of the version's code artifact — must resolve to an ingested blob (describe_blob), or 'none' to declare no code artifact.
harness_artifact_digestNosha256:<64 hex> digest of the harness/evaluation artifact — must resolve to an ingested blob; omit when there is none.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
candidate_idNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and handles it well: it discloses the lineage invariant ("every candidate keeps its ancestry", no latest-wins), the referential constraint (parent_id must name an existing candidate), and digest enforcement (must resolve to an ingested blob or be 'none'). It is silent on permissions, idempotency, and duplicate handling, which are relevant for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, and the enforcement paragraph is the most actionable content. Some sentences are conceptual prose ("The thing doing the research becomes a durable object", "a hash of nothing is a claim with no referent") that add flavor rather than call guidance, slightly inflating length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the lineage model and validation rules an agent needs. It remains incomplete on operational behavior for a mutation — auth requirements, idempotency, and error semantics on invalid parent or unresolved digest are not stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parent_id ("omit for genesis"), capability_profile ("object or JSON-encoded"), and the digest fields are already documented. The description mostly restates these; its only marginal addition is the phrase explaining that 'none' means no referent, which is a modest value-add over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Register") and resource ("a researcher version") and frames the scope with the RSI Phase 0 label plus the lineage concept. It is legible apart from read-only siblings like get_candidate_lineage or list_candidates because it explicitly denotes a creation/registration action with ancestry.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear conditional for the key decision — "omit for a genesis candidate" vs supplying parent_id — which tells the agent when the parent link applies. It does not name alternative tools or state exclusions (e.g., when to read lineage instead of register), so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_trialA

Run a trial by calling the executor role.

Imports run_training from the bundle's code_ref and calls it with the trial config. The code_ref must be a path to a .py file exposing def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. Read the executor://contract resource for the full contract.

When the bundle carries data_refs, the config handed to run_training is extended with two injected keys: 'data_paths' ({split: resolved read-only path}) and 'data_ref_paths' ({data_ref_id: resolved read-only path}). The stored config_json keeps the designed config verbatim — the injection is runtime-only.

For long-running jobs, this returns quickly with status "running". Use get_trial_status to poll for completion. Note the return is not immediate — there is a consistent ~10 s handshake/settle before status "running" comes back; that settle window is what makes cancel-race behaviour reproducible.

Enforcement: commitment 6 — the bundle must be controlled. Rejects if the bundle is not fully captured. Concurrency: rejects if trial is already running (no double execution).

ParametersJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
programme_idYesID of the programme owning the trial — required; the programme_id returned by design_experiment/list_trials (orphan check). The programme must be status=active — completed/abandoned/archived programmes are closed, immutable records.
timeout_secondsNoPer-trial hard deadline override in seconds — the executor kills the process past it. None uses the server default ([executor] timeout_seconds). Bounded by [executor] max_timeout_seconds; the applied value is recorded in executor_output.timeout_seconds.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
statusNo
messageNo
trial_idNo
executor_outputNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses the runtime-only injection of 'data_paths'/'data_ref_paths' while config_json stays verbatim, the non-immediate ~10 s handshake/settle window, the enforcement rule (commitment 6, bundle must be controlled), and the no-double-execution concurrency guard. This is unusually rich behavioral context for a mutation-style executor call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and each paragraph carries distinct information (contract, injection, async behavior, enforcement, concurrency). It is longer than average, and the aside about the settle window making cancel-race behaviour reproducible is marginal, but no sentence is dead weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out; the description still explains the async return shape and the polling path. Combined with rejection conditions, timeout semantics, and the executor contract reference, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so trial_id, programme_id, and timeout_seconds are already fully documented in the schema, making 3 the baseline. The description does not add parameter-level detail beyond that (its detail concerns the bundle contract, not the arguments), so it neither compensates nor detracts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a trial by calling the executor role') and immediately defines the mechanism (imports run_training from the bundle's code_ref and calls it with the trial config). It is cleanly distinguishable from siblings like get_trial_status, wait_trial, and cancel_trial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: for long-running jobs this returns quickly with status 'running', and 'Use get_trial_status to poll for completion.' It also names the conditions under which the call is rejected (bundle not fully captured, trial already running), which is exactly the when/when-not guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_beliefB

Update the belief state for a programme.

Calls the optimizer role's tell, then reads the updated posterior. Enforcement: commitment 2 — memory precedes optimization. Enforcement: commitment 9 — belief is a tracked quantity.

ParametersJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
programme_idYesID of the target research programme.
observation_idYesID of the observation the belief update consumes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
belief_idNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose a genuinely useful trait — that the tool invokes the optimizer role's tell and then reads back the updated posterior — and hints at ordering enforcement, but it says nothing about permissions, reversibility, side effects, or idempotency for what is clearly a state-mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core sentence is efficient and front-loaded, but the two 'Enforcement: commitment N' lines are cryptic pointers to an external policy document that most agents cannot resolve, so they consume space without conveying actionable meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, for a state-mutating tool embedded in a long trial/observation workflow with no annotations, the description leaves the caller unclear about preconditions and the relationship to sibling tools like record_observation and run_trial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (programme_id, trial_id, observation_id) with short descriptions. The description adds no format, constraint, or relationship detail beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Update the belief state for a programme') and even sketches the mechanism ('Calls the optimizer role's tell, then reads the updated posterior'). However, it does nothing to distinguish this tool from adjacent workflow siblings such as record_observation or run_trial, so an agent must infer the boundary from the schema alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite statement, and no named alternative. The only usage signal is implicit ('for a programme' plus an observation_id parameter), which leaves the agent guessing whether this should be called after record_observation or after run_trial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_metric_directionA

Update the optimization direction for an existing programme.

Use 'minimize' for loss-like metrics (perplexity, error rate) or 'maximize' for quality metrics (accuracy, F1). The optimizer uses this to rank best trials and steer the search.

This is needed for programmes created before the metric_direction parameter was added to create_programme.

ParametersJSON Schema
NameRequiredDescriptionDefault
programme_idYesID of the target research programme.
metric_directionYes'minimize' (loss-like: perplexity, error) | 'maximize' (quality: accuracy, F1) — steers search and best-trial ranking.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
programme_idNo
metric_directionNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden. It does disclose a real behavioral consequence – the value steers optimizer search and best-trial ranking – which is useful. However, it says nothing about permissions, whether the value can be changed repeatedly, or any side effects on already-recorded trials, so the mutation profile remains partly opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and resource, then the enum semantics, then the rationale – a sensible ordering with no filler sentences. It loses a point because the 'minimize/maximize' explanation largely duplicates the schema's parameter description, so that block does not fully earn its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter mutation tool with an output schema, the description covers purpose, value semantics, and rationale sufficiently. The only material gap is the absence of any auth/permission or reversibility note, which matters more given there are no annotations to supply it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are already documented, including the enum meanings for metric_direction. The description restates that same 'minimize = loss-like / maximize = quality' mapping rather than adding new format, default, or constraint information, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb (Update) and a specific resource/field (optimization direction for an existing programme), so it is immediately distinguishable from create_programme and the other programme-management siblings. Nothing about the intent is ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing sentence gives an explicit when-to-use condition – programmes created before metric_direction existed on create_programme – which both routes the agent away from the sibling and explains the migration scenario. It also implicitly frames the normal path (set direction at creation) rather than calling this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_archiveA

Verify an archive's integrity.

For sealed archives: recomputes SHA-256 of the .db file, compares to the recorded hash.

For open archives: verifies structural integrity (DB can be opened, row counts match the registry).

Returns: {verified: bool, recorded_hash, computed_hash (sealed only)}

ParametersJSON Schema
NameRequiredDescriptionDefault
archive_idYesID of the target archive.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
sealedNo
verifiedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it names the exact mechanism (SHA-256 recompute for sealed, open/row-count checks for open archives), which is substantive behavioral disclosure. It still omits whether the call is side-effect-free, permission requirements, and failure/error behavior, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by two tightly scoped branches and a return summary; every sentence earns its place and the sealed/open split is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the tool's dual behavior thoroughly and even restates the return shape (which the output schema already provides, but with the useful sealed-only qualifier). Missing only safety/permission context and error semantics, which no annotations supply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema description coverage, so the schema already documents archive_id. The description adds no syntax, format, or selection detail for the parameter, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (archive) plus the exact property being checked (integrity), which cleanly separates it from retrieval siblings like get_archive and list_archives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sealed-vs-open breakdown implies when each code path applies, but there is no explicit statement of when an agent should call verify_archive versus get_archive, check_invariants, or verify_data, and no prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_dataA

Verify data provenance by re-computing the hash.

Returns: {"verified": bool, "recorded_hash": ..., "computed_hash": ...}

If verified is false, the data was modified after preparation — a provenance violation that invalidates any trial using this data.

ParametersJSON Schema
NameRequiredDescriptionDefault
data_ref_idYesID of the target DataRef (from prepare_data).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
verifiedNo
computed_hashNo
recorded_hashNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does so well: it explains the hash recomputation, what a false result means (data modified post-preparation), and the downstream consequence for trials. It stops short of stating auth needs, idempotency, or that the operation is a pure read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the return shape, then the failure consequence — a sensible order with no filler. The Returns line partially duplicates the existing output schema, which is the only minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with an output schema and no annotations, this is nearly complete: purpose, result interpretation, and consequence are all covered. The main gap is the lack of explicit routing against verification siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is already self-describing (the DataRef ID from prepare_data), so baseline 3 applies. The description adds no format or sourcing detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (data provenance) plus the mechanism (re-computing the hash), which distinguishes it from the similarly-named sibling verify_archive. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: referencing the DataRef from prepare_data and noting that a failed check invalidates trials hints at context, but the description never states when to call this versus verify_archive or other verification siblings. No explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_trialA

Wait for a trial to reach a terminal state — one call instead of a polling loop.

Polls get_trial_status internally every poll_seconds until the trial is completed/failed/retryable/abandoned or timeout_seconds elapses (server-side cap: 60s). Returns the final status payload plus 'waited_seconds' and 'timed_out'. Caps the get_trial_status amplification loop: a 60s wait replaces ~30 round-trips.

ParametersJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
poll_secondsNoInterval between server-side status checks.
programme_idYesID of the target research programme.
timeout_secondsNoMax seconds to wait server-side (capped at 60 — long trials still need client polling, just far less of it).

Output Schema

ParametersJSON Schema
NameRequiredDescription
hintNo
reasonNo
statusNo
progressNo
trial_idNo
cancelledNo
started_atNo
eta_secondsNo
finished_atNo
artifact_pathNo
elapsed_secondsNo
executor_outputNo
duration_secondsNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so well: it discloses the internal polling of get_trial_status, the enumeration of terminal states (completed/failed/retryable/abandoned), the server-side 60s cap, the amplification bound (~30 round-trips), and the extra return fields 'waited_seconds' and 'timed_out'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded paragraphs with no filler; the purpose leads and mechanism/limits follow. Slight overlap between the prose statement of the 60s cap and the schema's own description of it keeps this from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a wait/polling tool with no annotations but a full output schema, the description supplies everything the agent needs: terminal-state enumeration, polling cadence, timeout cap, and the additional response fields. Nothing material for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters and the 60s cap on timeout_seconds. The description explains how poll_seconds and timeout_seconds drive the internal loop, which adds mild semantic value, but it is not information the schema lacks — baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('wait for a trial to reach a terminal state') plus the mechanism it replaces (a polling loop). An agent can distinguish it from get_trial_status, run_trial, and cancel_trial without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly positions the tool as a replacement for client-side polling of get_trial_status and names the condition that limits it ('long trials still need client polling'). It stops short of an explicit rule for when to call get_trial_status or cancel_trial instead, but the routing context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 46 tool updatesv0.1.28
    • First observedabandon_hypothesis
    • First observedacknowledge_violation
    • First observedarchive_pending_programmes
    • First observedassess_programme
    • First observedcancel_trial
    • First observedcapture_bundle
    • First observedcapture_bundle_from_code_hash
    • First observedcapture_pending_artifacts
    • First observedcheck_invariants
    • First observedclose_programme
    • First observedconclude_hypothesis
    • First observedcorrect_trial_status
    • First observedcreate_evaluation_contract
    • First observedcreate_programme
    • First observeddescribe_blob
    • First observeddesign_experiment
    • First observedformulate_hypothesis
    • First observedget_archive
    • First observedget_archived_programme
    • First observedget_blob
    • First observedget_candidate
    • First observedget_candidate_lineage
    • First observedget_candidate_scorecard
    • First observedget_evaluation_contract
    • First observedget_incumbent
    • First observedget_next_experiment
    • First observedget_trial_status
    • First observedlist_active_programmes
    • First observedlist_archives
    • First observedlist_candidates
    • First observedlist_hypotheses
    • First observedlist_programmes
    • First observedlist_promotion_decisions
    • First observedlist_trials
    • First observedmark_retryable
    • First observedprepare_data
    • First observedread_resource
    • First observedrecord_observation
    • First observedrecord_promotion_decision
    • First observedregister_candidate
    • First observedrun_trial
    • First observedupdate_belief
    • First observedupdate_metric_direction
    • First observedverify_archive
    • First observedverify_data
    • First observedwait_trial

TDQS

A3.7/5.0

Scored across 46 tools

Disambiguation4/5

Most tools target clearly distinct resources and lifecycle stages (programme/hypothesis/trial/candidate/contract/archive/blob). A few boundaries blur: list_active_programmes vs list_programmes, describe_blob vs get_blob, and the terminal-status cluster (mark_retryable, correct_trial_status, cancel_trial) require reading descriptions to separate. Descriptions do resolve the ambiguity, but an agent could misselect without care.

Naming Consistency4/5

Strong overall verb_noun snake_case convention (create_programme, conclude_hypothesis, record_observation, cancel_trial). Minor deviations: the get_ vs describe_ vs read_ prefix split (describe_blob/get_blob, read_resource) is semantically inconsistent, and capture_bundle vs capture_bundle_from_code_hash is a long variant. Still predictable and readable.

Tool Count2/5

46 tools is heavy for a single server, exceeding the 25+ threshold for concern. While the domain is genuinely broad (research loop + archives + integrity + blob store), operations like capture_pending_artifacts and archive_pending_programmes are migration one-offs that could be folded, and the trial-status family could be consolidated.

Completeness5/5

The surface covers the full research lifecycle: programme creation/closing, hypothesis formulate/conclude/abandon, experiment design/run/observe, candidate registration and lineage, evaluation contracts, promotion decisions, data prep/verify, archiving/verification, integrity auditing, and a resource/blob read path. Few obvious gaps remain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server that enforces evidence-graded, phase-gated, peer-reviewed research workflows for AI agents to conduct rigorous decision-making.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI-assisted scientific research workflow management through MCP, including project creation, ideation, experiment execution, and artifact handling, with integration for ChatGPT, Codex, and Claude Code.
    Apache 2.0
  • A
    license
    C
    quality
    A
    maintenance
    An append-only research operations framework and read-only MCP that tracks research plans, approvals, observations, claims, failures, revisions, and contributions with source-grounded evidence, providing search, evidence fetch, and audit capabilities without direct ledger writes.
    30
    64 PyPI
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables automated scientific paper analysis, citation credibility verification, consensus ratio calculation, and multi-hop research queries through the Model Context Protocol, integrating with MCP-compliant clients.
    8
    -