Skip to main content
Glama

capture_bundle

Seal a DESIGNED trial's code, environment, seeds, and splits before running to pre-register assumptions and prevent retrofitting.

Instructions

Seal the auxiliary bundle for a DESIGNED trial (commitment 6).

ORDERING: call AFTER design_experiment and BEFORE run_trial. The seal is pre-registration — the bundle fixes the auxiliary assumptions (code/env/seeds/splits) before any observation, so they cannot be retro-fitted to results. Trials that have left 'designed' are rejected. The post-run counterpart is executed_code.json — what actually ran, captured at finalization from the strace read-trace.

code_ref MUST be a path to a Python file that exposes: def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. The config is the same dict passed to design_experiment. Read the executor://contract resource for the full contract.

data_refs is an optional list of DataRef IDs (from prepare_data). When provided, the bundle records structured data provenance. When omitted, splits is used (backward-compatible).

extra_code_refs is an optional list of additional .py file paths that the trial depends on but cannot be discovered by AST import analysis — e.g. scripts invoked via subprocess.run(). These are captured into code_snippets and stored in code_hash_extra_json so the bundle is fully self-contained and rerunnable from archive (commitment 2 + 5 — the bundle must contain ALL code needed to reproduce, not just the executor).

Enforcement: commitment 1 — the loop is the unit (trial must exist). Enforcement: commitment 6 — the bundle must be controlled (code_ref validated). Concurrency: rejects if trial already has a bundle (no double capture).

Structured params (seeds, splits, data_refs, extra_code_refs) may be sent as JSON-encoded strings; seeds also accepts a bare int.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedsYesSeed set sealed into the bundle (pre-registration — fixes the assumption before any observation); list or JSON-encoded.
splitsYesSplit spec sealed into the bundle — which data each split used; object or JSON-encoded.
env_refYesEnvironment reference sealed into the bundle (pre-registration). Accepts a venv/conda directory (mounted read-only; its bin/python runs the trial) or a Python executable path (its venv root is mounted). Non-path values like 'python:3.12' are recorded as provenance but do not change the interpreter — the executor default runs.
code_refYesPath to a Python file exposing run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}} — see executor://contract.
trial_idYesID of the target trial.
data_refsNoDataRef IDs from prepare_data — structured data provenance (falls back to splits when omitted); list or JSON-encoded.
baseline_refNoReference to the baseline the trial compares against.
extra_code_refsNoAdditional code files to seal into the bundle; list or JSON-encoded.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
statusNo
code_refNo
warningsNo
bundle_idNo
code_hashNo
code_hash_extraNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full load and delivers: rejection on non-designed trials, no double-capture on a trial that already has a bundle, enforcement commitments (loop is the unit; code_ref validated), and the required code_ref callable contract. Runtime safety profile and constraints are fully disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and ordering, then structured by headings (Enforcement, Concurrency). Slightly long with some restatement of 'pre-registration' and repeated commitment framing, but each block earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values and doesn't. It instead completes the operational picture: sequencing, error conditions, contract resource pointer, and JSON-encoding conventions for structured params. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds real meaning beyond it: code_ref must expose run_training(config) -> dict with specific return keys, extra_code_refs exists for subprocess-invoked files undetectable by AST analysis, and data_refs falls back to splits. This explanation of intent and edge cases exceeds the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Seal the auxiliary bundle') scoped to a DESIGNED trial, and immediately distinguishes the action from run_trial and design_experiment by naming the ordering. An agent knows exactly what this does and where it sits in the loop.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit ORDERING clause (AFTER design_experiment, BEFORE run_trial), explicit exclusions ('Trials that have left designed are rejected'), and names the post-run counterpart (executed_code.json). When/when-not/alternatives are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.