capture_bundle
Seal a DESIGNED trial's code, environment, seeds, and splits before running to pre-register assumptions and prevent retrofitting.
Instructions
Seal the auxiliary bundle for a DESIGNED trial (commitment 6).
ORDERING: call AFTER design_experiment and BEFORE run_trial. The seal is pre-registration — the bundle fixes the auxiliary assumptions (code/env/seeds/splits) before any observation, so they cannot be retro-fitted to results. Trials that have left 'designed' are rejected. The post-run counterpart is executed_code.json — what actually ran, captured at finalization from the strace read-trace.
code_ref MUST be a path to a Python file that exposes: def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. The config is the same dict passed to design_experiment. Read the executor://contract resource for the full contract.
data_refs is an optional list of DataRef IDs (from prepare_data). When provided, the bundle records structured data provenance. When omitted, splits is used (backward-compatible).
extra_code_refs is an optional list of additional .py file paths that the trial depends on but cannot be discovered by AST import analysis — e.g. scripts invoked via subprocess.run(). These are captured into code_snippets and stored in code_hash_extra_json so the bundle is fully self-contained and rerunnable from archive (commitment 2 + 5 — the bundle must contain ALL code needed to reproduce, not just the executor).
Enforcement: commitment 1 — the loop is the unit (trial must exist). Enforcement: commitment 6 — the bundle must be controlled (code_ref validated). Concurrency: rejects if trial already has a bundle (no double capture).
Structured params (seeds, splits, data_refs, extra_code_refs) may be sent as JSON-encoded strings; seeds also accepts a bare int.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| seeds | Yes | Seed set sealed into the bundle (pre-registration — fixes the assumption before any observation); list or JSON-encoded. | |
| splits | Yes | Split spec sealed into the bundle — which data each split used; object or JSON-encoded. | |
| env_ref | Yes | Environment reference sealed into the bundle (pre-registration). Accepts a venv/conda directory (mounted read-only; its bin/python runs the trial) or a Python executable path (its venv root is mounted). Non-path values like 'python:3.12' are recorded as provenance but do not change the interpreter — the executor default runs. | |
| code_ref | Yes | Path to a Python file exposing run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}} — see executor://contract. | |
| trial_id | Yes | ID of the target trial. | |
| data_refs | No | DataRef IDs from prepare_data — structured data provenance (falls back to splits when omitted); list or JSON-encoded. | |
| baseline_ref | No | Reference to the baseline the trial compares against. | |
| extra_code_refs | No | Additional code files to seal into the bundle; list or JSON-encoded. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| code_ref | No | ||
| warnings | No | ||
| bundle_id | No | ||
| code_hash | No | ||
| code_hash_extra | No |