Skip to main content
Glama

capture_bundle_from_code_hash

Capture a new trial bundle from a code_hash content address instead of a file path, reusing archived code without filesystem access.

Instructions

Capture a bundle using a content address (code_hash) instead of a file path.

This is the rerun-from-archive path: the LLM reads an archived programme, gets the code_hash from the bundle, and captures a new bundle for a new trial using the same code content. No filesystem access is required — the code is loaded from code_snippets by hash, materialized as a real file next to the execution wrapper, and imported at run_trial time.

NOTE: code_hash is a "sha256:..." content address returned by a prior capture_bundle — it is NOT a bundle_id. To re-use an existing bundle's code, pass its code_hash field.

The code_hash must already exist in code_snippets (captured by a prior capture_bundle or prepare_data call). If it doesn't, the tool returns an error.

The bundle's code_ref is set to "code://{code_hash}" — a content address, not a file path. This is the carrier/content separation (Rule 5.4) made explicit: the bundle references the ICE directly.

code_hash_extra is an optional list of additional content addresses for locally-imported or subprocess-dispatched modules captured alongside the primary. Each must already exist in code_snippets. At run_trial time, these are materialized as real files under a _deps/ dir on sys.path, reconstructing each file's path suffix — so from src.mod import x resolves through the real import machinery (commitment 2 + 5).

All other parameters are identical to capture_bundle.

Structured params (seeds, splits, data_refs, code_hash_extra) may be sent as JSON-encoded strings; seeds also accepts a bare int.

Returns: {"bundle_id": "bundle-...", "status": "captured", "code_hash": ...}

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedsYesSeed set sealed into the bundle (pre-registration); list or JSON-encoded.
splitsYesSplit spec sealed into the bundle; object or JSON-encoded.
env_refYesEnvironment reference sealed into the bundle (pre-registration). Accepts a venv/conda directory (mounted read-only; its bin/python runs the trial) or a Python executable path (its venv root is mounted). Non-path values like 'python:3.12' are recorded as provenance but do not change the interpreter — the executor default runs.
trial_idYesID of the target trial.
code_hashYes'sha256:...' content address from a prior capture_bundle — NOT a bundle_id; must already exist in code_snippets.
data_refsNoDataRef IDs from prepare_data — structured data provenance; list or JSON-encoded.
baseline_refNoReference to the baseline the trial compares against.
code_hash_extraNoContent hashes of additional code files to seal; list or JSON-encoded.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
statusNo
code_refNo
warningsNo
bundle_idNo
code_hashNo
code_hash_extraNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full behavioral burden, and it does: the prerequisite that code_hash must already exist in code_snippets or the tool errors, the side effect of materializing a real file next to the execution wrapper at run_trial time, the absence of filesystem access, and the code_ref format. Failure mode, side effects, and lifecycle timing are all disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and information-dense, with the routing distinction first and prerequisites later. A few sentences lean on internal jargon that adds little for tool selection ('the carrier/content separation (Rule 5.4) made explicit', 'commitment 2 + 5'), which is the only drag on an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out (though the description still gives the shape). Prerequisites, error behavior, parameter roles, and the run_trial-time consequences are all covered for an 8-parameter, 5-required tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds real meaning the schema lacks: code_hash is a content address and NOT a bundle_id, code_hash_extra entries are materialized under a `_deps/` dir on sys.path so `from src.mod import x` resolves, and structured params accept JSON-encoded strings. It also clarifies the resulting code_ref is 'code://{code_hash}' rather than a path.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb and resource and pins the distinguishing mechanism: 'Capture a bundle using a content address (code_hash) instead of a file path.' That contrast with the sibling capture_bundle is stated up front, so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames the use case ('This is the rerun-from-archive path: the LLM reads an archived programme, gets the code_hash from the bundle, and captures a new bundle for a new trial') and notes no filesystem access is required. The alternative (plain capture_bundle) is implied by 'instead of a file path' and 'all other parameters are identical to capture_bundle', but the when-not condition is never stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.