Skip to main content
Glama

capture_pending_artifacts

Backfills artifact files for legacy trials missing stored copies by scanning disk, saving gzip-compressed data to SQLite, and marking oversized or lost files. Dry runs preview pending IDs.

Instructions

Capture artifact files for trials that predate artifact capture.

Scans all trials with an artifact_path but no rows in trial_artifacts. Reads files from disk and stores them in SQLite (gzip-compressed). Files over 50MB are skipped and marked oversized. Files that no longer exist on disk are marked lost.

dry_run=true reports the pending trial ids without reading or storing anything.

cleanup: when True, staging files are deleted after successful capture (the SQLite copy becomes authoritative). Default False keeps the originals — conservative for migration of old trials.

Under normal operation this is a no-op — _finalize_trial already captures artifacts automatically. This tool only catches trials that slipped through (finalized before artifact capture was implemented).

Returns: {captured: [...], oversized: [...], lost: [...], trials_scanned: int, trials_with_existing: int}

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
cleanupNoDelete staging files after successful capture (the SQLite copy becomes authoritative).
dry_runYesWhen true, report which trials would be captured without mutating. Required — pass false to apply.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
lostNo
capturedNo
oversizedNo
trials_scannedNo
trials_with_existingNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does: storage target (SQLite, gzip-compressed), >50MB files skipped and marked oversized, missing files marked lost, cleanup deletes staging files making SQLite authoritative, and dry_run performs no mutation. These are exactly the side effects an agent needs before invoking a mutating backfill.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in one line, then introduces the scan criteria, failure modes, dry_run, and cleanup. Slight redundancy with the schema descriptions for dry_run/cleanup and with the output schema's return list, but every section is scannable and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers scope, discovery criteria, failure handling, mutation semantics, and the no-op caveat for a mutating tool with no annotations. An output schema exists, so repetition of the return shape is harmless rather than a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented, capping the ceiling. The description still adds meaning beyond the schema by explaining the consequence of cleanup ('the SQLite copy becomes authoritative') and the conservative default rationale for migrating old trials.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('capture artifact files') and narrows scope precisely: trials with an artifact_path but no rows in trial_artifacts. The backfill framing distinguishes it from the normal capture path used by _finalize_trial, so an agent won't confuse it with capture_bundle or run_trial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it and when not to: 'Under normal operation this is a no-op — _finalize_trial already captures artifacts automatically. This tool only catches trials that slipped through.' It also documents the dry_run=false condition required to apply changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.