Skip to main content
Glama

get_provenance

Read-onlyIdempotent

Retrieve the full provenance manifest for a dataset output: input checksums, parameters, CRS decisions, engine and version, and deterministic check results. Get the exact operation details in one step.

Instructions

Return the manifest of the ONE operation that wrote this dataset.

Its inputs with their digests, the exact parameters, the CRS decisions with reasons, engine and version, and every deterministic check with its result. One step, in full. For the whole chain that led here — the operations behind the inputs, and the ones behind those — the catalogue operation get_lineage walks it and returns each step summarised instead of one step complete; run it through run_operation. It is not a tool of its own on purpose: two exposed tools answering the same question about the same file is the overlap that makes an agent pick wrong.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
output_pathYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint and idempotentHint already in annotations, the description adds meaningful behavioral detail: the manifest's exact contents (inputs with digests, parameters, CRS decisions, engine/version, deterministic checks). It also explains the design rationale for not exposing a separate get_lineage tool, which helps an agent understand the API shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main purpose is front-loaded in the first sentence, and the detailed manifest contents follow logically. The later paragraphs about get_lineage and design rationale are slightly longer than strictly needed but still earn their place by preventing agent mis-selection. Overall it is concise without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details are not required in the description. The description covers the operation's scope, manifest contents, and alternative routing. It does not mention any prerequisites (e.g., that the output_path must correspond to an existing dataset), but that is reasonably implicit and covered by the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries full responsibility for clarifying output_path. It only refers to 'this dataset' and 'same file', which implies a file path but does not state the expected format, validity constraints, or how the path should be resolved. This is a notable gap for a tool with a single undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return the manifest of the ONE operation that wrote this dataset.' It explicitly contrasts with get_lineage (which summarises the whole chain), making the tool's scope unmistakable and differentiating it from the sibling run_operation family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear 'when to use' (one step, in full) and 'when not to use' (for the whole chain, use get_lineage via run_operation). It names the alternative and the condition that selects it, leaving no ambiguity for an agent choosing between tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.