Skip to main content
Glama

Get example corpus provenance

get_corpus_provenance
Read-onlyIdempotent

Check a sample file's trustworthiness by retrieving its provenance sidecar, including sources, evidence confidence, validation results, and SHA-256 hash.

Instructions

Return the provenance sidecar of one example file.

Use this to know how far to trust a file: the sources it was derived
from, the confidence of the evidence (``verified``, ``derived`` or
``assumed``), the validation ladder result per rung, what the builder
renamed or dropped to fit the edition, and the file's SHA-256.

Delegates to :func:`pain001.corpus.provenance`.

Args:
    scenario_id: The scenario.
    version: The message type.
    variant: The overlay id of a bank variant, else ``None``.

Returns:
    The parsed sidecar as a dict, or ``{"error": ...}`` when there is
    no such file or no corpus.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
variantNoAn overlay id for the bank variant; omit for the generic file.
versionYesThe message type, e.g. 'pain.001.001.09'.
scenario_idYesThe market scenario (from list_corpus_files).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
buildNo
errorNo
twinsNo
familyNo
sha256No
countryNo
variantNo
scenarioNo
provenanceNo
validationNo
constraintsNo
descriptionNo
message_typeNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.0.71
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "A scenario's provenance record, as the corpus ships it.",
      +  "properties": {
      +    "build": {
      +      "additionalProperties": true,
      +      "title": "Build",
      +      "type": "object"
      +    },
      +    "constraints": {
      +      "items": {},
      +      "title": "Constraints",
      +      "type": "array"
      +    },
      +    "country": {
      +      "title": "Country",
      +      "type": "string"
      +    },
      +    "description": {
      +      "title": "Description",
      +      "type": "string"
      +    },
      +    "error": {
      +      "title": "Error",
      +      "type": "string"
      +    },
      +    "family": {
      +      "title": "Family",
      +      "type": "string"
      +    },
      +    "message_type": {
      +      "title": "Message Type",
      +      "type": "string"
      +    },
      +    "provenance": {
      +      "additionalProperties": true,
      +      "title": "Provenance",
      +      "type": "object"
      +    },
      +    "scenario": {
      +      "title": "Scenario",
      +      "type": "string"
      +    },
      +    "sha256": {
      +      "title": "Sha256",
      +      "type": "string"
      +    },
      +    "twins": {
      +      "additionalProperties": true,
      +      "title": "Twins",
      +      "type": "object"
      +    },
      +    "validation": {
      +      "additionalProperties": true,
      +      "title": "Validation",
      +      "type": "object"
      +    },
      +    "variant": {
      +      "anyOf": [
      +        {
      +          "type": "string"
      +        },
      +        {
      +          "type": "null"
      +        }
      +      ],
      +      "title": "Variant"
      +    }
      +  },
      +  "title": "CorpusProvenanceResult",
      +  "type": "object"
      +}
  2. Addedv0.0.67

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark this as read-only, idempotent, and non-destructive, and the description adds meaningful behavioral detail beyond that: it returns a parsed dict, returns an error dict when the file or corpus is missing, and delegates to pain001.corpus.provenance. It also discloses the provenance contents the agent should expect, which is valuable behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and use case, and the provenance contents list is informative rather than filler. The Args and Returns sections duplicate schema information and could be trimmed, but the overall size is reasonable and every section serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter read-only tool, the description is complete: it states what the tool returns, what the return looks like, when an error is returned, and what the provenance data means. Combined with full schema coverage and supportive annotations, an agent has everything needed to invoke and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents scenario_id, version, and variant, including the version example and the variant's 'omit for generic file' guidance. The description's Args section repeats this information without adding new parameter semantics, so it earns the baseline score and no more.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return the provenance sidecar of one example file.' This is clearer than the title and makes the tool's artifact obvious. However, it does not explicitly differentiate itself from sibling tools such as get_corpus_file or get_corpus_coverage, so it stops short of full sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this to know how far to trust a file' provides a clear, concrete usage context and explains what kind of information the result carries. It does not name alternative tools or state when not to use this tool, so there is clear context but no exclusions or alternates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.