Skip to main content
Glama

Dataset version manifest

dataset_manifest
Read-onlyIdempotent

The manifest of a dataset version (requires auth): schema, totals, and the content-addressed Parquet parts (sha256, bytes, records), each with a download URL valid 15 minutes — the way to fetch a whole version. Parts are shared across versions: pass have, the sha256s of the parts you already hold, and those come without a URL. The bytes of the parts that get a URL count toward your egress; the rest count nothing. The manifest is signed by the origin (Ed25519, keys at /.well-known/witan-keys) over its content without URLs, so it verifies either way.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
haveNosha256s of parts you already hold (up to 100): listed without a URL, no egress
slugYes
versionNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / have
      Added value: +{
      +  "description": "sha256s of parts you already hold (up to 100): listed without a URL, no egress",
      +  "items": {
      +    "pattern": "^[0-9a-f]{64}$",
      +    "type": "string"
      +  },
      +  "maxItems": 100,
      +  "type": "array"
      +}
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only/idempotent/closed-world, but the description adds substantially more: auth is required, download URLs expire in 15 minutes, only URL-bearing parts count toward egress, and the manifest is Ed25519-signed with keys at /.well-known/witan-keys. None of this is recoverable from the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, and each subsequent sentence carries distinct operational detail (URL TTL, egress accounting, signature verification). It is dense and long, but nearly every clause earns its place; the signing detail is the only part that flirts with excess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully shoulders the burden by describing exactly what the manifest returns (schema, totals, parts with sha256/bytes/records/URL) plus auth, expiry, egress, and verification behavior. An agent has what it needs to call and interpret this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate. It does a strong job on `have`: explaining that passing held sha256s returns those parts without a URL and with no egress cost, which is the key semantic beyond the pattern constraint. `slug` and `version` are left unexplained (e.g. what omitting `version` resolves to), leaving a residual gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific resource (the manifest of a dataset version) and enumerates its contents (schema, totals, content-addressed Parquet parts). More importantly it differentiates itself from siblings like query_dataset/read_dataset by declaring its role: 'the way to fetch a whole version.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when this tool is the right call (fetching an entire version) and prescribes the intended calling pattern with `have` to skip already-held parts. It stops short of naming an explicit alternative (e.g. query_dataset) for partial reads, so no exclusion is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.