Skip to main content
Glama

contact_sheet

Generate one PNG grid of video frames with burned-in timestamps and a text list of times to preview multiple moments at once.

Instructions

Look at frames: one PNG grid with burned-in timestamps (+ text with times).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nNo
t0No
t1No
cropNo[x0,y0,x1,y1] source px; null=full frame
timesNoexact times; else t0/t1/n
videoYesvideo path (relative to workdir ok)
max_widthNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It does disclose the output shape (one PNG grid with burned-in timestamps, plus text with times), which is genuinely useful behavioral context. However, it says nothing about permissions, cost of processing a video, or how frames are chosen, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short, front-loaded sentence with no wasted padding. The cryptic parenthetical '(+ text with times)' is slightly awkward but the overall size is appropriate for the content given.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter video-processing tool with no annotations, no output schema, and 43% parameter coverage, this description is far too sparse. It gives a hint of the return artifact but omits the meaning of most parameters and any usage context, so an agent cannot reliably invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%: n, t0, t1 and max_width have no schema descriptions, and the description does not compensate by explaining what they mean (frame count, start/end time, output width). The mention of 'timestamps'/'times' loosely touches the times param but adds no syntax or format detail, so the undocumented parameters remain ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description hints at the resource (frames from a video) and the output artifact (a single PNG grid with burned-in timestamps plus accompanying time text), which is more than a tautology. But the verb 'Look at' is vague and doesn't state that the tool generates/extracts a contact sheet, nor does it differentiate from siblings like probe, find_text, or chart_extract. Purpose is only partially conveyed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as probe, ocr_region, or track_text, and no mention of prerequisites or exclusions. The agent must infer that this is the frame-sampling tool purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.