Skip to main content
Glama

Extract frames

extract_frames
Read-only

Extract still frames from a video as JPEG images to inspect its content. Specify exact times or a frame count to spread evenly; returns images and file paths.

Instructions

Extract still frames from a video as JPEG images, so that you can look at the video. Give exact times, or a count of frames to spread evenly. Returns the images and their file paths.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNoThe number of frames to spread evenly. Default 6. Ignored when times is set.
inputYesA local file path or an http(s) URL of a video.
timesNoThe times of the frames, in seconds.
widthNoThe frame width in pixels. Default 512.
include_imagesNoReturn the images in the result. Default true. Set false to get only the file paths.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description adds that it returns both the images and their file paths. It doesn't mention the 24-frame cap, but with annotations covering the safety profile this is solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, then usage modes, then return format. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete enough to call correctly: input source, selection methods, and return shape are covered. No output schema exists, but the description states returns are images plus paths. Minor gap on the count maximum (24) and default (6), which live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents count, times, width, and include_images in detail. The description's note that times and count are mutually exclusive selection methods is a helpful synthesis but largely repeats the schema's 'Ignored when times is set.' Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Extract) + resource (still frames from a video) + output format (JPEG). 'So that you can look at the video' states the intent, which clearly distinguishes it from siblings like video_metadata or detect_shots.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells you the two selection modes (exact times vs. even count), which effectively routes usage. It doesn't name siblings as alternatives or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.