Skip to main content
Glama

video-tools-mcp

An MCP server that gives AI agents local video tools. Claude, Cursor, and other MCP clients can read video metadata, find shots and cuts, extract frames to look at, and check keyframes. The tools run FFmpeg on your machine. You do not need an API key, and the server does not upload your video anywhere.

Tool

What it does

video_metadata

Reads the duration, container, codecs, resolution, frame rate, rotation, bitrate, and audio facts.

detect_shots

Finds the camera cuts and returns each shot with a start and an end time.

extract_frames

Extracts JPEG frames at given times, or spread evenly, and returns the images to the agent.

check_keyframes

Lists the keyframe times and checks the interval for seeking, trimming, and HLS or DASH streaming.

video_context

Builds one JSON document with the media facts, shots, and keyframes in the open Video Context schema.

Each tool accepts a local file path or an http(s) URL.

Requirements

  • Node.js 20 or later.

  • FFmpeg and ffprobe on your PATH.

    • macOS: brew install ffmpeg

    • Ubuntu or Debian: sudo apt install ffmpeg

    • Windows: winget install ffmpeg

Related MCP server: video-reader-mcp

Install

Claude Code

claude mcp add video-tools -- npx -y github:video-context/video-tools-mcp

Claude Desktop, Cursor, Windsurf, and other MCP clients

Add this to the MCP configuration file of your client:

{
  "mcpServers": {
    "video-tools": {
      "command": "npx",
      "args": ["-y", "github:video-context/video-tools-mcp"]
    }
  }
}

For Claude Desktop, the file is claude_desktop_config.json. For Cursor, the file is .cursor/mcp.json.

Example prompts

  • "Look at ~/Movies/demo.mp4 and tell me what happens in it."

  • "Find the cuts in launch.mp4 and list each shot with its length."

  • "Get 8 frames from this video and write alt text for each one."

  • "Will ad.mp4 stream well with HLS? Check the keyframes."

  • "Make a Video Context JSON file for every video in this folder."

Example output

video_context returns a document like this one (illustrative):

{
  "schema_version": "0.1",
  "video_id": "launch.mp4",
  "duration": 15,
  "media": {
    "container": "mov,mp4,m4a,3gp,3g2,mj2",
    "video": { "codec": "h264", "width": 1280, "height": 720, "fps": 30, "pixel_format": "yuv420p" }
  },
  "shots": [
    { "start": 0, "end": 3 },
    { "start": 3, "end": 6 }
  ],
  "keyframes": [0, 3, 6],
  "provenance": [{ "source": "ffmpeg", "tracks": ["shots", "keyframes"] }]
}

How it works

  • video_metadata runs ffprobe -show_format -show_streams.

  • detect_shots runs the FFmpeg select='gt(scene,0.3)' filter on a small copy of each frame. Change threshold to find more or fewer cuts.

  • check_keyframes reads the packet flags with ffprobe. It does not decode the video, so it is fast.

  • extract_frames seeks to each time and saves one JPEG. The files go to your temporary folder.

These tools find technical facts. They do not understand the content. To search many videos by meaning, or to get transcripts, on-screen text, objects, and summaries, see Video Context.

Use the tools in code

import { buildContext, detectShots, extractFrames } from 'video-tools-mcp'

const { shots } = await detectShots('launch.mp4', { threshold: 0.25 })
const frames = await extractFrames('launch.mp4', { times: shots.map((s) => s.start) })

Develop

npm install
npm test

The tests make a short video with FFmpeg, so FFmpeg must be installed.

License

MIT

Available Tools

5 tools
check_keyframesCheck keyframesA
Read-only

List the keyframe (I-frame) times of a video and check the keyframe interval. Use it to find out if a video seeks fast, trims at exact times, and streams well with HLS or DASH.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesA local file path or an http(s) URL of a video.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so safety is covered. The description adds the outcome-oriented purpose (seek speed, trimming, streaming), but doesn't disclose return format details like whether times are timestamps, frame indices, or interval statistics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences: the first states the action and secondary check, the second enumerates concrete use cases. Front-loaded and no wasted words, though the sentence structure could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single-parameter read-only tool with full schema coverage and no output schema, the description provides sufficient context about what the tool does and why. Missing return-value details, but these are relatively minor for a diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single parameter whose description fully specifies accepted input (local path or http(s) URL). The description adds nothing about parameter semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (keyframe I-frame times) plus a secondary action (check interval). Distinguishes itself from siblings like video_metadata and detect_shots by focusing specifically on keyframe structure.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when to use it – diagnosing seek speed, exact-time trimming, and HLS/DASH streaming behavior. Doesn't explicitly name alternatives or exclusions, but the use cases are concrete enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_shotsDetect shotsA
Read-only

Find the camera cuts in a video and return the shots between them, with start and end times in seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesA local file path or an http(s) URL of a video.
min_shotNoIgnore cuts closer than this many seconds to the last cut. Default 0.5.
thresholdNoThe scene-change score that counts as a cut. Default 0.3. Lower finds more cuts.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so safety is already covered. The description adds useful output semantics beyond that: it returns the shots between cuts with start and end times in seconds. It doesn't discuss performance or edge cases, but it goes beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the action and result. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple read-only nature and full schema coverage, the description is adequate: purpose and return shape are clear. However, it lacks any guidance on when to choose this tool over its siblings, leaving a gap in a multi-tool video analysis context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented with defaults and semantics. The description does not add any parameter-level detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find cuts / return shots) and resource (video), and clarifies the output form (shots with start/end times). It is distinguishable from siblings like extract_frames or check_keyframes, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as check_keyframes or video_context, nor any prerequisites or exclusions. The purpose is clear but the routing context is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_framesExtract framesA
Read-only

Extract still frames from a video as JPEG images, so that you can look at the video. Give exact times, or a count of frames to spread evenly. Returns the images and their file paths.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNoThe number of frames to spread evenly. Default 6. Ignored when times is set.
inputYesA local file path or an http(s) URL of a video.
timesNoThe times of the frames, in seconds.
widthNoThe frame width in pixels. Default 512.
include_imagesNoReturn the images in the result. Default true. Set false to get only the file paths.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description adds that it returns both the images and their file paths. It doesn't mention the 24-frame cap, but with annotations covering the safety profile this is solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, then usage modes, then return format. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete enough to call correctly: input source, selection methods, and return shape are covered. No output schema exists, but the description states returns are images plus paths. Minor gap on the count maximum (24) and default (6), which live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents count, times, width, and include_images in detail. The description's note that times and count are mutually exclusive selection methods is a helpful synthesis but largely repeats the schema's 'Ignored when times is set.' Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Extract) + resource (still frames from a video) + output format (JPEG). 'So that you can look at the video' states the intent, which clearly distinguishes it from siblings like video_metadata or detect_shots.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells you the two selection modes (exact times vs. even count), which effectively routes usage. It doesn't name siblings as alternatives or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

video_contextVideo contextB
Read-only

Build one JSON document for a video in the open Video Context schema (https://github.com/video-context/video-schema): media facts, shots, and keyframes. For semantic search, transcripts, on-screen text, and objects across many videos, see Video Context: https://videocontextapi.com

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesA local file path or an http(s) URL of a video.
thresholdNoThe scene-change score for shot detection. Default 0.3.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds that it produces a single JSON document and cites the output schema, useful structural context, but says nothing about cost, processing time, or error behavior for a video-analysis operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what the tool produces, followed by the disambiguation clause. The external URL marketing line is slightly promotional but does carry the routing distinction, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should at least sketch the JSON shape it says it builds; instead it defers to an external GitHub link. For a tool whose whole purpose is assembling a structured document, this is a meaningful gap, though annotations and the cited schema cover part of the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (input path/URL, threshold with default 0.3) are fully documented in the schema. The description adds nothing about them; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb (Build) and artifact (one JSON document for a video) and names the schema it conforms to, with a link. It also distinguishes itself from a differently-scoped external service for semantic search/transcripts, though that reference points outside the sibling set rather than to its actual siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contrasts itself with a semantic-search API, implying 'use this for deterministic per-video media facts/shots/keyframes', but it never references the actual sibling tools (detect_shots, extract_frames, check_keyframes, video_metadata), which is where routing guidance matters most. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

video_metadataVideo metadataB
Read-only

Read the duration, container, codecs, resolution, frame rate, rotation, bitrate, and audio facts of a video.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesA local file path or an http(s) URL of a video.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this as a safe read, and the description usefully discloses what information is exposed. However, it says nothing about behavior with remote URLs (network fetch, failures, redirects) or what happens on unsupported containers, so it adds only moderate context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the enumerated field list is dense but every element is informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the returned facts, and annotations cover the safety profile. For a one-parameter read tool this is nearly complete, though it omits any mention of error/remote-fetch behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, and schema description coverage is 100% — the schema already explains that 'input' is a local path or http(s) URL. The description adds no additional parameter semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('video'), and enumerates exactly which facts are returned (duration, codecs, resolution, frame rate, etc.), which distinguishes it from frame/shot-oriented siblings. It stops short of explicitly naming a sibling to contrast with, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of the alternatives (detect_shots, extract_frames, check_keyframes, video_context). Usage is only implied by the fact that it reads metadata.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedcheck_keyframes
    • First observeddetect_shots
    • First observedextract_frames
    • First observedvideo_context
    • First observedvideo_metadata

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation4/5

Each tool targets a distinct aspect: metadata, shot detection, frame extraction, keyframe analysis, and context aggregation. There is mild overlap since video_metadata and check_keyframes both report technical video facts, and video_context aggregates the same data the individual tools produce, but purposes remain distinguishable.

Naming Consistency4/5

All names use snake_case, which is consistent. However, the prefix convention is mixed: detect_shots, extract_frames, and check_keyframes use verb_noun, while video_metadata and video_context use noun_noun, a minor deviation that keeps things readable.

Tool Count5/5

Five tools is well-scoped for a focused video inspection server, with each tool earning its place covering a distinct inspection task plus one aggregation helper.

Completeness4/5

The surface covers core video inspection: media facts, shots, frames, keyframes, and aggregated context output. Minor gaps exist (e.g. no audio/subtitle extraction or transcript tooling), but these are explicitly deferred to an external service, so agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers