video-tools-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@video-tools-mcpfind the cuts in ~/Movies/demo.mp4 and list each shot with its length"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
video-tools-mcp
An MCP server that gives AI agents local video tools. Claude, Cursor, and other MCP clients can read video metadata, find shots and cuts, extract frames to look at, and check keyframes. The tools run FFmpeg on your machine. You do not need an API key, and the server does not upload your video anywhere.
Tool | What it does |
| Reads the duration, container, codecs, resolution, frame rate, rotation, bitrate, and audio facts. |
| Finds the camera cuts and returns each shot with a start and an end time. |
| Extracts JPEG frames at given times, or spread evenly, and returns the images to the agent. |
| Lists the keyframe times and checks the interval for seeking, trimming, and HLS or DASH streaming. |
| Builds one JSON document with the media facts, shots, and keyframes in the open Video Context schema. |
Each tool accepts a local file path or an http(s) URL.
Requirements
Node.js 20 or later.
FFmpeg and ffprobe on your
PATH.macOS:
brew install ffmpegUbuntu or Debian:
sudo apt install ffmpegWindows:
winget install ffmpeg
Related MCP server: video-reader-mcp
Install
Claude Code
claude mcp add video-tools -- npx -y github:video-context/video-tools-mcpClaude Desktop, Cursor, Windsurf, and other MCP clients
Add this to the MCP configuration file of your client:
{
"mcpServers": {
"video-tools": {
"command": "npx",
"args": ["-y", "github:video-context/video-tools-mcp"]
}
}
}For Claude Desktop, the file is claude_desktop_config.json. For Cursor, the file is .cursor/mcp.json.
Example prompts
"Look at
~/Movies/demo.mp4and tell me what happens in it.""Find the cuts in
launch.mp4and list each shot with its length.""Get 8 frames from this video and write alt text for each one."
"Will
ad.mp4stream well with HLS? Check the keyframes.""Make a Video Context JSON file for every video in this folder."
Example output
video_context returns a document like this one (illustrative):
{
"schema_version": "0.1",
"video_id": "launch.mp4",
"duration": 15,
"media": {
"container": "mov,mp4,m4a,3gp,3g2,mj2",
"video": { "codec": "h264", "width": 1280, "height": 720, "fps": 30, "pixel_format": "yuv420p" }
},
"shots": [
{ "start": 0, "end": 3 },
{ "start": 3, "end": 6 }
],
"keyframes": [0, 3, 6],
"provenance": [{ "source": "ffmpeg", "tracks": ["shots", "keyframes"] }]
}How it works
video_metadatarunsffprobe -show_format -show_streams.detect_shotsruns the FFmpegselect='gt(scene,0.3)'filter on a small copy of each frame. Changethresholdto find more or fewer cuts.check_keyframesreads the packet flags with ffprobe. It does not decode the video, so it is fast.extract_framesseeks to each time and saves one JPEG. The files go to your temporary folder.
These tools find technical facts. They do not understand the content. To search many videos by meaning, or to get transcripts, on-screen text, objects, and summaries, see Video Context.
Use the tools in code
import { buildContext, detectShots, extractFrames } from 'video-tools-mcp'
const { shots } = await detectShots('launch.mp4', { threshold: 0.25 })
const frames = await extractFrames('launch.mp4', { times: shots.map((s) => s.start) })Develop
npm install
npm testThe tests make a short video with FFmpeg, so FFmpeg must be installed.
Related
video-schema: the open JSON schema for video data, with converters from Google Video Intelligence, Amazon Rekognition, and Amazon Transcribe.
awesome-video-understanding: a list of video understanding APIs, models, tools, and datasets.
Free video tools in the browser: scene detection, frame extraction, a keyframe checker, and a metadata viewer.
License
MIT
Available Tools
5 toolscheck_keyframesCheck keyframesARead-only
List the keyframe (I-frame) times of a video and check the keyframe interval. Use it to find out if a video seeks fast, trims at exact times, and streams well with HLS or DASH.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | A local file path or an http(s) URL of a video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so safety is covered. The description adds the outcome-oriented purpose (seek speed, trimming, streaming), but doesn't disclose return format details like whether times are timestamps, frame indices, or interval statistics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences: the first states the action and secondary check, the second enumerates concrete use cases. Front-loaded and no wasted words, though the sentence structure could be slightly tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single-parameter read-only tool with full schema coverage and no output schema, the description provides sufficient context about what the tool does and why. Missing return-value details, but these are relatively minor for a diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter whose description fully specifies accepted input (local path or http(s) URL). The description adds nothing about parameter semantics, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (keyframe I-frame times) plus a secondary action (check interval). Distinguishes itself from siblings like video_metadata and detect_shots by focusing specifically on keyframe structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use it – diagnosing seek speed, exact-time trimming, and HLS/DASH streaming behavior. Doesn't explicitly name alternatives or exclusions, but the use cases are concrete enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_shotsDetect shotsARead-only
Find the camera cuts in a video and return the shots between them, with start and end times in seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | A local file path or an http(s) URL of a video. | |
| min_shot | No | Ignore cuts closer than this many seconds to the last cut. Default 0.5. | |
| threshold | No | The scene-change score that counts as a cut. Default 0.3. Lower finds more cuts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so safety is already covered. The description adds useful output semantics beyond that: it returns the shots between cuts with start and end times in seconds. It doesn't discuss performance or edge cases, but it goes beyond what the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the action and result. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple read-only nature and full schema coverage, the description is adequate: purpose and return shape are clear. However, it lacks any guidance on when to choose this tool over its siblings, leaving a gap in a multi-tool video analysis context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented with defaults and semantics. The description does not add any parameter-level detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find cuts / return shots) and resource (video), and clarifies the output form (shots with start/end times). It is distinguishable from siblings like extract_frames or check_keyframes, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as check_keyframes or video_context, nor any prerequisites or exclusions. The purpose is clear but the routing context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_framesExtract framesARead-only
Extract still frames from a video as JPEG images, so that you can look at the video. Give exact times, or a count of frames to spread evenly. Returns the images and their file paths.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | The number of frames to spread evenly. Default 6. Ignored when times is set. | |
| input | Yes | A local file path or an http(s) URL of a video. | |
| times | No | The times of the frames, in seconds. | |
| width | No | The frame width in pixels. Default 512. | |
| include_images | No | Return the images in the result. Default true. Set false to get only the file paths. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds that it returns both the images and their file paths. It doesn't mention the 24-frame cap, but with annotations covering the safety profile this is solid added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then usage modes, then return format. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete enough to call correctly: input source, selection methods, and return shape are covered. No output schema exists, but the description states returns are images plus paths. Minor gap on the count maximum (24) and default (6), which live only in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents count, times, width, and include_images in detail. The description's note that times and count are mutually exclusive selection methods is a helpful synthesis but largely repeats the schema's 'Ignored when times is set.' Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Extract) + resource (still frames from a video) + output format (JPEG). 'So that you can look at the video' states the intent, which clearly distinguishes it from siblings like video_metadata or detect_shots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells you the two selection modes (exact times vs. even count), which effectively routes usage. It doesn't name siblings as alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_contextVideo contextBRead-only
Build one JSON document for a video in the open Video Context schema (https://github.com/video-context/video-schema): media facts, shots, and keyframes. For semantic search, transcripts, on-screen text, and objects across many videos, see Video Context: https://videocontextapi.com
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | A local file path or an http(s) URL of a video. | |
| threshold | No | The scene-change score for shot detection. Default 0.3. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds that it produces a single JSON document and cites the output schema, useful structural context, but says nothing about cost, processing time, or error behavior for a video-analysis operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the tool produces, followed by the disambiguation clause. The external URL marketing line is slightly promotional but does carry the routing distinction, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should at least sketch the JSON shape it says it builds; instead it defers to an external GitHub link. For a tool whose whole purpose is assembling a structured document, this is a meaningful gap, though annotations and the cited schema cover part of the burden.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (input path/URL, threshold with default 0.3) are fully documented in the schema. The description adds nothing about them; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (Build) and artifact (one JSON document for a video) and names the schema it conforms to, with a link. It also distinguishes itself from a differently-scoped external service for semantic search/transcripts, though that reference points outside the sibling set rather than to its actual siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contrasts itself with a semantic-search API, implying 'use this for deterministic per-video media facts/shots/keyframes', but it never references the actual sibling tools (detect_shots, extract_frames, check_keyframes, video_metadata), which is where routing guidance matters most. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_metadataVideo metadataBRead-only
Read the duration, container, codecs, resolution, frame rate, rotation, bitrate, and audio facts of a video.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | A local file path or an http(s) URL of a video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes this as a safe read, and the description usefully discloses what information is exposed. However, it says nothing about behavior with remote URLs (network fetch, failures, redirects) or what happens on unsupported containers, so it adds only moderate context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the enumerated field list is dense but every element is informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the returned facts, and annotations cover the safety profile. For a one-parameter read tool this is nearly complete, though it omits any mention of error/remote-fetch behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% — the schema already explains that 'input' is a local path or http(s) URL. The description adds no additional parameter semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('video'), and enumerates exactly which facts are returned (duration, codecs, resolution, frame rate, etc.), which distinguishes it from frame/shot-oriented siblings. It stops short of explicitly naming a sibling to contrast with, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of the alternatives (detect_shots, extract_frames, check_keyframes, video_context). Usage is only implied by the fact that it reads metadata.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
check_keyframes - First observed
detect_shots - First observed
extract_frames - First observed
video_context - First observed
video_metadata
TDQS
Scored across 5 tools
Each tool targets a distinct aspect: metadata, shot detection, frame extraction, keyframe analysis, and context aggregation. There is mild overlap since video_metadata and check_keyframes both report technical video facts, and video_context aggregates the same data the individual tools produce, but purposes remain distinguishable.
All names use snake_case, which is consistent. However, the prefix convention is mixed: detect_shots, extract_frames, and check_keyframes use verb_noun, while video_metadata and video_context use noun_noun, a minor deviation that keeps things readable.
Five tools is well-scoped for a focused video inspection server, with each tool earning its place covering a distinct inspection task plus one aggregation helper.
The surface covers core video inspection: media facts, shots, frames, keyframes, and aggregated context output. Minor gaps exist (e.g. no audio/subtitle extraction or transcript tooling), but these are explicitly deferred to an external service, so agents can work around them.
Maintenance
Related MCP Connectors
FFmpeg as a service for AI agents: typed video editing tools, async jobs, downloadable outputs.
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
15 media & data tools for AI agents: search, transcribe, subtitles, voiceover, translate & more.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables local media processing (video/audio) using FFmpeg and FFprobe, allowing frame extraction, audio conversion, and metadata retrieval through natural language.5-
- AlicenseNot gradedqualityFmaintenanceExtracts ffprobe metadata, subtitles, scenes, and timelines from video files without frame-by-frame LLM vision, providing evidence-first reading for AI agents.1,063 npm3MIT
- AlicenseAqualityCmaintenanceEnables coding agents to turn videos from social platforms or local files into a small set of distinct frames plus a manifest, so they can answer visual questions by reading image paths. Also provides metadata lookup for captions and authors without downloading the video.2MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to download, convert, and visually analyze videos by providing timestamped frames and locally transcribed speech, all without API keys.MIT