mcp-video-frames
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-video-framesSample one frame per second from /home/user/video.mp4"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-video-frames — Video MCP Server
An MCP server that lets image-only models look at video: frames with timecodes, plus measurable video properties.
Available Tools
The server provides these tools:
view_frames- Sample a time range and return frames as images, each preceded by its timecodevideo_info- Container metadata, optionally with a whole-file scene/silence/loudness scan
The model you are talking to must accept image input. Without it, the returned frames are dropped or degraded to text.
This is a local server: it is not hosted anywhere, and a client has to be pointed at it. See Self-Hosted below.
Related MCP server: Video MCP Server
Self-Hosted
Installs nothing into your system Python: dependencies go into the project's own .venv/, and the client config points at that interpreter directly.
Clone the repository
git clone https://github.com/Azzy-H/mcp-video-frames.git
cd mcp-video-framesGenerate the virtual environment and the client config
python setup_mcp.pysetup_mcp.py only does three things:
creates
.venv/in this directory (run it with any Python 3.10+ interpreter);installs the two runtime dependencies from
requirements.txtinto that.venv/;writes
mcp-config.json, with the paths already filled in.
Options:
python setup_mcp.py --max-images 12 # match your client's image limit
python setup_mcp.py --print # print the config only; write nothing
python setup_mcp.py --ffmpeg <path> # ffmpeg is not on PATHIf ffmpeg is missing the script offers to download a prebuilt build into .venv/ (Windows), to point at one you installed, or to skip. Nothing is ever installed system-wide.
Check the environment before wiring it up
.venv/Scripts/python run_server.py doctor # Windows
.venv/bin/python run_server.py doctor # macOS / LinuxThis reports the ffmpeg/ffprobe paths and versions, whether the required filters are present, and whether the cache directory is writable — naming whatever is missing rather than leaving you to guess. It finds a venv-local ffmpeg on its own, so nothing has to be exported first.
test_server.py goes further and speaks real MCP to the server over stdio, so the whole path can be verified before any client is configured:
.venv/Scripts/python test_server.py connect <video file path> # Windows
.venv/bin/python test_server.py connect <video file path> # macOS / LinuxIt checks the handshake, the tool list, the 2n + 1 content-block ordering, whether a single-frame request comes back as PNG, and that an over-limit request is refused — and that stdout stays clean for the JSON-RPC stream.
Add this MCP server
Merge the generated mcp-config.json into your client's config. If an mcpServers section already exists, just add the video-frames entry to it. run_server.py puts src/ on sys.path itself, so no PYTHONPATH is needed and nothing has to be pip installed anywhere.
Replace /path/to/mcp-video-frames with the actual path to this repository on your system. On Windows the interpreter is .venv\Scripts\python.exe.
Uninstall: delete .venv/ and remove that config entry. Nothing else was written.
Requirements
Python 3.10 or newer
ffmpegandffprobe.setup_mcp.pyinstalls (or downloads) them into the project's own.venv/, and the server looks for them beside its own interpreter first, then onPATH— so a venv-local install needs noFFMPEG_PATHto be exported.ffmpegis required; withoutffprobe, metadata falls back to parsing ffmpeg's own header output (fewer fields), anddoctorsays so explicitlyA client model that accepts image input
Tools
view_frames
Samples a time range and returns an interleaved timecode + image sequence. Set start and end to the same value to look at a single moment.
Parameter | Type | Default | Notes |
| string | — | absolute path (required) |
| number | 0 | start in seconds |
| number |
| end in seconds |
| number | 1.0 | sampling interval, 0.05–600 s |
| integer | 12 | frames for this call; may be raised up to the image limit |
| integer | 768 | long-edge size in pixels, 64–4096 |
|
|
| |
| integer | 85 | jpeg quality, 1–100 |
The result is 2n + 1 content blocks:
text: {"n":0,"t":0.0}
image: <frame 0>
text: {"n":1,"t":1.0}
image: <frame 1>
...
text: {"video":"...","frames":12,"range":[0.0,11.0],"interval":1.0,
"format":"jpeg","max_edge":768,"clamped":false,"cached":9,
"limits":{"max_images_per_call":20,"max_image_bytes":10485760,
"max_total_bytes":41943040}}The block order is load-bearing: an image without its timecode cannot be reasoned about in time. clamped reports whether the range was trimmed to the video's length, cached how many frames came from disk, and limits repeats the budget in force so the numbers do not have to be discovered by being rejected.
A range defaults to jpeg because at equal edge length it is about an order of magnitude smaller, which fits many more frames inside a client's byte budget. Ask for a single frame as png when detail matters.
Over a limit the call fails: nothing is truncated and the interval is never widened. Silently coarsening the sampling grid would invalidate the caller's assumption about time precision, so the whole call fails instead. If any single frame fails validation, no frames are returned.
video_info
Returns video metadata. detail="basic" reads container headers only and returns immediately; detail="full" adds an exhaustive whole-file scan for scene cuts, silences and EBU R128 loudness.
{
"path": "", "sha256": "", "size_bytes": 0,
"duration": 0.0, "fps": 0.0, "width": 0, "height": 0,
"video_codec": "", "pix_fmt": "",
"color_range": "", "color_space": "", "is_hdr": false,
"audio_tracks": [], "subtitle_tracks": [], "chapters": [],
"detail": "full",
"scene_cuts": [1.23, 5.67],
"silences": [{"start": 12.0, "end": 14.5, "duration": 2.5}],
"loudness": {"integrated_lufs": -23.0, "true_peak_dbfs": -1.2, "lra": 5.0},
"scan": {"ffmpeg_version": "", "computed_at": "", "cached": false}
}With detail="basic" the scene_cuts / silences / loudness / scan fields are absent rather than empty — "no value" and "not scanned" mean different things. The first full call scans the whole file; later calls reuse the cache. force=true recomputes.
Configuration
Environment variables take precedence over defaults.
Variable | Default | Purpose |
| 20 | maximum images per call (the one that matters most) |
|
| cache root |
| auto-detected | checked before |
| 10485760 | per-frame byte limit |
| 41943040 | per-result total byte limit |
| 8192 | per-frame edge limit |
| 5368709120 | cache size limit |
| 14 | frame age limit |
| 0.3 | scene-change sensitivity |
MCP_VIDEO_MAX_IMAGES caps how many frames one call may return. Set it to match your client: limits differ between clients, and while too low is merely conservative, too high makes the entire batch unusable. max_frames defaults to 12 and may be raised only up to this limit.
The effective limits are also stated in the tool descriptions and repeated in every view_frames result, so a caller can see them without being rejected first.
CLI
The same package ships a CLI for troubleshooting and batch work:
mcp-video-frames doctor # environment self-check
mcp-video-frames info video.mp4 # basic metadata
mcp-video-frames scan video.mp4 # pre-build the full scan
mcp-video-frames scan --dir ./videos # process a whole directory
mcp-video-frames frames video.mp4 --start 0 --end 5 # frames, as content-block JSON
mcp-video-frames cache --stats # cache usage
mcp-video-frames cache --prune # evict cache entriesWithout installing the project (Self-Hosted above), use the launcher for the same subcommands:
.venv/Scripts/python run_server.py doctor # Windows
.venv/bin/python run_server.py info video.mp4 # macOS / LinuxRunning it with no arguments prints help; as an MCP server it takes no arguments. When stdin is a pipe (as an MCP client provides) a bare invocation starts the server, which is how nothing has to be configured for the stdio case.
doctor prints its JSON report and then a one-line human verdict ("Self-check passed" / "Self-check failed"), so it is the one subcommand whose stdout is not pure JSON. The CLI and the MCP tools call the same core functions, so behaviour cannot drift between what works in a terminal and what works for a model.
Cache
Location: depends on how the server was started.
Deployment | Cache root |
|
|
A plain package install with no |
|
The setup script configures the project-local .cache/ deliberately, together with cwd. A relative path resolves against the client's working directory, so without cwd the cache would land wherever the client happened to start — less predictable than the temp default, not more. It is already in .gitignore.
MCP_VIDEO_CACHE_DIR overrides either default, and the resolved path is printed to stderr at startup.
Everything in the cache is reproducible — deleting the directory is always safe, and only costs a re-probe.
Entries are keyed by content hash (size + mtime + the first and last 1 MB), and frames are stored individually by absolute timestamp, so overlapping views reuse work: look at 0–20 s, then 10–30 s, and 10–20 s is already on disk.
If you left the cache in the system temp directory, do not rely on the operating system to clean it up — the platforms differ too much:
Platform | Temp-directory cleanup |
Windows | Essentially none. Storage Sense runs only under disk pressure and skips files in use |
Linux |
|
macOS |
|
So the server cleans up after itself, evicting whole video entries by LRU once MCP_VIDEO_CACHE_MAX_BYTES (default 5 GB) is exceeded, plus a 14-day frame age limit (MCP_VIDEO_FRAME_MAX_AGE_DAYS).
The tiers follow recomputation cost: frames are cheap and go first, while whole-file scan results (video_info(detail="full")) are expensive to rebuild and are kept preferentially.
Eviction runs asynchronously after a tool call returns — it never blocks a response, and there is no full scan at startup.
Multiple clients share one cache. Reads and writes are guarded by a file lock, writes go through a temp file plus os.replace, and a frame that has been evicted between listing and reading is regenerated rather than reported as an error.
License
MIT — see LICENSE.
Available Tools
2 toolsvideo_infoA
Metadata for a video. detail="basic" reads container headers only and returns immediately. detail="full" additionally scans the whole file for scene cuts, silence intervals and EBU R128 loudness; that scan is exhaustive and therefore slow on long files, but it is cached and reused. Use force=true to recompute. This tool returns no images, so the image limits do not apply to it, but your client's per-call time limit does: a full scan of a long video takes time proportional to its length, and may exceed the limit. Raise that limit if your client allows it.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | ||
| video | Yes | ||
| detail | No | basic |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it excels: it discloses that full scans are 'exhaustive and therefore slow,' that results are 'cached and reused,' that force recomputes, and that a full scan may 'exceed the limit' of the per-call time budget. It also states a non-behavior of the tool (no images, so image limits don't apply). This is model-level transparency for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but each clause earns its place: purpose, the two modes with their performance profiles, caching, force, image-limit exemption, and timeout risk. It is front-loaded with the definition before moving to behavior and caveats, and no sentence is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema already documents the return shape, so the description does not need to list response fields. What remains is well covered: mode selection, performance, caching, time limits, and the boundary against an image-returning sibling. The only meaningful omission is the format of the video input itself, which matters for call success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for two of the three parameters: detail is fully explained with concrete consequences of each enum value, and force is tied to the caching behavior ('Use force=true to recompute'). The required video parameter is never described — its format, source, or accepted identifiers are left to inference from the parameter name alone — which is the only gap preventing a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (video) and the core action (reading metadata) through the opening clause and the detailed account of what each detail level retrieves — container headers, scene cuts, silence intervals, and loudness. It differentiates from the sibling view_frames by explicitly noting it 'returns no images,' allowing an agent to tell them apart. The verb is implied rather than stated (a noun phrase opens the description), keeping this just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear internal routing guidance: basic for immediate header-only reads, full when deep analysis is wanted, and force=true when the cached scan must be recomputed. It also warns when the tool is not appropriate — when a call risks exceeding the client time limit — and notes that image limits are irrelevant since no images are returned. It does not name the sibling explicitly or state 'use view_frames when you need images,' so it falls just short of explicit when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
view_framesA
Sample a video and return frames as images, each preceded by its timecode — the only thing tying an image to a moment. Set start == end to look at a single moment (returns a lossless PNG); the interval is ignored in that case. Range requests default to jpeg, which is about an order of magnitude smaller and fits far more frames in the same budget. Frames are cached by absolute timestamp, so overlapping ranges reuse work. max_frames defaults to 12; raising it up to the image limit trades away detail — lower max_edge or use format="jpeg" to afford more frames. All of these limits are set by whoever deployed this server, not by you. Server limits for this deployment: at most 20 images per call, each at most 10 MB, 40 MB per result. Frames are never silently dropped or resampled, so the interval you ask for is the interval you get. Requests past a limit fail with the numbers and concrete alternatives — the server never truncates a range or coarsens your interval.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| start | No | ||
| video | Yes | ||
| format | No | ||
| quality | No | ||
| interval | No | ||
| max_edge | No | ||
| max_frames | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers richly: frames are cached by absolute timestamp, overlapping ranges reuse work, frames are never silently dropped or resampled, the interval requested is the interval returned, and requests past limits fail with numbers and alternatives. It also discloses the exact server limits (20 images, 10 MB each, 40 MB per result) and the lossless PNG behavior for start==end. This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core purpose, then moves through key behaviors and limits. Every sentence adds information, but it is a long paragraph that could benefit from light structuring (e.g., separating the caching note from the limits note). Still, no sentence is wasted, and the most important facts come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and a single sibling, the description covers the purpose, the key parameter interactions, the failure mode, the caching behavior, and the deployment-specific limits. The output schema exists, so return-value details are already structured. An agent has everything needed to call this tool correctly and to reason about trade-offs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for most parameters: start/end/interval semantics, format trade-offs, max_frames default and trade-offs, and max_edge as a lever. It does not explicitly explain quality, but the schema already provides min/max/default for it, and the description's trade-off framing covers the main decision parameters. Slight gap on quality, hence 4 rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Sample a video and return frames as images, each preceded by its timecode.' This clearly distinguishes it from the sibling video_info, which presumably returns metadata rather than frames. The timecode detail and the start==end single-moment behavior further sharpen the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: set start == end for a single moment, use jpeg for smaller payloads, raise max_frames only if you can trade detail, and lower max_edge or use jpeg to afford more frames. It also states that server limits are fixed by the deployer, not the agent, and that requests past a limit fail rather than being silently truncated. This is strong when-to-use and how-to-trade-off guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
video_info - First observed
view_frames
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one extracts frames from videos, the other retrieves metadata. No overlap in functionality; each addresses a separate need.
Both names follow a verb_noun pattern: 'view_frames' and 'video_info'. The verbs differ ('view' vs 'video') but the structure is consistent. Minor deviation in verb choice, but pattern is clear.
With only 2 tools, the set is minimal but each tool is substantial and well-scoped, covering the core functions of frame extraction and metadata retrieval. Slightly thin for a full video analysis workflow, but reasonable for a focused server.
The set covers the primary operations (extract frames, get info) but lacks additional common video operations such as transcoding, clipping, or audio extraction. The domain is video analysis, and while the two tools address key needs, there are notable gaps that agents might encounter.
Maintenance
Related MCP Connectors
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.
Video analysis AI: transcripts, summaries, visual scenes/shots, clips, answers in natural language.
Give AI random access to video: timestamped contact sheets + zoom into any start/end range.
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to watch YouTube videos by extracting frames at scene changes and visual references, pairing each frame with the exact words spoken at that timestamp. Provides dense frame-transcript interleaving for any model.9 npm2MIT
- FlicenseNot gradedqualityDmaintenanceBridges Claude and video content by extracting keyframes and transcribing audio, enabling Claude to analyze video files.-
- AlicenseNot gradedqualityCmaintenanceExtracts ffprobe metadata, subtitles, scenes, and timelines from video files without frame-by-frame LLM vision, providing evidence-first reading for AI agents.7 npm3MIT
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to watch and analyze videos by extracting scene-aware frames, transcribing speech, and generating shot timelines, all constrained by a token budget to fit LLM context limits.1MIT