video-understanding-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@video-understanding-mcpTranscribe /home/user/video.mp4 and save transcript to /home/user/transcripts"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
video-understanding-mcp
Local stdio MCP server for extracting bounded, reproducible evidence from video.
Early development: Milestones 1 and 2 provide safe video inspection and offline transcription. Timestamped frame extraction is planned next.
The server provides:
video_probe, backed byffprobevideo_transcribe, backed by FFmpeg and whisper.cpptimestamped transcript JSON and Markdown written to caller-selected roots
absolute-path and allowed-root enforcement
symlink-escape protection
bounded, cancellable subprocess execution
content-addressed probe and transcript caching
vu-doctorreadiness checks for FFmpeg, FFprobe, and whisper.cppsynthetic media fixtures and MCP stdio contract tests
Requirements
Node.js 22 or newer
ffmpegffprobewhisper-clia local whisper.cpp GGML model
Install whisper.cpp and bootstrap a model
On macOS with Homebrew:
brew install whisper-cpp
npm run bootstrap:model
export VU_WHISPER_MODEL_PATH="$HOME/.local/share/video-understanding-mcp/models/ggml-large-v3-turbo-q5_0.bin"The bootstrap script downloads the official
ggml-large-v3-turbo-q5_0.bin
model and verifies its Git LFS SHA-256 digest,
394221709cd5ad1f40c46e6031ca61bce88931e6e088c188294c6d5a55ffa7e2,
before reporting the configuration value. A failed checksum leaves no accepted
model. Pass a different destination directory when required:
npm run bootstrap:model -- /absolute/model/directoryThe server never downloads a model or makes any other network request while processing media.
Related MCP server: klaket-mcp
Build and run
git clone https://github.com/kwacky1/video-understanding-mcp.git
cd video-understanding-mcp
npm install
npm run build
VU_ALLOWED_READ_ROOTS="$HOME/Videos" \
VU_ALLOWED_WRITE_ROOTS="$HOME/Documents/transcripts" \
VU_WHISPER_MODEL_PATH="$HOME/.local/share/video-understanding-mcp/models/ggml-large-v3-turbo-q5_0.bin" \
npm startVU_ALLOWED_READ_ROOTS is a path-delimited list. If it is omitted, the server
allows reads only within its current working directory.
Optional configuration:
Variable | Default | Purpose |
| current directory | Readable filesystem roots |
| readable roots | Writable transcript roots |
|
| Maximum accepted input size |
|
| FFprobe executable |
|
| FFmpeg executable |
|
| whisper.cpp executable |
| none | Absolute path to a local GGML model |
| platform cache directory | Content-addressed cache root |
Both root variables are path-delimited lists (: on macOS/Linux, ; on
Windows). Output directories must already exist, resolve inside a configured
writable root, and be writable.
Transcription
video_transcribe accepts:
{
"path": "/absolute/path/demo.mp4",
"output_dir": "/absolute/path/transcripts",
"language": "en"
}It extracts the first audio stream to a temporary 16 kHz mono WAV, runs
whisper-cli locally, and removes scratch files on success, failure, or
cancellation. The result includes timestamped segments and durable
transcript_json_path and transcript_markdown_path values.
The transcript cache key includes the input SHA-256, transcription schema,
canonical parameters, whisper executable fingerprint, and model SHA-256.
Repeating the same call reuses cached transcript artefacts and returns
cache_hit: true; changing the input, language, model, or executable
configuration invalidates the cache.
Doctor
npm run doctor
npm run doctor -- --jsonThe doctor exits non-zero until all three executables are ready. Whisper is reported now so Milestone 2 prerequisites are visible before transcription is added.
Licence
Available Tools
2 toolsvideo_probeProbe local videoA
Inspect a local media file with ffprobe and return normalised metadata. The path must be absolute and inside an allowed read root.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to a local media file |
Output Schema
| Name | Required | Description |
|---|---|---|
| streams | Yes | |
| cache_hit | Yes | |
| container | Yes | |
| duration_ms | Yes | |
| input_sha256 | Yes | |
| schema_version | Yes | |
| ffprobe_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must stand alone. It does convey a read-only inspection via ffprobe, but it does not explicitly state that the file is unmodified, require specific permissions, or warn about failure modes. It is minimally transparent but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place: the first defines the operation, the second clarifies the precondition. The wording is direct, efficiently front-loaded, and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so the return value is already documented. With only one parameter and clear purpose, the description covers the key constraint (path root). It does not address branch conditions for failure or permission, but it is sufficiently complete for a simple read-only inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes 'path' as an absolute path to a local media file, but the description adds an important constraint: the path must be inside an allowed read root. This goes beyond the schema and adds meaningful semantic guidance for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action (inspect), resource (local media file), method (ffprobe), and output (normalised metadata). This plainly differentiates it from the sibling video_transcribe, which would perform transcription rather than metadata inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to choose this tool over video_transcribe. The description mentions an execution constraint (absolute path, allowed read root) but does not address selection among alternatives or any prerequisite/context beyond that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
video_transcribeTranscribe local videoA
Transcribe the first audio stream of a local media file with whisper.cpp. Processing is offline and writes timestamped JSON and Markdown into an allowed output directory.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to a local media file | |
| language | No | Whisper language code, or auto for detection | en |
| output_dir | Yes | Existing absolute directory for durable transcript files; the server does not create it |
Output Schema
| Name | Required | Description |
|---|---|---|
| engine | Yes | |
| language | Yes | |
| segments | Yes | |
| cache_hit | Yes | |
| duration_ms | Yes | |
| input_sha256 | Yes | |
| schema_version | Yes | |
| transcript_json_path | Yes | |
| transcript_markdown_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that processing is offline and that the tool writes timestamped JSON and Markdown files, which conveys privacy implications and side effects. It does not detail overwrite behavior or error cases, but the key operational behavior is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The primary action and constraints are front-loaded, and every clause adds meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the essential context: local file input, whisper.cpp engine, offline processing, and durable outputs. An output schema exists, so return-value documentation is not required here. Minor ambiguity around 'allowed output directory' is acceptable since the schema already notes the server does not create it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters well. The description adds no parameter-level detail, but it doesn't need to because the input schema covers path, output_dir, and language semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Transcribe the first audio stream of a local media file with whisper.cpp.' This is clearly distinct from the sibling video_probe, which implies probing metadata rather than producing transcripts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool when you need an offline transcription of a local media file written to timestamped output files. It does not explicitly say 'use video_probe instead for metadata' or list exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
video_probe - First observed
video_transcribe
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: video_probe inspects metadata, while video_transcribe performs audio transcription. There is no overlap or ambiguity between them.
Both tool names follow the same verb_noun snake_case pattern (video_probe, video_transcribe), with the action first and the domain prefix consistent. The naming is uniform and predictable.
With only two tools, the server feels thin for the broad domain of video understanding. The count is borderline but not extreme, as both operations are relevant and non-trivial.
The tool surface is severely incomplete for a server named 'video-understanding'. It only covers two basic operations: metadata inspection and speech transcription. There are no tools for common video-understanding tasks such as scene detection, object recognition, action classification, or even audio extraction.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Transcribe YouTube via Whisper. Summaries, chapters, semantic-search across your corpus.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Extract YouTube transcripts, search what was said, and read on-screen frames with cited timestamps.
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to query local video timelines by extracting speech, frame captions, and on-screen text into a SQLite store, exposing search and retrieval tools via MCP.PolyForm Noncommercial 1.0.0
- AlicenseAqualityBmaintenanceLet AI agents watch videos: local transcripts, speaker labels, scenes, chapters and exact-moment search from any video URL or file. Fully local, no API keys.42AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceDownloads and analyzes videos from platforms like Instagram, YouTube, and TikTok locally, returning keyframes and optional transcripts.MIT
- AlicenseNot gradedqualityAmaintenanceTurns a YouTube video or allowlisted local video into a timestamped transcript, chronological timeline, and retrievable image resources for transparent media preprocessing.1MIT