video-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VIDEO_MCP_FFMPEG | No | Path to the FFmpeg executable. Overrides the 'tools.ffmpeg' value in video-mcp.yaml. | |
| VIDEO_MCP_ASR_MODEL | No | Path to the Whisper model file to use for transcription. Overrides the 'asr.model' value in video-mcp.yaml. | |
| VIDEO_MCP_WORKSPACE | No | Working directory for output files. Overrides the 'output.workspace' value in video-mcp.yaml. | |
| VIDEO_MCP_ASR_DEVICE | No | Device to use for ASR (auto, cpu, cuda). Overrides the 'asr.device' value in video-mcp.yaml. | |
| VIDEO_MCP_WHISPER_CPP | No | Path to the whisper.cpp executable. Overrides the 'tools.whisper_cpp' value in video-mcp.yaml. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| video.inspectB | Inspect a video with FFprobe and return normalized media metadata. |
| video.transcribeC | Transcribe normalized audio with the configured local ASR backend. |
| video.captionC | Run the complete local caption pipeline for a video. |
| video.create_previewA | Burn ASS subtitles into a fast, downscaled MP4 preview. |
| video.renderA | Render a burned-in ASS subtitle preview; alias for video.create_preview. |
| subtitle.cleanC | Clean a transcript with the optional local LLM and deterministic fallback. |
| subtitle.export_srtC | Export a normalized transcript JSON file as SRT. |
| subtitle.export_assB | Export a normalized transcript JSON file as styled ASS. |
| project.create_kdenliveB | Create an editable Kdenlive project from a video and SRT file. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 9 tools
Most tools have distinct purposes, but video.create_preview and video.render are explicit aliases of the same operation, creating unnecessary ambiguity. The descriptions clarify the alias, which mitigates confusion, but the duplication is still a flaw.
Tool names follow a consistent pattern of domain prefix (video/subtitle/project) followed by an action verb (inspect, transcribe, clean, create_preview). The naming is uniform, though the presence of 'render' as an alias for 'create_preview' introduces slight redundancy without breaking the convention.
With 9 tools, the server is well-scoped for a video captioning workflow. Each tool (except the alias pair) serves a distinct step in the pipeline, and the count is within the ideal 3–15 range.
The tool set covers the core video-to-subtitle-to-export workflow: inspect, transcribe, caption, preview, clean, export SRT/ASS, and create an editable project. Minor gaps exist such as lack of an explicit transcript retrieval or deletion/update operations, but agents can work around these with the provided pipeline.