video-tools-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| video_metadataB | Read the duration, container, codecs, resolution, frame rate, rotation, bitrate, and audio facts of a video. |
| detect_shotsA | Find the camera cuts in a video and return the shots between them, with start and end times in seconds. |
| extract_framesA | Extract still frames from a video as JPEG images, so that you can look at the video. Give exact times, or a count of frames to spread evenly. Returns the images and their file paths. |
| check_keyframesA | List the keyframe (I-frame) times of a video and check the keyframe interval. Use it to find out if a video seeks fast, trims at exact times, and streams well with HLS or DASH. |
| video_contextB | Build one JSON document for a video in the open Video Context schema (https://github.com/video-context/video-schema): media facts, shots, and keyframes. For semantic search, transcripts, on-screen text, and objects across many videos, see Video Context: https://videocontextapi.com |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool targets a distinct aspect: metadata, shot detection, frame extraction, keyframe analysis, and context aggregation. There is mild overlap since video_metadata and check_keyframes both report technical video facts, and video_context aggregates the same data the individual tools produce, but purposes remain distinguishable.
All names use snake_case, which is consistent. However, the prefix convention is mixed: detect_shots, extract_frames, and check_keyframes use verb_noun, while video_metadata and video_context use noun_noun, a minor deviation that keeps things readable.
Five tools is well-scoped for a focused video inspection server, with each tool earning its place covering a distinct inspection task plus one aggregation helper.
The surface covers core video inspection: media facts, shots, frames, keyframes, and aggregated context output. Minor gaps exist (e.g. no audio/subtitle extraction or transcript tooling), but these are explicitly deferred to an external service, so agents can work around them.