video-understanding-mcp
Related Servers
Alternatives to video-understanding-mcp
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to process local videos into timestamped, citable text documents and then query them through tools for listing videos, retrieving transcripts, and fetching specific segments, all fully offline.MIT
- AlicenseAqualityBmaintenanceEnables AI agents to analyze videos by producing timestamped transcripts, keyframe contact sheets, and metadata reports from local files or online URLs, fully offline and without API keys.6MIT
- AlicenseNot gradedqualityAmaintenanceProvides local, offline transcription, keyframe extraction, OCR, and pre-publish review of audio, video, and image files, enabling AI agents to see and hear media without cloud or API keys.36 npmApache 2.0
- AlicenseNot gradedqualityAmaintenanceTurns a YouTube video or allowlisted local video into a timestamped transcript, chronological timeline, and retrievable image resources for transparent media preprocessing.1MIT
- AlicenseAqualityAmaintenanceEnables AI clients to watch local video files or YouTube/Bilibili and other supported URLs, receiving timestamped transcripts, subtitles, searchable text and keyframe contact sheets. Runs fully offline with local speech recognition and bundled ffmpeg, so no API key or cloud upload is needed.684 PyPI9MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to query local video timelines by extracting speech, frame captions, and on-screen text into a SQLite store, exposing search and retrieval tools via MCP.PolyForm Noncommercial 1.0.0
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: video_probe inspects metadata, while video_transcribe performs audio transcription. There is no overlap or ambiguity between them.
Both tool names follow the same verb_noun snake_case pattern (video_probe, video_transcribe), with the action first and the domain prefix consistent. The naming is uniform and predictable.
With only two tools, the server feels thin for the broad domain of video understanding. The count is borderline but not extreme, as both operations are relevant and non-trivial.
The tool surface is severely incomplete for a server named 'video-understanding'. It only covers two basic operations: metadata inspection and speech transcription. There are no tools for common video-understanding tasks such as scene detection, object recognition, action classification, or even audio extraction.