video-reader-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@video-reader-mcpread video sample.mp4 including subtitles and scene changes"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Merged into anymd (2026-09-25). anymd reads PDFs, Office files, EPUB, web pages, images and video into clean Markdown for AI agents:
npx -y @sylphx/anymd. This repository is archived.
Cue
Video answers with timestamp-level proof
Cue gives agents a local video timeline they can search and cite. The default read returns container metadata, streams, chapters, and embedded subtitles. Scenes and frames are separate.
npx -y @sylphx/cueFor Claude Code:
claude mcp add cue -- npx -y @sylphx/cueRead a local video
{
"sources": [{ "path": "/absolute/path/to/demo.mp4" }]
}That call uses the fast profile. It returns container metadata, streams,
chapters, and embedded subtitles. It does not detect scenes, extract frames,
or run speech recognition.
Then ask:
“Which chapter covers the pricing change?”
Cue returns chapter and subtitle locators with timestamp_ms and the source
hash. It does not invent a transcript. A quote search matches embedded
subtitles only. Render or crop a frame afterwards with video_evidence, once
you have a timestamp.
Related MCP server: video-to-llm
Jobs Cue is built for
Ask your agent | Cue returns |
“Find this quote.” | a timestamped match in embedded subtitles, when those subtitles exist |
“Summarize this meeting.” | chapters and embedded subtitles |
“Where is the code shown?” | one frame from |
“What changed in this demo?” | scene boundaries when |
“Give me the useful moments.” | chapters, embedded subtitles, warnings, and gaps |
Tool surface
Tool | Purpose |
| Default |
| Search embedded subtitle cues. It does not run speech recognition. |
| Named follow-up: render, crop, or OCR one frame at a timestamp |
Predictable defaults
Omitted
profileisfast: container metadata, streams, chapters, and embedded subtitles. No scenes, frames, or speech recognition.profilequalityadds ffmpeg scene detection. It does not extract frames or run speech recognition.include_keyframesandinclude_transcriptstay off. Setting either one on the shipped server returns a warning and an empty array.OCR stays on
video_evidence.No cloud video API or frame-by-frame vision model is required.
Missing ffprobe or embedded subtitles is reported as a gap, not guessed around.
Why agents trust it
Every claim can point back to timestamp_ms, a stream index, a subtitle range,
or the source hash. A frame index is present only after video_evidence.
Companion MCP tools
Product | Job |
PDF answers with page-level proof | |
Image facts and pixel evidence | |
Repository architecture and impact | |
Exact code-chunk retrieval | |
Web research with source excerpts |
Each product is independent. Install only the tools your agent needs.
Development
bun install
bun run build
bun test
cargo test
bun run benchmark:public-proof
bun run benchmark:release-gateLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Video analysis AI: transcripts, summaries, visual scenes/shots, clips, answers in natural language.
Extract structured insights from videos, podcasts, articles, and PDFs with multi-model AI
Any video URL to LLM-ready transcript. ASR built in, no captions needed. TikTok, X, TED and more.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to query local video timelines by extracting speech, frame captions, and on-screen text into a SQLite store, exposing search and retrieval tools via MCP.PolyForm Noncommercial 1.0.0
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to process local videos into timestamped, citable text documents and then query them through tools for listing videos, retrieving transcripts, and fetching specific segments, all fully offline.MIT
- AlicenseAqualityCmaintenanceEnables coding agents to turn videos from social platforms or local files into a small set of distinct frames plus a manifest, so they can answer visual questions by reading image paths. Also provides metadata lookup for captions and authors without downloading the video.2MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to download, convert, and visually analyze videos by providing timestamped frames and locally transcribed speech, all without API keys.MIT