Skip to main content
Glama

llm-video-mcp

CI License: MIT

Let any coding agent (Claude Code / Cursor / Kimi Code …) actually watch a video — within a token budget. An MCP server that downloads or ingests a video, extracts scene-aware frames with ffmpeg, transcribes speech with faster-whisper (VAD-gated), builds a shot timeline, and reverse-derives all extraction parameters from your token budget.

A video link tells your agent nothing, and piping raw frames eats the context window in seconds. Give the agent a token budget instead — llm-video-mcp derives everything else.

Install & connect (30 seconds)

Prerequisites: ffmpeg on PATH (winget install ffmpeg / brew install ffmpeg / sudo apt install ffmpeg), Python ≥ 3.10.

Add to your MCP host config (e.g. Claude Code / Claude Desktop):

{
  "mcpServers": {
    "llm-video-mcp": {
      "command": "uvx",
      "args": ["llm-video-mcp"]
    }
  }
}

Or from a clone: pip install . then llm-video-mcp (stdio). Then just tell your agent: "看看这个视频第 2 分钟讲了什么" — it calls the tools itself.

Related MCP server: markupR MCP Server

What the agent gets (real output)

analyze_video(path_or_url, token_budget=8000) on a real 60 s clip (腾讯视频, game-commentary video, faster-whisper tiny, zh) — unedited excerpts:

{
  "plan": {"token_budget": 8000, "max_frames": 28, "frame_width": 768,
           "grid": 4, "scene_threshold": 0.3, "estimated_tokens": 8000},
  "frame_count": 7,
  "transcript": {"source": "faster-whisper:tiny", "language": "zh",
                 "text": "…最後是中教材 就打到了…拿下了 Skog 的一個異血…"},
  "timeline": {"shot_count": 28, "cuts_per_min": 27.0,
               "shots": [{"start_sec": 0.0, "end_sec": 2.0, "motion": "static"},
                          {"start_sec": 2.0, "end_sec": 2.76, "motion": "static"}, …]}
}

Same pipeline on a 10 s synthetic test video at token_budget=8000: frames are packed into 2×2 contact sheets (grid: 4) and the estimate stays exactly within budget (estimated_tokens: 8000 ≤ 8000). Artifacts land in ~/.cache/llm-video-mcp/<content-hash>/: frames/*.jpg, frames.json, transcript.json, timeline.json.

Performance (measured 2026-07-20, CPU only — Intel Core Ultra X7 358H): a 5 min video at token_budget=8000 analyzes end-to-end in 41 s (whisper tiny) / 63 s (whisper base, model already cached) — scene frames + VAD-gated transcript + shot timeline, estimated_tokens: 7918 ≤ 8000.

Before / after

Without llm-video-mcp

With llm-video-mcp

You paste a video link into the chat. The agent sees a URL string and a title — and summarizes from guesswork.

analyze_video(path_or_url, token_budget=8000) downloads the video and returns scene-aware frames, a transcript, and a shot timeline shaped to the budget.

You hand-tune --max-frames 20 and either blow the context window or miss the scene that mattered.

token_budget reverse-derives frame count, resolution, contact-sheet packing, and transcript length. estimated_tokens ≤ token_budget, unit-tested and monotonic.

The video has no speech; the transcriber hallucinates a voiceover anyway.

Whisper is VAD-gated: silent videos honestly report no_speech (regression-tested). Subtitle tracks are preferred when present, no model download needed.

You want a closer look at one moment, so you re-run the whole extraction.

get_frames_at(path, timestamp) pulls frames around that moment only; full analyses are content-hash cached, so repeat calls return instantly.

When not to use this tool

  • Sub-second action precision — frame extraction (extractor.py:14) uses ffmpeg scene detection with a 1 frame/sec density floor; a 100 ms moment that is not a scene cut may not be captured. Use get_frames_at to target a specific timestamp, or reach for a frame‑accurate tool if every millisecond matters.

  • Real‑time / live streams — the pipeline downloads or reads a complete file (yt‑dlp / ffmpeg). It is not a streaming analyzer. Feed it a finished video, not a live feed.

  • Very long videos on CPU — faster‑whisper transcription (transcriber.py:81) is the bottleneck. At the measured throughput (5 min → 41 s @ tiny / 63 s @ base, CPU — see performance note above), a 2‑hour lecture means a multi‑minute wait. Use get_transcript for text‑only needs and skip frame extraction, or expect a long first analysis.

  • Frame‑dense visual detail — the budget controller (budget.py:85) deliberately reduces frame count, resolution, and packs frames into contact sheets to fit token_budget. Fast‑scrolling code or slides changing every second lose detail. If every frame is critical, a token‑budgeted sampler is the wrong shape.

  • Platform availability drift — yt‑dlp support for specific sites changes upstream. The supported‑sources table (dated 2026‑07‑20) is a point‑in‑time snapshot; login‑walled platforms other than 抖音/西瓜 (cookie‑bootstrapped) are out of scope.

The three tools

Tool

Purpose

analyze_video(path_or_url, token_budget, model?, lang?)

Full analysis: scene-aware frames + transcript + timeline, shaped to fit token_budget. Cached by content hash — repeat calls return instantly.

get_frames_at(path_or_url, timestamp, window?, count?)

On-demand frames around one moment — incremental viewing without re-analyzing.

get_transcript(path_or_url, start?, end?)

Text only. Prefers embedded/sidecar subtitles; whisper fallback is VAD-gated so silent videos honestly report no_speech instead of hallucinating.

Token budget controller (the core differentiator)

You give it a budget; it reverse-derives everything:

token_budget ──┬─► max_frames        (frame count cap)
               ├─► frame_width       (768 → 512 → 384 → 256)
               ├─► grid              (none → 2×2 → 3×3 contact sheets)
               ├─► scene_threshold   (fewer, more distinct frames when tight)
               └─► transcript_max_chars  (truncation)

Guarantees (unit-tested): estimated_tokens ≤ token_budget, and the estimate strictly decreases as the budget tightens. Built-in frame-token estimates for claude / gpt / gemini / kimi (model= param), overridable with a custom pricing config.

Supported video sources

Verified end-to-end on 2026-07-20 (yt-dlp 2026.07.04, real URLs, mainland-CN network):

  • Bilibili — works, no login needed

  • 腾讯视频 (v.qq.com) — works, no login needed

  • 优酷 (youku.com) — works, no login needed

  • 微博视频 (weibo.com) — works, no login needed

  • YouTube — works with a proxy (export https_proxy=… http_proxy=…)

  • 抖音 / 西瓜视频 — work after running llm-video-mcp bootstrap-cookies once (see below)

其他平台:请将视频下载为本地文件后使用本地路径输入。Local files are always a first-class input — any container/codec ffmpeg can read.

bootstrap-cookies: fresh cookies without touching your browser

ByteDance platforms (抖音/西瓜) require fresh (non-login) cookies. One command handles it:

llm-video-mcp bootstrap-cookies

It launches a separate Chrome instance with its own throwaway --user-data-dir, keeps a douyin.com page open for ~90 s to accumulate fresh cookies, closes only that instance gracefully (your browser is never touched — graceful close also flushes the cookie DB to disk), and exports ~/.cache/llm-video-mcp/cookies.txt. Pass it to the tools via cookies=...; douyin's __ac_signature cookie is bound to the browser UA, so the tools automatically replay the cookies with a User-Agent matching your local Chrome. Verified working 2026-07-20 on both douyin short links and long links.

vs. claude-real-video

Inspired by claude-real-video (MIT) — thanks to the author for proving the demand. We re-implemented everything from scratch; differences:

claude-real-video

llm-video-mcp

Agent integration

skill + folder of artifacts

native MCP server (3 structured tools)

Token economics

manual --max-frames guesswork

token_budget → parameters derived automatically

Timeline (shots / pan-zoom / cuts-per-min)

Pro paywall ($19+)

free, open source

On-demand re-viewing

re-run CLI

get_frames_at incremental queries

Development

pip install -e .[dev]
pytest tests -q      # synthetic test videos are generated by ffmpeg on the fly
ruff check src tests

Traction so far

Honest numbers, because this project is young:

  • v0.1.0 released 2026-07-20 — the full pipeline (scene frames, VAD-gated transcript, timeline, budget controller) works end-to-end today, verified on real Bilibili / Tencent Video / Youku / Weibo / YouTube / Douyin URLs.

  • The maintainer is still the only contributor. Zero stars, zero external PRs so far. You would genuinely be among the first.

  • GitHub traffic (14 days ending 2026-07-25): 29 unique cloners, 3 unique visitors — people are cloning it, but nobody has crossed the gap from "cloned" to "contributed" yet.

  • The contribution queue is ready: browse the open good-first-issue list — each issue is small, has hard acceptance criteria, and is testable offline.

The sibling project AgentCron used the same good-first-issue design and received its first 3 external PRs this way; the first-contribution case study documents what made those tasks approachable.

Support the project

If llm-video-mcp saved your agent from a hallucinated video summary — or you just want more coding agents that can actually watch things — a star on GitHub helps the next person find it.

Only download videos you have the right to download. MIT licensed.

A
license - permissive license
Not graded
quality - not tested
A
maintenance

Maintenance

1Releases (12mo)
Commit activity

Related MCP Connectors

  • Video scene understanding for AI agents via the Primate Vision API.

  • Agentic video editing on real footage: cut, caption, reframe, score, and export at full quality.

  • A real timeline video editor for AI agents: journaled edits, FFmpeg/MLT rendering, exports

  • Video analysis AI: transcripts, summaries, visual scenes/shots, clips, answers in natural language.

View all MCP Connectors

Related MCP Servers

View all related MCP servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/shkyyy18/llm-video-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server