gemini-video-mcp
This server lets MCP clients (like Claude) analyze public YouTube and Instagram videos as well as local video files through Google Gemini, combining audio and visuals to answer questions.
Analyze YouTube videos via URLs (
watch?v=,youtu.be/,/shorts/,/live/,/embed/) — no yt-dlp needed.Analyze Instagram reels and video posts (
/reel/,/p/,/tv/, etc.) — downloads via yt-dlp, uploads to Gemini, then analyzes.Analyze local video files by absolute path — upload once, reuse for follow-up questions.
Ask specific questions about a video with a
prompt.Get a full structured analysis when no prompt is given: summary, chapters with timestamps, key statements, visual details, and closing assessment.
Analyze a specific time range using
start/end(supports"12:30","750s", or milliseconds).Choose processing modes:
agentic(model navigates the video itself, ~90% fewer tokens on long videos),static(frame-by-frame, exact trimming), orauto(default, picks the best mode).Request high visual detail with
detail: "hoch"for reading on-screen text, code, or fine visual details.Manage uploaded local video copies via
manage_local_video: upload, check status/expiry, delete from Google.See token usage on every call, with automatic fallback through multiple Gemini models on quota errors (429/503).
Uses the Google Gemini API to process videos, with support for agentic navigation, static frame-by-frame analysis, adjustable resolution, and token usage reporting.
Allows analysis of public YouTube videos, including full-video summaries with chapters and timestamps, targeted time-range queries, and combined audio/visual understanding of on-screen content.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-video-mcpAnalyze this video and summarize key points with timestamps: https://www.youtube.com/watch?v=abc123"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
gemini-video-mcp
An MCP server that lets Claude — or any other MCP client — understand public
YouTube and Instagram videos as well as your own local video files, both what
is said and what is shown, through the Google Gemini API. It exposes
analyze_video over stdio, plus manage_local_video for the upload of local
files.
Why
Language models cannot watch YouTube videos. The usual workaround is to paste a transcript by hand, which throws away everything that only appears on screen, or to ask a second app and copy the answer back. This server removes that detour: the client asks a question about a URL and gets an answer grounded in the audio and the video track. The same call works for a YouTube link and an Instagram reel; the difference in how they are fetched is handled internally.
Related MCP server: yt-analysis-mcp
What it does
YouTube videos and Instagram reels and video posts
Your own video files by local path, uploaded once and reused for follow-up questions
Full analysis with chapters and timestamps across the whole video
Targeted analysis of a specific time range
Audio and visuals evaluated together — on-screen text, code, diagrams and demos are part of the answer, no separate transcript needed
Token usage reported on every call, so the cost of each request is visible
The agentic finding
This is why the project exists rather than just calling generateContent.
Gemini can process YouTube videos two ways. The classic endpoint,
models:generateContent, pulls the video into the context window frame by
frame. The POST /v1beta/interactions endpoint in agentic mode lets the model
navigate the video itself.
Same 15-minute video, same question:
|
| |
Tokens | 85,480 | 7,416 |
Answer | incomplete | more complete coverage |
Roughly 91 % fewer tokens, with better coverage. This server therefore uses
interactions throughout. The processing field that selects the mode does not
exist on generateContent at all — sending it there returns
400 INVALID_ARGUMENT: Unknown name "processing".
Measurements taken with this server on a 15-minute video:
Call | Mode | Tokens |
Whole video, specific question | agentic | 8,203 |
Range 10:30–12:30 | static, | 11,293 |
Range 10:30–12:30 | static, | 35,922 |
Worth noting: low picked up the same on-screen-only GitHub repository name
that high did, for a third of the cost. High resolution is therefore a
deliberate choice (detail: "hoch"), never automatic.
On a short clip the two modes cost about the same. Measured on a 51-second Instagram reel, same question:
Mode | Tokens | Answer |
agentic | 5,570 | full breakdown, separating what is said from what is shown |
static, | 5,288 | a shorter timestamped list |
So agentic stays the default even for short videos: 5 % more tokens bought a
noticeably richer answer, and the gap widens enormously as videos get longer.
Requirements
Node 22 or newer (developed and tested on 24.16)
Your own Gemini API key from Google AI Studio
yt-dlp — only for Instagram. YouTube works without it.
Install yt-dlp with pipx install yt-dlp, brew install yt-dlp,
winget install yt-dlp, or grab a binary from its releases page. If it is not
on your PATH, set YTDLP_PATH to the full path of the executable.
Installation
git clone https://github.com/DerYUYU/gemini-video-mcp.git
cd gemini-video-mcp
npm installCreate a .env file next to package.json (see .env.example):
GEMINI_API_KEY=your-key-here.env is gitignored and must never be committed.
Register with your MCP client
For Claude Code:
claude mcp add gemini-video -- node /path/to/gemini-video-mcp/src/index.jsFor clients that use a configuration file (Claude Desktop and others):
{
"mcpServers": {
"gemini-video": {
"command": "node",
"args": ["/path/to/gemini-video-mcp/src/index.js"]
}
}
}Replace /path/to/gemini-video-mcp with the absolute path to your clone. The
key can also be passed here instead of via .env:
{
"mcpServers": {
"gemini-video": {
"command": "node",
"args": ["/path/to/gemini-video-mcp/src/index.js"],
"env": { "GEMINI_API_KEY": "your-key-here" }
}
}
}Usage
The client calls the tool; you normally just ask in plain language. The arguments below are what the tool receives.
A simple question about a video
{
"url": "https://www.youtube.com/watch?v=VIDEO_ID",
"prompt": "Which tools does the speaker recommend? Include timestamps."
}Analysing a specific range — auto switches to the static mode here and
cuts exactly to the requested window:
{
"url": "https://www.youtube.com/watch?v=VIDEO_ID",
"start": "10:30",
"end": "12:30",
"prompt": "What is shown on screen during this part?"
}Higher visual detail — roughly three times the tokens per video second, so only worth it on a short range:
{
"url": "https://www.youtube.com/watch?v=VIDEO_ID",
"start": "14:00",
"end": "14:40",
"detail": "hoch",
"prompt": "Read the code in the terminal line by line."
}An Instagram reel — same call, same arguments. The download and upload happen internally:
{
"url": "https://www.instagram.com/reel/POST_ID/",
"prompt": "What happens in this clip? Include timestamps."
}Leaving out prompt requests a full structured analysis: summary, chapters
with timestamps covering the entire video, key statements, what is shown
visually, and a closing assessment.
Example output
Every answer starts with a header stating mode, resolution, range, model and token usage. Abbreviated real output from a question about a speaker's Linux setup:
**Video:** https://www.youtube.com/watch?v=VIDEO_ID
**Modus:** agentic
**Tokenverbrauch:** 8.203 Tokens
**Modell:** gemini-3.8-flash
---
* **Distribution [ca. 10:45 – 10:57]:**
Er nutzt das aktuelle **CachyOS**, eine auf **Arch Linux** basierende
Distribution [...]
* **Fingerabdrucksensor [ca. 14:27 – 15:02]:** Der Sensor wird unter Linux zwar
prinzipiell unterstützt, ist im Alltag jedoch extrem unzuverlässig [...]See Response language for why this example is in German.
Tool reference: analyze_video
Parameter | Type | Required | Default | Description |
| string | yes | — | URL of a public video. YouTube: |
| string | no | full analysis | The question to ask. Omitted, a complete structured analysis with chapters and timestamps is requested. |
|
| no |
| Processing mode, see below. |
| string | number | no | — | Start of a range: |
| string | number | no | — | End of the range, same formats. |
|
| no |
|
|
Modes
agentic — the model navigates the video itself and picks the relevant
passages. It has no hard time offsets. The right choice for whole videos and for
any question about content.
static — the video is processed second by second. This allows exact
trimming via start/end and the high resolution setting, but costs roughly
100 (low) to 300 (high) tokens per video second. Without a range it gets
expensive fast on long videos.
auto (default) decides as follows:
Situation | Result |
|
|
Range of 5 minutes or less |
|
Range longer than 5 minutes |
|
No range |
|
An explicit mode is honoured, except that detail: "hoch" always requires
static. Any such override is reported in the answer under "Hinweise" (notes).
Tool reference: manage_local_video
Parameter | Type | Required | Description |
|
| yes |
|
| string | yes | Absolute path of the local video file. For |
How Instagram works differently
For YouTube, the video URL is handed to Gemini directly and Google fetches the video itself. Gemini cannot do that for Instagram, so the server takes a detour:
yt-dlpdownloads the video to a temporary directorythe file is uploaded to the Gemini Files API
the server waits until Google reports the file as
ACTIVE— analysing it earlier failsthe analysis runs as usual
the uploaded file is deleted at Google and the temporary file is removed, including when a step in between fails
What this means in practice:
It is slower. A short reel takes a few seconds to download and several more to upload and process, before the analysis even starts. The YouTube path has none of that overhead.
It needs yt-dlp. YouTube does not.
Public posts only. The server holds no cookies, no session and no login, by design. Private accounts, stories and anything behind a login are out of scope and are reported as such, not worked around.
Video posts only. Image posts and image carousels are rejected with a clear message rather than a silent failure.
Instagram support in yt-dlp can break. Instagram changes its delivery paths regularly and actively works against downloaders, so extraction that works today may fail after a site change. A
yt-dlp -Uusually picks up the fix. The server says so in the error message instead of reporting a generic failure.Downloads are capped at
MAX_VIDEO_MB(250 MB by default), enforced both by yt-dlp during the download and by the server before the upload.Cleanup runs even when a step fails, but it cannot run if the process is killed outright. Should that happen, uploaded files expire at Google after 48 hours on their own.
TikTok is deliberately not supported, even though yt-dlp could handle it.
Local video files
Pass an absolute path instead of a URL. Windows paths with backslashes and a drive letter work, forward slashes too, and surrounding quotes from Explorer's "Copy as path" are stripped.
Before anything is uploaded the server checks that the file exists, that its
extension is one Gemini supports (.mp4, .mpeg, .mpg, .mov, .avi,
.flv, .webm, .wmv, .3gp) and that it is not larger than 2 GB. The Files
API documentation states 2 GB per file; the video documentation says 2 GB on the
free tier and 20 GB on paid tiers. The server defaults to the smaller figure,
counted in decimal gigabytes (2,000,000,000 bytes). LOKAL_MAX_MB raises it if
your tier allows more.
Unlike Instagram, the uploaded copy is not deleted after the answer. Uploading a 1.7 GB file again for every question would be absurd. Instead:
The copy is tagged at Google with a hash of path, size and modification time. A later call with the same, unchanged file reuses it, even after a server restart. The path itself never leaves your machine.
If the file has changed, it is uploaded again and older copies of the same path are deleted.
The Files API keeps files for 48 hours. A copy with less than an hour left is not used any more; the file is uploaded again.
Two calls for the same file at the same time share one upload.
manage_local_videowithaction: "delete"removes all copies of a path at Google right away.action: "status"shows whether and until when a copy exists. The local file is never touched.
Long videos
A 45-minute file of 1.7 GB takes a while to upload, depending on your upload
bandwidth, and Google then needs time to process it. The server waits up to
30 minutes for the processing (LOKAL_UPLOAD_TIMEOUT_MS); if it does not finish,
the copy is deleted. The recommended flow:
manage_local_videowithaction: "upload"and the path. This only uploads and waits forACTIVE; nothing is analysed yet.analyze_videowith the path and nostart/end: overview with chapters and timestamps.analyze_videoagain per chapter, withstartandendfrom the overview. Every call reuses the copy from step 1.manage_local_videowithaction: "delete"when you are done, or let the copy expire after 48 hours.
About timeouts in Claude Code: a single MCP tool call has a hard limit of about
28 hours by default (MCP_TOOL_TIMEOUT), which is not the issue. The issue is
the idle timeout: a stdio tool call that sends neither a response nor a progress
notification for 30 minutes is aborted (CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT,
Claude Code 2.1.203 or later). The server sends a progress notification every
minute while it works, provided the client asks for progress. If a call is
aborted anyway, the upload keeps running inside the server process, and the next
call for the same file waits for it instead of starting a second one. For a
guaranteed margin, set a per-server timeout in .mcp.json, which also acts as
a floor for the idle timeout:
{
"mcpServers": {
"gemini-video": {
"command": "node",
"args": ["/path/to/gemini-video-mcp/src/index.js"],
"timeout": 3600000
}
}
}If the server process is killed in the middle of an upload, the SDK's resumable
upload is never finalised, so no usable file should appear at Google. This
follows from the upload protocol and has not been tested by killing a live
upload. Anything that does get stuck shows up under action: "status" and can
be deleted.
Keep in mind that on the free tier Google may use submitted content to improve its products. Decide before uploading a private video.
Model fallback chain
The free tier's daily limit applies per model, so an exhausted model does not mean the API is unusable. When a request comes back with HTTP 429 (quota exhausted) or 503 (model overloaded), the server retries the identical request against the next model in the chain instead of surfacing an error.
Default chain, all of which support agentic video processing:
gemini-3.8-flashgemini-3.7-flashgemini-3.6-flashgemini-3.5-flash-lite
Rules:
Only 429 and 503 trigger a switch. An invalid key, a private video or a malformed argument fails the same way on every model, so retrying would just burn quota. Those errors are reported immediately.
Each model is tried at most once. When the whole chain is exhausted, the error says so and points out that the quota resets daily.
The switch is never silent. The answer always names the model that responded, and when a fallback was used, the notes say which models were skipped and why.
For Instagram the file is uploaded once. The fallback reuses the same uploaded file reference, so a model switch costs no extra upload.
Set GEMINI_MODEL to choose the starting model — it always stays first in the
chain. Set GEMINI_MODEL_FALLBACKS to a comma-separated list to replace the
models tried after it.
Limits and cost
Public videos only. Private, unlisted, deleted or region-blocked videos fail. For Instagram this also covers private accounts and stories.
Instagram costs extra time, not extra Gemini quota. The download and the file upload do not consume analysis tokens, but they do add wall-clock time.
Free tier: 20 requests per day, per model — not per account, and not the frequently cited 8 hours of YouTube footage per day. Because the limit is per model, the server falls back through a chain of models automatically (see below), so the practical ceiling is a multiple of 20 requests per day.
Token cost in static mode is about 100 tokens per video second at
lowand about 300 athigh. Agentic mode does not scale this way; the measurements above are representative.An API key bills separately from any Google AI subscription. A paid Gemini app plan does not cover API usage.
Long videos take time. Several minutes is normal in agentic mode. The default timeout is 10 minutes, adjustable via
GEMINI_TIMEOUT_MS.start/endonly take effect in static mode. In agentic mode the requested range is written into the prompt but not hard-trimmed.Timestamps come from the model and can be off by a few seconds.
Video only. No image or PDF tooling.
Response language
The model answers in the language of the question you ask. The default prompt
and the tool's own output labels and notes ship in German, so a call without an
explicit prompt returns a German analysis — as in the example above. Pass your
own prompt in English to get an English answer.
Three deviations from Google's documentation
All three were verified against the live API and cost real debugging time. If you are building against this API yourself, they are worth knowing.
1. Time offsets are second-strings, not millisecond numbers.
processing.start_offset and end_offset accept only strings with an s
suffix, such as "630s". Passing raw milliseconds returns
400 Invalid input at 'input[0].processing'. This server accepts "12:30",
"750s" and milliseconds from the caller and converts internally.
2. The field is resolution, not media_resolution.
media_resolution is the generateContent name. On the interactions
endpoint the setting sits directly on the video input as resolution.
3. The two most common errors are signalled the wrong way round.
An unreachable video returns 403 The caller does not have permission — with no
mention of the video at all. An invalid API key returns 400 with an empty
message body. Classifying these by their text alone reports a private video as a
key problem. This server resolves the ambiguity by issuing a models.list call
with the same key, which takes about 0.1 s, costs no tokens, and only runs on
the error path.
Project layout
File | Purpose |
| MCP server, stdio transport, tool registration |
| Input validation, mode selection, orchestration, result formatting |
| Wrapper around |
| Platform detection, YouTube vs Instagram vs local file vs rejected |
| Local files: validation, upload once and reuse, status, delete |
| Instagram download via yt-dlp, error translation |
| Gemini Files API: upload, wait for |
| Default prompt and time-range addendum |
| Time parsing, normalisation to seconds |
| YouTube URL validation |
src/analyze.js has no MCP dependencies and can be exercised directly from a
Node script.
Configuration reference
Variable | Required | Default | Purpose |
| yes | — | Key from Google AI Studio |
| no |
| Starting model, always first in the fallback chain |
| no | see chain above | Comma-separated models to fall back to on HTTP 429 or 503 |
| no |
| Timeout for the analysis request, in milliseconds |
| no |
| Full path to the yt-dlp executable (Instagram only) |
| no |
| Timeout for the Instagram download |
| no |
| How long to wait for Gemini to process an uploaded Instagram file |
| no |
| Size ceiling for a downloaded Instagram video |
| no |
| How long to wait for Gemini to process an uploaded local file |
| no |
| Size ceiling for a local file, in decimal megabytes |
License
MIT — see LICENSE.
Available Tools
1 toolanalyze_videoYouTube-Video analysierenARead-only
Analysiert ein oeffentliches YouTube-Video mit Google Gemini und beantwortet eine Frage dazu. Ohne "prompt" liefert das Tool eine vollstaendige Analyse mit Kapiteln, Zeitstempeln und visuellen Details. Bild und Ton werden gemeinsam ausgewertet. Die Antwort weist immer den Tokenverbrauch aus. Funktioniert nur mit oeffentlichen Videos (keine privaten oder nicht gelisteten).
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | Ende des Ausschnitts, gleiche Formate wie "start". | |
| url | Yes | URL des oeffentlichen YouTube-Videos, z. B. https://www.youtube.com/watch?v=VIDEOID | |
| mode | No | "agentic" (Default ueber auto): Gemini navigiert selbst durchs Video, ~90 % guenstiger und vollstaendiger bei langen Videos. "static": Frame-fuer-Frame, erlaubt exakte Zeit-Offsets und hohe Aufloesung, aber teuer. "auto": agentic, ausser bei engem Ausschnitt. | |
| start | No | Beginn des Ausschnitts: "12:30", "1:02:30", "750s" oder Millisekunden als Zahl. | |
| detail | No | "normal" (Default) oder "hoch". "hoch" erzwingt den statischen Modus mit hoher Aufloesung (~300 statt ~100 Tokens pro Videosekunde) -- sinnvoll fuer Bildschirmtexte, Code oder feine Bilddetails, nur zusammen mit einem kurzen Ausschnitt empfehlenswert. | |
| prompt | No | Die konkrete Frage an das Video. Ohne Angabe wird eine vollstaendige strukturierte Analyse mit Kapiteln und Zeitstempeln angefordert. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though readOnlyHint=true already signals a safe read-only operation, the description adds meaningful behavioral context: it names the model (Gemini), explains that both image and audio are evaluated together, and states that token usage is always reported. It also discloses the public-video restriction, which is useful beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it starts with the core purpose, then adds default behavior, multimodal processing, output characteristic, and a key constraint. Every sentence carries useful information without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main behavioral aspects and typical outputs, including the default analysis format and token reporting. Since there is no output schema and no sibling tools, the description sufficiently supports tool selection and invocation, though it could slightly expand on what the answer looks like when a prompt is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema provides detailed meaning for all parameters including mode, detail, start, end, prompt, and url. The description adds little parameter-level semantics beyond what the schema already states, such as the default full analysis when prompt is omitted, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: it analyzes a public YouTube video with Google Gemini and answers a question about it. It clearly distinguishes its behavior from generic video tools by stating that it processes image and audio together and can produce a full structured analysis with chapters, timestamps, and visual details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool is appropriate: for public YouTube videos only, and it explicitly excludes private and unlisted videos. It also explains the default behavior when no prompt is supplied, which helps the agent decide whether to call the tool with or without parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
analyze_video
TDQS
Scored across 1 tool
With only one tool in the server, there is no possibility of confusion or overlap. The tool's purpose is clearly defined and distinct from any other hypothetical tool.
The single tool uses a clear snake_case verb_noun pattern ('analyze_video'), which is internally consistent. There are no other tools to deviate from the pattern.
The server has a narrow, focused purpose (analyzing YouTube videos with Gemini), so one tool is slightly under the typical 3-15 range but still appropriate and not excessive. It covers the core functionality without unnecessary fragmentation.
For the stated domain of video analysis, the tool is complete: it handles both open-ended analysis and specific questions, works with public videos, and reports token usage. There are no obvious missing operations or dead ends.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
SubDownload exposes YouTube as an MCP-native data source. Connect via OAuth and your AI agent can summarize videos, fetch full transcripts (even for videos with no captions, via AI ASR), search across channels, and save everything into a private knowledge base. Works with Claude, ChatGPT, Cursor, and 40+ MCP clients. Free credits on signup, no card required.
Related MCP Servers
- AlicenseBqualityFmaintenanceMCP (Model Context Protocol) server that utilizes the Google Gemini Vision API to interact with YouTube videos. It allows users to get descriptions, summaries, answers to questions, and extract key moments from YouTube videos.422 npm6MIT
- FlicenseAqualityDmaintenanceEnables analysis of YouTube videos using the Gemini API to generate summaries and answer specific questions via direct URLs. It supports standard videos and shorts, allowing users to interact with video content without requiring manual downloads.510-
- AlicenseAqualityDmaintenanceEnables YouTube content browsing, video searching, and metadata retrieval via the YouTube Data API v3. It also facilitates fetching video transcripts for summarization and analysis within MCP-compatible AI clients.720 npm1MIT
- AlicenseNot gradedqualityAmaintenanceGives MCP clients access to YouTube video transcripts and metadata. It lets AI agents read and summarize video content from a URL.8 npmMIT