yt-transcript-mcp
Fetches YouTube video transcripts as clean, token-efficient markdown with metadata, language selection, timestamp options, and local caching for repeated requests.
yt-transcript-mcp
YouTube transcripts as token-efficient AI context. One fetch, cached forever.
Agent-first: returns structured JSON by default. Zero dependencies on yt-dlp, ffmpeg, or API keys.
Works with Claude Desktop, ChatGPT Desktop, Cursor, Windsurf, and any MCP client.
Why?
When AI browses YouTube for a transcript, it processes the entire page: navigation, ads, recommendations, scripts. That's 75,000-150,000 tokens of noise to extract maybe 6,000 tokens of actual content.
This tool fetches only the transcript.
Tokens | Speed | Repeat queries | |
AI browses YouTube | 75-150k | 20-90s | Same cost every time |
ytfetch-mcp | 6-12k | 1-3s | Instant (cached) |
~50 KB per video in cache. A year of daily use stays under 120 MB.
Related MCP server: YouTube Transcript MCP
Demo
You say:
Fetch the transcript from https://www.youtube.com/watch?v=dQw4w9WgXcQ
Default response (compact JSON, segments only):
{
"is_error": false,
"video_id": "dQw4w9WgXcQ",
"title": "Never Gonna Give You Up",
"channel": "Rick Astley",
"published": "2009-10-25",
"language": "en",
"caption_type": "manual",
"segment_count": 56,
"transcript_duration_seconds": 213.5,
"content_hash": "a1b2c3...",
"cache_hit": false,
"warnings": [],
"segments": [
{"text": "We're no strangers to love", "start": 18.0, "end": 21.4},
{"text": "You know the rules and so do I", "start": 21.4, "end": 24.8}
]
}Structured, machine-readable, one transcript representation. No HTML, no noise, no wasted tokens.
Install
One line. No git clone needed.
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"yt-transcript": {
"command": "uvx",
"args": ["ytfetch-mcp"]
}
}
}Restart Claude Desktop. Done.
ChatGPT Desktop
Same config in your Codex MCP settings, or add manually:
Field | Value |
Command |
|
Arguments |
|
Tip: Find your uvx path with
which uvx. Restart the app after config changes.
Cursor / Windsurf / VS Code
Paste the same JSON block into your MCP server config.
Requires
uv (includes uvx): curl -LsSf https://astral.sh/uv/install.sh | sh
Parameters
Parameter | Description | Default |
| YouTube URL (required, any format) | |
| Language codes in priority order |
|
|
|
|
|
|
|
|
|
|
| Override title | |
| Override channel | |
| Override date (YYYY-MM-DD) | |
|
|
|
Output modes
| What you get |
| Array of |
| Single readable string (clean or timestamped) |
| Both representations |
Markdown format (format=markdown) always renders readable text regardless of output mode.
Error handling
Every error returns a structured response with a machine-readable code and a retryable flag so agents can branch automatically:
{
"is_error": true,
"error_code": "VIDEO_UNAVAILABLE",
"error_message": "Video is unavailable, private, or removed.",
"retryable": false,
"retry_count": 0
}Error code | Meaning | Retryable |
| Not a YouTube URL or malformed video ID | No |
| Transcripts disabled for this video | No |
| No transcript in requested languages | No |
| Video unavailable, private, age-restricted, or unplayable | No |
| YouTube is blocking your IP | No |
| Video requires Proof-of-Origin token | No |
| YouTube rate limit (429) | Yes |
Provenance
Every response includes provenance so you know exactly where the data comes from:
caption_type:manual,auto-generated, orunknownmetadata_sources: per-field tracking ({"title": "oembed", "published": "pytubefix"})content_hash: SHA256 of the segments array for reproducibilitywarnings:AUTO_GENERATED(speech recognition, may contain errors),LANGUAGE_FALLBACK(got a different language than requested),METADATA_FETCH_FAILED(some metadata unavailable)
Cache
Transcripts cached locally in ~/.cache/yt-transcript/. Keyed by video ID + language preference. Second fetch: instant, zero network.
Cache entries validated on load (version, types, segments, metadata)
Legacy or corrupted entries silently skipped
Cache write failures never block transcript delivery
CLI
Also works standalone, no MCP client needed:
uvx ytfetch-mcp # starts the MCP server
uv run yt_transcript.py https://youtu.be/ABC123 # CLI mode, saves .md fileCLI flags: --date, --title, --channel, --lang, --out, --no-clean, --no-cache.
Roadmap
Summary mode -- condensed output for lower token cost
Token budget (
max_tokens) -- fit any context windowBatch URLs -- multiple videos in one call
Chapter/topic filtering -- return only relevant sections
Remote HTTP transport -- expose as streamable HTTP MCP server
Schema.org metadata -- replace pytubefix for publish date
MCP outputSchema / structured content
Support
If this saves you time or tokens:
MIT \u00a9 Bj\u00f6rn Walther
Available Tools
1 toolfetch_transcriptA
Fetch a YouTube video transcript. Returns compact JSON (default) or markdown. Defaults to segments-only for token efficiency. Cached by language preference. Retries only transient errors. Includes provenance, caption type, and verifiable content hash.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video URL (watch, youtu.be, shorts, embed) | |
| title | No | Manual title override | |
| format | No | Response format. JSON is compact (no indent). | json |
| output | No | Transcript representation. 'segments' (default): array of {text,start,end}. 'text': readable string. 'both': both. Markdown always renders text. | segments |
| channel | No | Manual channel override | |
| languages | No | Comma-separated language codes in priority order (default: sv,en) | sv,en |
| published | No | Manual date (YYYY-MM-DD) | |
| bypass_cache | No | Force fresh fetch | |
| include_timestamps | No | Include HH:MM:SS per line in text output |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does substantial work: it reveals caching behavior by language preference, retry policy (only transient errors), default output representation, and that responses include provenance, caption type, and a verifiable content hash. The only notable gaps are error/edge-case behavior (e.g., missing transcript) and any rate-limit or auth constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six short sentences, each carrying distinct information: purpose, format options, default rationale, caching, retry behavior, and response contents. The core purpose is front-loaded and there is zero redundancy or filler — every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description is remarkably complete: it covers purpose, formats, default behaviors, caching, retry semantics, and key response elements. Remaining gaps — error handling for missing transcripts and any rate-limit/authentication context — are minor given how much behavioral ground is already covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, establishing a baseline of 3. The description adds genuine value beyond the schema by explaining the 'why' behind defaults — 'Defaults to segments-only for token efficiency' maps to the output parameter, and 'Cached by language preference' illuminates the interplay between languages and bypass_cache. This rationale helps the agent reason about parameter choices.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Fetch a YouTube video transcript' uses a specific verb and clearly identifiable resource. It is further enriched by specifying return formats (compact JSON or markdown), default output mode, and content guarantees (provenance, caption type, content hash), leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives implied usage context through 'Defaults to segments-only for token efficiency' and 'Cached by language preference,' which signal this is an efficient default path for transcript retrieval. However, there is no explicit when-to-use guidance, no exclusions, and no alternative tools to route toward — though this is partly mitigated by there being no sibling tools listed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v1.2.0- First observed
fetch_transcript
TDQS
With only one tool, there is no possibility of confusing it with another. The tool's purpose is narrowly and clearly defined.
The single tool name 'fetch_transcript' follows a clear verb_noun convention. There are no other names to contradict this pattern.
One tool is borderline for a server, but it is reasonable for a narrowly focused YouTube transcript service. It feels thin if broader video or transcript management is expected, though the scope appears intentionally minimal.
The tool covers the core need of fetching a transcript with useful options like markdown output, caching, and provenance. Minor gaps exist around listing available transcripts or explicitly choosing languages, but these are workable and not fatal for the stated purpose.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents. No signup.
Provide token-optimized, structured YouTube data to enhance your LLM applications. Access efficien…
Transcribe YouTube via Whisper. Summaries, chapters, semantic-search across your corpus.
Any video URL to LLM-ready transcript. ASR built in, no captions needed. TikTok, X, TED and more.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceEnables AI assistants to fetch and process YouTube video transcripts in multiple formats and languages, with built-in caching and rate limiting for efficient video content analysis.-
- AlicenseNot gradedqualityCmaintenanceEnables AI models to extract transcripts from YouTube videos in multiple languages with zero local setup. It supports all YouTube URL formats and features smart caching via Cloudflare Workers for fast responses.461MIT
- AlicenseAqualityBmaintenanceExtract YouTube transcripts for AI agents, RAG pipelines, and LLM workflows. Supports any YouTube URL. Returns clean text or timestamped segments. No API keys required.14MIT
- AlicenseAqualityBmaintenanceFast, minimal YouTube 'watch' engine for AI agents that extracts clean transcripts, searches, and slices any YouTube video without requiring an API key.4MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/bjornwalther/yt-transcript'
If you have feedback or need assistance with the MCP directory API, please join our Discord server