youtube-transcript-mcp
Provides tools to fetch YouTube video transcripts from any YouTube URL, list available caption languages, retrieve video metadata, and parse YouTube URLs, with support for formats like text, markdown, SRT, and VTT.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@youtube-transcript-mcpget the transcript of https://youtu.be/dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
youtube-transcript-mcp
An MCP server that returns the transcript of a YouTube video from any YouTube URL — no API key, no browser, no OAuth. It knows that a meaningful share of videos have no transcript at all, and says so precisely instead of failing with a vague error or returning empty text.
Any reference shape. watch URLs,
youtu.be, Shorts, live, embeds, music, nocookie, mobile, legacy/v/, attribution links, oEmbed, URL-encoded links, Markdown links, prose around a link, or a bare 11-character id.Honest failure.
NO_TRANSCRIPT,TRANSCRIPT_GENERATING,PRIVATE_VIDEO,AGE_RESTRICTED,MEMBERS_ONLY,LIVE_ENDED,BOT_CHECK,RATE_LIMITED,LANGUAGE_UNAVAILABLE… each with a hint and, where useful, the list of languages that do exist.Resilient fetching. Several InnerTube clients, the watch page, the transcript panel endpoint and an optional local
yt-dlpfallback, tried in order, because a single endpoint gets you empty bodies or 429s depending on the IP you come from.Language aware. Preference lists, manual-vs-auto-generated awareness, and optional machine translation into a language the video does not offer.
Install
git clone https://github.com/0xCaFeBEef/free-youtube-transcript-mcp.git youtube-transcript-mcp
cd youtube-transcript-mcp
pnpm install # or: npm install
pnpm buildRequires Node.js 20+ (uses global fetch).
Wire it into an MCP client
Claude Desktop, DSH, Cursor, or anything else that speaks MCP over stdio:
{
"mcpServers": {
"youtube-transcript": {
"command": "node",
"args": ["/absolute/path/to/youtube-transcript-mcp/dist/index.js"],
"env": { "YTA_YTDLP": "auto" }
}
}
}Or keep the source checked out and let the client run it directly with tsx:
{
"mcpServers": {
"youtube-transcript": {
"command": "npx",
"args": ["tsx", "/absolute/path/to/youtube-transcript-mcp/src/index.ts"]
}
}
}Try it from a terminal first
node dist/index.js --probe "https://youtu.be/dQw4w9WgXcQ"
node dist/index.js --probe "<url>" --lang hi --format markdown --timestamps
node dist/index.js --list-languages "<url>"
node dist/index.js --info "<url>"
node dist/index.js --helpRelated MCP server: YouTube Transcript MCP
Tools
Tool | What it does |
| The transcript. |
| Every caption track: language, display name, whether it is auto-generated, plus YouTube's translation targets. |
| Title, author, duration, and whether captions exist at all. Cheap pre-flight for batch work. |
| Offline parse: video id, playlist id, start offset, how it was recognised. Costs no request. |
output is one of text (default), markdown (timestamped bullets with deep links),
segments (JSON with millisecond timings), srt or vtt.
URLs that work
All of these resolve to the same video:
https://www.youtube.com/watch?v=VIDEO_ID
https://www.youtube.com/watch?v=VIDEO_ID&list=RDVIDEO_ID&index=3&t=42s&si=abc123
https://www.youtube.com/watch?feature=share&v=VIDEO_ID
https://www.youtube.com/watch/VIDEO_ID
https://www.youtube.com/v/VIDEO_ID /vi/VIDEO_ID /e/VIDEO_ID
https://www.youtube.com/embed/VIDEO_ID (also youtube-nocookie.com)
https://www.youtube.com/shorts/VIDEO_ID
https://www.youtube.com/live/VIDEO_ID
https://youtu.be/VIDEO_ID?t=1m30s
https://m.youtube.com/watch?v=VIDEO_ID
https://music.youtube.com/watch?v=VIDEO_ID
https://www.youtube.co.uk/watch?v=VIDEO_ID
https://www.youtube.com/attribution_link?a=xyz&v=VIDEO_ID
https://www.youtube.com/oembed?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DVIDEO_ID
https://www.youtube.com/watch%3Fv%3DVIDEO_ID (double-encoded)
VIDEO_ID (bare id)
<https://youtu.be/VIDEO_ID> [title](https://youtu.be/VIDEO_ID) here: https://youtu.be/VIDEO_ID!Timestamps in t, start, start_s or #t= (accepting 42, 42s, 1m30s,
1h2m3s, 01:30, 1:02:03) are parsed and reported, but a transcript is always for
the whole video — a start offset does not trim it.
Playlists, channels, @handles, search pages, feeds, community posts and /clip/
links are not one video, so they are rejected with an explanation rather than a
guess. A list= parameter travelling with a video URL is noted and ignored.
When there is no transcript
This is a normal outcome, not an error condition to retry.
Code | Meaning | Retry? |
| No captions exist: none uploaded, none generated. Or only auto-captions exist and you asked not to include them. | no |
| YouTube has queued auto-captions but not produced them yet (common hours after upload). | later |
| Bad id, deleted, or never public. | no |
| Requires the owner's or a member's session. | with cookies |
| Blocked for this session or region. | with cookies / proxy |
| Recording gone, or a scheduled premiere. 24/7 streams carry no captions. | no |
| Captions exist, but not in the language you asked for. | different language |
| YouTube is throttling this IP (HTTP 403/429). Back off; cookies help. | yes, later |
| Every fetch path was refused. | yes |
Failures come back as isError: true with a JSON body:
{
"ok": false,
"error": "None of the requested languages (xx) is available for this video.",
"code": "LANGUAGE_UNAVAILABLE",
"retryable": false,
"hint": "Call youtube_list_languages to see what exists, or request one of those.",
"details": {
"requested": ["xx"],
"available": [
{ "language": "en", "name": "English", "generated": false },
{ "language": "en", "name": "English (auto-generated)", "generated": true }
]
}
}Languages
languages is a priority list: ["hi", "en"] prefers Hindi and falls back to English,
noting the substitution in notes. Matching is exact tag first, then base language
(en matches en-US), and human-written tracks always beat auto-generated ones at
the same distance. includeAutoCaptions: false refuses speech-recognised tracks
outright — useful when wording accuracy matters.
translateTo asks YouTube to machine-translate. It works even when the video has no
track in the target language: the tool picks the original-language track and
translates that, and says so in notes.
Configuration
All optional, all environment variables.
Variable | Default | Purpose |
|
| Interface language and region sent to YouTube; affects track display names. |
| – | Raw |
| – | Netscape cookie file (the format browser extensions and yt-dlp emit); parsed for HTTP and passed to yt-dlp. |
|
| Timeout per HTTP request. |
|
| Retries for transport-level failures, with backoff. |
|
| Transcript cap per call; truncation is reported, not hidden. |
|
| In-memory transcript cache lifetime. |
|
|
|
|
| Path to the binary. |
|
| Override the provider order. |
| desktop Chrome UA | Sent on every request. |
How it fetches, and why it sometimes cannot
InnerTube
playerAPI (Android, then iOS, MWEB, TV-embedded, web). The mobile clients still return ordinary caption URLs.The watch page, whose
ytInitialPlayerResponsealso lists tracks — but those URLs carry aexp=xpo,xpeproof-of-origin experiment flag that makes/api/timedtextanswer with an empty body. The flag is stripped before use.The
get_transcriptpanel endpoint, using the pre-encoded params the watch page embeds.Local
yt-dlp, if installed, as the most refusal-resistant path.
The first provider that lists tracks wins; if the download then fails, the providers
that were skipped are consulted for another copy of the same track. Payloads are
requested as json3, then srv3, then vtt, and parsed by sniffing when the answer
arrives in a different format than asked.
Practical consequences worth knowing:
Datacenter and shared IPs get throttled. Empty 200-body caption responses and HTTP 429 are both anti-abuse responses, not bugs in the video. The tool backs off instead of retrying harder, and reports
RATE_LIMITED/BOT_CHECK.Cookies unlock access-limited videos (
YTA_COOKIES/YTA_COOKIES_FILE).Auto-generated captions roll: each cue can repeat the previous line. Repeats are folded before you see them.
Music videos often carry placeholder captions (
[♪♪♪]); those are real cues, kept.
Development
pnpm test # 40+ unit tests: URL parsing, wire formats, selection, rendering
pnpm test:live # opt-in tests that hit YouTube (RUN_LIVE=1)
pnpm typecheckNotes
Unofficial API. Scraping transcripts can conflict with YouTube's Terms of Service; use it considerately, cache responses, and do not hammer the endpoints. Captions are the uploader's content — respect whatever licence applies to the video.
Available Tools
4 toolsyoutube_get_transcriptGet YouTube transcriptARead-onlyIdempotent
Return the transcript of a YouTube video, from any YouTube URL. Handles watch / youtu.be / shorts / live / embed / music / nocookie links and bare video ids. If the video has no captions it fails with code NO_TRANSCRIPT instead of inventing text; if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list. Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id. | |
| output | No | Output shape: text (default, readable lines), markdown (timestamped bullets), segments (JSON with millisecond timings), or srt / vtt subtitle files. | |
| maxChars | No | Truncate the transcript to roughly this many characters. | |
| languages | No | Language preferences in priority order, e.g. ["en"] or ["hi", "en"]. Defaults to ["en"]. Falls back to the closest available track and says so in notes. | |
| timestamps | No | Prefix each line of text output with its [mm:ss] offset. | |
| translateTo | No | Ask YouTube to machine-translate the selected captions into this language (for example es). Only works when the video offers translations. | |
| includeAutoCaptions | No | Include YouTube auto-generated (speech-recognised) captions. Default true; set false to require human-written captions only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the description's job is to add behavioral context. It does detail failure codes and URL handling, which is helpful. However, there is an internal contradiction: the description says 'if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list' while the languages parameter description says 'Falls back to the closest available track and says so in notes.' This inconsistency undermines trust and clarity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise for the amount of information it conveys. It front-loads the primary purpose and then covers failure modes and usage tips in a structured manner. It is not verbose, though it could be tightened by removing the contradictory language note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (7 parameters) and the lack of an output schema, the description covers URL formats, output options, failure modes, and truncation behavior (via maxChars in schema). However, the language fallback/failure contradiction leaves a gap that could confuse an agent. Overall, it is fairly complete but not perfect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed descriptions for all 7 parameters, so the baseline is 3. The description adds a small amount of guidance (e.g., explaining the difference between output=segments and timestamps=true) but does not significantly enhance parameter understanding beyond what the schema already provides. It also repeats the contradictory language behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Return the transcript of a YouTube video, from any YouTube URL.' It enumerates the URL formats it accepts and explicitly contrasts failure modes (NO_TRANSCRIPT, LANGUAGE_UNAVAILABLE). This unambiguously differentiates it from sibling tools like youtube_parse_url or youtube_video_info, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage guidance for output selection ('Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines') but does not explicitly state when to prefer this tool over its siblings. While the purpose is clear, there is no direct 'use this instead of X when...' guidance, though the context makes it obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_list_languagesList transcript languagesARead-onlyIdempotent
List every caption track a video offers, marking which are auto-generated. Use it after LANGUAGE_UNAVAILABLE, or to decide what youtube_get_transcript can return. hasCaptions=false means the video has no transcript at all.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds useful output semantics: it flags auto-generated tracks and explains that hasCaptions=false means no transcript exists at all, which is beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core function, the usage trigger, and a key result field meaning. The most important action is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter, read-only tool with strong annotations and a fully documented schema, the description covers what the tool returns and when to invoke it. No critical information is missing for an agent to select and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single url parameter at 100%, including a thorough list of accepted URL formats. The description adds no additional parameter-level meaning, so the schema carries the burden and the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "List every caption track a video offers," and clarifies the auto-generated marking. It differentiates from youtube_get_transcript by positioning this as the survey of available tracks rather than the transcript content itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger conditions: use after LANGUAGE_UNAVAILABLE or to decide what youtube_get_transcript can return. It names the relevant sibling tool and gives the agent a clear routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_parse_urlParse a YouTube URLARead-onlyIdempotent
Offline parse of a YouTube reference: video id, playlist id, start offset and how it was recognised. Makes no network requests. Useful for validating a link and explaining playlist, channel or clip links.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Any YouTube URL, or a bare 11-character video id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds valuable behavioral detail: it is offline and 'Makes no network requests.' It also reveals what kind of recognition output the tool produces, which complements the safe-read/idempotent hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler. The core capability and offline guarantee are front-loaded, and the use-case sentence earns its place by clarifying when the tool is helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, offline, read-only parser with full schema coverage and helpful annotations, the description is complete. It explains what the tool parses, how it behaves, and why an agent would choose it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single url parameter well. The description adds context about valid link forms (playlist, channel, clip) but does not substantially extend the parameter meaning beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Offline parse of a YouTube reference.' It clearly enumerates the outputs (video id, playlist id, start offset, recognition method) which distinguishes it from transcript, language, and video-info siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the intended use: 'validating a link and explaining playlist, channel or clip links.' This gives clear context for when to call the tool, though it does not explicitly say when not to use it or name alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_video_infoInspect video and caption availabilityARead-onlyIdempotent
Cheap report for one video: title, author, duration, whether captions exist and in which languages. Call it before a batch of transcript requests, or to explain why a transcript is missing.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds 'cheap' (performance expectation) and specifies the exact data returned, plus its diagnostic role. This adds context beyond annotations without contradiction, though it omits details like pagination or rate limits, which are not critical for a read-only info tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and output, then the usage context. No wasted words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only info tool with full schema coverage and annotations covering safety, the description is complete. It tells the agent what it returns, when to use it, and why, leaving nothing essential missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% – the url parameter is thoroughly described in the schema, listing all supported URL formats. The description adds no additional parameter information, so the baseline 3 applies. It does not compensate for any schema gap because there is none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Cheap report for one video') with concrete output fields (title, author, duration, caption languages), clearly distinguishing it from siblings like youtube_get_transcript and youtube_parse_url. The purpose is unambiguous and the tool's role in the transcript workflow is explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Call it before a batch of transcript requests, or to explain why a transcript is missing.' This tells the agent when to use it, though it does not explicitly name the alternative tools or state when not to use it (e.g., 'if you need the transcript text, use youtube_get_transcript'). The guidance is sufficient but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
youtube_get_transcript - First observed
youtube_list_languages - First observed
youtube_parse_url - First observed
youtube_video_info
TDQS
Scored across 4 tools
Each tool targets a distinct concern: parsing URLs, fetching transcripts, listing caption languages, and summarizing video info. Although list_languages and video_info both mention captions, their purposes are clearly separated (exploring vs. explaining/planning).
All tool names follow the consistent pattern youtube_<verb>_<object>: parse_url, get_transcript, list_languages, video_info. The style is uniform and predictable across the set.
Four tools is a well-scoped size for a focused YouTube transcript server. Each tool serves a clear function without redundancy or bloat, covering the primary workflow.
The surface covers the core transcript workflow completely: parse/validate references, fetch transcripts, list available languages, and get video-level metadata to explain missing captions. There are no dead ends for the stated purpose.
Maintenance
Related MCP Connectors
Fetch the full transcript of any YouTube video as clean text. No API key, no signup.
Transcripts of YouTube videos, playlists and channels with timestamps; SRT or WebVTT too.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Free YouTube transcripts, no API key: videos, channel lists, latest uploads, bulk download links.
Related MCP Servers
- AlicenseAqualityDmaintenanceRetrieves transcripts from YouTube videos with support for multiple languages, timestamp control, and language detection. Enables video content analysis, summarization, and quote extraction without manually downloading or watching videos.272 npm15MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI models to extract transcripts from YouTube videos in multiple languages with zero local setup. It supports all YouTube URL formats and features smart caching via Cloudflare Workers for fast responses.100 npm1MIT
- FlicenseAqualityCmaintenanceEnables fetching YouTube video transcripts with metadata, including timed captions in multiple formats (JSON, SRT, VTT, CSV, TXT) and preprocessing options.41-
- AlicenseNot gradedqualityDmaintenanceEnables retrieval of transcripts from YouTube videos, supporting multiple URL formats, language selection, timestamps, and ad filtering.1,050 npmMIT