Skip to main content
Glama

youtube-transcript-mcp

An MCP server that returns the transcript of a YouTube video from any YouTube URL — no API key, no browser, no OAuth. It knows that a meaningful share of videos have no transcript at all, and says so precisely instead of failing with a vague error or returning empty text.

  • Any reference shape. watch URLs, youtu.be, Shorts, live, embeds, music, nocookie, mobile, legacy /v/, attribution links, oEmbed, URL-encoded links, Markdown links, prose around a link, or a bare 11-character id.

  • Honest failure. NO_TRANSCRIPT, TRANSCRIPT_GENERATING, PRIVATE_VIDEO, AGE_RESTRICTED, MEMBERS_ONLY, LIVE_ENDED, BOT_CHECK, RATE_LIMITED, LANGUAGE_UNAVAILABLE … each with a hint and, where useful, the list of languages that do exist.

  • Resilient fetching. Several InnerTube clients, the watch page, the transcript panel endpoint and an optional local yt-dlp fallback, tried in order, because a single endpoint gets you empty bodies or 429s depending on the IP you come from.

  • Language aware. Preference lists, manual-vs-auto-generated awareness, and optional machine translation into a language the video does not offer.

Install

git clone https://github.com/0xCaFeBEef/free-youtube-transcript-mcp.git youtube-transcript-mcp
cd youtube-transcript-mcp
pnpm install        # or: npm install
pnpm build

Requires Node.js 20+ (uses global fetch).

Wire it into an MCP client

Claude Desktop, DSH, Cursor, or anything else that speaks MCP over stdio:

{
  "mcpServers": {
    "youtube-transcript": {
      "command": "node",
      "args": ["/absolute/path/to/youtube-transcript-mcp/dist/index.js"],
      "env": { "YTA_YTDLP": "auto" }
    }
  }
}

Or keep the source checked out and let the client run it directly with tsx:

{
  "mcpServers": {
    "youtube-transcript": {
      "command": "npx",
      "args": ["tsx", "/absolute/path/to/youtube-transcript-mcp/src/index.ts"]
    }
  }
}

Try it from a terminal first

node dist/index.js --probe "https://youtu.be/dQw4w9WgXcQ"
node dist/index.js --probe "<url>" --lang hi --format markdown --timestamps
node dist/index.js --list-languages "<url>"
node dist/index.js --info "<url>"
node dist/index.js --help

Related MCP server: YouTube Transcript MCP

Tools

Tool

What it does

youtube_get_transcript

The transcript. url, plus languages, includeAutoCaptions, output, timestamps, maxChars, translateTo.

youtube_list_languages

Every caption track: language, display name, whether it is auto-generated, plus YouTube's translation targets.

youtube_video_info

Title, author, duration, and whether captions exist at all. Cheap pre-flight for batch work.

youtube_parse_url

Offline parse: video id, playlist id, start offset, how it was recognised. Costs no request.

output is one of text (default), markdown (timestamped bullets with deep links), segments (JSON with millisecond timings), srt or vtt.

URLs that work

All of these resolve to the same video:

https://www.youtube.com/watch?v=VIDEO_ID
https://www.youtube.com/watch?v=VIDEO_ID&list=RDVIDEO_ID&index=3&t=42s&si=abc123
https://www.youtube.com/watch?feature=share&v=VIDEO_ID
https://www.youtube.com/watch/VIDEO_ID
https://www.youtube.com/v/VIDEO_ID      /vi/VIDEO_ID      /e/VIDEO_ID
https://www.youtube.com/embed/VIDEO_ID  (also youtube-nocookie.com)
https://www.youtube.com/shorts/VIDEO_ID
https://www.youtube.com/live/VIDEO_ID
https://youtu.be/VIDEO_ID?t=1m30s
https://m.youtube.com/watch?v=VIDEO_ID
https://music.youtube.com/watch?v=VIDEO_ID
https://www.youtube.co.uk/watch?v=VIDEO_ID
https://www.youtube.com/attribution_link?a=xyz&v=VIDEO_ID
https://www.youtube.com/oembed?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DVIDEO_ID
https://www.youtube.com/watch%3Fv%3DVIDEO_ID          (double-encoded)
VIDEO_ID                                              (bare id)
<https://youtu.be/VIDEO_ID>   [title](https://youtu.be/VIDEO_ID)   here: https://youtu.be/VIDEO_ID!

Timestamps in t, start, start_s or #t= (accepting 42, 42s, 1m30s, 1h2m3s, 01:30, 1:02:03) are parsed and reported, but a transcript is always for the whole video — a start offset does not trim it.

Playlists, channels, @handles, search pages, feeds, community posts and /clip/ links are not one video, so they are rejected with an explanation rather than a guess. A list= parameter travelling with a video URL is noted and ignored.

When there is no transcript

This is a normal outcome, not an error condition to retry.

Code

Meaning

Retry?

NO_TRANSCRIPT

No captions exist: none uploaded, none generated. Or only auto-captions exist and you asked not to include them.

no

TRANSCRIPT_GENERATING

YouTube has queued auto-captions but not produced them yet (common hours after upload).

later

VIDEO_NOT_FOUND

Bad id, deleted, or never public.

no

PRIVATE_VIDEO / MEMBERS_ONLY

Requires the owner's or a member's session.

with cookies

AGE_RESTRICTED / REGION_RESTRICTED

Blocked for this session or region.

with cookies / proxy

LIVE_ENDED / LIVE_NOT_STARTED

Recording gone, or a scheduled premiere. 24/7 streams carry no captions.

no

LANGUAGE_UNAVAILABLE

Captions exist, but not in the language you asked for. details.available lists what does.

different language

BOT_CHECK / RATE_LIMITED

YouTube is throttling this IP (HTTP 403/429). Back off; cookies help.

yes, later

PROVIDER_FAILED

Every fetch path was refused. attempts shows what each one said.

yes

Failures come back as isError: true with a JSON body:

{
  "ok": false,
  "error": "None of the requested languages (xx) is available for this video.",
  "code": "LANGUAGE_UNAVAILABLE",
  "retryable": false,
  "hint": "Call youtube_list_languages to see what exists, or request one of those.",
  "details": {
    "requested": ["xx"],
    "available": [
      { "language": "en", "name": "English", "generated": false },
      { "language": "en", "name": "English (auto-generated)", "generated": true }
    ]
  }
}

Languages

languages is a priority list: ["hi", "en"] prefers Hindi and falls back to English, noting the substitution in notes. Matching is exact tag first, then base language (en matches en-US), and human-written tracks always beat auto-generated ones at the same distance. includeAutoCaptions: false refuses speech-recognised tracks outright — useful when wording accuracy matters.

translateTo asks YouTube to machine-translate. It works even when the video has no track in the target language: the tool picks the original-language track and translates that, and says so in notes.

Configuration

All optional, all environment variables.

Variable

Default

Purpose

YTA_HL / YTA_GL

en / US

Interface language and region sent to YouTube; affects track display names.

YTA_COOKIES

–

Raw Cookie header, for age-restricted or your own private videos.

YTA_COOKIES_FILE

–

Netscape cookie file (the format browser extensions and yt-dlp emit); parsed for HTTP and passed to yt-dlp.

YTA_TIMEOUT_MS

20000

Timeout per HTTP request.

YTA_RETRIES

2

Retries for transport-level failures, with backoff.

YTA_MAX_CHARS

200000

Transcript cap per call; truncation is reported, not hidden.

YTA_CACHE_TTL_MS

300000

In-memory transcript cache lifetime.

YTA_YTDLP

auto

auto uses a local yt-dlp as fallback, always puts it first, never disables it.

YTA_YTDLP_PATH

yt-dlp

Path to the binary.

YTA_PROVIDERS

innertube,watch-page,yt-dlp

Override the provider order.

YTA_USER_AGENT

desktop Chrome UA

Sent on every request.

How it fetches, and why it sometimes cannot

  1. InnerTube player API (Android, then iOS, MWEB, TV-embedded, web). The mobile clients still return ordinary caption URLs.

  2. The watch page, whose ytInitialPlayerResponse also lists tracks — but those URLs carry a exp=xpo,xpe proof-of-origin experiment flag that makes /api/timedtext answer with an empty body. The flag is stripped before use.

  3. The get_transcript panel endpoint, using the pre-encoded params the watch page embeds.

  4. Local yt-dlp, if installed, as the most refusal-resistant path.

The first provider that lists tracks wins; if the download then fails, the providers that were skipped are consulted for another copy of the same track. Payloads are requested as json3, then srv3, then vtt, and parsed by sniffing when the answer arrives in a different format than asked.

Practical consequences worth knowing:

  • Datacenter and shared IPs get throttled. Empty 200-body caption responses and HTTP 429 are both anti-abuse responses, not bugs in the video. The tool backs off instead of retrying harder, and reports RATE_LIMITED / BOT_CHECK.

  • Cookies unlock access-limited videos (YTA_COOKIES / YTA_COOKIES_FILE).

  • Auto-generated captions roll: each cue can repeat the previous line. Repeats are folded before you see them.

  • Music videos often carry placeholder captions ([♪♪♪]); those are real cues, kept.

Development

pnpm test          # 40+ unit tests: URL parsing, wire formats, selection, rendering
pnpm test:live     # opt-in tests that hit YouTube (RUN_LIVE=1)
pnpm typecheck

Notes

Unofficial API. Scraping transcripts can conflict with YouTube's Terms of Service; use it considerately, cache responses, and do not hammer the endpoints. Captions are the uploader's content — respect whatever licence applies to the video.

Available Tools

4 tools
youtube_get_transcriptGet YouTube transcriptA
Read-onlyIdempotent

Return the transcript of a YouTube video, from any YouTube URL. Handles watch / youtu.be / shorts / live / embed / music / nocookie links and bare video ids. If the video has no captions it fails with code NO_TRANSCRIPT instead of inventing text; if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list. Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id.
outputNoOutput shape: text (default, readable lines), markdown (timestamped bullets), segments (JSON with millisecond timings), or srt / vtt subtitle files.
maxCharsNoTruncate the transcript to roughly this many characters.
languagesNoLanguage preferences in priority order, e.g. ["en"] or ["hi", "en"]. Defaults to ["en"]. Falls back to the closest available track and says so in notes.
timestampsNoPrefix each line of text output with its [mm:ss] offset.
translateToNoAsk YouTube to machine-translate the selected captions into this language (for example es). Only works when the video offers translations.
includeAutoCaptionsNoInclude YouTube auto-generated (speech-recognised) captions. Default true; set false to require human-written captions only.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the description's job is to add behavioral context. It does detail failure codes and URL handling, which is helpful. However, there is an internal contradiction: the description says 'if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list' while the languages parameter description says 'Falls back to the closest available track and says so in notes.' This inconsistency undermines trust and clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise for the amount of information it conveys. It front-loads the primary purpose and then covers failure modes and usage tips in a structured manner. It is not verbose, though it could be tightened by removing the contradictory language note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 parameters) and the lack of an output schema, the description covers URL formats, output options, failure modes, and truncation behavior (via maxChars in schema). However, the language fallback/failure contradiction leaves a gap that could confuse an agent. Overall, it is fairly complete but not perfect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed descriptions for all 7 parameters, so the baseline is 3. The description adds a small amount of guidance (e.g., explaining the difference between output=segments and timestamps=true) but does not significantly enhance parameter understanding beyond what the schema already provides. It also repeats the contradictory language behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Return the transcript of a YouTube video, from any YouTube URL.' It enumerates the URL formats it accepts and explicitly contrasts failure modes (NO_TRANSCRIPT, LANGUAGE_UNAVAILABLE). This unambiguously differentiates it from sibling tools like youtube_parse_url or youtube_video_info, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage guidance for output selection ('Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines') but does not explicitly state when to prefer this tool over its siblings. While the purpose is clear, there is no direct 'use this instead of X when...' guidance, though the context makes it obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

youtube_list_languagesList transcript languagesA
Read-onlyIdempotent

List every caption track a video offers, marking which are auto-generated. Use it after LANGUAGE_UNAVAILABLE, or to decide what youtube_get_transcript can return. hasCaptions=false means the video has no transcript at all.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds useful output semantics: it flags auto-generated tracks and explains that hasCaptions=false means no transcript exists at all, which is beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core function, the usage trigger, and a key result field meaning. The most important action is front-loaded and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter, read-only tool with strong annotations and a fully documented schema, the description covers what the tool returns and when to invoke it. No critical information is missing for an agent to select and call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single url parameter at 100%, including a thorough list of accepted URL formats. The description adds no additional parameter-level meaning, so the schema carries the burden and the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: "List every caption track a video offers," and clarifies the auto-generated marking. It differentiates from youtube_get_transcript by positioning this as the survey of available tracks rather than the transcript content itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger conditions: use after LANGUAGE_UNAVAILABLE or to decide what youtube_get_transcript can return. It names the relevant sibling tool and gives the agent a clear routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

youtube_parse_urlParse a YouTube URLA
Read-onlyIdempotent

Offline parse of a YouTube reference: video id, playlist id, start offset and how it was recognised. Makes no network requests. Useful for validating a link and explaining playlist, channel or clip links.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAny YouTube URL, or a bare 11-character video id.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds valuable behavioral detail: it is offline and 'Makes no network requests.' It also reveals what kind of recognition output the tool produces, which complements the safe-read/idempotent hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler. The core capability and offline guarantee are front-loaded, and the use-case sentence earns its place by clarifying when the tool is helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, offline, read-only parser with full schema coverage and helpful annotations, the description is complete. It explains what the tool parses, how it behaves, and why an agent would choose it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single url parameter well. The description adds context about valid link forms (playlist, channel, clip) but does not substantially extend the parameter meaning beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb and resource: 'Offline parse of a YouTube reference.' It clearly enumerates the outputs (video id, playlist id, start offset, recognition method) which distinguishes it from transcript, language, and video-info siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the intended use: 'validating a link and explaining playlist, channel or clip links.' This gives clear context for when to call the tool, though it does not explicitly say when not to use it or name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

youtube_video_infoInspect video and caption availabilityA
Read-onlyIdempotent

Cheap report for one video: title, author, duration, whether captions exist and in which languages. Call it before a batch of transcript requests, or to explain why a transcript is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds 'cheap' (performance expectation) and specifies the exact data returned, plus its diagnostic role. This adds context beyond annotations without contradiction, though it omits details like pagination or rate limits, which are not critical for a read-only info tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and output, then the usage context. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only info tool with full schema coverage and annotations covering safety, the description is complete. It tells the agent what it returns, when to use it, and why, leaving nothing essential missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% – the url parameter is thoroughly described in the schema, listing all supported URL formats. The description adds no additional parameter information, so the baseline 3 applies. It does not compensate for any schema gap because there is none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Cheap report for one video') with concrete output fields (title, author, duration, caption languages), clearly distinguishing it from siblings like youtube_get_transcript and youtube_parse_url. The purpose is unambiguous and the tool's role in the transcript workflow is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'Call it before a batch of transcript requests, or to explain why a transcript is missing.' This tells the agent when to use it, though it does not explicitly name the alternative tools or state when not to use it (e.g., 'if you need the transcript text, use youtube_get_transcript'). The guidance is sufficient but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedyoutube_get_transcript
    • First observedyoutube_list_languages
    • First observedyoutube_parse_url
    • First observedyoutube_video_info

TDQS

A4.4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct concern: parsing URLs, fetching transcripts, listing caption languages, and summarizing video info. Although list_languages and video_info both mention captions, their purposes are clearly separated (exploring vs. explaining/planning).

Naming Consistency5/5

All tool names follow the consistent pattern youtube_<verb>_<object>: parse_url, get_transcript, list_languages, video_info. The style is uniform and predictable across the set.

Tool Count5/5

Four tools is a well-scoped size for a focused YouTube transcript server. Each tool serves a clear function without redundancy or bloat, covering the primary workflow.

Completeness5/5

The surface covers the core transcript workflow completely: parse/validate references, fetch transcripts, list available languages, and get video-level metadata to explain missing captions. There are no dead ends for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers