Skip to main content
Glama
JADRT22
by JADRT22

yt-transcript

Give your AI agent eyes for YouTube. A lightweight MCP server + CLI that fetches the transcript, description, and comments of any YouTube video — so your agent can read, summarize, and analyze video content.

Works with opencode, Claude Code, Freebuff, Cursor, and any MCP-compatible client — or as a plain CLI any agent can call.

Why this one?

YouTube aggressively blocks scripted access ("confirm you're not a bot", HTTP 429). Most transcript tools break the first time YouTube pushes back. This one ships a 4-layer fallback cascade baked in:

Layer

What it does

1. youtube-transcript-api

Fast path, no API key needed

2. yt-dlp web client

Baseline fallback

3. yt-dlp android client

Bypasses the "not a bot" check on most videos

4. yt-dlp + browser cookies

Cures hard IP blocks / 429s (uses your Firefox/Chrome cookies)

Plus:

  • Disk cache (~/.cache/yt-transcript/) — second call for the same video is ~10x faster and consumes zero rate limit

  • Description + comments — likes, authors, replies; great context for richer summaries

  • Language aware — pick preferred languages; lists everything available per video

  • Zero config, zero API keys

Related MCP server: VidLens

Install

git clone https://github.com/JADRT22/yt-transcript.git
cd yt-transcript
python3 -m venv .venv
.venv/bin/pip install -e .

MCP setup

Point your client at the server binary:

/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp

In ~/.config/opencode/opencode.json (global) or opencode.json (project):

{
  "mcp": {
    "yt-transcript": {
      "type": "local",
      "command": ["/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"],
      "enabled": true
    }
  }
}
claude mcp add yt-transcript -- /absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp

Create a mcp.json in the project root where the agent runs:

{
  "mcpServers": {
    "yt-transcript": {
      "command": "/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"
    }
  }
}

Any agent that runs shell commands can read the output directly:

/absolute/path/to/yt-transcript/.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"

MCP tools

Tool

Description

get_transcript_tool(url, languages?)

Full transcript of the video

list_transcripts_tool(url)

Available captions (languages, auto-generated?)

get_video_info_tool(url)

Title, channel, duration, upload date

get_video_details_tool(url, max_comments?)

Full description + comments (author, text, likes)

Example prompts once connected:

Summarize this video: https://www.youtube.com/watch?v=...

What are people saying in the comments of this video? [URL]

Compare what the video claims with what the description promises: [URL]

CLI usage

# Full transcript (default languages: en, es, pt)
.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"

# Metadata only
.venv/bin/yt-transcript VIDEO_ID --info

# Description + comments (JSON)
.venv/bin/yt-transcript VIDEO_ID --desc --max-comments 50

# Available captions
.venv/bin/yt-transcript VIDEO_ID --list

# JSON output / other languages
.venv/bin/yt-transcript VIDEO_ID --json
.venv/bin/yt-transcript VIDEO_ID -l de,fr

Accepts youtube.com/watch?v=..., youtu.be/..., youtube.com/shorts/..., or the bare 11-character ID.

Configuration

Environment variable

Default

Purpose

YT_TRANSCRIPT_COOKIES_BROWSER

firefox

Browser for the cookie fallback (chrome, chromium, brave, ...; none disables)

YT_TRANSCRIPT_CACHE

~/.cache/yt-transcript

Cache directory (none disables)

To force a fresh fetch for one video, delete its file (SHA-256-named) from the cache directory.

How it works

  1. extract_video_id normalizes any YouTube URL format into a video ID.

  2. Transcripts: captions API first; on any failure, yt-dlp downloads the .vtt through the client cascade and converts it to plain text (each attempt only counts if the caption file is actually written — yt-dlp's simulation mode silently skips it, and --no-simulate guards that).

  3. Metadata/details: yt-dlp --print / --dump-json --write-comments through the same cascade.

  4. Everything is cached on disk; agents can retry cheaply.

Limitations

  • Videos with no captions at all can't be transcripted (no Whisper/ASR — this tool stays lightweight on purpose).

  • Browser-cookie access requires the target browser installed locally; in logged-in sessions YouTube sometimes serves empty format lists to yt-dlp, which is exactly why cookies are the last layer (and always paired with --ignore-no-formats-error).

  • Cache is content-blind: if a video's description/comments change, clear the cache entry.

License

MIT © 2026 JADRT22

Available Tools

4 tools
get_transcript_toolB

Returns the full transcript of a YouTube video.

Args: url: Video URL (youtube.com/watch?v=..., youtu.be/..., shorts/...) or ID. languages: Preferred language codes, comma-separated (e.g. "en,es,pt"). Defaults to the server default.

Returns: Transcript text with a metadata header, or an error message.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
languagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It partially discharges this by disclosing the return shape (transcript text with a metadata header, or an error message), which tells the agent how failures surface. It says nothing about permission requirements, rate limits, or fallback behavior when a preferred language is unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then cleanly sectioned Args/Returns. Every sentence carries information. Slightly verbose in the Returns block given an output schema exists, but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Both parameters are documented despite 0% schema coverage, and the return contract is stated even though an output schema exists. For a simple two-parameter read tool this is nearly complete; only the sibling disambiguation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it enumerates the accepted URL forms (watch?v=, youtu.be, shorts/, bare ID) and explains that 'languages' is a comma-separated preference list defaulting to the server default. That is meaningful detail absent from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Returns the full transcript of a YouTube video.' An agent knows exactly what it produces. It does not, however, differentiate itself from sibling list_transcripts_tool or get_video_info_tool, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus list_transcripts_tool or the other siblings. The adjective 'full' implies a distinction from a listing tool, but that inference is left entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_video_details_toolA

Returns title, channel, duration, FULL description and comments of a video.

Useful for understanding context, mentioned links and audience reception.

Args: url: Video URL or ID. max_comments: maximum number of comments (default 50).

Returns: JSON with description and a comments list (author, text, likes).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_commentsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the response shape ('JSON with description and a comments list (author, text, likes)') and that comments are truncated by max_comments, which is useful. However, it never states that this is a read-only operation, nor mentions rate limits, pagination, or failure modes for invalid URLs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence is front-loaded and the Args/Returns sections are scannable. It is slightly padded by restating the default comment count and by describing return values that the existing output schema already covers, but nothing is wasteful enough to impede use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, two-parameter read tool with an output schema already present, the description covers purpose, parameters, and response contents adequately. The main gap is the absence of any routing guidance against its three siblings, which the sibling set makes relevant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must document parameters itself, and it does: url is clarified as 'Video URL or ID' (not obvious from a plain string type) and max_comments is described as a maximum with default 50. Both parameters receive meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource ('Returns title, channel, duration, FULL description and comments of a video') and enumerates the specific fields returned. It only weakly separates itself from the sibling get_video_info_tool, relying on the word 'FULL' and the inclusion of comments rather than an explicit contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers an implied use case ('Useful for understanding context, mentioned links and audience reception') but gives no explicit when-to-use conditions, prerequisites, or alternatives to the sibling tools get_transcript_tool / get_video_info_tool. An agent must infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_video_info_toolC

Returns video metadata: title, channel, duration and upload date.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only lists output fields. It says nothing about permissions, error behavior for invalid URLs, rate limits, or whether any network/auth preconditions apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word earns its place. It is terse to the point of under-specification, but structurally sound.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the returned values need not be re-explained, and the description partially duplicates them anyway. Still, with zero annotation coverage and an undocumented required parameter, more guidance was warranted for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single url parameter has no schema description. The description does not compensate by explaining the expected format (full URL vs. video ID) or any constraints, so the parameter is essentially undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (Returns) and resource (video metadata) and enumerates the fields returned, so the agent knows exactly what it produces. However, it offers no differentiation from the near-identical sibling get_video_details_tool, leaving a real ambiguity about which to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or how it relates to get_video_details_tool, get_transcript_tool, or list_transcripts_tool. The agent must guess between two similarly named metadata tools on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_transcripts_toolA

Lists the captions/transcripts available for a video (languages and type).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It discloses that the listing includes languages and type, which is useful, but says nothing about the video source, whether captions may be absent, permissions, or rate limits. For a low-risk read it is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the purpose and the returned fields are packed into one clause. Appropriately sized, though the parenthetical could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and a one-parameter read-only listing tool needs little more than this. The main omission is any routing against the three sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'url' parameter is undocumented in both places. The parameter name is largely self-evident for a video tool, but the description does nothing to clarify whether it expects a video page URL or an ID, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Lists') and resource ('captions/transcripts available for a video'), and even indicates what the list contains (languages and type). It does not, however, distinguish itself from the sibling get_transcript_tool, which an agent must infer is the fetch counterpart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a discovery step (list what's available before fetching), but never states when to use this versus get_transcript_tool or get_video_info_tool. No prerequisites, auth needs, or exclusions are given, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.2.0
    • First observedget_transcript_tool
    • First observedget_video_details_tool
    • First observedget_video_info_tool
    • First observedlist_transcripts_tool

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation3/5

get_video_details_tool and get_video_info_tool overlap heavily, both returning title, channel and duration; the distinction (details adds description/comments, info adds upload date) is only implicit. get_transcript_tool and list_transcripts_tool are clearly separable, but the two metadata tools risk misselection.

Naming Consistency5/5

All four tools follow a consistent verb_noun_tool pattern (get_video_details_tool, get_transcript_tool, list_transcripts_tool, get_video_info_tool) with uniform snake_case. Naming is highly predictable.

Tool Count4/5

Four tools are well-scoped for a YouTube transcript/metadata server, each targeting a distinct operation. Slightly thin, and the count is partly inflated by two near-duplicate metadata tools that could be merged.

Completeness4/5

The surface covers listing available transcripts, fetching a transcript, retrieving metadata, and gathering description/comments, which spans the core domain lifecycle. No update/delete is needed for a read-only extractor, so coverage is solid with only minor gaps (e.g. no bulk or search operation).

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    B
    maintenance
    Enables AI agents to search, analyze, and extract insights from YouTube videos including transcripts, visual frames, and benchmarks without requiring API keys. Supports semantic search across playlists, sentiment analysis, and visual content indexing with automatic fallback chains for reliable access.
    41
    81 npm
    35
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Extract YouTube transcripts for AI agents, RAG pipelines, and LLM workflows. Supports any YouTube URL. Returns clean text or timestamped segments. No API keys required.
    1
    4
    MIT