yt-transcript
Provides tools to fetch YouTube video transcripts, list available captions, retrieve video metadata (title, channel, duration, upload date), and get full descriptions plus comments (authors, text, likes) for analysis and summarization.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@yt-transcriptsummarize this video: https://youtu.be/dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
yt-transcript
Give your AI agent eyes for YouTube. A lightweight MCP server + CLI that fetches the transcript, description, and comments of any YouTube video — so your agent can read, summarize, and analyze video content.
Works with opencode, Claude Code, Freebuff, Cursor, and any MCP-compatible client — or as a plain CLI any agent can call.
Why this one?
YouTube aggressively blocks scripted access ("confirm you're not a bot", HTTP 429). Most transcript tools break the first time YouTube pushes back. This one ships a 4-layer fallback cascade baked in:
Layer | What it does |
1. | Fast path, no API key needed |
2. yt-dlp web client | Baseline fallback |
3. yt-dlp android client | Bypasses the "not a bot" check on most videos |
4. yt-dlp + browser cookies | Cures hard IP blocks / 429s (uses your Firefox/Chrome cookies) |
Plus:
Disk cache (
~/.cache/yt-transcript/) — second call for the same video is ~10x faster and consumes zero rate limitDescription + comments — likes, authors, replies; great context for richer summaries
Language aware — pick preferred languages; lists everything available per video
Zero config, zero API keys
Related MCP server: VidLens
Install
git clone https://github.com/JADRT22/yt-transcript.git
cd yt-transcript
python3 -m venv .venv
.venv/bin/pip install -e .MCP setup
Point your client at the server binary:
/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcpIn ~/.config/opencode/opencode.json (global) or opencode.json (project):
{
"mcp": {
"yt-transcript": {
"type": "local",
"command": ["/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"],
"enabled": true
}
}
}claude mcp add yt-transcript -- /absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcpCreate a mcp.json in the project root where the agent runs:
{
"mcpServers": {
"yt-transcript": {
"command": "/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"
}
}
}Any agent that runs shell commands can read the output directly:
/absolute/path/to/yt-transcript/.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"MCP tools
Tool | Description |
| Full transcript of the video |
| Available captions (languages, auto-generated?) |
| Title, channel, duration, upload date |
| Full description + comments (author, text, likes) |
Example prompts once connected:
Summarize this video: https://www.youtube.com/watch?v=...
What are people saying in the comments of this video? [URL]
Compare what the video claims with what the description promises: [URL]CLI usage
# Full transcript (default languages: en, es, pt)
.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"
# Metadata only
.venv/bin/yt-transcript VIDEO_ID --info
# Description + comments (JSON)
.venv/bin/yt-transcript VIDEO_ID --desc --max-comments 50
# Available captions
.venv/bin/yt-transcript VIDEO_ID --list
# JSON output / other languages
.venv/bin/yt-transcript VIDEO_ID --json
.venv/bin/yt-transcript VIDEO_ID -l de,frAccepts youtube.com/watch?v=..., youtu.be/..., youtube.com/shorts/..., or the bare 11-character ID.
Configuration
Environment variable | Default | Purpose |
|
| Browser for the cookie fallback ( |
|
| Cache directory ( |
To force a fresh fetch for one video, delete its file (SHA-256-named) from the cache directory.
How it works
extract_video_idnormalizes any YouTube URL format into a video ID.Transcripts: captions API first; on any failure, yt-dlp downloads the
.vttthrough the client cascade and converts it to plain text (each attempt only counts if the caption file is actually written — yt-dlp's simulation mode silently skips it, and--no-simulateguards that).Metadata/details:
yt-dlp --print/--dump-json --write-commentsthrough the same cascade.Everything is cached on disk; agents can retry cheaply.
Limitations
Videos with no captions at all can't be transcripted (no Whisper/ASR — this tool stays lightweight on purpose).
Browser-cookie access requires the target browser installed locally; in logged-in sessions YouTube sometimes serves empty format lists to yt-dlp, which is exactly why cookies are the last layer (and always paired with
--ignore-no-formats-error).Cache is content-blind: if a video's description/comments change, clear the cache entry.
License
MIT © 2026 JADRT22
Available Tools
4 toolsget_transcript_toolB
Returns the full transcript of a YouTube video.
Args: url: Video URL (youtube.com/watch?v=..., youtu.be/..., shorts/...) or ID. languages: Preferred language codes, comma-separated (e.g. "en,es,pt"). Defaults to the server default.
Returns: Transcript text with a metadata header, or an error message.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| languages | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It partially discharges this by disclosing the return shape (transcript text with a metadata header, or an error message), which tells the agent how failures surface. It says nothing about permission requirements, rate limits, or fallback behavior when a preferred language is unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then cleanly sectioned Args/Returns. Every sentence carries information. Slightly verbose in the Returns block given an output schema exists, but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Both parameters are documented despite 0% schema coverage, and the return contract is stated even though an output schema exists. For a simple two-parameter read tool this is nearly complete; only the sibling disambiguation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it enumerates the accepted URL forms (watch?v=, youtu.be, shorts/, bare ID) and explains that 'languages' is a comma-separated preference list defaulting to the server default. That is meaningful detail absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Returns the full transcript of a YouTube video.' An agent knows exactly what it produces. It does not, however, differentiate itself from sibling list_transcripts_tool or get_video_info_tool, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus list_transcripts_tool or the other siblings. The adjective 'full' implies a distinction from a listing tool, but that inference is left entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_details_toolA
Returns title, channel, duration, FULL description and comments of a video.
Useful for understanding context, mentioned links and audience reception.
Args: url: Video URL or ID. max_comments: maximum number of comments (default 50).
Returns: JSON with description and a comments list (author, text, likes).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_comments | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the response shape ('JSON with description and a comments list (author, text, likes)') and that comments are truncated by max_comments, which is useful. However, it never states that this is a read-only operation, nor mentions rate limits, pagination, or failure modes for invalid URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose sentence is front-loaded and the Args/Returns sections are scannable. It is slightly padded by restating the default comment count and by describing return values that the existing output schema already covers, but nothing is wasteful enough to impede use.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, two-parameter read tool with an output schema already present, the description covers purpose, parameters, and response contents adequately. The main gap is the absence of any routing guidance against its three siblings, which the sibling set makes relevant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must document parameters itself, and it does: url is clarified as 'Video URL or ID' (not obvious from a plain string type) and max_comments is described as a maximum with default 50. Both parameters receive meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb+resource ('Returns title, channel, duration, FULL description and comments of a video') and enumerates the specific fields returned. It only weakly separates itself from the sibling get_video_info_tool, relying on the word 'FULL' and the inclusion of comments rather than an explicit contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers an implied use case ('Useful for understanding context, mentioned links and audience reception') but gives no explicit when-to-use conditions, prerequisites, or alternatives to the sibling tools get_transcript_tool / get_video_info_tool. An agent must infer selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_info_toolC
Returns video metadata: title, channel, duration and upload date.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it only lists output fields. It says nothing about permissions, error behavior for invalid URLs, rate limits, or whether any network/auth preconditions apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every word earns its place. It is terse to the point of under-specification, but structurally sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the returned values need not be re-explained, and the description partially duplicates them anyway. Still, with zero annotation coverage and an undocumented required parameter, more guidance was warranted for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single url parameter has no schema description. The description does not compensate by explaining the expected format (full URL vs. video ID) or any constraints, so the parameter is essentially undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Returns) and resource (video metadata) and enumerates the fields returned, so the agent knows exactly what it produces. However, it offers no differentiation from the near-identical sibling get_video_details_tool, leaving a real ambiguity about which to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or how it relates to get_video_details_tool, get_transcript_tool, or list_transcripts_tool. The agent must guess between two similarly named metadata tools on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_transcripts_toolA
Lists the captions/transcripts available for a video (languages and type).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses that the listing includes languages and type, which is useful, but says nothing about the video source, whether captions may be absent, permissions, or rate limits. For a low-risk read it is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the purpose and the returned fields are packed into one clause. Appropriately sized, though the parenthetical could be slightly tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and a one-parameter read-only listing tool needs little more than this. The main omission is any routing against the three sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'url' parameter is undocumented in both places. The parameter name is largely self-evident for a video tool, but the description does nothing to clarify whether it expects a video page URL or an ID, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and resource ('captions/transcripts available for a video'), and even indicates what the list contains (languages and type). It does not, however, distinguish itself from the sibling get_transcript_tool, which an agent must infer is the fetch counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a discovery step (list what's available before fetching), but never states when to use this versus get_transcript_tool or get_video_info_tool. No prerequisites, auth needs, or exclusions are given, leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.0- First observed
get_transcript_tool - First observed
get_video_details_tool - First observed
get_video_info_tool - First observed
list_transcripts_tool
TDQS
Scored across 4 tools
get_video_details_tool and get_video_info_tool overlap heavily, both returning title, channel and duration; the distinction (details adds description/comments, info adds upload date) is only implicit. get_transcript_tool and list_transcripts_tool are clearly separable, but the two metadata tools risk misselection.
All four tools follow a consistent verb_noun_tool pattern (get_video_details_tool, get_transcript_tool, list_transcripts_tool, get_video_info_tool) with uniform snake_case. Naming is highly predictable.
Four tools are well-scoped for a YouTube transcript/metadata server, each targeting a distinct operation. Slightly thin, and the count is partly inflated by two near-duplicate metadata tools that could be merged.
The surface covers listing available transcripts, fetching a transcript, retrieving metadata, and gathering description/comments, which spans the core domain lifecycle. No update/delete is needed for a read-only extractor, so coverage is solid with only minor gaps (e.g. no bulk or search operation).
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
YouTube data for AI agents: channels, videos, transcripts, comments, search. Video research.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Related MCP Servers
- AlicenseBqualityNot gradedmaintenanceYouTube intelligence layer for AI agents. 41 tools across 10 modules ; search, explore, transcripts, comments, visual search, analytics, and more. Zero config.4181 npm-
- AlicenseCqualityBmaintenanceEnables AI agents to search, analyze, and extract insights from YouTube videos including transcripts, visual frames, and benchmarks without requiring API keys. Supports semantic search across playlists, sentiment analysis, and visual content indexing with automatic fallback chains for reliable access.4181 npm35MIT
- AlicenseAqualityBmaintenanceExtract YouTube transcripts for AI agents, RAG pipelines, and LLM workflows. Supports any YouTube URL. Returns clean text or timestamped segments. No API keys required.14MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to fetch transcripts, metadata, and download videos/audio from YouTube without API keys.33MIT