yt-transcript-mcp
Fetches and cleans transcripts from YouTube videos using yt-dlp, supporting both manual and auto-generated subtitles.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@yt-transcript-mcpget transcript for https://www.youtube.com/watch?v=dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
YouTube Transcript MCP Server
A Model Context Protocol (MCP) server that fetches and cleans transcripts from YouTube videos using yt-dlp.
Features
Robust Transcript Fetching: Uses
yt-dlpto retrieve subtitles.Attempts to fetch manual subtitles first (high quality).
Falls back to auto-generated subtitles if manual ones are unavailable.
Smart Cleaning:
Removes VTT formatting, tags, and timestamps.
Deduplicates repeated lines common in auto-generated captions.
Cleans up common YouTube auto-caption artifacts.
MCP Integration: Fully compatible with Claude Desktop and other MCP clients.
Related MCP server: YouTube Transcript MCP Server
Tools
get_youtube_transcript
Fetches and cleans the transcript for a given YouTube video URL.
Arguments:
url(string) - The full URL of the YouTube video.Returns: A clean string containing the video transcript.
Installation
This project uses uv for dependency management.
Prerequisites
uv installed.
A recent version of Python.
Setup
Clone the repository:
git clone https://github.com/warshanks/yt-transcript-mcp.git cd yt-transcript-mcpInstall dependencies:
uv sync
Configuration
Claude Desktop
To use this server with Claude Desktop, add the following to your claude_desktop_config.json:
{
"mcpServers": {
"yt-transcript": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/yt-transcript-mcp",
"run",
"server.py"
]
}
}
}Replace /absolute/path/to/yt-transcript-mcp with the actual path to your cloned repository.
Development
To run the server locally for testing or development:
uv run server.pyLicense
MIT
Available Tools
1 toolget_youtube_transcriptA
Fetches and cleans the transcript for a YouTube video URL using yt-dlp. Attempts to get manual subtitles first, falling back to auto-generated subtitles.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description notes the fallback strategy from manual to auto-generated subtitles, which is a behavioral trait. However, it does not explain what 'cleans' entails, potential failure modes, or the output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, directly addressing the core function and the fallback behavior without extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter) and the presence of an output schema, the description is largely complete. It covers the primary functionality and fallback, though it could mention limitations like availability or language support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a parameter named 'url' with no description. The tool description clarifies that this parameter expects a YouTube video URL, adding essential meaning that the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches and cleans YouTube transcripts using yt-dlp, with a specific verb and resource. It also specifies fallback behavior, making it distinct and purposeful even without siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While there are no sibling tools to differentiate from, the description implies usage for YouTube video transcript retrieval. It does not explicitly state when not to use it, but the context is clear enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
get_youtube_transcript
TDQS
Scored across 1 tool
With only a single tool, there is no potential for confusion between tools. The purpose of get_youtube_transcript is clearly defined and specific to fetching YouTube transcripts.
The tool name follows the standard verb_noun pattern (get_youtube_transcript), which is consistent and descriptive. Even though there is only one tool, the naming convention is appropriate.
The server's scope is narrowly defined as retrieving YouTube transcripts, and a single tool covers this purpose effectively. While the count is below the typical 3-15 range, it is not excessive or insufficient for such a focused domain.
The tool fully addresses the domain of fetching and cleaning YouTube transcripts, including fallback from manual to auto-generated subtitles. There are no obvious missing operations for a transcript-fetching service.
Maintenance
Related MCP Connectors
Clean YouTube transcripts for agents: single videos, channels, playlists, plus AI caption cleanup.
Fetch the full transcript of any YouTube video as clean text. No API key, no signup.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Search YouTube, read video metadata, and fetch transcripts with language preferences
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables fetching, searching, and analyzing YouTube video transcripts in multiple languages using yt-dlp. Supports timestamp filtering, language detection, and transcript summaries with robust error handling for production use.4MIT
- AlicenseAqualityDmaintenanceRetrieves transcripts from YouTube videos with support for multiple languages, timestamp control, and language detection. Enables video content analysis, summarization, and quote extraction without manually downloading or watching videos.264 npm15MIT
- AlicenseAqualityDmaintenanceExtracts clean text transcripts from YouTube videos using their subtitles and returns them as plain text.15 npmMIT
- AlicenseNot gradedqualityDmaintenanceFetches YouTube subtitles via yt-dlp, cleans them into plain text, and provides tools for transcript retrieval, file management, and session-based storage with paging.MIT