yt-transcript-mcp
Fetches and cleans transcripts from YouTube videos using yt-dlp, supporting both manual and auto-generated subtitles.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@yt-transcript-mcpget transcript for https://www.youtube.com/watch?v=dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
YouTube Transcript MCP Server
A Model Context Protocol (MCP) server that fetches and cleans transcripts from YouTube videos using yt-dlp.
Features
Robust Transcript Fetching: Uses
yt-dlpto retrieve subtitles.Attempts to fetch manual subtitles first (high quality).
Falls back to auto-generated subtitles if manual ones are unavailable.
Smart Cleaning:
Removes VTT formatting, tags, and timestamps.
Deduplicates repeated lines common in auto-generated captions.
Cleans up common YouTube auto-caption artifacts.
MCP Integration: Fully compatible with Claude Desktop and other MCP clients.
Related MCP server: YouTube Transcript MCP Server
Tools
get_youtube_transcript
Fetches and cleans the transcript for a given YouTube video URL.
Arguments:
url(string) - The full URL of the YouTube video.Returns: A clean string containing the video transcript.
Installation
This project uses uv for dependency management.
Prerequisites
uv installed.
A recent version of Python.
Setup
Clone the repository:
git clone https://github.com/warshanks/yt-transcript-mcp.git cd yt-transcript-mcpInstall dependencies:
uv sync
Configuration
Claude Desktop
To use this server with Claude Desktop, add the following to your claude_desktop_config.json:
{
"mcpServers": {
"yt-transcript": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/yt-transcript-mcp",
"run",
"server.py"
]
}
}
}Replace /absolute/path/to/yt-transcript-mcp with the actual path to your cloned repository.
Development
To run the server locally for testing or development:
uv run server.pyLicense
MIT
Available Tools
1 toolget_youtube_transcriptA
Fetches and cleans the transcript for a YouTube video URL using yt-dlp. Attempts to get manual subtitles first, falling back to auto-generated subtitles.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description notes the fallback strategy from manual to auto-generated subtitles, which is a behavioral trait. However, it does not explain what 'cleans' entails, potential failure modes, or the output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, directly addressing the core function and the fallback behavior without extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter) and the presence of an output schema, the description is largely complete. It covers the primary functionality and fallback, though it could mention limitations like availability or language support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a parameter named 'url' with no description. The tool description clarifies that this parameter expects a YouTube video URL, adding essential meaning that the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches and cleans YouTube transcripts using yt-dlp, with a specific verb and resource. It also specifies fallback behavior, making it distinct and purposeful even without siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While there are no sibling tools to differentiate from, the description implies usage for YouTube video transcript retrieval. It does not explicitly state when not to use it, but the context is clear enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only a single tool, there is no potential for confusion between tools. The purpose of get_youtube_transcript is clearly defined and specific to fetching YouTube transcripts.
The tool name follows the standard verb_noun pattern (get_youtube_transcript), which is consistent and descriptive. Even though there is only one tool, the naming convention is appropriate.
The server's scope is narrowly defined as retrieving YouTube transcripts, and a single tool covers this purpose effectively. While the count is below the typical 3-15 range, it is not excessive or insufficient for such a focused domain.
The tool fully addresses the domain of fetching and cleaning YouTube transcripts, including fallback from manual to auto-generated subtitles. There are no obvious missing operations for a transcript-fetching service.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Clean YouTube transcripts for agents: single videos, channels, playlists, plus AI caption cleanup.
Fetch the full transcript of any YouTube video as clean text. No API key, no signup.
Free YouTube transcript fetcher: clean text, timestamped, SRT, VTT, Markdown, or JSON. No API key.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables fetching, searching, and analyzing YouTube video transcripts in multiple languages using yt-dlp. Supports timestamp filtering, language detection, and transcript summaries with robust error handling for production use.4MIT
- AlicenseAqualityFmaintenanceRetrieves transcripts from YouTube videos with support for multiple languages, timestamp control, and language detection. Enables video content analysis, summarization, and quote extraction without manually downloading or watching videos.212315MIT
- AlicenseAqualityDmaintenanceExtracts clean text transcripts from YouTube videos using their subtitles and returns them as plain text.122MIT
- AlicenseNot gradedqualityDmaintenanceFetches YouTube subtitles via yt-dlp, cleans them into plain text, and provides tools for transcript retrieval, file management, and session-based storage with paging.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/warshanks/yt-transcript-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server