Skip to main content
Glama

YouTube Transcript MCP Server

A Model Context Protocol (MCP) server that fetches and cleans transcripts from YouTube videos using yt-dlp.

Features

  • Robust Transcript Fetching: Uses yt-dlp to retrieve subtitles.

    • Attempts to fetch manual subtitles first (high quality).

    • Falls back to auto-generated subtitles if manual ones are unavailable.

  • Smart Cleaning:

    • Removes VTT formatting, tags, and timestamps.

    • Deduplicates repeated lines common in auto-generated captions.

    • Cleans up common YouTube auto-caption artifacts.

  • MCP Integration: Fully compatible with Claude Desktop and other MCP clients.

Related MCP server: YouTube Transcript MCP Server

Tools

get_youtube_transcript

Fetches and cleans the transcript for a given YouTube video URL.

  • Arguments: url (string) - The full URL of the YouTube video.

  • Returns: A clean string containing the video transcript.

Installation

This project uses uv for dependency management.

Prerequisites

  • uv installed.

  • A recent version of Python.

Setup

  1. Clone the repository:

    git clone https://github.com/warshanks/yt-transcript-mcp.git
    cd yt-transcript-mcp
  2. Install dependencies:

    uv sync

Configuration

Claude Desktop

To use this server with Claude Desktop, add the following to your claude_desktop_config.json:

{
  "mcpServers": {
    "yt-transcript": {
      "command": "uv",
      "args": [
        "--directory",
        "/absolute/path/to/yt-transcript-mcp",
        "run",
        "server.py"
      ]
    }
  }
}

Replace /absolute/path/to/yt-transcript-mcp with the actual path to your cloned repository.

Development

To run the server locally for testing or development:

uv run server.py

License

MIT

Available Tools

1 tool
get_youtube_transcriptA

Fetches and cleans the transcript for a YouTube video URL using yt-dlp. Attempts to get manual subtitles first, falling back to auto-generated subtitles.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description notes the fallback strategy from manual to auto-generated subtitles, which is a behavioral trait. However, it does not explain what 'cleans' entails, potential failure modes, or the output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, directly addressing the core function and the fallback behavior without extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter) and the presence of an output schema, the description is largely complete. It covers the primary functionality and fallback, though it could mention limitations like availability or language support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only a parameter named 'url' with no description. The tool description clarifies that this parameter expects a YouTube video URL, adding essential meaning that the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches and cleans YouTube transcripts using yt-dlp, with a specific verb and resource. It also specifies fallback behavior, making it distinct and purposeful even without siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

While there are no sibling tools to differentiate from, the description implies usage for YouTube video transcript retrieval. It does not explicitly state when not to use it, but the context is clear enough to guide an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

With only a single tool, there is no potential for confusion between tools. The purpose of get_youtube_transcript is clearly defined and specific to fetching YouTube transcripts.

Naming Consistency5/5

The tool name follows the standard verb_noun pattern (get_youtube_transcript), which is consistent and descriptive. Even though there is only one tool, the naming convention is appropriate.

Tool Count4/5

The server's scope is narrowly defined as retrieving YouTube transcripts, and a single tool covers this purpose effectively. While the count is below the typical 3-15 range, it is not excessive or insufficient for such a focused domain.

Completeness5/5

The tool fully addresses the domain of fetching and cleaning YouTube transcripts, including fallback from manual to auto-generated subtitles. There are no obvious missing operations for a transcript-fetching service.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Fetches YouTube subtitles via yt-dlp, cleans them into plain text, and provides tools for transcript retrieval, file management, and session-based storage with paging.
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/warshanks/yt-transcript-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server