Skip to main content
Glama

šŸŽ¬ YouTube Subtitle MCP Server

A Model Context Protocol (MCP) server for fetching YouTube video subtitles/transcripts with support for multiple output formats (SRT, VTT, TXT, JSON).

License TypeScript MCP

✨ Features

  • šŸŽ„ Fetch subtitles from any public YouTube video

  • šŸ“ Multiple output formats: SRT, VTT, TXT, JSON

  • šŸŒ Multi-language subtitle support

  • ⚔ Two deployment modes: stdio (local) and HTTP (server)

  • šŸ”§ Zero-configuration setup with npx

  • šŸ“Š Complete timestamp information

  • šŸš€ Production-ready with TypeScript

  • šŸ”„ Built with youtubei.js for reliable and stable subtitle extraction

Related MCP server: YouTube Transcript MCP Server

šŸš€ Quick Start

Zero configuration required! Simply use npx:

npx -y youtube-subtitle-mcp

Or install globally:

npm install -g youtube-subtitle-mcp
youtube-subtitle-mcp

🌐 Method 2: HTTP Mode (For Server Deployment)

# Clone repository
git clone https://github.com/guangxiangdebizi/youtube-subtitle-mcp.git
cd youtube-subtitle-mcp

# Install dependencies
npm install

# Build
npm run build

# Start HTTP server
npm run start:http

Server will start at http://localhost:3000

šŸ“¦ Installation

For Development

# Clone repository
git clone https://github.com/guangxiangdebizi/youtube-subtitle-mcp.git
cd youtube-subtitle-mcp

# Install dependencies
npm install

# Build
npm run build

For Production

npm install -g youtube-subtitle-mcp

šŸ”§ Configuration

Add to your MCP client configuration file:

Claude Desktop / Cursor Configuration:

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "youtube-subtitle": {
      "command": "npx",
      "args": ["-y", "youtube-subtitle-mcp"]
    }
  }
}

For local development:

{
  "mcpServers": {
    "youtube-subtitle": {
      "command": "node",
      "args": ["C:/path/to/youtube-subtitle-mcp/build/index.js"]
    }
  }
}

🌐 HTTP Mode Configuration

{
  "mcpServers": {
    "youtube-subtitle": {
      "type": "streamableHttp",
      "url": "http://localhost:3000/mcp",
      "timeout": 600
    }
  }
}

šŸ› ļø Tool: fetch_youtube_subtitles

Parameters

Parameter

Type

Required

Description

Example

url

string

āœ… Yes

YouTube video URL or video ID

https://www.youtube.com/watch?v=dQw4w9WgXcQ

format

string

āŒ No

Output format (default: JSON)

SRT, VTT, TXT, JSON

lang

string

āŒ No

Language code (default: auto)

zh-Hans, en, ja

Supported URL Formats

  • Standard: https://www.youtube.com/watch?v=VIDEO_ID

  • Short: https://youtu.be/VIDEO_ID

  • Embed: https://www.youtube.com/embed/VIDEO_ID

  • Direct ID: VIDEO_ID

Language Codes

  • zh-Hans - Simplified Chinese

  • zh-Hant - Traditional Chinese

  • en - English

  • ja - Japanese

  • ko - Korean

  • es - Spanish

  • fr - French

  • de - German

šŸ“ Usage Examples

Example 1: Fetch JSON Format Subtitles (Default)

Please fetch subtitles from this video:
https://www.youtube.com/watch?v=dQw4w9WgXcQ

Example 2: Fetch SRT Format Subtitles

Please fetch subtitles in SRT format from:
https://www.youtube.com/watch?v=dQw4w9WgXcQ

Example 3: Fetch Specific Language Subtitles

Please fetch Simplified Chinese subtitles in VTT format from:
https://www.youtube.com/watch?v=dQw4w9WgXcQ
Language code: zh-Hans

Example 4: Fetch Plain Text Content

Please fetch plain text subtitles from:
https://youtu.be/dQw4w9WgXcQ
Format: TXT

šŸ“Š Output Format Examples

JSON Format

[
  {
    "text": "Hello world",
    "start": 0,
    "end": 2000,
    "duration": 2000
  },
  {
    "text": "Welcome to YouTube",
    "start": 2000,
    "end": 5000,
    "duration": 3000
  }
]

SRT Format

1
00:00:00,000 --> 00:00:02,000
Hello world

2
00:00:02,000 --> 00:00:05,000
Welcome to YouTube

VTT Format

WEBVTT

00:00:00.000 --> 00:00:02.000
Hello world

00:00:02.000 --> 00:00:05.000
Welcome to YouTube

TXT Format

Hello world
Welcome to YouTube
This is a subtitle example

šŸ—ļø Project Structure

youtube-subtitle-mcp/
ā”œā”€ā”€ src/
│   ā”œā”€ā”€ index.ts              # stdio mode entry (recommended for local use)
│   ā”œā”€ā”€ httpServer.ts         # HTTP mode entry (for server deployment)
│   └── tools/
│       ā”œā”€ā”€ fetchYoutubeSubtitles.ts  # Main tool implementation
│       ā”œā”€ā”€ formatters.ts     # Format converters (SRT/VTT/TXT/JSON)
│       └── utils.ts          # Utility functions
ā”œā”€ā”€ build/                    # Compiled JavaScript (generated)
ā”œā”€ā”€ package.json
ā”œā”€ā”€ tsconfig.json
└── README.md

šŸ“š Development

Build

npm run build

Watch Mode

npm run watch

Start stdio Mode

npm run start:stdio

Start HTTP Mode

npm run start:http

Custom Port (HTTP Mode)

PORT=8080 npm run start:http

🐳 Docker Deployment

Using Docker

FROM node:18-alpine
WORKDIR /app
COPY package*.json ./
RUN npm ci --only=production
COPY . .
RUN npm run build
EXPOSE 3000
CMD ["npm", "run", "start:http"]
# Build image
docker build -t youtube-subtitle-mcp .

# Run container
docker run -d -p 3000:3000 --name youtube-mcp youtube-subtitle-mcp

Using Docker Compose

version: '3.8'
services:
  youtube-subtitle-mcp:
    build: .
    ports:
      - "3000:3000"
    environment:
      - PORT=3000
    restart: unless-stopped
docker-compose up -d

šŸš€ Deployment

PM2 (Process Manager)

# Install PM2
npm install -g pm2

# Start service
pm2 start build/httpServer.js --name youtube-subtitle-mcp

# View status
pm2 status

# View logs
pm2 logs youtube-subtitle-mcp

# Restart
pm2 restart youtube-subtitle-mcp

# Stop
pm2 stop youtube-subtitle-mcp

šŸ” API Reference

Health Check (HTTP Mode)

Endpoint: GET /health

Response:

{
  "status": "healthy",
  "transport": "streamable-http",
  "activeSessions": 0,
  "name": "youtube-subtitle-mcp",
  "version": "1.0.0",
  "timestamp": "2024-01-01T12:00:00.000Z"
}

MCP Endpoint (HTTP Mode)

Endpoint: POST /mcp

Headers:

  • Content-Type: application/json

  • Mcp-Session-Id: <session-id> (after initialization)

Request Body: JSON-RPC 2.0 format

āš ļø Notes

  1. Public videos only: Cannot fetch subtitles from private or restricted videos

  2. Subtitles required: Video must have available subtitles (auto-generated or uploaded)

  3. Language codes: If specified language doesn't exist, an error will be returned

  4. Network required: Requires stable network connection to access YouTube

  5. Terms of Service: Please comply with YouTube's Terms of Service, use for research/analysis purposes only

ā“ FAQ

Q: Server started but client can't connect?

A: Check the following:

  1. Verify server is running: visit http://localhost:3000/health

  2. Check URL in configuration file: http://localhost:3000/mcp

  3. Ensure firewall isn't blocking port 3000

  4. Restart Cursor/Claude Desktop

Q: Why can't I fetch subtitles?

A: Possible reasons:

  • Video has no subtitles

  • Video is private or region-restricted

  • Specified language code doesn't exist

  • Network connection issue

  • Server network can't access YouTube

Q: Do you support auto-generated subtitles?

A: Yes, supports YouTube auto-generated subtitles.

Q: Can I batch process multiple videos?

A: Current version processes one video at a time. For batch processing, loop the tool call on the client side.

Q: What's the timestamp unit?

A: Timestamps in JSON format are in milliseconds (ms).

Q: Can I deploy on a remote server?

A: Yes! Just:

  1. Install Node.js on remote server

  2. Clone project and run npm install

  3. Start server with npm run start:http

  4. Use remote URL in client configuration: http://your-server:3000/mcp

  5. Recommend using HTTPS and reverse proxy (e.g., Nginx) for security

Q: How to change port?

A: Two methods:

  1. Environment variable: PORT=8080 npm run start:http

  2. Create .env file: add PORT=8080

šŸ¤ Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

šŸ“„ License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

šŸ‘¤ Author

Xingyu Chen

šŸ™ Acknowledgments

šŸ“® Support

If you have any questions or suggestions, please create an Issue.


Made with ā¤ļø by Xingyu Chen

Available Tools

1 tool
fetch_youtube_subtitlesFetch YouTube SubtitlesA

Fetch subtitles/transcripts from YouTube videos. Supports multiple output formats (SRT, VTT, TXT, JSON) and language selection. Returns complete subtitle content with timestamps.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video URL or video ID. Supported formats: https://www.youtube.com/watch?v=xxx, https://youtu.be/xxx, or direct video ID
formatNoOutput format. SRT: subtitle file format (with sequence numbers), VTT: WebVTT format, TXT: plain text (text only), JSON: structured JSON (with timestamps)JSON
langNoSubtitle language code (optional). Examples: zh-Hans (Simplified Chinese), zh-Hant (Traditional Chinese), en (English). Auto-detect if not specified

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions output formats and language auto-detection, which are useful, but lacks details on error handling, rate limits, authentication needs, or whether the operation is read-only (implied by 'fetch'). More behavioral context would improve transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with core functionality and efficiently lists key features in a single, well-structured sentence. Every element ('fetch subtitles/transcripts', 'output formats', 'language selection', 'returns content') adds value without redundancy, making it appropriately sized and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description adequately covers the tool's purpose and basic features, but lacks completeness for a tool with 3 parameters and potential behavioral complexities. It does not explain return values in detail or address edge cases, leaving gaps in contextual understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds marginal value by summarizing key parameters ('multiple output formats', 'language selection'), but does not provide additional semantics beyond what's in the schema. Baseline is 3, but the concise mention of parameters in context slightly elevates it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('fetch subtitles/transcripts'), target resource ('YouTube videos'), and key capabilities ('multiple output formats', 'language selection', 'complete subtitle content with timestamps'). It uses precise verbs and distinguishes what the tool does without redundancy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context through features like format and language selection, but provides no explicit guidance on when to use this tool versus alternatives (e.g., for different video platforms or content types). Since no sibling tools are listed, this is less critical, but general best practices are not addressed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • First observedfetch_youtube_subtitles

TDQS

A3.8/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool has a clear, distinct purpose focused on fetching YouTube subtitles.

Naming Consistency5/5

Since there is only one tool, naming consistency is inherently perfect. The tool name 'fetch_youtube_subtitles' follows a clear verb_noun pattern, but no comparison is needed.

Tool Count2/5

A single tool is too few for a server with the apparent scope of YouTube subtitle management, as it lacks operations like searching for videos, handling errors, or managing subtitle files. This minimal set may cause agent failures due to incomplete functionality.

Completeness2/5

The server is severely incomplete for its domain; it only provides fetching subtitles but lacks essential operations such as searching for videos, uploading or editing subtitles, or handling multiple videos. This creates significant gaps in coverage for subtitle-related workflows.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers