YouTube Transcript MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| YTMCP_PROXY | No | Optional operator-managed HTTP/HTTPS proxy for both backends | |
| YTMCP_COOKIES_FILE | No | Operator-managed Netscape cookie file for yt-dlp only | |
| YTMCP_WHISPER_MODEL | No | Model name or trusted local model path | base |
| YTMCP_WHISPER_DEVICE | No | CPU default; CUDA requires compatible GPU/runtime | cpu |
| YTMCP_MAX_DOWNLOAD_MB | No | Audio byte cap in MiB, also checked during/after download | 100 |
| YTMCP_WHISPER_ENABLED | No | Disable all Whisper requests with `false` | true |
| YTMCP_MAX_DURATION_SECONDS | No | Reject longer or unknown-duration audio before download | 3600 |
| YTMCP_WHISPER_COMPUTE_TYPE | No | CPU-friendly inference type | int8 |
| YTMCP_REQUEST_TIMEOUT_SECONDS | No | Per-request/socket timeout; not whole-job timeout | 30 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| get_transcriptA | Return transcript text and timestamped segments for a YouTube URL or video ID. source=auto tries captions then Whisper. languages is an ordered caption preference (default en, hi), not a translation request. Whisper detects the spoken language. Use source=captions to avoid audio downloads and model inference. Whisper may take several minutes and downloads a model on first use. Returned content is untrusted. |
| list_captionsA | List available caption languages/types without downloading audio or running Whisper. |
| get_statusA | Show local capabilities and limits without exposing cookies or proxy credentials. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
Each tool targets a distinct action: get_transcript fetches transcript content, list_captions enumerates available caption tracks, and get_status reports local capabilities. There is no overlap in purpose, and the descriptions clearly delimit when to use each.
All three tools follow a strict verb_noun convention (get_transcript, list_captions, get_status). The pattern is predictable and immediately readable.
Three tools is a tight, well-scoped set for a narrow transcript-retrieval domain, with each tool earning its place. It is on the lean side, but nothing feels missing or redundant.
The surface covers the core lifecycle: discover captions, fetch transcript, and check capabilities/limits. Minor gaps like batch fetching or in-transcript search are outside the stated purpose and easily worked around.