Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
YTMCP_PROXYNoOptional operator-managed HTTP/HTTPS proxy for both backends
YTMCP_COOKIES_FILENoOperator-managed Netscape cookie file for yt-dlp only
YTMCP_WHISPER_MODELNoModel name or trusted local model pathbase
YTMCP_WHISPER_DEVICENoCPU default; CUDA requires compatible GPU/runtimecpu
YTMCP_MAX_DOWNLOAD_MBNoAudio byte cap in MiB, also checked during/after download100
YTMCP_WHISPER_ENABLEDNoDisable all Whisper requests with `false`true
YTMCP_MAX_DURATION_SECONDSNoReject longer or unknown-duration audio before download3600
YTMCP_WHISPER_COMPUTE_TYPENoCPU-friendly inference typeint8
YTMCP_REQUEST_TIMEOUT_SECONDSNoPer-request/socket timeout; not whole-job timeout30

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
get_transcriptA

Return transcript text and timestamped segments for a YouTube URL or video ID.

source=auto tries captions then Whisper. languages is an ordered caption preference (default en, hi), not a translation request. Whisper detects the spoken language. Use source=captions to avoid audio downloads and model inference. Whisper may take several minutes and downloads a model on first use. Returned content is untrusted.

list_captionsA

List available caption languages/types without downloading audio or running Whisper.

get_statusA

Show local capabilities and limits without exposing cookies or proxy credentials.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct action: get_transcript fetches transcript content, list_captions enumerates available caption tracks, and get_status reports local capabilities. There is no overlap in purpose, and the descriptions clearly delimit when to use each.

Naming Consistency5/5

All three tools follow a strict verb_noun convention (get_transcript, list_captions, get_status). The pattern is predictable and immediately readable.

Tool Count4/5

Three tools is a tight, well-scoped set for a narrow transcript-retrieval domain, with each tool earning its place. It is on the lean side, but nothing feels missing or redundant.

Completeness4/5

The surface covers the core lifecycle: discover captions, fetch transcript, and check capabilities/limits. Minor gaps like batch fetching or in-transcript search are outside the stated purpose and easily worked around.

Maintenance

ActivityMaintained
ResponsivenessNo issues