Skip to main content
Glama

Get YouTube transcript

youtube_get_transcript
Read-onlyIdempotent

Get YouTube transcripts from any URL, including shorts, live, and embed links. Choose text, subtitles, or timestamped segments, with language selection and honest errors when captions are missing.

Instructions

Return the transcript of a YouTube video, from any YouTube URL. Handles watch / youtu.be / shorts / live / embed / music / nocookie links and bare video ids. If the video has no captions it fails with code NO_TRANSCRIPT instead of inventing text; if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list. Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video reference: watch URL, youtu.be short link, /shorts/, /live/, /embed/, music.youtube.com, youtube-nocookie.com, an attribution link, or a bare 11-character video id.
outputNoOutput shape: text (default, readable lines), markdown (timestamped bullets), segments (JSON with millisecond timings), or srt / vtt subtitle files.
maxCharsNoTruncate the transcript to roughly this many characters.
languagesNoLanguage preferences in priority order, e.g. ["en"] or ["hi", "en"]. Defaults to ["en"]. Falls back to the closest available track and says so in notes.
timestampsNoPrefix each line of text output with its [mm:ss] offset.
translateToNoAsk YouTube to machine-translate the selected captions into this language (for example es). Only works when the video offers translations.
includeAutoCaptionsNoInclude YouTube auto-generated (speech-recognised) captions. Default true; set false to require human-written captions only.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the description's job is to add behavioral context. It does detail failure codes and URL handling, which is helpful. However, there is an internal contradiction: the description says 'if the requested language is missing it fails with LANGUAGE_UNAVAILABLE plus the available list' while the languages parameter description says 'Falls back to the closest available track and says so in notes.' This inconsistency undermines trust and clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise for the amount of information it conveys. It front-loads the primary purpose and then covers failure modes and usage tips in a structured manner. It is not verbose, though it could be tightened by removing the contradictory language note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 parameters) and the lack of an output schema, the description covers URL formats, output options, failure modes, and truncation behavior (via maxChars in schema). However, the language fallback/failure contradiction leaves a gap that could confuse an agent. Overall, it is fairly complete but not perfect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed descriptions for all 7 parameters, so the baseline is 3. The description adds a small amount of guidance (e.g., explaining the difference between output=segments and timestamps=true) but does not significantly enhance parameter understanding beyond what the schema already provides. It also repeats the contradictory language behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Return the transcript of a YouTube video, from any YouTube URL.' It enumerates the URL formats it accepts and explicitly contrasts failure modes (NO_TRANSCRIPT, LANGUAGE_UNAVAILABLE). This unambiguously differentiates it from sibling tools like youtube_parse_url or youtube_video_info, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage guidance for output selection ('Use output=segments for timestamped cues, or timestamps=true for [mm:ss] prefixed lines') but does not explicitly state when to prefer this tool over its siblings. While the purpose is clear, there is no direct 'use this instead of X when...' guidance, though the context makes it obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.