Skip to main content
Glama

get_youtube_transcript

Fetch the full text transcript of a YouTube video. Accepts a YouTube URL (watch, youtu.be, shorts, embed) or a bare 11-character video ID, and returns the transcript as plain text with title and metadata. Optionally pass a BCP-47 language code to select a specific caption track.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
langNoOptional BCP-47 language code, e.g. "en", "zh-CN", "es". Omit for the default track.
videoYesYouTube video URL or 11-character video ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It discloses the return format ('plain text with title and metadata'), input normalization across four URL formats plus bare ID, and how the lang parameter changes caption-track selection. This covers the core observable behavior well, though it omits edge-case behavior such as what happens when a video has no captions or an invalid ID is supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose is front-loaded in the first clause, the second sentence covers both input flexibility and output format, and the third covers the optional parameter. There is no fluff, no repetition of what the schema already states.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, it is important that the description states the return shape, and it does ('plain text with title and metadata'). Both parameters are fully documented in the schema and invocation behavior is clear. The only missing context is failure-mode behavior (unavailable transcripts, invalid IDs), which is a minor gap for a simple read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — both video and lang are already documented with clear descriptions including BCP-47 examples. The tool description adds marginal value by explicitly enumerating the URL variants (watch, youtu.be, shorts, embed), but the schema already carries the semantic load, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pair: 'Fetch the full text transcript of a YouTube video.' It is unambiguous about what is retrieved and explicitly states the output form (plain text with title and metadata), leaving no room to confuse it with a summary, metadata-only call, or video download tool even with no siblings present.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

With no sibling tools to route between, the description supplies clear usage context instead: it enumerates the accepted input forms (watch, youtu.be, shorts, embed URLs or bare 11-character ID) and explains when to set the optional language parameter versus omitting it for the default track. It provides clear invocation context, though there are no explicit exclusions or when-not-to-use statements since no alternatives exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.4/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The tool's purpose is clear and unambiguous by default.

Naming Consistency5/5

The single tool name get_youtube_transcript follows a clear verb_noun pattern and is consistent with standard naming conventions. There are no mixed styles or conflicting verb choices.

Tool Count4/5

One tool is minimal but well-matched to the server's narrow purpose of fetching YouTube transcripts. While slightly below the typical 3–15 tool range, the server does not feel under-scoped or artificially padded.

Completeness5/5

The tool fully covers the core domain need by accepting various YouTube URL formats or a bare video ID, returning transcript text, and supporting an optional language parameter. For a transcript-only server, there are no significant missing operations.