Skip to main content
Glama

agdata

YouTube transcript of a video

youtube-transcript
Read-only

YouTube transcript API for AI agents: send a video URL or 11-character id and get its published subtitles as timed JSON segments, plain text, SRT or WebVTT, with title, channel, duration, views and publish date. Choose one of ten subtitle languages or any. No speech-to-text. A video without subtitles in that language is not charged.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
langNosubtitle language; any takes English if the video has it, else the first tracken
videoYesone video: an 11-character video id, or a watch, shorts, live, embed or youtu.be URL of one video (not a channel, playlist or search)
formatNojson: timed segments in transcript; text, srt or vtt: the transcript as one stringjson

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which already mark this read-only, open-world and non-destructive), the description adds meaningful behavioral context: the no-speech-to-text constraint and the cost rule that a video lacking the requested language is not charged. It does not cover rate limits or pagination, but the added cost/limitation disclosure is real value over structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: capability, inputs, outputs, constraints and cost are each covered in one tight sentence with little waste. Slightly packed, but every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the returned fields (segments, title, channel, duration, views, publish date) and the available formats, so an agent knows what it gets back. Only edge behavior (errors, empty results) is left unaddressed, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 100% the baseline is 3, and the description goes further by restating that video accepts a URL or 11-char id, that lang spans ten languages plus 'any', and that format yields timed JSON or a single SRT/VTT/text string. This adds format semantics beyond the raw enum lists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('send a video URL or 11-character id and get its published subtitles'), and its scope as a transcript-only tool is clear against siblings like youtube (metadata) and youtube-comments. An agent can tell it apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description discloses a key when-not ('No speech-to-text') and the schema routes input away from channels/playlists/search, giving clear context. However, it never names an alternative tool or the condition that would send the agent to youtube comments or plain youtube instead, so routing is inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources