Skip to main content
Glama
retracn

automationnation-mcp

YouTube transcripts

get_youtube_transcripts
Read-only

Fetch YouTube video transcripts and captions from URLs, IDs, Shorts, channels, or playlists, including title, channel, language, and optional timestamps for summarizing, quoting, or fact-checking.

Instructions

Get the transcript (captions) of YouTube videos, with title, channel, duration, caption language and word count, and optional timestamps. Accepts video URLs or IDs, Shorts, youtu.be links, and channel or playlist URLs (it takes their latest videos). Manual captions in the requested language come first, then auto-generated ones; translation into another language is best effort. Use it to summarise, quote, fact-check or search inside videos. Videos without captions come back with an error status and aren't billed. Cost on your Apify account: $1.50 per 1,000 transcripts ($1.20 on Gold and above).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
videosYesYouTube video URLs or 11-character IDs, Shorts or youtu.be links, or channel/playlist URLs. Up to 10.
languageNoPreferred caption language code, e.g. en, es, de, fr, ja.en
translate_toNoOptional language code to machine-translate the transcript into (best effort; YouTube rate-limits translation).
max_charactersNoCut each transcript after this many characters to protect your context window.
include_timestampsNoReturn the transcript as timestamped lines ("[1:23] text") instead of plain text.
max_videos_per_channelNoFor channel or playlist URLs: how many of the latest videos to take.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: caption-priority ordering, best-effort translation with YouTube rate-limiting, error status and no billing for captionless videos, and explicit pricing ($1.50/1,000, $1.20 Gold+). This is exactly the extra behavioral context annotations cannot supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph, but it is front-loaded with the core purpose and input forms, and every sentence (language priority, billing, pricing) carries information. Slightly long and run-on, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers return shape (title, channel, duration, caption language, word count, timestamps) and failure behavior, so an agent has everything needed to call it and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds real meaning: it clarifies the accepted URL/ID/Shorts/youtu.be/channel/playlist forms and that channel or playlist URLs yield their latest videos, which ties directly to max_videos_per_channel.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get the transcript (captions) of YouTube videos') and enumerates the returned fields and accepted input forms. It is clearly distinguishable from the sibling search_google_videos tool because it is about captions of a known video, not video discovery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names intended uses ('summarise, quote, fact-check or search inside videos') and states the fallback behavior (manual captions first, then auto-generated; translation best effort). It does not, however, say when NOT to use it or point to an alternative sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.