YouTube Transcript AI
Server Details
Fetch the full transcript of any YouTube video as clean text. No API key, no signup.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP
- URL
- Repository
- zxl777/youtube-transcript-mcp
- GitHub Stars
- 0
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion between tools. The tool's purpose is clear and unambiguous by default.
The single tool name get_youtube_transcript follows a clear verb_noun pattern and is consistent with standard naming conventions. There are no mixed styles or conflicting verb choices.
One tool is minimal but well-matched to the server's narrow purpose of fetching YouTube transcripts. While slightly below the typical 3–15 tool range, the server does not feel under-scoped or artificially padded.
The tool fully covers the core domain need by accepting various YouTube URL formats or a bare video ID, returning transcript text, and supporting an optional language parameter. For a transcript-only server, there are no significant missing operations.
Available Tools
1 toolget_youtube_transcriptAInspect
Fetch the full text transcript of a YouTube video. Accepts a YouTube URL (watch, youtu.be, shorts, embed) or a bare 11-character video ID, and returns the transcript as plain text with title and metadata. Optionally pass a BCP-47 language code to select a specific caption track.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Optional BCP-47 language code, e.g. "en", "zh-CN", "es". Omit for the default track. | |
| video | Yes | YouTube video URL or 11-character video ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses the return format ('plain text with title and metadata'), input normalization across four URL formats plus bare ID, and how the lang parameter changes caption-track selection. This covers the core observable behavior well, though it omits edge-case behavior such as what happens when a video has no captions or an invalid ID is supplied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose is front-loaded in the first clause, the second sentence covers both input flexibility and output format, and the third covers the optional parameter. There is no fluff, no repetition of what the schema already states.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, it is important that the description states the return shape, and it does ('plain text with title and metadata'). Both parameters are fully documented in the schema and invocation behavior is clear. The only missing context is failure-mode behavior (unavailable transcripts, invalid IDs), which is a minor gap for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — both video and lang are already documented with clear descriptions including BCP-47 examples. The tool description adds marginal value by explicitly enumerating the URL variants (watch, youtu.be, shorts, embed), but the schema already carries the semantic load, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource pair: 'Fetch the full text transcript of a YouTube video.' It is unambiguous about what is retrieved and explicitly states the output form (plain text with title and metadata), leaving no room to confuse it with a summary, metadata-only call, or video download tool even with no siblings present.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
With no sibling tools to route between, the description supplies clear usage context instead: it enumerates the accepted input forms (watch, youtu.be, shorts, embed URLs or bare 11-character ID) and explains when to set the optional language parameter versus omitting it for the default track. It provides clear invocation context, though there are no explicit exclusions or when-not-to-use statements since no alternatives exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- First observed
get_youtube_transcript
Related MCP Connectors
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Clean YouTube transcripts for agents: single videos, channels, playlists, plus AI caption cleanup.
Free YouTube transcripts, no API key: videos, channel lists, latest uploads, bulk download links.
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents. No signup.
Related MCP Servers
- AlicenseAqualityDmaintenanceExtracts clean text transcripts from YouTube videos using their subtitles and returns them as plain text.15 npmMIT
- AlicenseAqualityBmaintenanceExtract YouTube transcripts for AI agents, RAG pipelines, and LLM workflows. Supports any YouTube URL. Returns clean text or timestamped segments. No API keys required.14MIT
- AlicenseAqualityFmaintenanceRetrieves transcripts from YouTube videos with support for multiple languages, timestamp control, and language detection. Enables video content analysis, summarization, and quote extraction without manually downloading or watching videos.275 npm15MIT
- AlicenseAqualityCmaintenanceFetches YouTube video subtitles and transcripts with support for multiple languages and output formats (SRT, VTT, TXT, JSON).17 npmApache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.