TranscriptYT
Server Details
YouTube transcripts for AI agents: text, JSON, SRT, or VTT in 150+ languages, with AI fallback.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP ยท MCP 2025-11-25
- URL
TDQS
Scored across 3 tools
get_transcript, get_usage, and list_languages serve clearly distinct purposes: fetching content, checking account credits, and listing available caption tracks. No overlap or ambiguity.
All tools follow a consistent snake_case verb_noun pattern: get_transcript, get_usage, list_languages. The convention is predictable and readable.
With only 3 tools, each earns its place for a focused transcript retrieval API. The count is well-scoped and avoids unnecessary bloat.
The toolset covers the core workflow: list available languages, fetch the transcript in multiple formats, and check credit usage. No obvious gaps for the stated domain.
Available Tools
3 toolsget_transcriptGet YouTube transcriptARead-onlyInspect
Fetch the transcript of a public YouTube video. Accepts a full URL or an 11-character video ID. Defaults to Markdown (title + text), which is the most compact format for reading or summarizing. Use json for timestamped segments, srt/vtt for subtitle files. Costs 1 credit on success.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube URL or video ID | |
| format | No | Output format | md |
| language | No | Preferred caption language (BCP-47, e.g. en, es, pt-BR) | |
| timestamps | No | Prefix lines with timestamps (text/md) | |
| translate_to | No | Translate the transcript into this language |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds real behavioral context beyond them, notably 'Costs 1 credit on success' (billing model) and the 'public YouTube video' scope constraint that implies private/unavailable videos fail. It does not describe failure modes for videos lacking captions, so it falls just short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action, then inputs, then format guidance, then cost. Every sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete enough for a read-only, open-world retrieval tool: it explains input forms, output shapes, format trade-offs, and credit cost, and needs no return-value documentation since none of the fields are hidden. It omits edge-case behavior (e.g., videos without captions) but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so a baseline of 3 applies. The description goes beyond the schema's terse 'Output format' label by explaining what each value yields (Markdown = title + text, json = timestamped segments, srt/vtt = subtitle files) and by clarifying the dual URL/ID input, adding genuine semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Fetch the transcript of a public YouTube video.' The scope qualifier 'public' and the input forms (full URL or 11-character video ID) let an agent immediately distinguish this from unrelated sibling tools like get_usage and list_languages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance on format selection: 'Defaults to Markdown... Use json for timestamped segments, srt/vtt for subtitle files.' This tells the agent which format maps to which task. It stops short of naming an alternative tool to use instead (none exists among the siblings), so no explicit when-not is possible here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageCheck creditsARead-onlyInspect
Show the remaining TranscriptYT credits and plan for this API key. Free.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds two useful bits beyond the schema - that the result is keyed to the current API key and that the call costs nothing - but says nothing about rate limits, error behavior, or freshness of the credit figure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the resource and the cost caveat are both front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only status tool with no output schema, the description names the returned entities (remaining credits and plan) and the credential scope, which is enough to call it correctly. It could say a bit more about what happens when the key is invalid, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter detail, since there is nothing to parameterize.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Show) and resource (remaining credits and plan) scoped to 'this API key', so an agent immediately knows what it returns. It does not name or contrast sibling tools, though get_transcript and list_languages operate on entirely different resources so confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied - the 'Free' tag hints this is a cheap pre-flight check, but the description never states when to call it (e.g., before consuming credits) or any alternative. Adequate but no explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_languagesList caption languagesARead-onlyInspect
List the caption tracks available for a YouTube video. Free โ does not use credits.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube URL or video ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds a genuinely non-structured behavioral fact โ that the call is free and does not consume credits โ which matters in a billing-aware API where get_usage is a sibling. It stops short of describing return shape or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler: the action first, then the cost attribute. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only listing tool with no output schema and a fully documented parameter, the description plus annotations cover what an agent needs to call it correctly. Only the relation to get_transcript is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single documented parameter ('YouTube URL or video ID'), so the schema already carries full parameter meaning. The description adds no format or syntax detail beyond that baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the caption tracks available for a YouTube video'), which is clear and matches the title. It does not explicitly contrast itself with the sibling get_transcript, though the resource is distinct enough that an agent can differentiate them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the use case (discovering available caption tracks) and adds a cost note, but never states when to use this instead of get_transcript or that it is typically a prerequisite step. Usage is inferable rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
get_transcript - First observed
get_usage - First observed
list_languages
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
Transcripts of YouTube videos, playlists and channels in any language: text, SRT, VTT or JSON.
Transcripts of YouTube videos, playlists and channels with timestamps; SRT or WebVTT too.
๐ฏ The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free.
Related MCP Servers
AlicenseAqualityBmaintenanceEnables AI agents to retrieve YouTube transcripts from individual videos, channels, and search results, supporting multiple output formats such as plain text, SRT, and VTT.328 npmMIT- AlicenseAqualityBmaintenanceExtract YouTube transcripts for AI agents, RAG pipelines, and LLM workflows. Supports any YouTube URL. Returns clean text or timestamped segments. No API keys required.14MIT
- FlicenseNot gradedqualityCmaintenancePay-per-success YouTube transcript extractor for AI agents and RAG pipelines. Timestamps, SRT, plain text. Failed videos are never charged. $0.005/video.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to extract YouTube video metadata and generate high-quality multilingual transcriptions with voice activity detection, supporting 99 languages with translation capabilities and intelligent caching.56MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.