YouTube Transcript + YouTube Search MCP
Server Details
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
- Status
- Healthy
- Uptime
- 85.9% over 22 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- tubeagentkit/youtube-mcp
- GitHub Stars
- 15
- Server Listing
- YouTube Transcript + Search MCP
TDQS
Scored across 7 tools
Each tool has a clearly distinct scope: transcript retrieval, YouTube-wide search, in-channel search, channel latest videos, full channel uploads, playlist videos, and credit balance. Descriptions explicitly cross-reference related tools and clarify boundaries, such as using list_channel_videos for full history versus get_channel_latest_videos for recent uploads. No two tools appear to do the same thing.
All tool names use consistent snake_case with a verb_noun or verb_resource_modifier pattern (get_, list_, search_). The verbs are predictable and directly indicate the operation, and resources are named clearly. No naming convention mixing is present.
Seven tools is well-scoped for a YouTube transcript and search server. Each tool earns its place by covering a specific retrieval mode or support function without redundancy. The count is neither thin nor bloated.
The surface covers the core domain thoroughly: video transcript retrieval, global search, channel search, channel upload browsing, latest channel videos, playlist traversal, and credit checking. Pagination is supported via continuation tokens, and search results provide metadata needed before transcript retrieval. No obvious lifecycle or workflow gap prevents common agent tasks.
Available Tools
7 toolsget_channel_latest_videosGet Channel Latest VideosARead-onlyInspect
Use this when the user asks what a YouTube channel has posted recently, or wants a channel's details (name, description, links) with its newest uploads. Accepts an @handle, channel URL, or UC... ID. For a channel's full upload history use list_channel_videos; to find a topic inside one channel use search_channel_videos.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | Yes | Channel @handle, URL, or UC... id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds only the 'newest uploads' scope; it says nothing about how many videos are returned, pagination, rate limits, or quota cost, so it does not go meaningfully beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler: use-case first, accepted input next, sibling routing last. The trigger is front-loaded before the alternatives.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one parameter, no output schema, and full annotation coverage, an agent has everything needed to call correctly. The description even sketches the return payload ('name, description, links with newest uploads'), compensating for the absent output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'channel' parameter, and the description repeats the same accepted formats ('@handle, channel URL, or UC... ID') without adding new semantics. Baseline 3 applies when the schema already documents the parameter fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource — fetching a channel's newest uploads plus its details (name, description, links). It explicitly distinguishes itself from list_channel_videos (full history) and search_channel_videos (topic within a channel), so an agent can route without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger condition ('user asks what a YouTube channel has posted recently') and names two alternative tools with the exact conditions that select them. Nothing about when/when-not is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_creditsGet CreditsARead-onlyInspect
Use this when the user asks how many credits their getyoutubetranscript.com account has left, or after another tool reports the account has no credits remaining. Returns plan and top-up credit balances. Calling it does not use credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered structurally. The description adds genuinely useful behavioral context beyond that: 'Calling it does not use credits' resolves the obvious agent concern for a billing-adjacent tool, and it names what is returned (plan and top-up balances).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the usage trigger and reserving the last line for the non-obvious behavioral fact. Each sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with no output schema, the description supplies the trigger, the returned values (plan and top-up balances), and the cost behavior. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to compensate for regarding parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('how many credits their getyoutubetranscript.com account has left') and explicitly distinguishes its domain from every sibling tool, which are all transcript/video retrieval tools. An agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives two explicit trigger conditions: when the user asks about remaining credits, and after another tool reports zero credits remaining. Nothing is left to inference about when to reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_youtube_transcriptGet YouTube TranscriptARead-onlyInspect
Use this when the user shares a YouTube video link or ID, or asks to summarize, explain, quote, translate, take notes on, or chat about a specific YouTube video, lecture, podcast, or talk. Returns the full spoken text (captions) as one block, plus the title and channel. Accepts watch, youtu.be, Shorts, and live URLs. No per-line timestamps, so do not promise a timestamped breakdown. Do not use for non-YouTube videos or for files the user uploads. If a video has no captions, this returns an error: tell the user instead of guessing what the video says.
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | Language code, e.g. 'en', 'es'. Defaults to 'en'. | |
| video_url | Yes | YouTube URL (full or short) or an 11-character video ID | |
| send_metadata | No | Include title/author/thumbnail metadata. Defaults to true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, non-destructive, open-world profile. The description goes well beyond them by disclosing the return shape (one block of text plus title and channel), the absence of per-line timestamps, the accepted URL variants, and the exact failure mode for caption-less videos with instructions to report rather than guess.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the trigger condition and followed by behavior, constraints, and exclusion. No sentence is filler; each carries a distinct operational instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return payload and its limitations (no timestamps, single text block). Annotations cover safety, so nothing an agent needs to call or interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds genuine meaning: it enumerates the accepted URL forms (watch, youtu.be, Shorts, live) beyond the schema's generic 'full or short URL or 11-character ID'. It is silent on the language and send_metadata parameters, so it falls short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (retrieve the full spoken transcript of a YouTube video) and immediately scopes it to YouTube, which separates it from every search/list sibling (search_youtube, list_channel_videos, etc.). An agent can tell what it returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions ('user shares a YouTube video link or ID, or asks to summarize, explain, quote, translate, take notes on, or chat about'), explicit exclusions ('non-YouTube videos or files the user uploads'), and a fallback behavior when captions are missing. This is complete routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_channel_videosList Channel VideosARead-onlyInspect
Use this when the user wants to browse a YouTube channel's uploads beyond its latest videos, for example to see a creator's older episodes or find a video they remember from that channel. Returns the channel's uploads (its Videos tab) one page at a time. Pass continuation from a previous response for the next page.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Channel @handle, URL, or UC... id. Required unless `continuation` is set. | |
| continuation | No | Opaque token from a previous response - fetches the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safe-read (readOnly, non-destructive) and open-world. The description adds pagination semantics ('one page at a time', pass continuation for next page), which is valuable context beyond the annotations. It stops short of return-shape detail, but with no output schema that's a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads the when-to-use and the second covers the what-it-returns and how to paginate. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param read-only tool with no output schema, the description covers purpose, trigger context, and pagination. It could mention result ordering or that continuation is required for subsequent pages, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema. The description reiterates the continuation-pagination behavior but adds no format or constraint detail beyond what the schema provides. Baseline 3 for a fully-covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list/browse) and resource (a channel's uploads), and explicitly scopes it to the Videos tab beyond 'latest videos'. This clearly distinguishes it from get_channel_latest_videos and search_channel_videos.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case ('older episodes', 'find a video they remember') and contrasts with latest videos implicitly, but does not name the sibling tools or state when-not-to-use (e.g., exclude search_channel_videos).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_playlist_videosList Playlist VideosARead-onlyInspect
Use this when the user shares a YouTube playlist, or wants to work through a course, lecture series, or podcast playlist video by video (for example to build study notes for each lecture). Returns the videos in playlist order, one page at a time. Pass continuation from a previous response for the next page.
| Name | Required | Description | Default |
|---|---|---|---|
| playlist | No | Playlist URL or id. Required unless `continuation` is set. | |
| continuation | No | Opaque token from a previous response - fetches the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, non-destructive, open-world behavior, so the bar is lower; the description still adds real operational context by disclosing pagination ('one page at a time') and the continuation-token workflow. Ordering semantics ('in playlist order') is also non-obvious value. It does not mention rate limits or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the usage trigger, then behavior, then the pagination mechanic. No filler or repetition; each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only paginated list tool with no output schema, the description covers when to call it, ordering, and paging — the essentials an agent needs. It could briefly indicate what fields each returned video carries, but nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including the 'required unless continuation is set' rule. The description's instruction to pass the continuation token largely restates the schema, so it adds workflow framing rather than new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Returns the videos in playlist order') and scopes it to playlists, which implicitly separates it from the channel- and search-oriented siblings. It stops short of naming an alternative tool, so the differentiation is inferred rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete triggering conditions ('when the user shares a YouTube playlist, or wants to work through a course, lecture series, or podcast playlist video by video') with a worked example. It does not state when NOT to use it or name a sibling to prefer instead, leaving that to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_channel_videosSearch Channel VideosARead-onlyInspect
Use this when the user wants videos from one specific YouTube channel about a topic, for example 'what has @hubermanlab said about sleep' or 'find the MIT OpenCourseWare videos on recursion'. Searches only inside that channel. Pass continuation from a previous response for the next page.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Query to search within the channel. Required unless `continuation` is set. | |
| channel | No | Channel @handle, URL, or UC... id. Required unless `continuation` is set. | |
| continuation | No | Opaque token from a previous response - fetches the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the one bit of useful behavior beyond that: pagination via a continuation token. It says nothing about result limits, ordering, or error behavior when the channel is not found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the usage scenario, then scope, then pagination. No wasted clauses, though the example-heavy opening is near the limit of what's needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read tool with no output schema and full annotation coverage, the description covers selection, scope, and pagination adequately. Return shape and result-count behavior are the only unaddressed items, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented, and the description's note about passing `continuation` for the next page restates what the schema says. Baseline 3 applies since the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (search videos) and a hard scope constraint: 'Searches only inside that channel.' This cleanly distinguishes it from search_youtube and the list_* siblings, which have no within-channel query semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete trigger examples ('what has @hubermanlab said about sleep') and states the scoping condition, which implicitly rules out global search. It does not name search_youtube or list_channel_videos as the alternative for the non-scoped case, so routing is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_youtubeSearch YouTubeARead-onlyInspect
Use this when the user wants to find YouTube videos or channels on a topic, for example 'find lectures on linear algebra' or 'popular videos about the Apollo missions', before reading or comparing them. Returns each result's title, video ID, link, channel, views, length, and upload date. Pass continuation from a previous response for the next page. Does not return what a video says: call get_youtube_transcript with a result's video ID for that.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search query. Required unless `continuation` is set. | |
| search_type | No | 'video' (default) or 'channel'. | |
| continuation | No | Opaque token from a previous response - fetches the next page. Do not construct this yourself. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful context beyond that: the exact fields returned per result, the pagination contract ('Pass continuation from a previous response for the next page'), and the negative capability that it does not return spoken content. It stops short of describing result ordering or rate/credit behavior, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, zero waste, and the trigger condition plus examples are front-loaded ahead of the return-value and alternative-tool notes. Every sentence carries distinct information: trigger, return shape, pagination, and boundary with the transcript sibling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by enumerating returned fields (title, video ID, link, channel, views, length, upload date) and the pagination mechanism. Combined with annotations covering the safety profile, an agent has everything needed to call this correctly and know what comes back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so query, search_type, and continuation are already documented in the schema, including the note not to construct continuation tokens. The description's mention of passing continuation adds a usage cue but no syntax or format detail beyond the schema. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('find YouTube videos or channels') with concrete query examples, and explicitly distinguishes itself from the sibling get_youtube_transcript by scope ('Does not return what a video says'). An agent can route between search and transcript without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition ('when the user wants to find... before reading or comparing them'), two concrete query examples, and names the alternative tool with the exact condition that selects it (call get_youtube_transcript with a result's video ID). Nothing about when to use or not use this tool is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
get_credits
6 tool updates
- First observed
get_channel_latest_videos - First observed
get_youtube_transcript - First observed
list_channel_videos - First observed
list_playlist_videos - First observed
search_channel_videos - First observed
search_youtube
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
YouTube data for AI agents: channels, videos, transcripts, comments, search. Video research.
💯 The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables YouTube content browsing, video searching, and metadata retrieval via the YouTube Data API v3. It also facilitates fetching video transcripts for summarization and analysis within MCP-compatible AI clients.721 npm1MIT
- AlicenseNot gradedqualityAmaintenanceGives MCP clients access to YouTube video transcripts and metadata. It lets AI agents read and summarize video content from a URL.8 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables interaction with YouTube via MCP, including search, transcript retrieval, trending videos, and AI community discovery.MIT
- AlicenseNot gradedqualityBmaintenanceA comprehensive MCP server providing YouTube transcript retrieval, video search, channel browsing, playlist extraction, and upload monitoring for AI agents.10MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.