youtube-research-mcp
Enables AI agents to research YouTube videos by collecting transcripts, metadata, comments, and channel information, with optional YouTube Data API for search and comment analysis.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@youtube-research-mcpGet the transcript and main ideas from https://youtu.be/abc"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
YouTube Research MCP
μ νλΈμ μλ λͺ¨λ μ 보λ₯Ό β μλ§, λκΈ, μ±λ β AI 리μμΉ μμ€λ‘.
AIμκ² μμ²λ§ νλ©΄ β μλ§ μμ§, λκΈ μ¬λ‘ λΆμ, μ¬λ¬ μμ λΉκ΅κΉμ§ μ λΆ AIκ° μ§μ μ²λ¦¬ν©λλ€. API ν€ μμ΄λ μλ§Β·μ±λ λΆμμ μ¦μ μ¬μ©ν μ μμ΅λλ€.
Related MCP server: mcp-server-youtube
μ΄λμ μΈ μ μλμ?
MCP(Model Context Protocol)λ₯Ό μ§μνλ AI ν΄λΌμ΄μΈνΈλΌλ©΄ μ΄λμλ μ¬μ©ν μ μμ΅λλ€.
ν΄λΌμ΄μΈνΈ | μ§μ μ¬λΆ |
β | |
β | |
β | |
β | |
MCP μ§μ ν΄λΌμ΄μΈνΈ μ 체 | β |
μ΄λ° κ² λ©λλ€
AI μ±ν μ°½μ κ·Έλ₯ λ§νλ―μ΄ μ λ ₯νλ©΄ λ©λλ€.
API ν€ μμ΄ λ°λ‘
μ΄ μμ ν΅μ¬λ§ μμ½ν΄μ€: https://www.youtube.com/watch?v=tTw1z10yMCI@fireship μ±λ μ΅κ·Ό μμ 5κ° λΆμν΄μ μμ¦ μ΄λ€ κΈ°μ μ£Όμ λ€λ£¨λμ§ μλ €μ€μ΄ 3κ° μμ λΉκ΅ν΄μ κ°μ μ΄λ€ μ£Όμ₯ νλμ§, 곡ν΅μ Β·μ°¨μ΄μ μ 리ν΄μ€:
https://www.youtube.com/watch?v=aaa
https://www.youtube.com/watch?v=bbb
https://www.youtube.com/watch?v=cccAPI ν€ μμΌλ©΄ λκΈ μ¬λ‘ κΉμ§
μ€λ νκ΅ μ£Όμ λ€λ£¬ μ νλ² μμ 3κ° λκΈκΉμ§ μ‘°μ¬ν΄μ μμ₯ μ΄μλ λΆμκΈ° 체ν¬ν΄μ€μ΄ μ€λ§νΈν° 리뷰 μμ β ν¬λ¦¬μμ΄ν° νκ°λ μ€μ λκΈ λ°μμ΄ μΌλ§λ λ€λ₯Έμ§ λΉκ΅ν΄μ€"AI μμ΄μ νΈ" μμ 4κ° κ²μν΄μ κ°μ μ΄λ€ μ£Όμ₯μΈμ§ λΉκ΅νκ³ , λκΈ λ°μλ μ 리ν΄μ€μ€μ λ‘ μ΄λ»κ² λ΅μ΄ λμ€λμ?
π€ μ λ ₯:
"μ΄ μ€λ§νΈν° 리뷰 μμ ν΅μ¬ μ₯λ¨μ μ 리νκ³ , λκΈμμ κ°μ₯ λ§μ΄ λ°λ³΅λλ λΆλ§ 3κ°μ§ μλ €μ€."
π€ AI μλ΅:
"μμμμ μ μμλ μΉ΄λ©λΌ μ±λ₯κ³Ό λ°°ν°λ¦¬λ₯Ό μ£Όμ μ₯μ μΌλ‘ κΌ½μμ΅λλ€. κ·Έλ¬λ μμ§λ λκΈ λΆμ κ²°κ³Ό μ€μ μ¬μ©μλ€μ΄ κ°μ₯ λ§μ΄ μΈκΈν λΆλ§μ β λ°μ΄ λ¬Έμ , β‘ νΉμ μ±μμμ νλ μ λλ, β’ μΆ©μ μλμμ΅λλ€. μμμ κΈμ μ νκ°μ μ€μ μ¬μ©μ κ²½ν μ¬μ΄μ μ¨λμ°¨κ° μμ΅λλ€."
κΈ°λ₯ μμ½
κΈ°λ₯ | API ν€ μμ΄ | API ν€ μμ λ |
μμ μλ§ μμ§ | β | β |
μμ λ©νλ°μ΄ν° μ‘°ν | β (yt-dlp κ²½μ ) | β |
μ±λ μ΅μ μμ λΆμ | β (yt-dlp κ²½μ ) | β |
ν€μλ κ²μ | β | β |
λκΈ μμ§ λ° μ¬λ‘ λΆμ | β | β |
API μ¬μ©λ μ‘°ν | β | β |
λκΈΒ·κ²μ κΈ°λ₯μ YouTube Data API ν€κ° νμν©λλ€. λ°κΈμ 무λ£, ν루 10,000 μ λ μ 곡.
μ€μΉ
μΆμ² β uvxλ‘ μ€μΉ μμ΄ λ°λ‘ μ¬μ©
uvλ§ μ€μΉνλ©΄ λ³λ νκ²½ μΈν μμ΄ λ°λ‘ μ°κ²°λ©λλ€.
# uv μ€μΉ (μμ§ μλ€λ©΄)
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
# λλ: winget install astral-sh.uv # Windowsλμ β pipμΌλ‘ μ€μΉ
pip install youtube-research-mcpMCP ν΄λΌμ΄μΈνΈ μ°κ²°
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json νμΌμ μΆκ°ν©λλ€.
νμΌμ΄ μμΌλ©΄ μλ‘ λ§λμΈμ. Claude Desktopμ λ¨Όμ ν λ² μ€νν΄μΌ ν΄λκ° μκΉλλ€.
API ν€ μμ΄ (μλ§ + μ±λ λΆμ)
{
"mcpServers": {
"youtube-research": {
"command": "uvx",
"args": ["youtube-research-mcp"]
}
}
}API ν€ μμ λ (μ 체 κΈ°λ₯)
{
"mcpServers": {
"youtube-research": {
"command": "uvx",
"args": ["youtube-research-mcp"],
"env": {
"YOUTUBE_API_KEY": "AIzaSy..."
}
}
}
}Cursor / Windsurf / κΈ°ν MCP ν΄λΌμ΄μΈνΈ
κ° ν΄λΌμ΄μΈνΈμ MCP μ€μ νμΌμ λμΌν λ°©μμΌλ‘ μΆκ°νλ©΄ λ©λλ€. commandμ argsλ λμΌν©λλ€.
μ€μ μ μ₯ ν ν΄λΌμ΄μΈνΈλ₯Ό μμ ν μ’ λ£νκ³ μ¬μμνλ©΄ μ μ©λ©λλ€.
pipμΌλ‘ μ€μΉνλ€λ©΄
"command": "youtube-research-mcp","args": []λ‘ μ€μ νμΈμ.
YouTube API ν€ λ°κΈ (μ ν μ¬ν)
κ²μΒ·λκΈ κΈ°λ₯μ νμν©λλ€. 무λ£μ΄λ©° ν루 10,000 μ λ β μΌλ° μ¬μ©μΌλ‘λ μμ§λμ§ μμ΅λλ€.
Google Cloud Console μ μ
μ νλ‘μ νΈ μμ±
APIs & Services β Library β
YouTube Data API v3κ²μ β EnableAPIs & Services β Credentials β + Create Credentials β API key
μμ±λ ν€ λ³΅μ¬
보μ μ€μ κΆμ₯: Edit API key β API restrictions β YouTube Data API v3λ§ νμ©
μ€μΉν΄λ μμ νκ°μ?
ν μ€ μμ½: λ€, μμ ν©λλ€.
μ΄ μλ²κ° νλ μΌ | |
β | YouTubeμμ μλ§κ³Ό λ©νλ°μ΄ν°λ₯Ό κ°μ Έμ΅λλ€ |
β | κ²°κ³Όλ¬Όμ λ΄ μ»΄ν¨ν°μ λ‘컬 SQLite νμΌμ μΊμν©λλ€ |
β | API ν€λ₯Ό μ 곡ν κ²½μ°μλ§ YouTube Data API v3λ₯Ό νΈμΆν©λλ€ |
β | μμ§ν λ°μ΄ν°λ₯Ό μΈλΆ μλ²λ‘ μ μ‘νμ§ μμ΅λλ€ |
β | LLMΒ·AI APIλ₯Ό νΈμΆνμ§ μμ΅λλ€ |
β | μΊμ λλ ν 리 μΈμ λ‘컬 νμΌμ μ κ·Όνμ§ μμ΅λλ€ |
β | μ Έ λͺ λ Ήμ΄ μ€νμ΄λ μμ€ν μ κ·Όμ νμ§ μμ΅λλ€ |
μλ§κ³Ό λκΈμλ safety_notice νλκ° ν¬ν¨λμ΄ ν둬ννΈ μΈμ μ
μ λ°©μ§ν©λλ€.
μ 체 μμ€ μ½λλ GitHubμ 곡κ°λμ΄ μμ΅λλ€.
μ€κ³ μμΉ
LLM νΈμΆ μμ β λ°μ΄ν° μμ§λ§ λ΄λΉν©λλ€. λΆμΒ·μμ½Β·νλ¨μ AI μ΄μμ€ν΄νΈκ° ν©λλ€.
ν둬ννΈ μΈμ μ λ°©μ΄ β μλ§κ³Ό λκΈμ
safety_noticeλ₯Ό ν¬ν¨ν΄ μΈλΆ μ½ν μΈ μμ λͺ μν©λλ€.API μκΈ νν μμ β μλ§Β·λκΈΒ·κ²μ κ²°κ³Όλ₯Ό SQLiteμ μΊμν΄ μ€λ³΅ νΈμΆμ μ°¨λ¨ν©λλ€.
ν€ μμ΄λ ν΅μ¬ κΈ°λ₯ μ¬μ© β μλ§ μμ§κ³Ό μ±λ λΆμμ yt-dlpλ‘ API ν€ μμ΄ λμν©λλ€.
λꡬ λͺ©λ‘
API ν€ μμ΄ μ¬μ© κ°λ₯
get_transcript β μμ μλ§ κ°μ Έμ€κΈ°
μ΄ μμ ν΅μ¬ λ΄μ©λ§ μμ½ν΄μ€:
https://www.youtube.com/watch?v=tTw1z10yMCIνλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| YouTube URL λλ video ID | νμ |
| μλ§ μΈμ΄ μ°μ μμ (μ: | μλ μ ν |
analyze_videos β μ¬λ¬ μμ ν λ²μ λΆμ
URL λͺ©λ‘μ μ£Όλ©΄ μλ§Β·λ©νλ°μ΄ν°λ₯Ό λ³λ ¬ μμ§ν©λλ€. API ν€κ° μμΌλ©΄ λκΈλ ν¨κ» μμ§ν©λλ€.
μ΄ 3κ° μμ λΆμν΄μ κ°μ μ΄λ€ μ£Όμ₯μΈμ§, 곡ν΅μ κ³Ό μ°¨μ΄μ μ 리ν΄μ€:
https://www.youtube.com/watch?v=aaa
https://www.youtube.com/watch?v=bbb
https://www.youtube.com/watch?v=cccνλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| URL λλ video ID λͺ©λ‘ | νμ |
| μλ§ μΈμ΄ μ°μ μμ | μλ μ ν |
| λκΈ ν¬ν¨ μ¬λΆ (β οΈ API ν€ νμ) |
|
| μμλΉ μ΅λ λκΈ μ |
|
| μλ§ μ΅λ κΈμ μ (0 = μ ν μμ) |
|
analyze_channel β μ±λ λΆμ
μ±λ νΈλ€(@μ±λλͺ
) λλ μ±λ IDλ‘ μ΅μ μμ Nκ°λ₯Ό μμ§Β·λΆμν©λλ€.
@ycombinator μ±λ μ΅κ·Ό μμ 5κ° λ³΄κ³ μ΄λ€ μ€ννΈμ
νΈλ λ λ€λ£¨λμ§ λΆμν΄μ€νλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| μ±λ νΈλ€ λλ ID | νμ |
| μμ§ν μμ μ (μ΅λ 8) |
|
| μ΅μ μμ κΈΈμ΄ (μ΄) |
|
| μ΅λ μμ κΈΈμ΄ (μ΄) |
|
| λκΈ ν¬ν¨ μ¬λΆ (β οΈ API ν€ νμ) |
|
get_capabilities β νμ¬ μ¬μ© κ°λ₯ν κΈ°λ₯ νμΈ
μ§κΈ μ΄λ€ κΈ°λ₯μ μΈ μ μμ΄?API ν€ νμ
search_videos β ν€μλλ‘ μμ κ²μ
"AI agent" κ΄λ ¨ μ΅μ μμ 5κ° κ²μν΄μ€νλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| κ²μμ΄ | νμ |
| μ΅λ κ²°κ³Ό μ |
|
| μ΄ν λ μ§ (YYYY-MM-DD) | μμ |
| μ΄μ λ μ§ (YYYY-MM-DD) | μμ |
| μΌμΈ μ μΈ μ¬λΆ |
|
get_video_comments β λκΈ μμ§
νλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| YouTube URL λλ video ID | νμ |
| μ΅λ λκΈ μ |
|
| μ λ ¬ λ°©μ ( |
|
| λλκΈ ν¬ν¨ μ¬λΆ |
|
collect_video_discussion β μλ§ + λκΈ ν λ²μ
ν¬λ¦¬μμ΄ν° μ£Όμ₯κ³Ό μμ²μ λ°μμ ν λ²μ λΉκ΅ν λ μ μ©ν©λλ€.
μ΄ μμ λ΄μ©μ΄λ λκΈ λ°μ κ°μ΄ λΆμν΄μ€:
https://www.youtube.com/watch?v=tTw1z10yMCIcollect_research_sources β κ²μ β μλ§ λ¬Άμ μμ§
"λ¬μ€νΈ vs κ³ λΉκ΅" μμ 5κ° κ²μν΄μ κ° μμμ΄ μ΄λ€ κ²°λ‘ λ΄λ¦¬λμ§ μ 리ν΄μ€νλΌλ―Έν° | μ€λͺ | κΈ°λ³Έκ° |
| κ²μμ΄ | νμ |
| μμ§ν μμ μ (μ΅λ 8) |
|
| μ΅μ μμ κΈΈμ΄ (μ΄) |
|
| μΌμΈ μ μΈ μ¬λΆ |
|
| μ΅μ μ‘°νμ |
|
collect_research_discussions β κ²μ β μλ§ + λκΈ λ¬Άμ μμ§
κ°μ₯ κ°λ ₯ν 리μμΉ λꡬ. κ²μΒ·μλ§Β·λκΈμ ν λ²μ λ³λ ¬ μμ§ν©λλ€.
"LLM νμΈνλ" μμ 3κ° κ²μν΄μ
ν¬λ¦¬μμ΄ν°λ€μ΄ 곡ν΅μΌλ‘ κ°μ‘°νλ κ², μλ‘ λ€λ₯Έ μ견, λκΈμμ λ°λ³΅λλ μ§λ¬Έ μ 리ν΄μ€get_quota_usage β API μ¬μ©λ μ‘°ν
μ€λ μ¬μ©ν YouTube API μΏΌν°μ λ¨μ μμ νμΈν©λλ€.
μΊμ λμ λ°©μ
κ°μ μμμ μ¬λ¬ λ² λΆμν΄λ μΆκ° API μΏΌν°κ° μλΉλμ§ μμ΅λλ€.
λ°μ΄ν° | μΊμ μ ν¨ κΈ°κ° |
μλ§ | 30μΌ |
λκΈ | 6μκ° |
κ²μ κ²°κ³Ό | 2μκ° |
μμ λ©νλ°μ΄ν° | μꡬ |
μΊμ μμΉ (OS μλ μ ν, YOUTUBE_RESEARCH_CACHE_DB νκ²½ λ³μλ‘ λ³κ²½ κ°λ₯):
macOS:
~/Library/Application Support/youtube-research-mcp/cache.dbWindows:
%APPDATA%\youtube-research-mcp\cache.dbLinux:
~/.local/share/youtube-research-mcp/cache.db
λΌμ΄μ μ€
Available Tools
10 toolsanalyze_channelC
Collect transcripts and comments from a specific YouTube channel's recent videos.
| Name | Required | Description | Default |
|---|---|---|---|
| channel_id_or_handle | Yes | ||
| max_videos | No | ||
| max_comments_per_video | No | ||
| languages | No | ||
| include_segments | No | ||
| include_comments | No | ||
| min_comment_length | No | ||
| min_like_count | No | ||
| force_refresh | No | ||
| min_duration_seconds | No | ||
| max_duration_seconds | No | ||
| published_after | No | ||
| published_before | No | ||
| exclude_shorts | No | ||
| min_view_count | No | ||
| max_transcript_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions data collection (transcripts, comments) but omits any behavioral details such as read-only nature, caching, rate limits, or authentication requirements. The description is too brief given the lack of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is under-specified for the tool's complexity. While concise, it fails to provide necessary context and is not well-structured for a tool with many parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the large parameter count and absence of schema descriptions or annotations, the description is inadequate. It does not explain defaults, required parameters, filtering logic, or any caveats, leaving the agent with insufficient information to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no meaning to the parameters. With 16 parameters, this is a critical gapβthe agent has no guidance on the role of each parameter beyond their names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Collect' and the resources 'transcripts and comments from a specific YouTube channel's recent videos', distinguishing it from single-video tools like get_transcript or get_video_comments. However, 'recent' is ambiguous given the date-range parameters, and sibling differentiation is not explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like analyze_videos or collect_video_discussion. No when-to-use or when-not-to-use context is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_videosC
Collect transcripts, metadata, and optional comments for specific videos.
| Name | Required | Description | Default |
|---|---|---|---|
| urls_or_video_ids | Yes | ||
| languages | No | ||
| max_comments_per_video | No | ||
| comment_order | No | relevance | |
| include_replies | No | ||
| include_comments | No | ||
| include_segments | No | ||
| max_transcript_chars | No | ||
| min_comment_length | No | ||
| min_like_count | No | ||
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states what is collected, but fails to mention read-only nature, authentication needs, rate limits, or behavior on private/missing videos. The description is too brief for a tool with 11 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the purpose. However, it is too minimal for the tool's complexity and does not include additional structured information that would aid an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 parameters, no annotations, no schema descriptions), the description is severely lacking. It does not hint at the output format or behavior, and fails to help an agent decide when to use this tool over siblings like get_transcript or get_video_comments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no information about any of the 11 parameters. Parameters like 'comment_order', 'force_refresh', and 'min_comment_length' are not explained, leaving the agent without necessary semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'collect' and the resource 'specific videos', listing the data types collected (transcripts, metadata, optional comments). It is specific and implies a combined functionality, though it does not explicitly differentiate from sibling tools like get_transcript or get_video_comments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites, exclusions, or contexts where other tools like search_videos or get_transcript would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collect_research_discussionsC
Search YouTube and collect transcript plus comments for each usable video.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_videos | No | ||
| max_comments_per_video | No | ||
| min_duration_seconds | No | ||
| max_duration_seconds | No | ||
| languages | No | ||
| published_after | No | ||
| published_before | No | ||
| comment_order | No | relevance | |
| include_replies | No | ||
| exclude_shorts | No | ||
| min_view_count | No | ||
| max_per_channel | No | ||
| include_segments | No | ||
| max_transcript_chars | No | ||
| min_comment_length | No | ||
| min_like_count | No | ||
| search_order | No | relevance | |
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only states basic function without mentioning rate limits, auth requirements, or what 'usable' means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (one sentence), but it sacrifices clarity and completeness. It would benefit from additional detail while remaining efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (19 parameters, output schema), the description is incomplete. It omits details about output format, filtering logic, and edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no semantic information about the 19 parameters. It fails to clarify the purpose of key parameters like query, max_videos, or filtering options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it searches YouTube and collects transcripts and comments, but 'usable video' is vague and does not differentiate from siblings like 'collect_video_discussion' or 'get_transcript'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidance is provided. The description does not indicate when to use this tool over alternatives like 'search_videos' or 'collect_research_sources'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collect_research_sourcesC
Search YouTube and collect transcript-ready sources for analysis.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_videos | No | ||
| min_duration_seconds | No | ||
| max_duration_seconds | No | ||
| languages | No | ||
| published_after | No | ||
| published_before | No | ||
| exclude_shorts | No | ||
| min_view_count | No | ||
| max_per_channel | No | ||
| include_segments | No | ||
| max_transcript_chars | No | ||
| search_order | No | relevance | |
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only says 'collect transcript-ready sources' without detailing quota usage, filtering logic, or what 'transcript-ready' entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Overly concise for a tool with 14 parameters; one sentence is insufficient to convey necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing context on output format, how sources are collected, and how it differs from similar tools. Incomplete for a complex tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain any of the 14 parameters beyond the implicit 'query'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (search, collect) and resource (YouTube, transcript-ready sources), and distinguishes from siblings like search_videos and collect_research_discussions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., search_videos, get_transcript). No when-not or context provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collect_video_discussionC
Collect one video's transcript plus comments for discussion analysis.
| Name | Required | Description | Default |
|---|---|---|---|
| url_or_video_id | Yes | ||
| languages | No | ||
| max_comments | No | ||
| comment_order | No | relevance | |
| include_replies | No | ||
| include_segments | No | ||
| max_transcript_chars | No | ||
| min_comment_length | No | ||
| min_like_count | No | ||
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. The description is vague, mentioning only that it collects transcript and comments. It does not disclose any behavioral traits such as caching, rate limits, or side effects, even though the schema includes a force_refresh parameter that implies caching behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise and front-loaded with the main purpose. However, given the tool's complexity with 10 parameters, it is too brief and omits necessary detail, making it under-specified rather than appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 10 parameters and an output schema, but the description only covers the basic function. It does not explain the output, how to use parameters, or provide any context for configuration. This is insufficient for an agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning no parameter descriptions exist in the schema. The description adds no meaning to the 10 parameters; it only states the general action. This fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'collect' and the resource 'one video's transcript plus comments', with the purpose 'for discussion analysis'. It is specific and distinguishes from sibling tools like get_transcript and get_video_comments by combining both.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives like get_transcript or get_video_comments. No prerequisites or context of use are mentioned, leaving the agent without direction for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_capabilitiesA
Return which tools are available based on current configuration.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It mentions 'based on current configuration' but does not elaborate on how configuration affects results, whether there are side effects, or any dynamic behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of a single sentence with ten words. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description does not need to explain return values. However, it lacks details on whether the result is cached or dynamic, but overall it is adequate for a simple discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the description adds no parameter-specific detail. With 0 parameters, the baseline is 4, and the description appropriately explains the tool's purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Return' and identifies the resource as 'which tools are available based on current configuration.' It clearly distinguishes from sibling tools like search_videos or get_quota_usage, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives, nor does it mention any prerequisites or context for its use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_quota_usageA
Return today's YouTube Data API quota usage and remaining estimate.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description indicates a read-only, non-destructive operation. It clearly states what is returned (quota usage and remaining estimate), but could mention potential limitations (e.g., estimate accuracy).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no redundant information, perfectly front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool is simple and has an output schema, the description could be more complete by advising use before other API calls or noting that this is the only quota-checking tool among siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so baseline is 4. The description adds no parameter info, which is acceptable since there are none to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and clearly identifies the resource ('today's YouTube Data API quota usage and remaining estimate'), which distinguishes it from sibling tools that focus on analysis and collection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking quota, but does not explicitly state when to use it or mention alternatives. While it is likely the only tool for this purpose, no guidance on exclusions or prerequisites is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptC
Fetch a YouTube transcript by URL or video ID.
| Name | Required | Description | Default |
|---|---|---|---|
| url_or_video_id | Yes | ||
| languages | No | ||
| preserve_formatting | No | ||
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states that the tool 'fetches' a transcript, implying a read operation, but fails to mention auth requirements, rate limits, or behavior when a transcript is unavailable. The description is too sparse for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it omits critical details that would fit in a few more sentences. It is front-loaded with the core action, but the brevity compromises completeness without earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters and an output schema, the description is incomplete. It does not explain the role of optional parameters like 'languages' or 'preserve_formatting', nor does it describe any nuances of transcript retrieval (e.g., auto-generated vs manual). The output schema exists but the description adds no context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description provides no additional meaning for any parameter. Although the schema defines parameters like 'languages' and 'preserve_formatting', the description only mentions the required input. The agent gets no guidance on how to use optional parameters effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Fetch', the resource 'YouTube transcript', and the input method 'by URL or video ID'. This distinguishes it from sibling tools like analyze_channel or search_videos, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal guidance on when to use this tool. It only mentions input options (URL or video ID) but does not specify when to use this vs alternatives, nor does it mention any prerequisites or limitations. This leaves the agent with insufficient context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_commentsC
Fetch top-level YouTube comments for one video.
| Name | Required | Description | Default |
|---|---|---|---|
| url_or_video_id | Yes | ||
| max_comments | No | ||
| order | No | relevance | |
| include_replies | No | ||
| min_comment_length | No | ||
| min_like_count | No | ||
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description only says 'Fetch', which implies a read operation, but no details on quotas, authentication, rate limits, or caching behavior are provided. Since no annotations exist, the description carries the full burden and fails to disclose these aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (one sentence) but lacks essential parameter explanations and usage context. It is just barely adequate but not optimally structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no parameter descriptions, and an output schema presumed present but not explained, the description is insufficient for the complexity. It fails to provide complete guidance for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the tool description adds no meaning to any of the 7 parameters. The agent has no guidance on what parameters like min_comment_length or force_refresh do.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (Fetch) and resource (top-level YouTube comments for one video). It is specific about scope (one video) but does not differentiate from sibling tools like collect_video_discussion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, limitations, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_videosC
Search YouTube videos and return structured metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_results | No | ||
| published_after | No | ||
| published_before | No | ||
| exclude_shorts | No | ||
| search_order | No | relevance | |
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so description carries full burden. It only states 'search' and 'return structured metadata', omitting behavioral traits like rate limits, pagination, caching, or any side effects. Very limited transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise but overly brief for a tool with 7 parameters and no schema descriptions. The front-loading is minimal; it could be structured better with key details upfront.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (7 parameters, no schema descriptions, no annotations) and the existence of an output schema, the description is incomplete. It does not explain optional parameters like published_after or search_order, leaving the agent unsupported.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Description adds no meaning to any of the 7 parameters. Schema description coverage is 0%, meaning the schema itself provides no parameter descriptions. The description does not compensate; it mentions only the main action, leaving all parameters uninterpretable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool searches YouTube videos and returns structured metadata. It distinguishes from siblings like analyze_channel or analyze_videos, which imply deeper analysis rather than search-and-return. However, it could be more specific about the metadata structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. No mention of prerequisites, exclusions, or context for choosing this over similar tools like collect_research_sources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
analyze_channel - First observed
analyze_videos - First observed
collect_research_discussions - First observed
collect_research_sources - First observed
collect_video_discussion - First observed
get_capabilities - First observed
get_quota_usage - First observed
get_transcript - First observed
get_video_comments - First observed
search_videos
TDQS
Scored across 10 tools
Several tools have overlapping purposes, such as analyze_channel, analyze_videos, collect_research_discussions, collect_research_sources, and collect_video_discussion, all dealing with transcripts and comments. This could confuse an agent about which tool to use.
Tool names use a mix of verbs (analyze, collect, get, search) and the verb_noun pattern is not consistently applied. For example, 'analyze_channel' and 'analyze_videos' are similar but 'get_capabilities' and 'search_videos' follow a different pattern.
With 10 tools, the server covers a reasonable scope for YouTube research without being overly complex. The count is appropriate for the domain.
The tool set covers the core workflow of searching, fetching transcripts and comments, and analyzing channels. Minor gaps exist, such as missing direct channel metadata retrieval or playlist support, but overall it is fairly complete.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents. No signup.
YouTube public video, comment, reply, channel, search, and speech-to-text transcript tools.
Related MCP Servers
- AlicenseBqualityFmaintenanceEnables AI assistants to analyze and summarize YouTube videos by extracting captions, subtitles, and comprehensive metadata including title, description, and duration in multiple languages.16661MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to search videos, read channels, browse playlists, fetch comments, and get transcripts from YouTube using the YouTube Data API v3 and InnerTube API for captions.2GPL 3.0
- AlicenseAqualityFmaintenanceEnables AI assistants to extract YouTube video metadata, subtitles, and top comments without downloading videos.36MIT
- AlicenseNot gradedqualityDmaintenanceExtracts captions, metadata, and descriptions from YouTube videos to enable AI assistants to summarize their content.5MIT