YouTube Toolbox
py-mcp-youtube-툴박스
YouTube와 상호작용할 수 있는 강력한 도구를 AI 어시스턴트에게 제공하는 MCP 서버로, 비디오 검색, 대본 추출, 댓글 검색 등이 포함됩니다.
개요
py-mcp-youtube-toolbox는 다음과 같은 YouTube 관련 기능을 제공합니다.
고급 필터링 옵션으로 YouTube 동영상 검색
비디오 및 채널에 대한 자세한 정보를 얻으세요
정렬 옵션을 사용하여 비디오 댓글 검색
여러 언어로 비디오 대본과 자막을 추출합니다.
주어진 비디오와 관련된 비디오를 찾으세요
지역별 인기 영상 보기
대본을 기반으로 비디오 콘텐츠 요약 생성
필터링, 검색 및 다중 비디오 기능을 갖춘 고급 대본 분석
Related MCP server: YouTube MCP Server
목차
필수 조건
Python : Python 3.12 이상 설치
YouTube API 키 :
Google Cloud Console 로 이동
새 프로젝트를 만들거나 기존 프로젝트를 선택하세요
YouTube 데이터 API v3 활성화:
"API 및 서비스" > "라이브러리"로 이동하세요.
"YouTube Data API v3"를 검색하여 활성화하세요.
자격 증명을 만듭니다.
"API 및 서비스" > "자격 증명"으로 이동하세요.
"자격 증명 만들기" > "API 키"를 클릭하세요.
API 키를 적어 두세요
설치
Git 복제
지엑스피1
구성
UV 패키지 관리자를 설치하세요:
curl -LsSf https://astral.sh/uv/install.sh | sh가상 환경을 만들고 활성화하세요.
uv venv -p 3.12
source .venv/bin/activate # On MacOS/Linux
# or
.venv\Scripts\activate # On Windows종속성 설치:
uv pip install -r requirements.txt환경 변수:
cp env.example .env
vi .env
# Update with your YouTube API key
YOUTUBE_API_KEY=your_youtube_api_keyDocker 사용
Docker 이미지를 빌드합니다.
docker build -t py-mcp-youtube-toolbox .컨테이너를 실행합니다.
docker run -e YOUTUBE_API_KEY=your_youtube_api_key py-mcp-youtube-toolbox로컬 사용
서버를 실행합니다:
mcp run server.pyMCP 검사기를 실행합니다.
mcp dev server.pyMCP 설정 구성
MCP 설정 파일에 서버 구성을 추가합니다.
클로드 데스크톱 앱
Smithery를 통해 자동으로 설치하려면:
npx -y @smithery/cli install @jikime/py-mcp-youtube-toolbox --client claude수동으로 설치하려면
~/Library/Application Support/Claude/claude_desktop_config.json엽니다.
mcpServers 개체에 다음을 추가합니다.
{
"mcpServers": {
"YouTube Toolbox": {
"command": "/path/to/bin/uv",
"args": [
"--directory",
"/path/to/py-mcp-youtube-toolbox",
"run",
"server.py"
],
"env": {
"YOUTUBE_API_KEY": "your_youtube_api_key"
}
}
}
}커서 IDE
~/.cursor/mcp.json 엽니다
mcpServers 개체에 다음을 추가합니다.
{
"mcpServers": {
"YouTube Toolbox": {
"command": "/path/to/bin/uv",
"args": [
"--directory",
"/path/to/py-mcp-youtube-toolbox",
"run",
"server.py"
],
"env": {
"YOUTUBE_API_KEY": "your_youtube_api_key"
}
}
}
}도커를 위해
{
"mcpServers": {
"YouTube Toolbox": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"-e", "YOUTUBE_API_KEY=your_youtube_api_key",
"py-mcp-youtube-toolbox"
]
}
}
}도구 문서
비디오 도구
search_videos: 고급 필터링 옵션(채널, 길이, 지역 등)을 사용하여 YouTube 동영상을 검색합니다.get_video_details: 특정 YouTube 동영상에 대한 자세한 정보(제목, 채널, 조회수, 좋아요 등)를 가져옵니다.get_video_comments: 정렬 옵션을 사용하여 YouTube 비디오에서 댓글 검색get_related_videos: 특정 YouTube 동영상과 관련된 동영상을 찾습니다.get_trending_videos: 지역별 YouTube 인기 영상 받기
채널 도구
get_channel_details: YouTube 채널에 대한 자세한 정보(이름, 구독자, 조회수 등)를 가져옵니다.
필사 도구
get_video_transcript: 지정된 언어로 된 YouTube 비디오에서 대본/캡션을 추출합니다.get_video_enhanced_transcript: 필터링, 검색 및 다중 비디오 기능을 갖춘 고급 대본 추출
프롬프트 도구
transcript_summary: 사용자 정의 옵션을 사용하여 대본을 기반으로 YouTube 비디오 콘텐츠 요약을 생성합니다.
리소스 도구
youtube://available-youtube-tools: 사용 가능한 모든 YouTube 도구 목록을 가져옵니다.youtube://video/{video_id}: 특정 비디오에 대한 자세한 정보를 가져옵니다.youtube://channel/{channel_id}: 특정 채널에 대한 정보를 가져옵니다.youtube://transcript/{video_id}?language={language}: 특정 비디오의 대본을 가져옵니다.
개발
로컬 테스트를 위해 포함된 클라이언트 스크립트를 사용할 수 있습니다.
# Example: Search videos
uv run client.py search_videos query="MCP" max_results=5
# Example: Get video details
uv run client.py get_video_details video_id=zRgAEIoZEVQ
# Example: Get channel details
uv run client.py get_channel_details channel_id=UCRpOIr-NJpK9S483ge20Pgw
# Example: Get video comments
uv run client.py get_video_comments video_id=zRgAEIoZEVQ max_results=10 order=time
# Example: Get video transcript
uv run client.py get_video_transcript video_id=zRgAEIoZEVQ language=ko
# Example: Get related videos
uv run client.py get_related_videos video_id=zRgAEIoZEVQ max_results=5
# Example: Get trending videos
uv run client.py get_trending_videos region_code=ko max_results=10
# Example: Advanced transcript extraction
uv run client.py get_video_enhanced_transcript video_ids=zRgAEIoZEVQ language=ko format=timestamped include_metadata=true start_time=100 end_time=200 query=에이전트 case_sensitive=true segment_method=equal segment_count=2
# Example: 특허
MIT 라이센스
Available Tools
8 toolsget_channel_detailsB
Get detailed information about a YouTube channel
| Name | Required | Description | Default |
|---|---|---|---|
| channel_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool retrieves 'detailed information' but doesn't specify what details are included (e.g., subscriber count, videos, metadata), whether it's a read-only operation, rate limits, authentication needs, or error handling. This leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with zero wasted words. It's front-loaded with the core purpose and appropriately sized for a simple tool, making it easy to parse quickly without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (1 parameter, no nested objects) and the presence of an output schema (which should define return values), the description is minimally adequate. However, with no annotations and low schema coverage, it lacks context on behavior and parameters that could help the agent use it effectively, leaving room for improvement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 1 parameter with 0% description coverage, so the schema provides no semantic context. The description doesn't add any parameter-specific information beyond implying a 'channel_id' is needed. It doesn't explain what a channel ID is, where to find it, or format requirements, resulting in minimal value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and resource 'detailed information about a YouTube channel', making the purpose immediately understandable. However, it doesn't differentiate this tool from sibling tools like 'get_video_details' or 'get_video_enhanced_transcript', which also retrieve YouTube content information, so it doesn't fully distinguish its specific scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'get_video_details' for video-specific data or 'search_videos' for broader searches, leaving the agent without context for tool selection. There's no indication of prerequisites or typical use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trending_videosB
Get trending videos on YouTube by region
| Name | Required | Description | Default |
|---|---|---|---|
| region_code | No | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool retrieves trending videos by region but lacks details on permissions, rate limits, data freshness, or response format. This is a significant gap for a tool that likely involves external API calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's appropriately sized and front-loaded, clearly stating the tool's purpose without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no annotations, but with an output schema), the description is minimally adequate. The output schema likely covers return values, reducing the burden, but the lack of behavioral details and incomplete parameter semantics leaves room for improvement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'by region,' which aligns with the 'region_code' parameter, adding some context beyond the schema. However, with 0% schema description coverage and 2 parameters, it doesn't fully explain 'max_results' or provide format details for 'region_code,' leaving gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get trending videos') and resource ('on YouTube by region'), providing a specific purpose. However, it doesn't differentiate this tool from sibling tools like 'search_videos' or 'get_related_videos', which could also retrieve videos in some context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when-not scenarios, prerequisites, or compare it to siblings like 'search_videos' for general queries or 'get_ideo_details' for specific videos, leaving usage ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_commentsC
Get comments for a YouTube video
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | ||
| max_results | No | ||
| order | No | relevance | |
| include_replies | No | ||
| page_token | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states 'Get comments' but doesn't clarify if this is a read-only operation, whether it requires API keys or permissions, if there are rate limits, or what the output format looks like. The description is too minimal to provide adequate behavioral context for a tool with 5 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that gets straight to the point with no wasted words. It's appropriately sized for a simple tool name and is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters, 0% schema description coverage, no annotations, but does have an output schema, the description is incomplete. It covers the basic purpose but misses parameter explanations, usage context, and behavioral details. The output schema helps with return values, but the description alone doesn't provide enough context for effective tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, meaning none of the 5 parameters have descriptions in the schema. The tool description only mentions 'video_id' implicitly ('for a YouTube video'), leaving the other 4 parameters (max_results, order, include_replies, page_token) completely undocumented. This fails to compensate for the schema's lack of coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('comments for a YouTube video'), making the purpose immediately understandable. It doesn't explicitly distinguish from sibling tools like 'get_video_details' or 'get_video_transcript', but the focus on comments is specific enough to imply differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_video_details' (which might include comments) or 'search_videos' (which might find videos with certain comments). It also lacks context about prerequisites, such as needing a valid video ID or authentication requirements.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_detailsB
Get detailed information about a YouTube video
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves 'detailed information' but doesn't specify what that includes (e.g., title, duration, view count), whether it requires authentication, rate limits, or error handling. For a read operation with zero annotation coverage, this leaves significant gaps in understanding how the tool behaves beyond its basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded with the core action and resource, making it easy to parse. Every part of the sentence earns its place by conveying essential information, achieving ideal conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (1 parameter, no nested objects) and the presence of an output schema (which should define return values), the description is minimally complete. However, with no annotations and low schema coverage, it lacks details on behavioral aspects like authentication or error cases. For a basic read tool, it's adequate but leaves room for improvement in contextual richness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 1 parameter with 0% description coverage, so the description must compensate. It implies the parameter is a 'video_id' for a YouTube video but doesn't clarify format (e.g., YouTube URL vs. ID string), validation, or examples. Since schema coverage is low, the description adds minimal value beyond what's inferred from the schema property name, meeting the baseline for adequate but incomplete parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('detailed information about a YouTube video'), making the purpose unambiguous. It distinguishes this tool from siblings like 'get_video_comments' or 'get_video_transcript' by focusing on general video metadata rather than specific aspects. However, it doesn't explicitly differentiate from 'get_video_enhanced_transcript' which might also provide detailed information, keeping it from a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings like 'get_related_videos', 'get_trending_videos', and 'search_videos', there's no indication that this is for retrieving metadata of a specific known video ID versus browsing or searching. No exclusions or prerequisites are mentioned, leaving usage context implied at best.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_enhanced_transcriptC
Advanced transcript extraction tool with filtering, search, and multi-video capabilities. Provides rich transcript data for detailed analysis and processing. Features: 1) Extract transcripts from multiple videos; 2) Filter by time ranges; 3) Search within transcripts; 4) Segment transcripts; 5) Format output in different ways; 6) Include video metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| video_ids | Yes | ||
| language | No | ko | |
| start_time | No | ||
| end_time | No | ||
| query | No | ||
| case_sensitive | No | ||
| segment_method | No | equal | |
| segment_count | No | ||
| format | No | timestamped | |
| include_metadata | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. While it lists features (multi-video extraction, filtering, search, segmentation, formatting, metadata inclusion), it doesn't describe important behavioral aspects: whether this is a read-only operation, potential rate limits, authentication requirements, error conditions, or what happens with invalid parameters. The feature list is helpful but incomplete for behavioral understanding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise with two sentences and a numbered feature list. However, the first sentence is somewhat redundant ('Advanced transcript extraction tool' and 'Provides rich transcript data' convey similar ideas). The feature list is helpful but could be more efficiently integrated. Overall, it's adequately structured but not optimally front-loaded with the most critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (10 parameters, no annotations, but with output schema), the description provides a reasonable feature overview but has significant gaps. The output schema existence means return values are documented elsewhere, but the description should better explain the tool's scope, limitations, and relationship to sibling tools. For a complex tool with many parameters and no annotations, more behavioral context would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage for 10 parameters, the description provides some parameter context by listing features that map to parameters: 'Extract transcripts from multiple videos' (video_ids), 'Filter by time ranges' (start_time, end_time), 'Search within transcripts' (query, case_sensitive), 'Segment transcripts' (segment_method, segment_count), 'Format output' (format), 'Include video metadata' (include_metadata). However, it doesn't explain the language parameter or provide details about parameter values, constraints, or interactions between parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Advanced transcript extraction tool with filtering, search, and multi-video capabilities' and specifies it 'Provides rich transcript data for detailed analysis and processing.' This is specific about the verb (extract) and resource (transcripts from videos) with additional capabilities. However, it doesn't explicitly distinguish it from its sibling 'get_video_transcript' - the 'enhanced' aspect is implied but not directly contrasted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. While it lists features like filtering, search, and multi-video capabilities, it doesn't indicate when these advanced features are needed versus the simpler 'get_video_transcript' sibling tool. There's no mention of prerequisites, performance considerations, or specific use cases that would help an agent choose between available transcript tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_transcriptB
Get transcript/captions for a YouTube video
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | ||
| language | No | ko |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action but lacks details on permissions, rate limits, error handling, or response format. For a tool with an output schema, some context is implied, but key behavioral traits are missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It is front-loaded and directly states the tool's purpose, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no annotations, but with an output schema), the description is minimally adequate. The output schema reduces the need to explain return values, but the lack of behavioral context and usage guidelines leaves gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'transcript/captions' and 'YouTube video', which loosely relates to the 'video_id' parameter, but adds minimal semantic value beyond the schema. With 0% schema description coverage, it partially compensates by hinting at the resource, but does not explain parameter roles or the 'language' parameter's purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('transcript/captions for a YouTube video'), making the tool's function immediately understandable. However, it does not distinguish this tool from its sibling 'get_video_enhanced_transcript', which could cause confusion about when to use each one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as 'get_video_enhanced_transcript' or other sibling tools. There is no mention of prerequisites, context, or exclusions, leaving the agent without usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_videosC
Search for YouTube videos with advanced filtering options
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_results | No | ||
| channel_id | No | ||
| order | No | ||
| video_duration | No | ||
| published_after | No | ||
| published_before | No | ||
| video_caption | No | ||
| video_definition | No | ||
| region_code | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool searches with filtering but doesn't mention key traits like whether it's read-only, has rate limits, requires authentication, returns paginated results, or what the output format entails. For a search tool with 10 parameters and no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Search for YouTube videos with advanced filtering options.' It's front-loaded with the core purpose and avoids unnecessary words, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (10 parameters, no annotations, but has an output schema), the description is incomplete. It covers the basic purpose but lacks usage guidelines, behavioral details, and parameter explanations. The presence of an output schema means return values are documented elsewhere, but for a tool with many parameters and no annotations, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning parameter titles (e.g., 'Query', 'Max Results') lack detailed descriptions in the schema. The description adds minimal value by hinting at 'advanced filtering options' but doesn't explain what parameters like 'order', 'video_duration', or 'region_code' mean or how to use them. It fails to compensate for the low schema coverage, leaving most parameters semantically unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search for YouTube videos with advanced filtering options.' It specifies the verb ('search'), resource ('YouTube videos'), and scope ('advanced filtering options'), which is specific and actionable. However, it doesn't explicitly differentiate from sibling tools like 'get_related_videos' or 'get_trending_videos,' which might also involve video retrieval, so it's not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions 'advanced filtering options' but doesn't specify scenarios where this is preferred over siblings like 'get_trending_videos' (for trending content) or 'get_related_videos' (for context-based retrieval). Without such context, the agent lacks clear usage rules.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v1.0.0- Changed
get_channel_details1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_channel_detailsOutput", + "type": "object" +}
- Changed
get_related_videos1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_related_videosOutput", + "type": "object" +}
- Changed
get_trending_videos1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_trending_videosOutput", + "type": "object" +}
- Changed
get_video_comments1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_video_commentsOutput", + "type": "object" +}
- Changed
get_video_details1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_video_detailsOutput", + "type": "object" +}
- Changed
get_video_enhanced_transcript1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_video_enhanced_transcriptOutput", + "type": "object" +}
- Changed
get_video_transcript1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "get_video_transcriptOutput", + "type": "object" +}
- Changed
search_videos1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "additionalProperties": true, + "title": "Result", + "type": "object" + } + }, + "required": [ + "result" + ], + "title": "search_videosOutput", + "type": "object" +}
8 tool updates
- First observed
get_channel_details - First observed
get_related_videos - First observed
get_trending_videos - First observed
get_video_comments - First observed
get_video_details - First observed
get_video_enhanced_transcript - First observed
get_video_transcript - First observed
search_videos
TDQS
Scored across 8 tools
Most tools have distinct purposes, but there is notable overlap between get_video_transcript and get_video_enhanced_transcript, which could cause confusion as both handle transcripts with the enhanced version being a superset. Other tools like get_video_details and get_channel_details are clearly differentiated, but the transcript duplication weakens clarity.
All tool names follow a consistent snake_case pattern with a verb-noun structure (e.g., get_channel_details, search_videos). The naming is predictable and uniform across all eight tools, making it easy for agents to understand and use them without confusion.
With 8 tools, the count is well within the typical 3-15 range for a focused domain like YouTube data retrieval. It feels slightly thin for comprehensive coverage but reasonable for core functionalities such as fetching videos, channels, comments, and transcripts.
The toolset covers key read operations like getting details, searching, and fetching comments/transcripts, but lacks obvious write or management capabilities (e.g., upload, update, delete) that might be expected in a full YouTube toolbox. This creates notable gaps for agents needing to perform more than retrieval tasks.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
An MCP server that integrates with Discord to provide AI-powered features.
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
Driflyte MCP server which lets AI assistants query topic-specific knowledge from web and GitHub.
Related MCP Servers
- AlicenseAqualityBmaintenanceAn MCP server for intelligent YouTube video analysis that provides token-optimized summaries, sentiment analysis, and entity extraction from transcripts. It enables AI assistants to perform video reporting, channel monitoring, and comprehensive YouTube searches through structured data tools.1052Apache 2.0
- AlicenseAqualityDmaintenanceMCP server that provides YouTube video data to AI agents, supporting search, metadata, comments, and transcripts without an API key.517 npmMIT
- AlicenseAqualityDmaintenanceMCP server that lets AI agents search YouTube and fetch transcripts.23MIT
- FlicenseNot gradedqualityDmaintenanceMCP server for YouTube that provides tools to fetch video metadata and transcripts, enabling natural language queries about YouTube videos.2-