youtube-mcp
youtube-mcp
专为频道所有者设计的 OAuth 认证 YouTube MCP。您可以编辑视频元数据、回复并审核评论、管理播放列表、查询频道分析,并通过 ComfyUI 桥接生成或设置 AI 缩略图。它超越了该领域常见的只读 Data API v3 封装器。
核心优势
大多数现有的 YouTube MCP 使用 Data API v3 的 API 密钥,仅能搜索视频、获取公开元数据,且为只读。本项目使用 OAuth 2.0(授权码 + PKCE),因此可以真正对您的频道进行写入操作:更新视频标题、描述和标签,回复评论,审核垃圾信息,管理播放列表。它还调用独立的 YouTube Analytics API 获取频道统计数据,并通过 ComfyUI 生成缩略图,在一次 MCP 调用中将其推送到 YouTube。
Claude, use generate_and_set_thumbnail on video abc123:
prompt: "cyberpunk hacker at keyboard, neon blue and pink, high contrast"ComfyUI 渲染 1280×720 的图像,youtube-mcp 获取字节流,并 POST 到 thumbnails.set。完成。
Related MCP server: yt-fetch
安装
# npx, no install
npx @miller-joe/youtube-mcp --help
# Docker
docker run -p 9120:9120 \
-e YOUTUBE_CLIENT_ID=... \
-e YOUTUBE_CLIENT_SECRET=... \
-e YOUTUBE_TOKEN_FILE=/token/token.json \
-v $PWD/token:/token \
ghcr.io/miller-joe/youtube-mcp:latest设置:Google Cloud 一次性配置(约 10 分钟)
Google 账号及 YouTube 频道。 请使用个人账号,不要使用可能会丢失的 Workspace 账号。
Google Cloud 项目,访问 https://console.cloud.google.com。随意命名(例如
youtube-mcp)。启用 API:
YouTube Data API v3
YouTube Analytics API
OAuth 同意屏幕: 选择“外部”,填写应用名称和支持邮箱。在“范围 (Scopes)”中添加:
youtube.uploadyoutube.force-sslyt-analytics.readonly
保持在 测试 (Testing) 模式。将您自己添加为 测试用户(必需)。作为项目所有者,您的刷新令牌不会过期。
创建 OAuth 客户端 ID: 应用类型 = 桌面应用 (Desktop app)。下载 JSON 文件。
运行交互式认证流程:
npx @miller-joe/youtube-mcp --auth --client-secret-file ./client_secret.json浏览器将打开,您登录绑定 YouTube 频道的 Google 账号并授予请求的范围。成功后,刷新令牌将保存到
~/.config/youtube-mcp/token.json。启动服务器:
npx @miller-joe/youtube-mcp --client-secret-file ./client_secret.json或者通过环境变量提供客户端凭据:
YOUTUBE_CLIENT_SECRET_FILE,或YOUTUBE_CLIENT_ID+YOUTUBE_CLIENT_SECRET。
连接 MCP 客户端
claude mcp add --transport http youtube http://localhost:9120/mcp或者将您的 MCP 网关指向 Streamable HTTP 端点。
配置
CLI 参数 | 环境变量 | 默认值 | 说明 |
|
| (无) | Google OAuth JSON 文件路径 |
|
| (无) | 客户端 ID(替代 secret 文件) |
|
| (无) | 客户端密钥(替代 secret 文件) |
|
|
| 刷新令牌存储路径 |
|
|
| 绑定主机(仅限 HTTP 模式) |
|
|
| 绑定端口(仅限 HTTP 模式) |
|
| (未设置) | 通过 stdio 而非 HTTP 进行 MCP 通信。当被 stdio 优先的 MCP 客户端(如 Claude Desktop, mcp-inspector)作为子进程启动时使用。 |
|
| (未设置,桥接禁用) | 用于桥接工具的 ComfyUI HTTP URL |
(无参数) |
|
| 桥接工具的默认检查点 |
传输方式
服务器默认使用 streamable HTTP(非常适合 Claude Code、MetaMCP、原生 fetch)。传入 --stdio(或设置 MCP_TRANSPORT=stdio)可切换到 stdio 模式,这是 Claude Desktop 和 MCP Inspector 等 stdio 优先客户端所期望的:
// claude_desktop_config.json
{
"mcpServers": {
"youtube": {
"command": "npx",
"args": ["-y", "@miller-joe/youtube-mcp", "--stdio"],
"env": {
"YOUTUBE_CLIENT_SECRET_FILE": "/path/to/client_secret.json",
"YOUTUBE_TOKEN_FILE": "/path/to/token.json"
}
}
}
}Stdio 模式会跳过 OAuth 令牌预检 — 即使没有存储的令牌,服务器也会启动,并在工具调用时抛出认证错误。在将 Claude Desktop 指向它之前,请先在 HTTP 模式下运行 youtube-mcp --auth --client-secret-file <path> 以生成刷新令牌。
工具
视频
list_my_videos: 分页列出已认证频道的上传视频。get_video: 获取单个视频的详细信息。update_video_metadata: 更新标题、描述、标签、类别、隐私设置。delete_video: 永久删除视频。需要confirm_video_title与当前标题完全匹配,以防止误删。
字幕
list_captions: 列出视频的字幕轨道(语言、名称、状态、草稿标记)。upload_caption: 上传 SRT 或 WebVTT 字幕轨道到视频。delete_caption: 删除字幕轨道。
Shorts
list_my_shorts: 在最近上传中查找 Shorts(按时长 ≤60s 过滤)。get_shorts_analytics: 仅限 Shorts 的 YouTube 分析查询 (creatorContentType==SHORTS)。
播放列表
create_playlist: 创建播放列表(默认为私有)。add_to_playlist: 将视频添加到现有播放列表。
评论
list_comments: 获取视频的顶级评论线程。reply_to_comment: 回复顶级评论。moderate_comment: 暂存、批准或拒绝评论。
分析
query_channel_analytics: 带有可选维度和过滤器的日期范围指标查询。
桥接 (当配置了 COMFYUI_URL 时)
generate_and_set_thumbnail: 通过 ComfyUI 生成缩略图并在一次调用中将其设置到视频上。
配额说明
YouTube Data API 免费层级 = 10,000 单位/天。关键操作成本:
videos.list,commentThreads.list: 每个 1 单位。videos.update,comments.insert,thumbnails.set: 每个 50 单位。videos.insert(上传): 1,600 单位,因此在免费层级下每天约可上传 6 个视频。
大多数创作者运营工作流都在免费额度内。
架构
┌────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ MCP client │────▶│ youtube-mcp │────▶│ YouTube APIs │
│ (Claude etc.) │◀────│ (this server) │◀────│ (Data/Analytics)│
└────────────────┘ └────────┬─────────┘ └─────────────────┘
│
│ (bridge tools only)
▼
┌──────────────────┐
│ ComfyUI │
│ (txt2img) │
└──────────────────┘OAuth 刷新令牌在本地缓存,并在过期前即时刷新。桥接工具在内部从 ComfyUI 下载图像字节,因此 ComfyUI 不需要公开可访问。
开发
git clone https://github.com/miller-joe/youtube-mcp
cd youtube-mcp
npm install
npm run dev
npm run build
npm test需要 Node 20+。
路线图
已发布:
视频:列表、获取、更新元数据、带标题匹配确认的删除功能。
字幕:上传、列表、删除。
Shorts:
list_my_shorts(时长过滤)和get_shorts_analytics(creatorContentType==SHORTS)。播放列表:创建、添加。
评论:列表、回复、审核。
分析:频道分析查询。
ComfyUI 缩略图桥接:
generate_and_set_thumbnail。
计划中:
视频上传 (
video_upload),支持断点续传。用于批量历史数据导出的 Reporting API。
许可证
MIT © Joe Miller
支持
如果这为您节省了时间,请考虑支持开发:
Available Tools
15 toolsadd_to_playlistB
Add a video to an existing playlist. Both playlist_id and video_id are YouTube IDs (not URLs).
| Name | Required | Description | Default |
|---|---|---|---|
| playlist_id | Yes | YouTube playlist ID | |
| video_id | Yes | YouTube video ID to add |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description minimally discloses that IDs are not URLs but does not mention whether adding a duplicate is safe, if the operation is idempotent, or any error conditions. A mutation tool should provide more behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. Front-loaded with the core action and key constraint. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite simplicity, the description omits important context for a mutation tool: no mention of success/failure indicators, duplicate handling, or permission requirements. Leaves an AI agent underinformed for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by clarifying that both parameters are YouTube IDs and not URLs, which disambiguates potential misinterpretation. This is a meaningful addition beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a video to an existing playlist') and the resource, distinguishing it from siblings like create_playlist. The clarification that IDs are YouTube IDs (not URLs) further sharpens purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, prerequisites (e.g., playlist existence), or edge cases (e.g., duplicate video). The description is purely operational without directional help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_playlistB
Create a new playlist on the authenticated channel. Default privacy is 'private'.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| description | No | ||
| privacy_status | No | private |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Only mentions default privacy but omits other behaviors like authentication, side effects, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with front-loaded purpose. No redundant info, though could include more detail without bloating.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers main purpose and key default, but lacks details about return value or success response. Adequate for a simple creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 0% coverage; description hints at privacy_status via default but does not explain title or description parameters. Inadequate for a 3-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'create', resource 'playlist', and context 'on authenticated channel'. It also notes the default privacy setting, distinguishing it from siblings like add_to_playlist.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., adding to existing playlist). Lacks when-not or scenario exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_captionA
Delete a caption track by ID. Use list_captions to find the track ID first.
| Name | Required | Description | Default |
|---|---|---|---|
| caption_id | Yes | Caption track ID to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states it deletes a caption track, indicating a destructive operation, but it does not mention authorization, rate limits, error behavior (e.g., if ID does not exist), or whether the deletion is irreversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundant information. The purpose is front-loaded, and the usage guidance follows naturally. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no nested objects), the description covers the essential purpose and prerequisite. However, it could be more complete by noting any irreversible nature or confirmation behavior, but overall it is adequate for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema desccription for caption_id is 'Caption track ID to delete.', which already conveys the parameter's purpose. The tool description adds no additional semantics beyond what the schema provides. Since schema description coverage is 100%, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'delete' and resource 'caption track', clearly stating the action and how to identify the track (by ID). It also distinguishes from siblings by referencing list_captions to find the ID.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises using list_captions first to obtain the track ID, providing a necessary prerequisite. It does not elaborate on when not to use the tool, but for a delete operation, this is reasonable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_videoA
Permanently delete a video. Requires confirm_video_title to match the video's current title exactly — guards against deleting the wrong video by ID. Deletion is irreversible.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to delete. | |
| confirm_video_title | Yes | Exact current title of the video. Must match what YouTube returns to proceed — prevents accidental deletion of the wrong video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses irreversible deletion and the confirmation guard, which is key behavioral context. Without annotations, this provides necessary transparency, though permissions or side effects are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with core purpose, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given simple tool with 2 fully-described parameters and no output schema, the description covers all essential aspects: purpose, safety guard, and irreversibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining why confirm_video_title is needed (prevents accidental deletion) and that deletion is irreversible. Schema descriptions are clear, but description enriches understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Permanently delete a video', using specific verb and resource. Distinguishes from siblings like 'delete_caption' which deletes a different resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance. While the safety guard implies careful usage, alternative tools or conditions are not mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_shorts_analyticsA
Query YouTube Analytics restricted to Shorts for the authenticated channel. Applies filters=creatorContentType==SHORTS on top of the usual start_date/end_date/metrics/dimensions knobs.
| Name | Required | Description | Default |
|---|---|---|---|
| start_date | Yes | YYYY-MM-DD start date (inclusive). | |
| end_date | Yes | YYYY-MM-DD end date (inclusive). | |
| metrics | No | Comma-separated YouTube Analytics metrics. | views,estimatedMinutesWatched,averageViewDuration,subscribersGained |
| dimensions | No | Optional dimensions, e.g. 'day' for a time series. | |
| sort | No | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions authentication ('for the authenticated channel') and the filter, but does not disclose behavioral traits like pagination, rate limits, or return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with main purpose. Efficiently communicates the core functionality without excessive verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides adequate context for a query tool but lacks details about return values and behavioral aspects that could be expected from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so baseline is 3. The description adds that the tool applies a Shorts filter on top of usual knobs, but does not elaborate on parameters like sort or max_results beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it queries YouTube Analytics restricted to Shorts for the authenticated channel, using a specific filter. It distinguishes itself from sibling tools like query_channel_analytics, which likely covers all content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for Shorts analytics by mentioning the filter, and the sibling set includes query_channel_analytics for general analytics. However, it does not explicitly state when not to use or provide alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_videoA
Fetch full details for one video by ID — snippet, status, statistics, duration.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID (the part after v= in the URL) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavior. It indicates a read operation by 'Fetch full details', but does not explicitly state it is read-only, idempotent, or what happens on failure (e.g., missing video). The listing of returned fields adds some transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the action and what is returned, with no extraneous words. It is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple single-parameter schema and no output schema, the description adequately covers the tool's purpose and output. It could mention error behavior or prerequisites (e.g., video must exist), but overall it provides enough context for this straightforward tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds practical guidance: 'the part after v= in the URL', which helps the agent correctly format the video_id parameter, going beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch full details for one video by ID', specifying the resource (video) and action (fetch), and lists the exact details returned (snippet, status, statistics, duration). This distinguishes it from sibling tools like delete_video or update_video_metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing full details for a single video but provides no explicit guidance on when to use this tool versus alternatives (e.g., list_my_videos for multiple videos) or conditions where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_captionsA
List caption tracks on a video with their language, name, status, and whether they are drafts.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to list captions for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states what is listed without mentioning behavioral traits like pagination, read-only nature, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that is clear and to the point, with no unnecessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lists fields returned, which compensates partially for missing output schema; but does not mention output format or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and description adds no additional meaning beyond the schema's description of video_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it lists caption tracks on a video and specifies fields returned (language, name, status, draft status). Distinguishes from sibling tools like delete_caption and upload_caption.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, but context of siblings (e.g., delete_caption) implies it is used for viewing before other actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_commentsB
List top-level comment threads on a video (newest first). Returns comment IDs, authors, text, and like counts.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to list comments from | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as authentication requirements, rate limits, or side effects. It only states ordering and return fields, leaving gaps in transparency for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no extraneous information. It efficiently conveys the core purpose and return value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description mentions return fields, it lacks context on pagination, authentication, rate limits, or whether replies are included. For a simple listing tool, it is adequate but has notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only video_id has a description). The description adds no additional parameter semantics beyond the schema, leaving max_results undefined in meaning and constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List'), resource ('top-level comment threads on a video'), ordering ('newest first'), and return fields ('comment IDs, authors, text, and like counts'). It distinguishes from sibling tools like moderate_comment and reply_to_comment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing comments, but does not provide explicit guidance on when to use this tool versus alternatives (e.g., for replies, nested comments). No conditions or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_my_shortsA
List your recent Shorts — scans the most recent uploads and filters to videos ≤60s. Useful when the Data API doesn't expose a direct Shorts filter.
| Name | Required | Description | Default |
|---|---|---|---|
| max_candidates | No | How many of the most recent uploads to scan. Shorts are detected by duration ≤ 60s after fetching. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains the scanning and filtering behavior, but does not disclose return format, performance implications, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words, efficiently conveying the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks information about the return value (fields, structure), pagination, or ordering, which is important given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The parameter is fully described in the schema; the description adds context about why it scans recent uploads, but adds minimal extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recent Shorts by scanning uploads and filtering by duration ≤60s, distinguishing it from siblings like list_my_videos.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions it is useful when the Data API lacks a direct Shorts filter, but does not specify when not to use it or suggest alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_my_videosA
List videos on the authenticated channel (newest first via the uploads playlist). Returns video IDs, titles, view counts, and privacy status.
| Name | Required | Description | Default |
|---|---|---|---|
| max_results | No | ||
| page_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently states the operation is a read-only list ('list videos'), specifies ordering, and lists return fields. It does not mention pagination behavior or rate limits, but for a straightforward list tool, this is adequate and does not mislead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the main purpose and includes key details (ordering, return fields). Every part earns its place; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does explain return values. However, it omits pagination behavior (how to use page_token) and does not differentiate from sibling tools. It is adequate for a simple list but could be more complete with pagination details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 2 parameters with 0% description coverage, meaning no parameter descriptions exist in the schema. The description adds no explanation for 'max_results' or 'page_token', failing to compensate for the schema gap. The only info is that the list is paginated (implied by page_token) but not stated explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list', the resource 'videos on the authenticated channel', and specifies ordering ('newest first via the uploads playlist'). It also enumerates the returned fields (IDs, titles, view counts, privacy status). This differentiates it from sibling tools like 'list_my_shorts' (shorts) and 'get_video' (single video).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing own videos but provides no explicit guidance on when to use this tool versus alternatives (e.g., 'list_my_shorts' for shorts or 'get_video' for details). No when-not-to-use or prerequisites are mentioned. The context is implied but not clarified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
moderate_commentA
Change the moderation status of a comment: heldForReview (hide pending approval), published (approve), or rejected (delete).
| Name | Required | Description | Default |
|---|---|---|---|
| comment_id | Yes | ||
| moderation_status | Yes | heldForReview hides until approved, published approves, rejected deletes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the effects of each status (hides, approves, deletes) but does not mention required permissions, reversibility, or side effects. This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the tool's purpose and the three possible statuses. It is front-loaded with the action and resource, with no extraneous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation tool with 2 parameters and no output schema, the description covers the core functionality. It explains the purpose, the status options, and their effects. It could mention success/error behavior or authentication needs, but it is fairly complete given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to the 'moderation_status' parameter by explaining the result of each enum value, which goes beyond the schema's brief enum description. However, it does not clarify 'comment_id', which is a simple identifier and may be self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Change the moderation status of a comment' with explicit enumeration of the three possible statuses and their effects. It distinguishes the tool from sibling tools like list_comments or reply_to_comment by focusing solely on moderation actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for moderating a comment but provides no explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites or exclusions. The purpose is clear from context, but the description lacks explicit when/when-not instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_channel_analyticsA
Query YouTube Analytics for the authenticated channel. Returns tabular data — useful for views/watch-time/retention/traffic-source reports. Date-ranged and optionally grouped by dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| start_date | Yes | YYYY-MM-DD (inclusive) | |
| end_date | Yes | YYYY-MM-DD (inclusive) | |
| metrics | No | Comma-separated metric names (see YouTube Analytics API). Defaults cover the most common creator-dashboard stats. | views,estimatedMinutesWatched,averageViewDuration,subscribersGained |
| dimensions | No | Comma-separated dimensions, e.g. 'day', 'video', 'country'. Omit for channel totals. | |
| filters | No | Filter expression, e.g. 'video==VIDEO_ID' to scope to one video, or 'country==US'. | |
| sort | No | Sort spec, e.g. '-views' for descending by views | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns tabular data and supports date-range and optional grouping, but does not mention authentication requirements, rate limits, error handling, or pagination behavior. The description provides some transparency but is incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the primary purpose, and contains no unnecessary words. Every sentence adds value, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters and no output schema, the description is brief. It does not explain the return format beyond 'tabular data', lacks details on error handling or edge cases, and does not cover how to interpret results. This leaves gaps for an agent using the tool in complex scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (86%), so the schema already documents parameters well. The description adds minimal value beyond the schema, only referencing date-range and optional grouping. It does not provide additional semantic meaning for parameters like 'filters' or 'sort'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it queries YouTube Analytics for the authenticated channel, specifying it returns tabular data for common metrics like views and watch-time. The verb 'Query' and resource 'YouTube Analytics for the authenticated channel' are specific, and it distinguishes from siblings like 'get_shorts_analytics' by being for general channel analytics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions the tool is useful for certain reports but does not explicitly say when to use it versus alternatives like 'get_shorts_analytics'. There is no guidance on when not to use it or exclusions, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_commentB
Reply to a top-level comment. Requires youtube.force-ssl scope.
| Name | Required | Description | Default |
|---|---|---|---|
| parent_id | Yes | Comment ID to reply to (top-level comment.id from list_comments) | |
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only notes scope requirement. Does not disclose mutation behavior, idempotency, error responses, or return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences conveying purpose and requirement with no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and minimal description; missing details on expected return, error handling, or post-conditions for a write operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers parent_id with clear description, but text param lacks description. Tool description adds no parameter information beyond schema, not compensating for the 50% coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states verb 'reply' and resource 'top-level comment', with scope requirement. Distinguishes from sibling tools like list_comments or moderate_comment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as moderate_comment or add_to_playlist. Only mentions required OAuth scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_video_metadataA
Update a video's metadata — title, description, tags, category, or privacy. Only provide fields you want changed.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | ||
| title | No | ||
| description | No | ||
| tags | No | ||
| category_id | No | YouTube category ID as a string (e.g. '22' = People & Blogs, '27' = Education, '28' = Science & Tech) | |
| privacy_status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It describes the update operation but does not disclose side effects, auth requirements, rate limits, or whether changes are irreversible. The partial update hint is helpful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy. The purpose is front-loaded, and the partial update instruction is concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema, the description covers the essential update semantics. It implies that only provided fields are changed, which is critical. Could mention that omitted fields remain unchanged, but it's already clear from 'Only provide fields you want changed.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (17% – only category_id has a description). The description compensates by listing the fields and emphasizing partial updates, adding meaning beyond the schema's property names and types. However, it does not explain formatting for tags or category_id further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it updates video metadata, listing specific fields (title, description, tags, category, privacy). This distinguishes it from sibling tools like delete_video or get_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Only provide fields you want changed,' indicating a partial update pattern. It does not explicitly state when to use vs. alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_captionA
Upload a caption track (SRT or WebVTT) to a video. Creates a new track — use a distinct name per language/track, or is_draft=true while iterating.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID the caption belongs to. | |
| language | Yes | BCP-47 language code, e.g. 'en', 'en-US', 'es', 'ja'. Must match a language the video supports. | |
| name | No | Caption track name shown in the player's caption menu. Empty string for the default track. | |
| caption_text | Yes | Caption content as a string (SRT or WebVTT format). Source this from a file or the model's output. | |
| format | No | Content type of caption_text: 'srt' (SubRip, application/x-subrip) or 'vtt' (WebVTT, text/vtt). | srt |
| is_draft | No | Draft captions aren't visible to viewers. Useful while reviewing auto-translations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'creates a new track' implying mutation, but lacks details on idempotency, conflict behavior (e.g., does it overwrite or fail if track with same name+language exists?), authentication needs, or rate limits. This is a significant gap for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two clear, front-loaded sentences with no redundant information. The first sentence states the primary purpose, and the second adds practical usage tips. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of siblings for deletion and listing, the description is fairly complete for a creation tool. It covers formats and draft usage. However, it lacks details on error handling, file size limits, or post-creation behavior, leaving some gaps for a comprehensive understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters have descriptions. The description adds value beyond the schema by providing usage context for `name` and `is_draft`, but for other parameters, it repeats schema info. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it uploads a caption track (SRT or WebVTT) to a video, specifying that it creates a new track. This distinguishes it from siblings like list_captions (listing) and delete_caption (deletion), providing specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable guidance: use a distinct `name` per language/track, or `is_draft=true` while iterating. However, it does not explicitly state when not to use it (e.g., for updates, consider deleting first, referencing the sibling delete_caption), so some exclusions are missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.1.0- First observed
add_to_playlist - First observed
create_playlist - First observed
delete_caption - First observed
delete_video - First observed
get_shorts_analytics - First observed
get_video - First observed
list_captions - First observed
list_comments - First observed
list_my_shorts - First observed
list_my_videos - First observed
moderate_comment - First observed
query_channel_analytics - First observed
reply_to_comment - First observed
update_video_metadata - First observed
upload_caption
TDQS
Scored across 15 tools
Each tool targets a distinct resource and action. Analytics tools are clearly differentiated (general vs Shorts-specific), and video list vs Shorts list is scoped by content type. Comments, captions, and playlists each have their own dedicated tools with no overlap.
All tool names follow a consistent verb_noun snake_case pattern (list_, get_, create_, update_, delete_, upload_, reply_to_, add_to_, moderate_, query_). There are no mixed conventions or cryptic abbreviations, making the naming predictable and easy to reason about.
15 tools cover a broad YouTube management surface without feeling bloated. The set is well-scoped for the server's purpose—video, playlist, caption, comment, and analytics operations—and each tool has a clear role within that scope.
The coverage has significant gaps, especially for playlists: there is no way to list, retrieve, delete, or remove items from playlists, creating dead ends after creation. Caption updates are also absent (only create and delete), although video CRUD and comment moderation are well covered.
Maintenance
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for Google Veo AI video generation
Hosted MCP for YouTube Studio: uploads, metadata, playlists, comments, analytics, captions.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
Related MCP Servers
- FlicenseAqualityDmaintenanceA Model Context Protocol server that enables Claude to interact with YouTube data and functionality through the Claude Desktop application.111-
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.-
- AlicenseAqualityDmaintenanceMCP server for ComfyUI — text-to-image, variations, img2img refine, upscale, image proxy, and workflow runner.15115 npm1MIT
- AlicenseAqualityAmaintenanceA comprehensive MCP server integrating YouTube Data, Analytics, and Reporting APIs, providing 40 tools for channel management, analytics, video publishing, transcripts, SEO, and comments.4021MIT