youtube-mcp
youtube-mcp
채널 소유자를 위한 OAuth 인증 기반 YouTube MCP입니다. 동영상 메타데이터 수정, 댓글 답글 및 관리, 재생목록 관리, 채널 분석 쿼리, 그리고 ComfyUI 브릿지를 통한 AI 썸네일 생성 및 설정이 가능합니다. 이 분야의 주류인 읽기 전용 Data API v3 래퍼를 넘어선 기능을 제공합니다.
소개
대부분의 기존 YouTube MCP는 Data API v3에 대해 API 키를 사용하며, 동영상 검색 및 공개 메타데이터 조회와 같은 읽기 전용 기능만 제공합니다. 이 서버는 OAuth 2.0(Authorization Code + PKCE)을 사용하여 채널에 직접 쓰기 작업을 수행할 수 있습니다. 동영상 제목, 설명, 태그 업데이트, 댓글 답글, 스팸 관리, 재생목록 관리가 가능합니다. 또한 별도의 YouTube Analytics API를 통해 채널 통계를 조회하며, ComfyUI를 통해 썸네일을 생성하고 단일 MCP 호출로 YouTube에 업로드할 수 있습니다.
Claude, use generate_and_set_thumbnail on video abc123:
prompt: "cyberpunk hacker at keyboard, neon blue and pink, high contrast"ComfyUI가 1280×720 이미지를 렌더링하면, youtube-mcp가 바이트를 가져와 thumbnails.set으로 POST합니다. 완료입니다.
Related MCP server: yt-fetch
설치
# npx, no install
npx @miller-joe/youtube-mcp --help
# Docker
docker run -p 9120:9120 \
-e YOUTUBE_CLIENT_ID=... \
-e YOUTUBE_CLIENT_SECRET=... \
-e YOUTUBE_TOKEN_FILE=/token/token.json \
-v $PWD/token:/token \
ghcr.io/miller-joe/youtube-mcp:latest설정: Google Cloud 1회성 작업 (~10분)
Google 계정 및 YouTube 채널. 잃어버릴 위험이 있는 워크스페이스 계정이 아닌 개인 계정을 사용하세요.
Google Cloud 프로젝트를 https://console.cloud.google.com 에서 생성합니다. 이름은 원하는 대로 지정하세요 (예:
youtube-mcp).API 활성화:
YouTube Data API v3
YouTube Analytics API
OAuth 동의 화면: 외부(External) 유형, 앱 이름, 지원 이메일을 설정합니다. 범위(Scopes)에 다음을 추가하세요:
youtube.uploadyoutube.force-sslyt-analytics.readonly
테스트 모드를 유지하세요. 본인을 테스트 사용자로 추가해야 합니다(필수). 프로젝트 소유자로서 리프레시 토큰은 만료되지 않습니다.
OAuth 클라이언트 ID 생성: 애플리케이션 유형 = 데스크톱 앱. JSON 파일을 다운로드하세요.
대화형 인증 흐름 실행:
npx @miller-joe/youtube-mcp --auth --client-secret-file ./client_secret.json브라우저가 열리면 YouTube 채널과 연결된 Google 계정으로 로그인하고 요청된 범위를 승인하세요. 성공하면 리프레시 토큰이
~/.config/youtube-mcp/token.json에 저장됩니다.서버 시작:
npx @miller-joe/youtube-mcp --client-secret-file ./client_secret.json또는 환경 변수를 통해 클라이언트 자격 증명을 제공하세요:
YOUTUBE_CLIENT_SECRET_FILE, 또는YOUTUBE_CLIENT_ID+YOUTUBE_CLIENT_SECRET.
MCP 클라이언트 연결
claude mcp add --transport http youtube http://localhost:9120/mcp또는 MCP 게이트웨이를 Streamable HTTP 엔드포인트로 지정하세요.
구성
CLI 플래그 | 환경 변수 | 기본값 | 참고 |
|
| (없음) | Google OAuth JSON 경로 |
|
| (없음) | 비밀 파일 대신 사용 |
|
| (없음) | 비밀 파일 대신 사용 |
|
|
| 리프레시 토큰 저장소 |
|
|
| 바인드 호스트 (HTTP 모드 전용) |
|
|
| 바인드 포트 (HTTP 모드 전용) |
|
| (설정 안 됨) | HTTP 대신 stdio를 통해 MCP 통신. stdio 우선 MCP 클라이언트(Claude Desktop, mcp-inspector)가 서브프로세스로 실행할 때 사용. |
|
| (설정 안 됨, 브릿지 비활성화) | 브릿지 도구용 ComfyUI HTTP URL |
(플래그 없음) |
|
| 브릿지 도구용 기본 체크포인트 |
전송 방식
서버는 기본적으로 Streamable HTTP를 사용합니다(Claude Code, MetaMCP, 원시 fetch에 적합). --stdio를 전달하거나 MCP_TRANSPORT=stdio를 설정하여 stdio 모드로 전환할 수 있습니다. 이는 Claude Desktop 및 MCP Inspector와 같은 stdio 우선 클라이언트가 기대하는 방식입니다:
// claude_desktop_config.json
{
"mcpServers": {
"youtube": {
"command": "npx",
"args": ["-y", "@miller-joe/youtube-mcp", "--stdio"],
"env": {
"YOUTUBE_CLIENT_SECRET_FILE": "/path/to/client_secret.json",
"YOUTUBE_TOKEN_FILE": "/path/to/token.json"
}
}
}
}Stdio 모드는 OAuth 토큰 사전 확인을 건너뜁니다. 서버는 저장된 토큰 없이도 부팅되며 도구 호출 시점에 인증 오류를 발생시킵니다. Claude Desktop을 연결하기 전에 HTTP 모드에서 youtube-mcp --auth --client-secret-file <path>를 한 번 실행하여 리프레시 토큰을 생성하세요.
도구
동영상
list_my_videos: 인증된 채널의 업로드 목록을 페이지 단위로 가져옵니다.get_video: 동영상 하나에 대한 상세 정보를 가져옵니다.update_video_metadata: 제목, 설명, 태그, 카테고리, 공개 범위를 수정합니다.delete_video: 동영상을 영구 삭제합니다. 잘못된 동영상 삭제를 방지하기 위해confirm_video_title이 현재 제목과 정확히 일치해야 합니다.
자막
list_captions: 동영상의 자막 트랙 목록(언어, 이름, 상태, 초안 플래그)을 가져옵니다.upload_caption: SRT 또는 WebVTT 자막 트랙을 동영상에 업로드합니다.delete_caption: 자막 트랙을 삭제합니다.
쇼츠(Shorts)
list_my_shorts: 최근 업로드에서 쇼츠를 찾습니다(길이 60초 이하로 필터링).get_shorts_analytics: 쇼츠로 제한된 YouTube 분석 쿼리(creatorContentType==SHORTS).
재생목록
create_playlist: 재생목록을 생성합니다(기본값 비공개).add_to_playlist: 기존 재생목록에 동영상을 추가합니다.
댓글
list_comments: 동영상의 최상위 댓글 스레드를 가져옵니다.reply_to_comment: 최상위 댓글에 답글을 답니다.moderate_comment: 댓글을 보류, 승인 또는 거부합니다.
분석
query_channel_analytics: 선택적 차원 및 필터가 포함된 기간별 지표를 조회합니다.
브릿지 (COMFYUI_URL이 구성된 경우)
generate_and_set_thumbnail: ComfyUI를 통해 썸네일을 생성하고 단일 호출로 동영상에 설정합니다.
할당량 참고
YouTube Data API 무료 티어 = 일일 10,000 유닛. 주요 작업 비용:
videos.list,commentThreads.list: 각 1 유닛.videos.update,comments.insert,thumbnails.set: 각 50 유닛.videos.insert(업로드): 1,600 유닛 (무료 티어에서 하루 약 6회 업로드 가능).
대부분의 크리에이터 운영 워크플로우는 무료 한도 내에서 충분히 처리됩니다.
┌────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ MCP client │────▶│ youtube-mcp │────▶│ YouTube APIs │
│ (Claude etc.) │◀────│ (this server) │◀────│ (Data/Analytics)│
└────────────────┘ └────────┬─────────┘ └─────────────────┘
│
│ (bridge tools only)
▼
┌──────────────────┐
│ ComfyUI │
│ (txt2img) │
└──────────────────┘OAuth 리프레시 토큰은 로컬에 캐시되며 만료 직전에 갱신됩니다. 브릿지 도구는 ComfyUI에서 이미지 바이트를 내부적으로 다운로드하므로 ComfyUI가 공개적으로 접근 가능할 필요는 없습니다.
개발
git clone https://github.com/miller-joe/youtube-mcp
cd youtube-mcp
npm install
npm run dev
npm run build
npm testNode 20+가 필요합니다.
로드맵
구현 완료:
동영상: 목록, 조회, 메타데이터 업데이트, 제목 일치 확인을 통한 삭제.
자막: 업로드, 목록, 삭제.
쇼츠:
list_my_shorts(길이 필터) 및get_shorts_analytics(creatorContentType==SHORTS).재생목록: 생성, 추가.
댓글: 목록, 답글, 관리.
분석: 채널 분석 쿼리.
ComfyUI 썸네일 브릿지:
generate_and_set_thumbnail.
계획 중:
재개 가능한 업로드를 지원하는 동영상 업로드(
video_upload).대량 기록 데이터 내보내기를 위한 보고 API.
라이선스
MIT © Joe Miller
지원
이 도구가 시간을 절약해 주었다면 개발 지원을 고려해 주세요:
Available Tools
15 toolsadd_to_playlistB
Add a video to an existing playlist. Both playlist_id and video_id are YouTube IDs (not URLs).
| Name | Required | Description | Default |
|---|---|---|---|
| playlist_id | Yes | YouTube playlist ID | |
| video_id | Yes | YouTube video ID to add |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description minimally discloses that IDs are not URLs but does not mention whether adding a duplicate is safe, if the operation is idempotent, or any error conditions. A mutation tool should provide more behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. Front-loaded with the core action and key constraint. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite simplicity, the description omits important context for a mutation tool: no mention of success/failure indicators, duplicate handling, or permission requirements. Leaves an AI agent underinformed for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by clarifying that both parameters are YouTube IDs and not URLs, which disambiguates potential misinterpretation. This is a meaningful addition beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a video to an existing playlist') and the resource, distinguishing it from siblings like create_playlist. The clarification that IDs are YouTube IDs (not URLs) further sharpens purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, prerequisites (e.g., playlist existence), or edge cases (e.g., duplicate video). The description is purely operational without directional help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_playlistB
Create a new playlist on the authenticated channel. Default privacy is 'private'.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| description | No | ||
| privacy_status | No | private |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Only mentions default privacy but omits other behaviors like authentication, side effects, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with front-loaded purpose. No redundant info, though could include more detail without bloating.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers main purpose and key default, but lacks details about return value or success response. Adequate for a simple creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 0% coverage; description hints at privacy_status via default but does not explain title or description parameters. Inadequate for a 3-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'create', resource 'playlist', and context 'on authenticated channel'. It also notes the default privacy setting, distinguishing it from siblings like add_to_playlist.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., adding to existing playlist). Lacks when-not or scenario exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_captionA
Delete a caption track by ID. Use list_captions to find the track ID first.
| Name | Required | Description | Default |
|---|---|---|---|
| caption_id | Yes | Caption track ID to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states it deletes a caption track, indicating a destructive operation, but it does not mention authorization, rate limits, error behavior (e.g., if ID does not exist), or whether the deletion is irreversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundant information. The purpose is front-loaded, and the usage guidance follows naturally. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no nested objects), the description covers the essential purpose and prerequisite. However, it could be more complete by noting any irreversible nature or confirmation behavior, but overall it is adequate for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema desccription for caption_id is 'Caption track ID to delete.', which already conveys the parameter's purpose. The tool description adds no additional semantics beyond what the schema provides. Since schema description coverage is 100%, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'delete' and resource 'caption track', clearly stating the action and how to identify the track (by ID). It also distinguishes from siblings by referencing list_captions to find the ID.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises using list_captions first to obtain the track ID, providing a necessary prerequisite. It does not elaborate on when not to use the tool, but for a delete operation, this is reasonable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_videoA
Permanently delete a video. Requires confirm_video_title to match the video's current title exactly — guards against deleting the wrong video by ID. Deletion is irreversible.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to delete. | |
| confirm_video_title | Yes | Exact current title of the video. Must match what YouTube returns to proceed — prevents accidental deletion of the wrong video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses irreversible deletion and the confirmation guard, which is key behavioral context. Without annotations, this provides necessary transparency, though permissions or side effects are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with core purpose, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given simple tool with 2 fully-described parameters and no output schema, the description covers all essential aspects: purpose, safety guard, and irreversibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining why confirm_video_title is needed (prevents accidental deletion) and that deletion is irreversible. Schema descriptions are clear, but description enriches understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Permanently delete a video', using specific verb and resource. Distinguishes from siblings like 'delete_caption' which deletes a different resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance. While the safety guard implies careful usage, alternative tools or conditions are not mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_shorts_analyticsA
Query YouTube Analytics restricted to Shorts for the authenticated channel. Applies filters=creatorContentType==SHORTS on top of the usual start_date/end_date/metrics/dimensions knobs.
| Name | Required | Description | Default |
|---|---|---|---|
| start_date | Yes | YYYY-MM-DD start date (inclusive). | |
| end_date | Yes | YYYY-MM-DD end date (inclusive). | |
| metrics | No | Comma-separated YouTube Analytics metrics. | views,estimatedMinutesWatched,averageViewDuration,subscribersGained |
| dimensions | No | Optional dimensions, e.g. 'day' for a time series. | |
| sort | No | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions authentication ('for the authenticated channel') and the filter, but does not disclose behavioral traits like pagination, rate limits, or return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with main purpose. Efficiently communicates the core functionality without excessive verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides adequate context for a query tool but lacks details about return values and behavioral aspects that could be expected from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so baseline is 3. The description adds that the tool applies a Shorts filter on top of usual knobs, but does not elaborate on parameters like sort or max_results beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it queries YouTube Analytics restricted to Shorts for the authenticated channel, using a specific filter. It distinguishes itself from sibling tools like query_channel_analytics, which likely covers all content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for Shorts analytics by mentioning the filter, and the sibling set includes query_channel_analytics for general analytics. However, it does not explicitly state when not to use or provide alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_videoA
Fetch full details for one video by ID — snippet, status, statistics, duration.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID (the part after v= in the URL) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavior. It indicates a read operation by 'Fetch full details', but does not explicitly state it is read-only, idempotent, or what happens on failure (e.g., missing video). The listing of returned fields adds some transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the action and what is returned, with no extraneous words. It is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple single-parameter schema and no output schema, the description adequately covers the tool's purpose and output. It could mention error behavior or prerequisites (e.g., video must exist), but overall it provides enough context for this straightforward tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds practical guidance: 'the part after v= in the URL', which helps the agent correctly format the video_id parameter, going beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch full details for one video by ID', specifying the resource (video) and action (fetch), and lists the exact details returned (snippet, status, statistics, duration). This distinguishes it from sibling tools like delete_video or update_video_metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing full details for a single video but provides no explicit guidance on when to use this tool versus alternatives (e.g., list_my_videos for multiple videos) or conditions where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_captionsA
List caption tracks on a video with their language, name, status, and whether they are drafts.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to list captions for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states what is listed without mentioning behavioral traits like pagination, read-only nature, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that is clear and to the point, with no unnecessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lists fields returned, which compensates partially for missing output schema; but does not mention output format or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and description adds no additional meaning beyond the schema's description of video_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it lists caption tracks on a video and specifies fields returned (language, name, status, draft status). Distinguishes from sibling tools like delete_caption and upload_caption.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, but context of siblings (e.g., delete_caption) implies it is used for viewing before other actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_commentsB
List top-level comment threads on a video (newest first). Returns comment IDs, authors, text, and like counts.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID to list comments from | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as authentication requirements, rate limits, or side effects. It only states ordering and return fields, leaving gaps in transparency for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no extraneous information. It efficiently conveys the core purpose and return value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description mentions return fields, it lacks context on pagination, authentication, rate limits, or whether replies are included. For a simple listing tool, it is adequate but has notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only video_id has a description). The description adds no additional parameter semantics beyond the schema, leaving max_results undefined in meaning and constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List'), resource ('top-level comment threads on a video'), ordering ('newest first'), and return fields ('comment IDs, authors, text, and like counts'). It distinguishes from sibling tools like moderate_comment and reply_to_comment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing comments, but does not provide explicit guidance on when to use this tool versus alternatives (e.g., for replies, nested comments). No conditions or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_my_shortsA
List your recent Shorts — scans the most recent uploads and filters to videos ≤60s. Useful when the Data API doesn't expose a direct Shorts filter.
| Name | Required | Description | Default |
|---|---|---|---|
| max_candidates | No | How many of the most recent uploads to scan. Shorts are detected by duration ≤ 60s after fetching. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains the scanning and filtering behavior, but does not disclose return format, performance implications, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words, efficiently conveying the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks information about the return value (fields, structure), pagination, or ordering, which is important given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The parameter is fully described in the schema; the description adds context about why it scans recent uploads, but adds minimal extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recent Shorts by scanning uploads and filtering by duration ≤60s, distinguishing it from siblings like list_my_videos.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions it is useful when the Data API lacks a direct Shorts filter, but does not specify when not to use it or suggest alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_my_videosA
List videos on the authenticated channel (newest first via the uploads playlist). Returns video IDs, titles, view counts, and privacy status.
| Name | Required | Description | Default |
|---|---|---|---|
| max_results | No | ||
| page_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently states the operation is a read-only list ('list videos'), specifies ordering, and lists return fields. It does not mention pagination behavior or rate limits, but for a straightforward list tool, this is adequate and does not mislead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the main purpose and includes key details (ordering, return fields). Every part earns its place; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does explain return values. However, it omits pagination behavior (how to use page_token) and does not differentiate from sibling tools. It is adequate for a simple list but could be more complete with pagination details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 2 parameters with 0% description coverage, meaning no parameter descriptions exist in the schema. The description adds no explanation for 'max_results' or 'page_token', failing to compensate for the schema gap. The only info is that the list is paginated (implied by page_token) but not stated explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list', the resource 'videos on the authenticated channel', and specifies ordering ('newest first via the uploads playlist'). It also enumerates the returned fields (IDs, titles, view counts, privacy status). This differentiates it from sibling tools like 'list_my_shorts' (shorts) and 'get_video' (single video).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing own videos but provides no explicit guidance on when to use this tool versus alternatives (e.g., 'list_my_shorts' for shorts or 'get_video' for details). No when-not-to-use or prerequisites are mentioned. The context is implied but not clarified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
moderate_commentA
Change the moderation status of a comment: heldForReview (hide pending approval), published (approve), or rejected (delete).
| Name | Required | Description | Default |
|---|---|---|---|
| comment_id | Yes | ||
| moderation_status | Yes | heldForReview hides until approved, published approves, rejected deletes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the effects of each status (hides, approves, deletes) but does not mention required permissions, reversibility, or side effects. This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the tool's purpose and the three possible statuses. It is front-loaded with the action and resource, with no extraneous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation tool with 2 parameters and no output schema, the description covers the core functionality. It explains the purpose, the status options, and their effects. It could mention success/error behavior or authentication needs, but it is fairly complete given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to the 'moderation_status' parameter by explaining the result of each enum value, which goes beyond the schema's brief enum description. However, it does not clarify 'comment_id', which is a simple identifier and may be self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Change the moderation status of a comment' with explicit enumeration of the three possible statuses and their effects. It distinguishes the tool from sibling tools like list_comments or reply_to_comment by focusing solely on moderation actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for moderating a comment but provides no explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites or exclusions. The purpose is clear from context, but the description lacks explicit when/when-not instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_channel_analyticsA
Query YouTube Analytics for the authenticated channel. Returns tabular data — useful for views/watch-time/retention/traffic-source reports. Date-ranged and optionally grouped by dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| start_date | Yes | YYYY-MM-DD (inclusive) | |
| end_date | Yes | YYYY-MM-DD (inclusive) | |
| metrics | No | Comma-separated metric names (see YouTube Analytics API). Defaults cover the most common creator-dashboard stats. | views,estimatedMinutesWatched,averageViewDuration,subscribersGained |
| dimensions | No | Comma-separated dimensions, e.g. 'day', 'video', 'country'. Omit for channel totals. | |
| filters | No | Filter expression, e.g. 'video==VIDEO_ID' to scope to one video, or 'country==US'. | |
| sort | No | Sort spec, e.g. '-views' for descending by views | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns tabular data and supports date-range and optional grouping, but does not mention authentication requirements, rate limits, error handling, or pagination behavior. The description provides some transparency but is incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the primary purpose, and contains no unnecessary words. Every sentence adds value, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters and no output schema, the description is brief. It does not explain the return format beyond 'tabular data', lacks details on error handling or edge cases, and does not cover how to interpret results. This leaves gaps for an agent using the tool in complex scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (86%), so the schema already documents parameters well. The description adds minimal value beyond the schema, only referencing date-range and optional grouping. It does not provide additional semantic meaning for parameters like 'filters' or 'sort'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it queries YouTube Analytics for the authenticated channel, specifying it returns tabular data for common metrics like views and watch-time. The verb 'Query' and resource 'YouTube Analytics for the authenticated channel' are specific, and it distinguishes from siblings like 'get_shorts_analytics' by being for general channel analytics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions the tool is useful for certain reports but does not explicitly say when to use it versus alternatives like 'get_shorts_analytics'. There is no guidance on when not to use it or exclusions, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_commentB
Reply to a top-level comment. Requires youtube.force-ssl scope.
| Name | Required | Description | Default |
|---|---|---|---|
| parent_id | Yes | Comment ID to reply to (top-level comment.id from list_comments) | |
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only notes scope requirement. Does not disclose mutation behavior, idempotency, error responses, or return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences conveying purpose and requirement with no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and minimal description; missing details on expected return, error handling, or post-conditions for a write operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers parent_id with clear description, but text param lacks description. Tool description adds no parameter information beyond schema, not compensating for the 50% coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states verb 'reply' and resource 'top-level comment', with scope requirement. Distinguishes from sibling tools like list_comments or moderate_comment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as moderate_comment or add_to_playlist. Only mentions required OAuth scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_video_metadataA
Update a video's metadata — title, description, tags, category, or privacy. Only provide fields you want changed.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | ||
| title | No | ||
| description | No | ||
| tags | No | ||
| category_id | No | YouTube category ID as a string (e.g. '22' = People & Blogs, '27' = Education, '28' = Science & Tech) | |
| privacy_status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It describes the update operation but does not disclose side effects, auth requirements, rate limits, or whether changes are irreversible. The partial update hint is helpful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy. The purpose is front-loaded, and the partial update instruction is concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema, the description covers the essential update semantics. It implies that only provided fields are changed, which is critical. Could mention that omitted fields remain unchanged, but it's already clear from 'Only provide fields you want changed.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (17% – only category_id has a description). The description compensates by listing the fields and emphasizing partial updates, adding meaning beyond the schema's property names and types. However, it does not explain formatting for tags or category_id further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it updates video metadata, listing specific fields (title, description, tags, category, privacy). This distinguishes it from sibling tools like delete_video or get_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Only provide fields you want changed,' indicating a partial update pattern. It does not explicitly state when to use vs. alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_captionA
Upload a caption track (SRT or WebVTT) to a video. Creates a new track — use a distinct name per language/track, or is_draft=true while iterating.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Video ID the caption belongs to. | |
| language | Yes | BCP-47 language code, e.g. 'en', 'en-US', 'es', 'ja'. Must match a language the video supports. | |
| name | No | Caption track name shown in the player's caption menu. Empty string for the default track. | |
| caption_text | Yes | Caption content as a string (SRT or WebVTT format). Source this from a file or the model's output. | |
| format | No | Content type of caption_text: 'srt' (SubRip, application/x-subrip) or 'vtt' (WebVTT, text/vtt). | srt |
| is_draft | No | Draft captions aren't visible to viewers. Useful while reviewing auto-translations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'creates a new track' implying mutation, but lacks details on idempotency, conflict behavior (e.g., does it overwrite or fail if track with same name+language exists?), authentication needs, or rate limits. This is a significant gap for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two clear, front-loaded sentences with no redundant information. The first sentence states the primary purpose, and the second adds practical usage tips. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of siblings for deletion and listing, the description is fairly complete for a creation tool. It covers formats and draft usage. However, it lacks details on error handling, file size limits, or post-creation behavior, leaving some gaps for a comprehensive understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters have descriptions. The description adds value beyond the schema by providing usage context for `name` and `is_draft`, but for other parameters, it repeats schema info. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it uploads a caption track (SRT or WebVTT) to a video, specifying that it creates a new track. This distinguishes it from siblings like list_captions (listing) and delete_caption (deletion), providing specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable guidance: use a distinct `name` per language/track, or `is_draft=true` while iterating. However, it does not explicitly state when not to use it (e.g., for updates, consider deleting first, referencing the sibling delete_caption), so some exclusions are missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.1.0- First observed
add_to_playlist - First observed
create_playlist - First observed
delete_caption - First observed
delete_video - First observed
get_shorts_analytics - First observed
get_video - First observed
list_captions - First observed
list_comments - First observed
list_my_shorts - First observed
list_my_videos - First observed
moderate_comment - First observed
query_channel_analytics - First observed
reply_to_comment - First observed
update_video_metadata - First observed
upload_caption
TDQS
Scored across 15 tools
Each tool targets a distinct resource and action. Analytics tools are clearly differentiated (general vs Shorts-specific), and video list vs Shorts list is scoped by content type. Comments, captions, and playlists each have their own dedicated tools with no overlap.
All tool names follow a consistent verb_noun snake_case pattern (list_, get_, create_, update_, delete_, upload_, reply_to_, add_to_, moderate_, query_). There are no mixed conventions or cryptic abbreviations, making the naming predictable and easy to reason about.
15 tools cover a broad YouTube management surface without feeling bloated. The set is well-scoped for the server's purpose—video, playlist, caption, comment, and analytics operations—and each tool has a clear role within that scope.
The coverage has significant gaps, especially for playlists: there is no way to list, retrieve, delete, or remove items from playlists, creating dead ends after creation. Caption updates are also absent (only create and delete), although video CRUD and comment moderation are well covered.
Maintenance
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for Google Veo AI video generation
Hosted MCP for YouTube Studio: uploads, metadata, playlists, comments, analytics, captions.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
Related MCP Servers
- FlicenseAqualityDmaintenanceA Model Context Protocol server that enables Claude to interact with YouTube data and functionality through the Claude Desktop application.111-
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.-
- AlicenseAqualityDmaintenanceMCP server for ComfyUI — text-to-image, variations, img2img refine, upscale, image proxy, and workflow runner.15115 npm1MIT
- AlicenseAqualityAmaintenanceA comprehensive MCP server integrating YouTube Data, Analytics, and Reporting APIs, providing 40 tools for channel management, analytics, video publishing, transcripts, SEO, and comments.4021MIT