YouTube MCP Server
Enables searching for YouTube videos through natural language commands and playing videos or search results as playlists in the browser.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@YouTube MCP Serverfind and play some lofi hip hop beats for studying"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
YouTube MCP Server
A Model Context Protocol (MCP) server that enables AI agents to search and play YouTube videos through natural language commands.
Features
searchVideos: Search YouTube for videos matching any query
playPlaylist: Open videos in your default browser as a playlist
Related MCP server: YouTube MCP Server
Prerequisites
Node.js 18+ installed
YouTube Data API v3 key from Google Cloud Console
Getting a YouTube API Key
Go to Google Cloud Console
Create a new project (or select existing)
Navigate to APIs & Services > Library
Search for "YouTube Data API v3" and enable it
Go to APIs & Services > Credentials
Click Create Credentials > API Key
Copy your API key
Installation
Clone the repository
Install dependencies:
npm installConfigure Environment Variables: Create a
.envfile in the root directory (copy from.env.example):cp .env.example .envEdit
.envand add yourYOUTUBE_API_KEY.Build the project:
npm run build
Add the following to your MCP configuration (e.g., .cursor/mcp.json or Claude Desktop config):
{
"mcpServers": {
"youtube-mcp": {
"command": "node",
"args": ["/path/to/your/project/youtubeMCP/dist/index.js"]
}
}
}Note: The server will automatically load the
YOUTUBE_API_KEYfrom the.envfile in the project directory. Alternatively, you can pass it directly in theenvobject in the JSON config above.
Usage Examples
Once configured, you can ask your AI agent:
"Search for 10 best English songs on YouTube"
"Play 5 relaxing piano music videos"
"Find Taylor Swift songs and play them"
"Search for coding tutorials"
Available Tools
searchVideos
Search YouTube for videos by query.
Parameters:
query(string, required): Search querymaxResults(number, optional): Number of results (1-50, default: 10)
playPlaylist
Play videos in the browser.
Parameters:
videoIds(string[], optional): Array of video IDs to playquery(string, optional): Search and play videos matching this querymaxResults(number, optional): Number of videos when using query (default: 10)
Development
# Watch mode for development
npm run dev
# Build
npm run build
# Start
npm startProject Structure
YouTubeMCP/
├── package.json
├── tsconfig.json
├── README.md
├── src/
│ ├── index.ts # MCP server entry point
│ ├── youtube-client.ts # YouTube API wrapper
│ └── tools/
│ ├── search.ts # searchVideos tool
│ └── play.ts # playPlaylist tool
└── dist/ # Compiled JavaScriptLicense
MIT
Available Tools
2 toolsplayPlaylistA
Play YouTube videos in the browser. Provide video IDs directly or a search query to find and play videos automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| videoIds | No | Array of YouTube video IDs to play as a playlist | |
| query | No | Search query to find and play videos (used if videoIds not provided) | |
| maxResults | No | Number of videos to play when using query (default: 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While it states the tool plays videos in the browser, it doesn't mention important behavioral aspects like whether this opens a new tab/window, requires browser permissions, has rate limits, or what happens if multiple instances are invoked. The description is insufficient for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences that each earn their place. The first sentence states the core purpose, and the second explains the parameter options. There's zero wasted language and it's front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters, 100% schema coverage, but no annotations or output schema, the description provides adequate basic context about what the tool does and parameter options. However, as a mutation tool (playing videos implies side effects), it should disclose more about behavioral expectations and potential constraints given the lack of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds marginal value by mentioning the alternative between videoIds and query parameters, but doesn't provide additional semantic context beyond what's in the schema. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Play YouTube videos in the browser') and resource ('YouTube videos'), distinguishing it from the sibling tool 'searchVideos' which presumably only searches without playing. It explicitly mentions both direct video ID input and search query functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use each parameter ('Provide video IDs directly or a search query'), but doesn't explicitly state when to choose this tool over the sibling 'searchVideos' or mention any prerequisites or exclusions. The guidance is helpful but lacks sibling differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchVideosC
Search for YouTube videos by query. Returns video IDs, titles, channels, and descriptions.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query for YouTube videos (e.g., "best english songs", "relaxing piano music") | |
| maxResults | No | Maximum number of results to return (1-50, default: 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. While it mentions what the tool returns, it doesn't cover important behavioral aspects like whether this is a read-only operation, potential rate limits, authentication requirements, error conditions, or pagination behavior. The description is minimal and lacks the depth needed for a tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with two clear sentences. The first sentence states the core functionality, and the second sentence specifies what information is returned. There's no wasted language or unnecessary elaboration. However, it could be slightly more structured by separating purpose from return values more explicitly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there are no annotations and no output schema, the description is insufficiently complete. For a search tool with 2 parameters, the description should provide more context about the search behavior, result format, limitations, and how it differs from the sibling 'playPlaylist' tool. The current description leaves too many behavioral questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with both parameters well-documented in the schema itself. The description doesn't add any meaningful parameter information beyond what's already in the schema - it mentions 'by query' which is already covered by the schema's query parameter description. This meets the baseline expectation when schema coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search for YouTube videos by query.' It specifies the verb (search) and resource (YouTube videos), and mentions the return fields (video IDs, titles, channels, descriptions). However, it doesn't explicitly differentiate from the sibling tool 'playPlaylist', which appears to serve a different function (playing playlists vs. searching videos).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention the sibling tool 'playPlaylist' or any other search-related tools that might exist. There's no context about when this search is appropriate or when other methods should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- First observed
playPlaylist - First observed
searchVideos
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: playPlaylist is for playing videos in the browser, while searchVideos is for searching and returning video metadata. There is no overlap in functionality, making it easy for an agent to choose the correct tool based on the task.
The naming is inconsistent: playPlaylist uses camelCase, while searchVideos uses a verb_noun pattern. This mix of conventions lacks a predictable pattern, which could confuse agents expecting uniformity across tools.
With only 2 tools, the server feels thin for a YouTube domain, which typically involves operations like listing playlists, managing subscriptions, or uploading videos. The current set is insufficient for comprehensive YouTube interactions, suggesting a significant under-scoping.
The tool surface is severely incomplete for a YouTube server. It lacks basic CRUD operations such as creating playlists, managing subscriptions, or handling user accounts. The existing tools cover only playback and search, leaving major gaps that will likely cause agent failures in broader tasks.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
Search YouTube, read video metadata, and fetch transcripts with language preferences
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceEnables AI assistants to search YouTube for videos, channels, and playlists while retrieving detailed analytics and metrics through the YouTube Data API v3. Supports advanced filtering options and provides comprehensive statistics for content discovery and analysis.1MIT
- FlicenseNot gradedqualityDmaintenanceEnables interaction with YouTube through the YouTube Data API, allowing users to search for videos, playlists, and channels, generate video titles using AI, and manage YouTube content through natural language commands.22-
- AlicenseAqualityCmaintenanceEnables AI models to interact with YouTube content including video details, transcripts, channel information, playlists, and search functionality through the YouTube Data API.714 npm10MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to search videos, read channels, browse playlists, fetch comments, and get transcripts from YouTube using the YouTube Data API v3 and InnerTube API for captions.2GPL 3.0