youtube-watchlater-mcp
Allows fetching YouTube Watch Later playlist and video subtitles using yt-dlp with browser cookies.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@youtube-watchlater-mcpshow my watch later playlist"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
youtube-watchlater-mcp
MCP server for fetching your YouTube Watch Later playlist via yt-dlp.
Why This Exists
The YouTube Watch Later playlist is invisible to the YouTube Data API.
Most YouTube MCP servers and integrations use the official YouTube Data API v3 with OAuth authentication. This works well for public playlists, subscriptions, search, and liked videos — but Watch Later is a special private playlist (WL) that Google has explicitly excluded from the API. Even with full OAuth scopes and a valid token, the API returns an empty result or a 403 Forbidden error for this playlist.
This means:
OAuth-based MCPs cannot access Watch Later — the endpoint simply does not exist in the API.
Scraping the YouTube web UI requires managing session cookies, handling anti-bot measures, and is fragile against layout changes.
yt-dlp with browser cookies is the only reliable approach: it reads the authentication cookies that your browser already holds from your active YouTube login, and uses them to fetch the playlist directly — exactly as your browser would.
This server wraps that mechanism as an MCP tool so AI assistants can read your Watch Later queue without you needing API keys, OAuth flows, or exposing credentials.
Related MCP server: YouTube Watch Later MCP Server
Requirements
Node.js 18+
yt-dlpinstalled and available on$PATHYouTube account logged in to a browser on your machine
Installation
1. Install yt-dlp:
# macOS
brew install yt-dlp
# Linux
sudo curl -L https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp -o /usr/local/bin/yt-dlp
sudo chmod a+rx /usr/local/bin/yt-dlp
# Windows (winget)
winget install yt-dlp2. Install Node.js dependencies:
npm installRunning
npm startThe server communicates over stdio and is intended to be connected to an MCP host (e.g. Claude Desktop), not run directly in a browser.
Claude Desktop Setup
Add to claude_desktop_config.json:
{
"mcpServers": {
"youtube-watchlater": {
"command": "node",
"args": ["/path/to/youtube-watchlater-mcp/server.mjs"]
}
}
}Tool: get_watch_later
Returns videos from your Watch Later playlist.
Parameter | Type | Default | Description |
|
|
| Browser to read YouTube cookies from |
| number (1–500) |
| Number of videos to return |
| string | — | Browser profile name, e.g. |
Example response:
{
"items": [
{
"videoId": "dQw4w9WgXcQ",
"title": "Never Gonna Give You Up",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"channel": "Rick Astley"
}
]
}Tool: get_subtitles
Downloads auto-generated subtitles for a YouTube video and returns the VTT content.
Parameter | Type | Default | Description |
| string | — | YouTube video URL or bare video ID (e.g. |
| string |
| Subtitle language code, e.g. |
|
|
| Browser to read YouTube cookies from |
| string | — | Browser profile name, e.g. |
Example response:
{
"lang": "en",
"videoId": "dQw4w9WgXcQ",
"subtitles": "WEBVTT\n..."
}How It Works
The server calls yt-dlp --cookies-from-browser <browser> --flat-playlist --dump-json against the WL playlist. Cookies are read directly from the local browser — no tokens or passwords are transmitted anywhere.
Available Tools
2 toolsget_subtitlesA
Downloads auto-generated subtitles for a YouTube video via yt-dlp and returns the VTT content.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Subtitle language code, e.g. "en" or "ru". | en |
| video | Yes | YouTube video URL or bare video ID (e.g. dQw4w9WgXcQ). | |
| browser | No | Which browser to read YouTube cookies from. | chrome |
| profile | No | Browser profile name, e.g. "Default" or "Profile 1". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that subtitles are auto-generated, that yt-dlp is used, and that the output is VTT content, which is useful context. However, it does not mention potential failures (e.g., subtitles unavailable, network errors), rate limits, or the need for YouTube cookies, which are notable behavioral aspects of this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the action and includes essential information (tool, method, output). No word is wasted, and it is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides a reasonable overview for a straightforward tool, but with no output schema, it could clarify what happens when subtitles are not found or how the VTT content is returned (e.g., as a file path, string). The complexity is moderate due to 4 parameters, but the description does not fully cover edge cases or the exact behavior of the browser/profile parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter has a clear description (e.g., 'Subtitle language code', 'YouTube video URL or bare video ID'). The tool description adds little beyond this, but since the schema is complete, the baseline of 3 is appropriate. It does not explain how parameters interact (e.g., browser/profile for cookie access) beyond what schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Downloads') and clearly identifies the resource ('auto-generated subtitles for a YouTube video'), method ('via yt-dlp'), and output ('VTT content'). This distinguishes it from the sibling tool get_watch_later, which has a different purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when subtitles for a YouTube video are needed, and the inclusion of browser/profile parameters hints at cookie-based authentication, but it does not explicitly state when to use this tool instead of alternatives or provide exclusionary conditions. The alternative (get_watch_later) is unrelated, so no direct comparison is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_watch_laterA
Returns videos from your YouTube Watch Later list using yt-dlp with cookies read from a local browser.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many items to return from Watch Later. | |
| browser | No | Which browser to read YouTube cookies from. | chrome |
| profile | No | Browser profile name, e.g. "Default" or "Profile 1" (Chrome/Brave/Edge). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It mentions using yt-dlp and reading cookies from a local browser, which is a key behavioral detail. However, it does not explain caveats like cookie availability, browser being closed, or failure scenarios. Some value is added, but it is not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that is front-loaded with the action and resource, then specifies the method. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list-retrieval tool, the description covers what it does and how (yt-dlp, local browser cookies). It lacks information about return format or error conditions, but given no output schema, it is reasonably complete. It could explicitly note that browser cookies must exist and the profile must match, which is only implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents limit, browser, and profile adequately. The description itself adds no additional parameter information, so it relies on the schema, giving a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Returns videos') and the specific resource ('your YouTube Watch Later list'), also mentioning the method (yt-dlp with cookies). This distinguishes it from the sibling tool get_subtitles, which is about subtitles, not watch later.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it is for retrieving Watch Later videos. However, it does not explicitly mention when to use this tool instead of alternatives or any exclusions. Given the only sibling is get_subtitles, the distinct purpose is clear, but explicit guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
get_subtitles - First observed
get_watch_later
TDQS
The two tools serve entirely different purposes: one retrieves the watch-later playlist, the other fetches subtitles for a specific video. There is no overlap or ambiguity between them.
Both tools follow a consistent verb_noun pattern with 'get_' prefix, making the action and resource clear. No mixed conventions or irregular naming.
With only two tools, the server feels thin for a YouTube-related service. The scope is narrow (watch later + subtitles), so the count is borderline but not extreme.
The watch-later domain lacks operations like removing or adding videos, and the subtitle tool is unrelated to the watch-later focus. The surface is minimal and leaves obvious gaps for a cohesive workflow.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
YouTube transcripts, subtitles, and video metadata as structured JSON via an Apify Actor.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Search YouTube, read video metadata, and fetch transcripts with language preferences
Transcribe YouTube via Whisper. Summaries, chapters, semantic-search across your corpus.
Related MCP Servers
- AlicenseAqualityCmaintenanceUses yt-dlp to download subtitles from YouTube and connects it to claude.ai via Model Context Protocol.11,014544MIT
- AlicenseNot gradedqualityDmaintenanceEnables secure access to your YouTube Watch Later playlist, allowing retrieval of video URLs added within a specified timeframe through a simple interface using OAuth2 authentication.11MIT
- AlicenseAqualityDmaintenanceEnables fetching, searching, and analyzing YouTube video transcripts in multiple languages using yt-dlp. Supports timestamp filtering, language detection, and transcript summaries with robust error handling for production use.4MIT
- AlicenseAqualityCmaintenanceFetches YouTube video subtitles and transcripts with support for multiple languages and output formats (SRT, VTT, TXT, JSON).119Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/tulaman/youtube-watchlater-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server