YouTube Video Summarizer MCP Server
Extracts video captions, subtitles, metadata, titles, descriptions, and duration from YouTube videos to enable comprehensive video analysis and summarization.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@YouTube Video Summarizer MCP Serversummarize this video: https://youtube.com/watch?v=dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
YouTube Video Summarizer MCP Server
An MCP (Model Context Protocol) server that enables AI assistants to analyze and summarize YouTube videos by extracting captions, descriptions, and metadata.
Features
Extract video captions/subtitles in multiple languages
Retrieve comprehensive video metadata (title, description, duration)
Provide structured data to AI assistants for comprehensive video summarization
Works with any MCP-compatible client through MCP integration
Support for multiple YouTube URL formats
Language-specific caption extraction
Related MCP server: YouTube DLP MCP Server
Integrating with MCP Clients
To add the MCP server to your MCP client, you can use either method:
Option 1: Using npx (No Installation Required)
Add the following to your MCP client configuration file:
{
"mcpServers": {
"youtube-video-summarizer": {
"command": "npx",
"args": ["-y", "youtube-video-summarizer-mcp"]
}
}
}The server automatically filters out any npm/npx output to ensure MCP protocol compliance.
Option 2: Global Installation (Recommended for Production)
Install the package globally:
npm install -g youtube-video-summarizer-mcpAdd the following to your MCP client configuration file:
{
"mcpServers": {
"youtube-video-summarizer": {
"command": "youtube-video-summarizer",
"args": []
}
}
}Available Tools
When integrated with an MCP client, the following commands become available:
get-video-info-for-summary-from-url: Extract video information and captions from a YouTube URL
get-video-captions: Get captions/subtitles for a specific video
get-video-metadata: Retrieve comprehensive video metadata
Usage Examples
Once integrated with your MCP client, you can use natural language to request video summaries:
"Can you summarize this YouTube video: https://youtube.com/watch?v=VIDEO_ID"
"What are the main points from this video's captions?"
"Extract the key information from this YouTube link"Installation
npm install -g youtube-video-summarizer-mcpDevelopment
git clone https://github.com/nabid-pf/youtube-video-summarizer-mcp.git
cd youtube-video-summarizer-mcp
npm install
npm run buildHow It Works
URL Parsing: Extracts video IDs from various YouTube URL formats
Caption Extraction: Uses youtube-caption-extractor to get subtitles
Metadata Retrieval: Fetches video title, description, and other details
MCP Integration: The Model Context Protocol (MCP) to communicate with AI assistants
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Available Tools
1 toolget-video-info-for-summary-from-urlC
Get details or explanation about a YouTube video, get captions or subtitles of Youtube video from a URL
| Name | Required | Description | Default |
|---|---|---|---|
| videoUrl | Yes | The URL or ID of the YouTube video | |
| languageCode | No | The language code of the video |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the tool can 'get details or explanation' and 'get captions or subtitles', but fails to specify permissions needed, rate limits, error conditions, or output format. For a tool that likely involves external API calls, this lack of transparency is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, consisting of a single sentence that efficiently communicates the core functionality. However, it could be slightly more structured by separating the two distinct purposes (details vs. captions) for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of interacting with YouTube videos (likely involving external APIs), no annotations, and no output schema, the description is incomplete. It lacks crucial information about behavioral traits, error handling, and what the return values entail, making it inadequate for confident tool invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters (videoUrl and languageCode). The description adds no additional meaning beyond what the schema provides, such as format examples or usage notes for the parameters, resulting in the baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('get details or explanation', 'get captions or subtitles') and identifies the resource ('YouTube video from a URL'). It distinguishes between two related but distinct functions (details/explanation vs. captions/subtitles), though without sibling tools to differentiate from, it cannot achieve a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, prerequisites, or exclusions. It merely lists what the tool does without context for application, leaving the agent to infer usage scenarios independently.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v1.8.7- First observed
get-video-info-for-summary-from-url
TDQS
With only one tool, there is no possibility of confusion or overlap between tools. The tool's purpose is clearly defined as retrieving video details and captions for summarization from a YouTube URL.
Since there is only one tool, naming consistency is inherently perfect. The tool name uses a clear verb-noun pattern (get-video-info-for-summary-from-url) that is descriptive and follows a logical structure.
A single tool is too few for a server named 'YouTube Video Summarizer MCP Server', as it suggests a broader summarization functionality. The tool only fetches video info and captions, lacking operations like generating summaries, processing text, or managing multiple videos, making the scope feel incomplete and thin.
The tool set is severely incomplete for the server's purpose. It only provides video details and captions, but does not include any summarization tools (e.g., generate-summary, analyze-captions) or related operations (e.g., list-videos, save-summary), leaving obvious gaps that will hinder agent workflows.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents. No signup.
YouTube transcripts, subtitles, and video metadata as structured JSON via an Apify Actor.
Search YouTube, read video metadata, and fetch transcripts with language preferences
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
Related MCP Servers
- FlicenseDqualityDmaintenanceEnables AI assistants to fetch and analyze transcripts from YouTube videos using video IDs or URLs, with support for multiple language preferences.1-
- AlicenseAqualityFmaintenanceEnables AI assistants to extract YouTube video metadata, subtitles, and top comments without downloading videos.36MIT
- AlicenseCqualityDmaintenanceEnables AI assistants to research YouTube videos by collecting captions, comments, and channel information for analysis and comparison.109MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to fetch YouTube video transcripts with precise timestamps, multi-language support, and time-range filtering.31MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nabid-pf/youtube-video-summarizer-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server