mcp-youtube-transcript
Servidor de transcripciones de YouTube de MCP
Un servidor de Protocolo de Contexto de Modelo que permite la recuperación de transcripciones de vídeos de YouTube. Este servidor proporciona acceso directo a las transcripciones de vídeo mediante una interfaz sencilla, lo que lo hace ideal para el análisis y procesamiento de contenido.
Tabla de contenido
Related MCP server: YouTube Transcript Server
Características
✨ Capacidades clave:
Extraer transcripciones de vídeos de YouTube
Soporte para múltiples idiomas
Formatear texto con modo continuo o de párrafo
Recuperar títulos y metadatos de vídeos
Segmentación automática de párrafos
Normalización de texto y decodificación de entidades HTML
Manejo robusto de errores
Detección de marcas de tiempo y superposición
Empezando
Prerrequisitos
Node.js 18 o superior
Instalación
Ofrecemos dos métodos de instalación:
Opción 1: Configuración manual (recomendada para producción)
Cree o edite el archivo de configuración de Claude Desktop:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonVentanas:
%APPDATA%\Claude\claude_desktop_config.json
Agregue la siguiente configuración:
{
"mcpServers": {
"youtube-transcript": {
"command": "npx",
"args": [
"-y",
"@sinco-lab/mcp-youtube-transcript"
]
}
}
}Script de configuración rápida para macOS:
# Create directory if it doesn't exist
mkdir -p ~/Library/Application\ Support/Claude
# Create or update config file
cat > ~/Library/Application\ Support/Claude/claude_desktop_config.json << 'EOL'
{
"mcpServers": {
"youtube-transcript": {
"command": "npx",
"args": [
"-y",
"@sinco-lab/mcp-youtube-transcript"
]
}
}
}
EOLOpción 2: Vía Herrería (Solo Desarrollo)
npx -y @smithery/cli install @sinco-lab/mcp-youtube-transcript --client claude⚠️ Nota : Este método no se recomienda para uso en producción ya que depende de los servicios de proxy de Smithery.
Uso
Configuración básica
Para utilizar con Claude Desktop/Cursor/cline, asegúrese de que su configuración coincida:
{
"mcpServers": {
"youtube-transcript": {
"command": "npx",
"args": ["-y", "@sinco-lab/mcp-youtube-transcript"]
}
}
}Pruebas
Con la aplicación Claude
Reinicie la aplicación Claude después de la instalación
Prueba con un comando simple:
https://www.youtube.com/watch?v=AJpK3YTTKZ4 Summarize this video
Ejemplo de salida:
Con MCP Inspector
# Clone and setup
git clone https://github.com/sinco-lab/mcp-youtube-transcript.git
cd mcp-youtube-transcript
npm install
npm run build
# Launch inspector
npx @modelcontextprotocol/inspector node "dist/index.js"
# Access http://localhost:6274 and try these commands:
# 1. List Tools: clink `List Tools`
# 2. Test get_transcripts with:
# url: "https://www.youtube.com/watch?v=AJpK3YTTKZ4"
# lang: "en" (optional)
# enableParagraphs: false (optional)Solución de problemas y mantenimiento
Comprobación de los registros de Claude
Para supervisar los registros de Claude, puede utilizar el siguiente comando:
tail -n 20 -f ~/Library/Logs/Claude/mcp*.logEsto mostrará las últimas 20 líneas del archivo de registro y continuará mostrando nuevas entradas a medida que se agreguen.
Nota : La aplicación Claude prefija automáticamente los archivos de registro del servidor MCP con
mcp-server-. Por ejemplo, los registros de nuestro servidor se escribirán enmcp-server-youtube-transcript.log.
Limpieza de la caché npx
Si encuentra problemas relacionados con el caché npx , puede limpiarlo manualmente usando:
rm -rf ~/.npm/_npxEsto eliminará los paquetes almacenados en caché y le permitirá comenzar de nuevo.
Referencia de API
obtener_transcripciones
Obtiene transcripciones de vídeos de YouTube.
Parámetros:
url(cadena, obligatoria): URL o ID del video de YouTubelang(cadena, opcional): Código de idioma (predeterminado: "en")enableParagraphs(booleano, opcional): Habilitar el modo de párrafo (predeterminado: falso)
Formato de respuesta:
{
"content": [{
"type": "text",
"text": "Video title and transcript content",
"metadata": {
"videoId": "video_id",
"title": "video_title",
"language": "transcript_language",
"timestamp": "processing_time",
"charCount": "character_count",
"transcriptCount": "number_of_transcripts",
"totalDuration": "total_duration",
"paragraphsEnabled": "paragraph_mode_status"
}
}]
}Desarrollo
Estructura del proyecto
├── src/
│ ├── index.ts # Server entry point
│ ├── youtube.ts # YouTube transcript fetching logic
├── dist/ # Compiled output
└── package.jsonComponentes clave
YouTubeTranscriptFetcher: Funcionalidad principal para obtener transcripcionesYouTubeUtils: Procesamiento de texto y utilidades
Características y capacidades
Manejo de errores:
URL/ID no válidos
Transcripciones no disponibles
Disponibilidad de idiomas
Errores de red
Limitación de velocidad
Procesamiento de texto:
Decodificación de entidades HTML
Normalización de la puntuación
Normalización espacial
Detección inteligente de párrafos
Contribuyendo
¡Agradecemos las contribuciones! No dudes en enviar problemas y solicitudes de incorporación de cambios.
Licencia
Este proyecto está licenciado bajo la licencia MIT: consulte el archivo de LICENCIA para obtener más detalles.
Proyectos relacionados
Available Tools
1 toolget_transcriptsA
Extract and process transcripts from a YouTube video.
Parameters:
url(string, required): YouTube video URL or ID.lang(string, optional, default 'en'): Language code for transcripts (e.g. 'en', 'uk', 'ja', 'ru', 'zh').enableParagraphs(boolean, optional, default false): Enable automatic paragraph breaks.
IMPORTANT: If the user does not specify a language code, DO NOT include the lang parameter in the tool call. Do not guess the language or use parts of the user query as the language code.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video URL or ID | |
| lang | No | Language code for transcripts, default 'en' (e.g. 'en', 'uk', 'ja', 'ru', 'zh') | en |
| enableParagraphs | No | Enable automatic paragraph breaks, default `false` |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively explains the tool's core function and includes important behavioral guidance about parameter handling (the IMPORTANT note about not guessing language). However, it doesn't mention potential limitations like video availability, transcript existence, rate limits, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear purpose statement followed by parameter documentation and important usage notes. Every sentence serves a purpose, though the parameter list slightly duplicates schema information. The IMPORTANT section is appropriately emphasized for critical behavioral guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides adequate coverage for the tool's basic function and parameters. However, it lacks information about return values, error handling, and operational constraints that would be helpful for an agent. The IMPORTANT note adds valuable context, but more behavioral transparency would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description repeats this information in a bulleted list without adding significant semantic context beyond what's in the schema. The IMPORTANT note about language parameter handling adds some value, but overall the description doesn't enhance parameter understanding beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('extract and process') and resource ('transcripts from a YouTube video'). It distinguishes itself from potential alternatives by focusing on transcript extraction rather than other video-related operations, though no sibling tools exist for direct comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage guidance through the IMPORTANT note about language parameter handling, but it doesn't explicitly state when to use this tool versus alternatives (e.g., when transcripts are needed vs. other video metadata). Since no sibling tools exist, this is less critical, but general context about appropriate use cases is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- First observed
get_transcripts
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined and distinct by default.
Since there is only one tool, naming consistency is inherently perfect. The tool name 'get_transcripts' follows a clear verb_noun pattern.
A single tool is too few for a server focused on YouTube transcripts, as it lacks operations like searching transcripts, managing multiple videos, or handling errors. This minimal scope limits functionality and agent workflows.
The tool set is severely incomplete for the domain of YouTube transcript processing. It only provides extraction (get_transcripts), missing essential operations such as searching within transcripts, listing available languages, or handling video metadata, which are common needs in this context.
Maintenance
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
Transcripts of YouTube videos, playlists and channels in any language: text, SRT, VTT or JSON.
Related MCP Servers
- AlicenseAqualityFmaintenanceA Model Context Protocol server that enables retrieval of transcripts from YouTube videos. This server provides direct access to video captions and subtitles through a simple interface.1656 npm595MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables retrieval of transcripts from YouTube videos with language-specific support.1656 npm1MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI assistants to extract transcripts from YouTube videos, allowing AI to analyze and work with video content directly.111 npm4MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables access to YouTube video content through transcripts, translations, summaries, and subtitle generation in various languages.55MIT