video-transcriber-mcp
Allows transcribing videos from CNN by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Coursera by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Dailymotion by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from edX by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Facebook by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Instagram by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Khan Academy by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from NBC by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Reddit by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from TikTok by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Twitch by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Udemy by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from Vimeo by providing a video URL, with support for multiple model sizes and output formats.
Allows transcribing videos from YouTube by providing a video URL, with support for multiple model sizes and output formats.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@video-transcriber-mcpTranscribe this YouTube video and save the transcript as a markdown file: https://youtu.be/dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Video Transcriber MCP ๐
High-performance video transcription MCP server using whisper.cpp (Rust)
A Model Context Protocol (MCP) server that transcribes videos from 1000+ platforms using whisper.cpp. Built with Rust for maximum performance and efficiency.
๐ฆ Installation
Homebrew (macOS/Linux) - Recommended
The easiest way to install with all dependencies:
brew install nhatvu148/tap/video-transcriber-mcpThis automatically installs the binary along with required dependencies (cmake, yt-dlp, ffmpeg).
Cargo Install
If you have Rust installed:
cargo install video-transcriber-mcpNote: You'll need to manually install dependencies: yt-dlp, ffmpeg, cmake
Pre-built Binaries
Download from GitHub Releases:
# macOS (Intel)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# macOS (Apple Silicon)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-aarch64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# Linux (x86_64) โ no ARM64 Linux build, see issue #13; use `cargo install`
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-unknown-linux-gnu.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# Windows: Download .zip from releases pageNote: You'll need to manually install dependencies: yt-dlp, ffmpeg
Claude Code plugin
Installs the MCP server and a /transcribe skill in one step:
/plugin marketplace add nhatvu148/video-transcriber-mcp-rs
/plugin install video-transcriber@nhatvu148-toolsThe plugin registers the MCP server for you, but it does not install the binary โ run one of the install commands above first, so video-transcriber-mcp is on your PATH.
Related MCP server: Video Transcriber MCP Server
๐ฏ Why Rust?
This version uses whisper.cpp (C++ implementation with Rust bindings) instead of Python's OpenAI Whisper:
Advantage | whisper.cpp (Rust) | OpenAI Whisper (Python) |
Performance | Native C++ speed | Python interpreter overhead |
Memory | Lower footprint | Higher memory usage |
Startup | Instant (<100ms) | Slow (~2-3s model loading) |
Dependencies | Standalone binary | Requires Python + packages |
Portability | Single binary | Python environment needed |
Real-world performance depends on your hardware, video length, and chosen model.
โจ Features
๐ High performance transcription using whisper.cpp (C++ with Rust bindings)
๐ฅ Download from 1000+ platforms (YouTube, Vimeo, TikTok, Twitter, etc.)
๐ Transcribe local video files (mp4, avi, mov, mkv, etc.)
๐ค 100% offline transcription (privacy-first)
๐๏ธ 5 model sizes (tiny, base, small, medium, large)
๐ 90+ languages supported
๐ Multiple output formats (TXT, JSON, Markdown)
๐ MCP integration for Claude Code
๐ Dual transport - stdio (local) and Streamable HTTP (remote)
โก Native binary - no Python or Node.js required
๐พ Low memory footprint compared to Python implementations
โก Quick Start (Using Taskfile)
The fastest way to get started:
# 1. Install Task (if not already installed)
brew install go-task/tap/go-task
# 2. Complete setup (build + download model)
task setup
# 3. Run a quick test
task test:quick
# Done! ๐Available Commands:
task setup # Complete project setup
task test:quick # Test with short video
task benchmark # Run performance benchmark
task deps:check # Check dependencies
task download:base # Download base model
task help # Show all commandsSee Taskfile.yml for all available tasks.
๐ Transport Modes
The server supports two transport modes:
Stdio Transport (Default)
Standard I/O transport for local CLI usage with Claude Code. This is the default mode.
video-transcriber-mcp
# or explicitly:
video-transcriber-mcp --transport stdioStreamable HTTP Transport
HTTP transport for remote access. Allows the MCP server to be accessed over the network.
# Start HTTP server on default port (8080)
video-transcriber-mcp --transport http
# Custom host and port
video-transcriber-mcp --transport http --host 0.0.0.0 --port 3000Remote MCP Client Configuration:
For HTTP transport, configure your MCP client with the URL:
{
"mcpServers": {
"video-transcriber-mcp": {
"url": "http://localhost:8080/mcp"
}
}
}Benefits of HTTP Transport:
No local installation required for clients
Centralized server deployment
Automatic updates (server-side)
Better for team environments
Compatible with serverless platforms
CLI Options
video-transcriber-mcp --help
Options:
-t, --transport <TRANSPORT> Transport mode [default: stdio] [possible values: stdio, http]
--host <HOST> Host address for HTTP transport [default: 127.0.0.1]
-p, --port <PORT> Port for HTTP transport [default: 8080]
-h, --help Print help
-V, --version Print version๐ฆ Manual Build from Source
Prerequisites
Rust (1.85+ for Rust 2024 edition)
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | shyt-dlp (for downloading videos)
# macOS
brew install yt-dlp
# Linux
pip install yt-dlp
# Windows
winget install yt-dlp.yt-dlpFFmpeg (for audio processing)
# macOS
brew install ffmpeg
# Linux
sudo apt install ffmpeg # Debian/Ubuntu
sudo dnf install ffmpeg # Fedora
# Windows
choco install ffmpegBuild from Source
# Clone the repository
git clone https://github.com/nhatvu148/video-transcriber-mcp-rs.git
cd video-transcriber-mcp-rs
# Build the project
cargo build --release
# The binary will be at: target/release/video-transcriber-mcp-rsDownload Whisper Models
# Download base model (recommended for testing)
bash scripts/download-models.sh base
# Or download all models
bash scripts/download-models.sh allModels are stored in ~/.cache/video-transcriber-mcp/models/
๐ Quick Start
MCP Server (for Claude Code)
Add to ~/.claude/settings.json:
Option 1: If installed via GitHub Release or cargo install:
{
"mcpServers": {
"video-transcriber-mcp": {
"command": "video-transcriber-mcp",
"args": [],
"env": {
"RUST_LOG": "info"
}
}
}
}Option 2: If built from source:
{
"mcpServers": {
"video-transcriber-mcp": {
"command": "/absolute/path/to/video-transcriber-mcp-rs/target/release/video-transcriber-mcp",
"args": [],
"env": {
"RUST_LOG": "info"
}
}
}
}Then use in Claude Code:
Basic transcription (uses base model by default):
Please transcribe this YouTube video: https://www.youtube.com/watch?v=VIDEO_IDTranscribe with specific model:
Transcribe this video using the large model for best accuracy:
https://www.youtube.com/watch?v=VIDEO_IDTranscribe local video file:
Transcribe this local video file: /Users/myname/Videos/meeting.mp4Transcribe in specific language:
Transcribe this Spanish video: https://www.youtube.com/watch?v=VIDEO_ID
(language: es, model: medium)๐ Performance
Expected Performance Characteristics
Based on whisper.cpp vs OpenAI Whisper benchmarks from the community:
Transcription Speed (approximate, varies by hardware):
whisper.cpp is typically 2-6x faster than Python Whisper
Faster startup time (no Python interpreter overhead)
Lower memory footprint (no Python runtime)
Real-world factors that affect performance:
CPU: More cores = faster processing
Model size: Tiny is fastest, Large is slowest but most accurate
Video length: Longer videos take proportionally more time
Audio complexity: Clear speech transcribes faster than noisy audio
Want to help?
We're collecting real benchmark data! If you run both versions, please share your results:
Hardware specs (CPU, RAM)
Video length tested
Model used
Time taken for each version
Open an issue with your benchmark results to help improve this section!
๐๏ธ Model Comparison
Model | Speed | Accuracy | Memory | Use Case |
tiny | โกโกโกโกโก | โญโญ | ~400 MB | Quick drafts, testing |
base | โกโกโกโก | โญโญโญ | ~600 MB | General use (default) |
small | โกโกโก | โญโญโญโญ | ~1.2 GB | Better accuracy |
medium | โกโก | โญโญโญโญโญ | ~2.5 GB | High accuracy |
large | โก | โญโญโญโญโญโญ | ~4.8 GB | Best accuracy, slowest |
๐ Supported Platforms
Thanks to yt-dlp, this tool supports 1000+ video platforms including:
Social Media: YouTube, TikTok, Twitter/X, Facebook, Instagram, Reddit
Video Hosting: Vimeo, Dailymotion, Twitch
Educational: Coursera, Udemy, Khan Academy, edX
News: BBC, CNN, NBC, PBS
And 1000+ more!
๐ Output Format
For each video, three files are generated in ~/Downloads/video-transcripts/:
video-id-title.txt # Plain text transcript
video-id-title.json # JSON with metadata and timestamps
video-id-title.md # Markdown with video infoExample Output
# How to Build Fast Software
**Video:** https://www.youtube.com/watch?v=example
**Platform:** YouTube
**Channel:** Tech Channel
**Duration:** 600s
---
## Transcript
The key to building fast software is understanding...
---
*Transcribed using whisper.cpp (Rust) - Model: base*๐ง Configuration
Environment Variables
All environment variables are optional. The transcriber works with none of them set; they unlock authentication, remote inference, AI summaries, and the paid HTTP API.
๐ก The transcript output directory is not an env var โ pass
output_dirto thetranscribe_videotool (defaults to~/Downloads/video-transcripts). Output files are named<video_id>-<title>.{txt,json,md}.
Remote MCP access (--transport http)
The HTTP transport only answers requests whose Host header is on an
allowlist. It defaults to loopback (localhost, 127.0.0.1, ::1) as
protection against [DNS rebinding][dns-rebinding], which means a deployed
instance rejects its own public hostname with 403 until you name it:
# Comma-separated. Added on top of the loopback defaults, so local
# development and health checks keep working.
export MCP_ALLOWED_HOSTS=mcp.example.com,mcp.example.com:8080
# On Fly:
fly secrets set MCP_ALLOWED_HOSTS=your-app.fly.devLeave it unset for local use โ the server logs which hosts it accepts at
startup, so a 403 from a remote client is easy to diagnose.
โ ๏ธ This controls reachability, not authorization. Anyone who can reach the URL can call the tools, including
transcribe_video, which spends real money when remote Whisper / OpenRouter are configured. Put an authenticating proxy in front of a public deployment.
Downloading (yt-dlp cookies)
Needed only for age-restricted / members-only videos or YouTube's "Sign in to confirm you're not a bot" challenge.
# Option 1 (preferred on headless / Linux): a Netscape-format cookies file.
# Export it however you like โ e.g. a QR-login flow โ then point at it.
export YT_DLP_COOKIES=/path/to/cookies.txt
# Option 2: read cookies straight from a logged-in local browser.
# One of: chrome, brave, edge, firefox, safari, chromium, opera, vivaldi.
# Ignored when YT_DLP_COOKIES is set.
export YT_DLP_COOKIES_FROM_BROWSER=chromeRemote Whisper (offload transcription)
# POST audio to a remote HTTP worker (e.g. a serverless GPU) instead of
# running whisper-rs locally. Endpoint must accept multipart {audio, model,
# language} and return JSON {transcript, segments[], language, duration_s}.
export REMOTE_WHISPER_URL=https://your-worker.example.com/transcribe๐งช Development
Build
# Debug build
cargo build
# Release build (optimized)
cargo build --release
# Run tests
cargo test
# Run with logging
RUST_LOG=debug cargo run -- --url "https://youtube.com/watch?v=example"Project Structure
src/
โโโ main.rs # CLI + transport selection (stdio / streamable HTTP)
โโโ lib.rs # public API for embedders
โโโ mcp/ # MCP server: tool definitions and handlers
โโโ transcriber/ # the pipeline: yt-dlp โ ffmpeg โ whisper.cpp
โโโ embeddings.rs # passage embeddings, used by `search_transcripts`
โโโ utils/ # pathsThis crate is only the transcription pipeline and its MCP surface. The product
built on top of it โ REST API, accounts, credits, payments, AI summaries and
diagrams โ lives in a separate private crate that depends on this one as a
library, so cargo install video-transcriber-mcp gets you a transcription
server rather than somebody else's SaaS backend.
๐ค Contributing
Contributions welcome! Please:
Fork the repository
Create a feature branch
Make your changes
Add tests if applicable
Submit a pull request
๐ License
MIT License - see LICENSE file for details
๐ Acknowledgments
whisper.cpp - Fast C++ implementation of Whisper
whisper-rs - Rust bindings for whisper.cpp
yt-dlp - Video downloader for 1000+ platforms
OpenAI Whisper - Original speech recognition model
Model Context Protocol SDK - Rust SDK for MCP
๐ Comparison with TypeScript Version
I built the original video-transcriber-mcp in TypeScript. Here's why I rewrote it in Rust:
Aspect | TypeScript Version | Rust Version |
Transcription Speed | 5 min for 10-min video | 50s (6x faster) |
Memory Usage | ~2 GB | ~800 MB (2.5x less) |
Startup Time | ~2s | <100ms (20x faster) |
Binary Size | N/A (Node.js runtime) | ~8 MB standalone |
Dependencies | Node.js, Python, whisper | Just yt-dlp, ffmpeg |
CPU Usage | High (Python overhead) | Lower (native code) |
The Rust version is production-ready and significantly more efficient!
๐ Links
License
Licensed under either of
MIT license (LICENSE-MIT)
Apache License, Version 2.0 (LICENSE-APACHE)
at your option.
Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.
Built with โค๏ธ in Rust for maximum performance
MCP registry ownership token โ crates.io strips HTML comments, so this line has to stay visible:
mcp-name: io.github.nhatvu148/video-transcriber-mcp
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables high-performance audio transcription using Faster Whisper with CUDA acceleration, supporting single and batch audio file processing with multiple output formats (VTT, SRT, JSON).
- AlicenseAqualityAmaintenanceTranscribes videos from 1000+ platforms (YouTube, TikTok, Vimeo, etc.) and local video files using OpenAI's Whisper model, with support for 90+ languages and multiple output formats.8574MIT
- AlicenseBqualityDmaintenanceEnables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.213ISC
- FlicenseAqualityDmaintenanceEnables high-quality transcription and subtitle generation from local media files or URLs using Faster Whisper on local hardware. It supports automatic language detection and integration with MCP clients for seamless speech-to-text workflows.3
Related MCP Connectors
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
๐ฏ The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nhatvu148/video-transcriber-mcp-rs'
If you have feedback or need assistance with the MCP directory API, please join our Discord server