groq-whisper-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@groq-whisper-mcptranscribe ~/Downloads/interview.mov"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
groq-whisper-mcp
MCP server for audio/video transcription via Groq-hosted Whisper. Word-level timestamps, automatic caching, cost estimation.
9x cheaper than OpenAI Whisper — same model, different API.
Provider | Model | Cost/hour |
Groq | whisper-large-v3-turbo | $0.04 |
Groq | whisper-large-v3 | $0.111 |
Groq | distil-whisper | $0.02 |
OpenAI | whisper-1 | $0.36 |
Prerequisites
Python 3.10+
ffmpeg (for audio extraction from video)
Related MCP server: Whisper Speech Recognition MCP Server
Installation
git clone https://github.com/bis-code/groq-whisper-mcp.git
cd groq-whisper-mcp
python3 -m venv venv
venv/bin/pip install mcp openaiClaude Code Configuration
Add to your project's .mcp.json:
{
"mcpServers": {
"whisper": {
"command": "/path/to/groq-whisper-mcp/venv/bin/python",
"args": ["-m", "server"],
"cwd": "/path/to/groq-whisper-mcp/src",
"env": {
"GROQ_API_KEY": "your-groq-api-key"
}
}
}
}Tools
transcribe_video
Transcribe a video file with word-level timestamps.
Parameter | Type | Required | Description |
| string | yes | Path to video file |
| string | no | Whisper model (default: |
| boolean | no | Bypass cache (default: |
Returns full text, word-level timestamps [{word, start, end}], and duration. Results are cached per-project.
estimate_transcription_cost
Estimate cost before transcribing.
Parameter | Type | Required | Description |
| string | yes | Path to video file |
| string | no | Whisper model (default: |
Returns duration, estimated cost, and comparison with OpenAI pricing.
How It Works
Extracts audio from video via ffmpeg (128kbps MP3, falls back to 64kbps if >25MB)
Sends to Groq's OpenAI-compatible API (same
openaiSDK, differentbase_url)Returns word-level timestamps with ~0.1s precision
Caches results as
whisper_words.jsonalongside the video
Running Tests
PYTHONPATH=src venv/bin/pytest tests/ -vLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
- mcpOAuthso.transcribe
Transcribe audio and video into speaker-labelled transcripts, subtitles, clips, and cited Q&A.
Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceThis service provides fast and reliable transcriptions for audio/video files and voice memos. It allows LLMs to interact with the text content of audio/video file.8MIT
- FlicenseNot gradedqualityCmaintenanceEnables high-performance audio transcription using Faster Whisper with CUDA acceleration, supporting single and batch audio file processing with multiple output formats (VTT, SRT, JSON).-
- AlicenseAqualityAmaintenanceTranscribes videos from 1000+ platforms (YouTube, TikTok, Vimeo, etc.) and local video files using OpenAI's Whisper model, with support for 90+ languages and multiple output formats.814 npm6MIT
- AlicenseAqualityDmaintenanceAutomatically transcribes Google Chat voice messages using Groq Whisper API, and also supports local audio file transcription.2MIT