Skip to main content
Glama

whisper-mcp

Local Whisper transcription, exposed as an MCP server — so any MCP client (Claude Desktop, Claude Code, etc.) can transcribe audio/video files, generate .srt subtitles, and burn captions into a video, directly as tool calls. No cloud API, no manual "run a script and go read the output file" step.

Built on faster-whisper for inference and ffmpeg for audio extraction / caption burn-in.

Claude Code calling the transcribe tool and returning a timestamped transcript

Architecture

MCP client (Claude, etc.) --stdio--> whisper-mcp server --> faster-whisper (Whisper model)
                                                        \--> ffmpeg (audio extract / burn-in)

The server keeps loaded Whisper models cached in memory for the life of the process, so repeated tool calls in one session don't re-pay model load time.

Related MCP server: Whisper MCP Server

Tools

Tool

Description

transcribe(path, model_size="small", device="auto")

Returns detected language + timestamped segments as structured data.

generate_srt(path, output_path=None, model_size="small", device="auto")

Transcribes and writes a .srt file.

burn_captions(video_path, srt_path, output_path=None)

Burns an .srt into a video via ffmpeg.

Requirements

  • Python 3.10+

  • ffmpeg on PATH

  • Optional: a CUDA-capable GPU (falls back to CPU automatically)

Install & run

pip install -e ".[dev]"
whisper-mcp

Or, zero-install:

uvx --from git+https://github.com/Chain-P/whisper-mcp whisper-mcp

Configure in an MCP client

Add to your client's MCP config (e.g. Claude Code's .mcp.json or Claude Desktop's claude_desktop_config.json):

{
  "mcpServers": {
    "whisper": {
      "command": "whisper-mcp"
    }
  }
}

Then ask the client to transcribe a file, e.g. "transcribe samples/podcast_clip.mp4 and give me the SRT."

Development

pip install -e ".[dev]"
ruff check .
pytest                        # unit tests only
pytest -m integration         # + real transcription against a sample file

Roadmap

  • Speaker diarization (pyannote.audio / whisperx) for multi-speaker labeling in the SRT output.

  • MCP progress notifications for long transcriptions.

  • A resource exposing recent transcript history.

License

MIT

Install Server
A
license - permissive license
A
quality
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables high-quality transcription and subtitle generation from local media files or URLs using Faster Whisper on local hardware. It supports automatic language detection and integration with MCP clients for seamless speech-to-text workflows.
    3
  • A
    license
    A
    quality
    F
    maintenance
    Provides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.
    3
    67
    3
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    Enables local media processing (video/audio) using FFmpeg and FFprobe, allowing frame extraction, audio conversion, and metadata retrieval through natural language.
    5
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to query local video timelines by extracting speech, frame captions, and on-screen text into a SQLite store, exposing search and retrieval tools via MCP.
    PolyForm Noncommercial 1.0.0

View all related MCP servers

Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.

  • MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Chain-P/whisper-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server