Skip to main content
Glama
ZahiriNatZuke

whisper-transcribe-mcp

whisper-transcribe-mcp

PyPI version CI License: MIT Python 3.10+

MCP server for audio transcription using faster-whisper (local, free, offline) or OpenAI Whisper API (cloud, requires API key). Works with Claude Desktop and Claude Code on macOS, Windows, and Linux.


Prerequisites

macOS

Option A — uv (recommended):

brew install uv
# or
curl -LsSf https://astral.sh/uv/install.sh | sh

Option B — Python: Python 3.10+ is included in macOS 12.3+. You can also install it with brew install python.


Windows

Option A — uv (recommended):

winget install astral-sh.uv

Or download the installer from astral.sh/uv.

Option B — Python: Download Python 3.10+ from python.org. During installation, check "Add Python to PATH".

No need to install ffmpeg or any compiler — everything is bundled in the package.


Linux

Option A — uv (recommended):

curl -LsSf https://astral.sh/uv/install.sh | sh

Option B — Python:

# Debian/Ubuntu
sudo apt install python3.12 python3.12-venv

# Fedora
sudo dnf install python3.12

# Arch
sudo pacman -S python

No additional system dependencies required.


Related MCP server: audio-transcription-mcp

Installation

uvx automatically downloads and installs the package in an isolated environment. Only requires uv to be installed.

# Local backend:
uvx "whisper-transcribe-mcp[local]"

# OpenAI backend:
uvx "whisper-transcribe-mcp[openai]"

# Both backends:
uvx "whisper-transcribe-mcp[all]"

Option B — pip

# Local backend:
pip install "whisper-transcribe-mcp[local]"

# OpenAI backend:
pip install "whisper-transcribe-mcp[openai]"

# Both backends:
pip install "whisper-transcribe-mcp[all]"

Use Cases

Case 1 — Local backend only (free, works offline)

Uses faster-whisper to transcribe locally. The model is downloaded from HuggingFace on first use (~74MB for base) and cached.

Install:

pip install "whisper-transcribe-mcp[local]"

Environment variables:

WHISPER_MODEL=base   # or tiny, small, medium, large-v3

Case 2 — OpenAI backend only (best accuracy, requires API key)

Uses OpenAI's whisper-1 model. Requires an API key and internet connection. No local model downloads.

Install:

pip install "whisper-transcribe-mcp[openai]"

Environment variables:

OPENAI_API_KEY=sk-...

Case 3 — Both backends (OpenAI if key present, local as fallback)

If OPENAI_API_KEY is set, OpenAI is used automatically. Otherwise falls back to local faster-whisper.

Install:

pip install "whisper-transcribe-mcp[all]"

Configuration

Claude Desktop

Config file location by operating system:

OS

Path

macOS

~/Library/Application Support/Claude/claude_desktop_config.json

Windows

%APPDATA%\Claude\claude_desktop_config.json

Linux

~/.config/Claude/claude_desktop_config.json

Add the entry inside "mcpServers":

Windows note: Claude Desktop runs in a restricted environment and may not have uvx in its PATH, and it may use a Python version (e.g. 3.14) for which ctranslate2 (a dependency of faster-whisper) does not yet have prebuilt wheels. Two fixes are required:

  1. Use the full path to uvx.exe instead of just uvx. Run where.exe uvx in PowerShell to find it (usually C:\Users\<YourUser>\.local\bin\uvx.exe).

  2. Force Python 3.12 via the --python 3.12 flag so that a compatible wheel is used.

Case 1 — Local:

macOS / Linux:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "uvx",
      "args": ["whisper-transcribe-mcp[local]"],
      "env": {
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Windows:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "C:\\Users\\<YourUser>\\.local\\bin\\uvx.exe",
      "args": ["--python", "3.12", "whisper-transcribe-mcp[local]"],
      "env": {
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Case 2 — OpenAI:

macOS / Linux:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "uvx",
      "args": ["whisper-transcribe-mcp[openai]"],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Windows:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "C:\\Users\\<YourUser>\\.local\\bin\\uvx.exe",
      "args": ["--python", "3.12", "whisper-transcribe-mcp[openai]"],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Case 3 — Both (OpenAI takes priority if key is set):

macOS / Linux:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "uvx",
      "args": ["whisper-transcribe-mcp[all]"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Windows:

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "C:\\Users\\<YourUser>\\.local\\bin\\uvx.exe",
      "args": ["--python", "3.12", "whisper-transcribe-mcp[all]"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Restart Claude Desktop after editing the file.


Claude Code

Works the same on macOS, Windows, and Linux. Requires uv installed.

Claude Code config file location:

OS

Global

Per project

macOS / Linux

~/.claude.json

.claude/settings.json (project root)

Windows

C:\Users\<user>\.claude.json

.claude\settings.json (project root)

The easiest way to add the server is via the Claude Code CLI, which updates the config file automatically:

# Case 1 — Local:
claude mcp add whisper-transcribe uvx -- "whisper-transcribe-mcp[local]"

# Case 2 — OpenAI:
claude mcp add whisper-transcribe uvx --env OPENAI_API_KEY=sk-... -- "whisper-transcribe-mcp[openai]"

# Case 3 — Both (OpenAI with local fallback):
claude mcp add whisper-transcribe uvx --env OPENAI_API_KEY=sk-... --env WHISPER_MODEL=base -- "whisper-transcribe-mcp[all]"

Windows + [all]: Add --python 3.12 before the package name to avoid ctranslate2 wheel issues. Edit ~/.claude.json directly and use "args": ["--python", "3.12", "whisper-transcribe-mcp[all]"].

To add it globally (available in all projects), use --scope user:

claude mcp add --scope user whisper-transcribe uvx -- "whisper-transcribe-mcp[local]"

Or edit ~/.claude.json directly and add inside "mcpServers":

{
  "mcpServers": {
    "whisper-transcribe": {
      "command": "uvx",
      "args": ["whisper-transcribe-mcp[local]"],
      "env": {
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Environment Variables

Variable

Default

Description

WHISPER_MODEL

base

Local model size: tiny, base, small, medium, large-v3

OPENAI_API_KEY

—

If set, activates the OpenAI backend instead of local

WHISPER_POST_PROCESS_MODEL

gpt-5.4-nano

OpenAI chat model used when post_process=true

Backend selection and fallback ([all] only)

When installed with [all], the backend is chosen at startup:

  • OPENAI_API_KEY set → OpenAI is used. If the API call fails at runtime (network error, invalid key, quota exceeded), the server automatically falls back to local faster-whisper and includes a "fallback_reason" field in the response.

  • OPENAI_API_KEY not set → local faster-whisper is used directly, no fallback attempted.


Security and Permissions

The server runs locally over stdio and only needs the access its tools imply:

Access

Why

File system read

transcribe_file reads the audio file path you pass it.

Temporary files

transcribe_base64 writes the decoded audio to a private temp file and always deletes it.

Network

Only for the OpenAI backend / post-processing, and for the first download of a local model from Hugging Face.

Environment variables

OPENAI_API_KEY, WHISPER_MODEL, WHISPER_POST_PROCESS_MODEL.

Tool inputs are validated before use: model_size must be a known Whisper model name, language must be an ISO 639 code, and extension must be a known audio format. OpenAI errors are returned as a short summary (error type and HTTP status); the full message goes to stderr. Releases are published from GitHub Actions with PyPI Trusted Publishing and PEP 740 provenance attestations.


Available Tools

transcribe_file

Transcribes an audio file by path (mp3, wav, m4a, ogg, flac, webm, etc.).

Parameters:

  • file_path (required): Absolute path to the audio file

  • language (optional): Language code (es, en, fr, etc.). Auto-detected if not provided.

  • model_size (optional): Local model size. Ignored with the OpenAI backend.

  • post_process (optional, default false): If true, passes the transcription through GPT-4.1 to fix spelling, grammar, and punctuation. Requires the openai package ([openai] or [all]).

  • post_process_prompt (optional): Custom system prompt for GPT post-processing. Use it to provide domain-specific context, proper nouns, or product names that Whisper may have misspelled. Falls back to a generic correction prompt if not provided.

Response (without post-processing):

{
  "text": "Full transcription...",
  "language": "en",
  "language_probability": 0.99,
  "segments": [
    { "start": 0.0, "end": 4.2, "text": "First segment..." }
  ],
  "backend": "local",
  "model": "base"
}

Response (with post_process: true):

{
  "text": "Corrected transcription...",
  "raw_text": "Original transcription from Whisper...",
  "post_process_model": "gpt-4.1",
  "language": "en",
  "language_probability": 0.99,
  "segments": [...],
  "backend": "local",
  "model": "base"
}

If post-processing fails, text retains the original transcription and a post_process_error field is added.


transcribe_base64

Transcribes audio provided as a base64-encoded string. Useful for programmatic integrations.

Parameters:

  • audio_base64 (required): Base64-encoded audio data

  • extension (optional, default mp3): File extension (mp3, wav, ogg, etc.)

  • language (optional): Language code

  • model_size (optional): Local model size

  • post_process (optional, default false): Same as in transcribe_file.

  • post_process_prompt (optional): Same as in transcribe_file.


list_models

Shows the active backend configuration, available local models, and the GPT model used for post-processing.


Local Model Sizes

Model

Size

Relative Speed

Notes

tiny

39 MB

~32x

Fastest, least accurate

base

74 MB

~16x

Good balance (default)

small

244 MB

~6x

Better accuracy

medium

769 MB

~2x

High accuracy

large-v3

1.5 GB

~1x

Best accuracy, slowest

Models are downloaded automatically from HuggingFace on first use and cached locally.


Troubleshooting

MCP not loading in Claude Desktop on Windows

Symptom: The server fails to start with a dependency resolution error like:

ctranslate2>=4.6.1 has no wheels with a matching platform tag (e.g., `win32`)
hint: You require CPython 3.14 (`cp314`), but we only found wheels for `ctranslate2` with: `cp39`, `cp310`, `cp311`, `cp312`, `cp313`

Cause: Two issues combined:

  1. Claude Desktop does not include the user's local bin in its PATH, so uvx must be referenced by full path.

  2. Claude Desktop's uvx may pick a Python version (e.g. 3.14) for which ctranslate2 — a native dependency of faster-whisper — does not yet have prebuilt wheels for Windows.

Fix: Use the full path to uvx.exe and force Python 3.12 explicitly:

"whisper-transcribe": {
  "command": "C:\\Users\\<YourUser>\\.local\\bin\\uvx.exe",
  "args": ["--python", "3.12", "whisper-transcribe-mcp[local]"],
  "env": { "WHISPER_MODEL": "base" }
}

To find your exact uvx.exe path, run in PowerShell:

where.exe uvx

Transcribing audio files in Claude Desktop

Symptom: Claude Desktop fails to transcribe an uploaded audio file. It may attempt to read the file as base64 and pass it to transcribe_base64, which then fails or hangs for files larger than ~50 KB.

Cause: Claude Desktop runs in a sandboxed Linux container. When you upload a file using the attachment button, it is stored at a path like /mnt/user-data/uploads/audio.mp3 — inside the container. The MCP server runs on your Windows machine and has no access to that container path. Claude's fallback of base64-encoding the file and passing it to transcribe_base64 fails in practice because even a small audio file produces hundreds of kilobytes of base64 text, which overflows the context window before the tool call can be made.

Fix: Do not use the attachment button to upload audio files. Instead, place the file anywhere on your Windows filesystem and reference its path directly in the message:

"Transcribe the file at C:\Users\YourUser\Downloads\audio.mp3"

The MCP server will read the file directly from Windows and send it to the transcription backend. This works for files of any size within the Whisper API limit (25 MB).


Development

This project uses uv for reproducible development environments:

uv sync --group dev
uv run ruff check .
uv run ruff format --check .
uv run pytest -q
uv run pre-commit install

See CONTRIBUTING.md for the contribution workflow.

Distribution and releases

The package is distributed through PyPI and described by server.json for the official MCP Registry. Version tags publish both destinations through GitHub OIDC, without long-lived publishing tokens. See docs/publishing.md for the release checklist and one-time repository setup.

The Registry entry describes the base PyPI package. Choose the [local], [openai], or [all] extra from the installation examples above so the transcription backend you need is installed.


License

MIT — see LICENSE

Available Tools

3 tools
list_modelsA

List available Whisper model sizes and current configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description states it lists information but does not explicitly confirm read-only or side effects. Minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with zero waste; front-loaded with action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No parameters, output schema exists to describe return values; description fully covers purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters in schema (0 count), baseline is 4. Description adds no param info, but none needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb 'List' and specific resource 'available Whisper model sizes and current configuration'. Distinguishes from sibling tools transcribe_base64 and transcribe_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage for checking available models, but no explicit when-to-use or alternatives given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribe_base64C

Transcribe audio provided as a base64-encoded string.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNoLanguage code. Auto-detected if not provided.
extensionNoAudio format of the data (mp3, wav, m4a, ogg, flac, webm, mp4, etc.).mp3
model_sizeNoLocal model size. Ignored when using the OpenAI backend.
audio_base64YesBase64-encoded audio data.
post_processNoIf True, passes the transcription through GPT to fix spelling, grammar, and punctuation. Requires the openai package.
post_process_promptNoCustom system prompt for post-processing. Use this to provide domain-specific context, proper nouns, or product names.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing beyond the verb. It does not mention input size ceilings (a real risk for base64 payloads), model/backend selection behavior, latency, or error conditions when the payload is malformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler that names the action and the input form. It is efficient, though its brevity is closer to under-specification than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and all six parameters are documented. However, for a tool with no annotations and a sibling that handles the file variant, the absence of routing guidance leaves the definition only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents language detection, extension formats, model_size backend interaction, and post-processing. The description adds no semantics beyond the schema, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transcribe audio') plus the input modality ('base64-encoded string'). The base64 modality implicitly separates it from transcribe_file, though the sibling is never named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this over transcribe_file or list_models, and no prerequisites such as input size limits or required backend configuration. The agent must infer routing from the phrase 'base64-encoded' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribe_fileC

Transcribe an audio file to text.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNoLanguage code (e.g. 'es', 'en', 'fr'). Auto-detected if not provided.
file_pathYesAbsolute path to the audio file (mp3, wav, m4a, ogg, flac, etc.)
model_sizeNoLocal model size: tiny, base, small, medium, large-v3 (plus the .en, turbo, and distil variants). Ignored when using the OpenAI backend. Defaults to the WHISPER_MODEL environment variable (default: 'base').
post_processNoIf True, passes the transcription through GPT to fix spelling, grammar, and punctuation. Requires the openai package.
post_process_promptNoCustom system prompt for post-processing. Use this to provide domain-specific context, proper nouns, or product names that Whisper may have misspelled. Falls back to a generic correction prompt if not provided.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and adds nothing behavioral: it doesn't mention the local-vs-OpenAI backend, whether it is a long-running operation, cost, or any side effects. The important behaviors (model_size, post_process requiring openai) exist only in the schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, which is appropriate for a simple call. It leans toward under-specification rather than being overly brief for its own sake.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema and full parameter descriptions cover return values and inputs, so the description needn't explain them. Still, for a tool with backend selection and optional GPT post-processing, the absence of any behavioral framing leaves the definition only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including language, model_size, and post_process_prompt is already documented in the schema with defaults and caveats. The description adds no parameter meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transcribe an audio file to text'), so the agent immediately knows the operation. However, it offers no distinction from the sibling transcribe_base64, which presumably does the same thing from a different input form.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus transcribe_base64 or list_models, and no prerequisites or exclusions are stated. The agent must infer selection purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.1.3
    • Changedtranscribe_base641 field changed
      • changedInput schema / properties / extension / description
        Previous value: -"File extension for the temp file (mp3, wav, m4a, ogg, etc.)."New value: +"Audio format of the data (mp3, wav, m4a, ogg, flac, webm, mp4, etc.)."
    • Changedtranscribe_file1 field changed
      • changedInput schema / properties / model_size / description
        Previous value: -"Local model size: tiny, base, small, medium, large-v3.\n        Ignored when using the OpenAI backend.\n        Defaults to the WHISPER_MODEL environment variable (default: 'base')."New value: +"Local model size: tiny, base, small, medium, large-v3 (plus the .en,\n        turbo, and distil variants). Ignored when using the OpenAI backend.\n        Defaults to the WHISPER_MODEL environment variable (default: 'base')."
  2. 3 tool updatesv1.1.1
    • First observedlist_models
    • First observedtranscribe_base64
    • First observedtranscribe_file

TDQS

A3.6/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: transcribe_file (local file input), transcribe_base64 (encoded string input), and list_models (metadata/discovery). The file-vs-base64 split is unambiguous because it is keyed to input source, not overlapping functionality.

Naming Consistency5/5

All three tools follow a consistent snake_case verb_noun pattern (transcribe_file, transcribe_base64, list_models). The convention is predictable and readable throughout.

Tool Count5/5

Three tools is well-scoped for a narrow, single-purpose transcription server; each tool covers a genuinely distinct need (file input, string input, model discovery) with no redundancy.

Completeness4/5

The core transcribe-and-inspect lifecycle is covered, but common Whisper capabilities like translation, language detection, or alternate output formats (SRT/VTT) are absent. These are minor gaps an agent can work around given the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers