local-whisper-speech-to-text
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-whisper-speech-to-texttranscribe this audio URL: https://example.com/interview.mp3"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
local-whisper-mcp
Local speech-to-text service, in a single process: the FastAPI REST API and the MCP
Streamable HTTP server share the same JobManager, the same job queue and the same
Whisper model cache.
No cloud API calls. Audio is downloaded, transcribed, and the job metadata is kept on local disk.
Copyright (C) 2026 Laurent Lemercier. Licensed under the GNU AGPL v3.0.
Contents
Related MCP server: faster-whisper-mcp
Features
Asynchronous submission: every transcription returns a
job_idimmediately. The REST call responds202 Accepted; the client then polls for state.Three audio sources: remote URL, multipart upload, or Base64.
Two interfaces: REST and MCP Streamable HTTP, on the same port, with no second process.
Webhooks: when a job ends (success or failure), the service POSTs the job state to a callback URL.
Transcription metrics: audio duration, processing duration, and real-time factor (
real_time_factor).SSRF protection: download and callback URLs are validated, including on every redirect.
On-disk job persistence, readable again after a restart.
Requirements
Python 3.11 or later.
FFmpeg is only required for the Docker containers. Locally, audio duration is determined by PyAV (a bundled native library) and, as a fallback, by Mutagen — no external binary is installed.
Disk space: model weights are downloaded on first use (roughly 500 MB for
small, roughly 1.5 GB formedium).
Installation
git clone https://github.com/laurentlemercier/local-whisper-mcp.git
cd local-whisper-mcp
python -m venv .venv
# Windows (PowerShell)
.venv\Scripts\Activate.ps1
# Linux / macOS
source .venv/bin/activate
pip install -r requirements.txtConfiguration
Every variable is read from the process environment at startup (app/config.py).
The application does not load a .env file: the .env.example in this repository
serves as reference only and is not applied automatically.
Variable | Default | Purpose |
| (empty) | Bearer token required by REST and MCP. Empty means no authentication. |
|
| Data root: audio inputs and job state. |
|
| Listening interface. |
|
| Listening port. |
|
| Device passed to |
|
| Quantisation ( |
|
| Maximum audio file size, in bytes (500 MiB). |
|
| Maximum time to download audio, in seconds. |
|
| Maximum time for the callback POST, in seconds. |
|
| Allow audio URLs pointing at private IP addresses. |
| (empty) | Comma-separated allow-list of hosts permitted for audio. |
|
| Allow callback URLs pointing at private IP addresses. |
| (empty) | Comma-separated allow-list of hosts permitted for callbacks. |
Getting started
# Set a token, then start
export API_TOKEN="a-long-random-token" # Linux / macOS
$env:API_TOKEN="a-long-random-token" # Windows PowerShell
uvicorn app.main:app --host 0.0.0.0 --port 8000REST: http://localhost:8000
Interactive OpenAPI documentation: http://localhost:8000/docs
Why
uvicorn app.main:appinstead ofpython -m app.run?app/run.pycallsmount_mcp(app), which mounts the MCP server on/mcpa second time even thoughapp/main.pyhas already mounted it. The second mount is unreachable: it creates an unused session manager. Going throughuvicorn app.main:appavoids the duplication. See Known limitations.
For development with automatic reload:
uvicorn app.main:app --reloadREST API
Every route except /health requires the Authorization: Bearer <API_TOKEN> header.
Method | Route | Description |
|
| Readiness probe. Unauthenticated. |
|
| Supported models and models currently loaded in memory. |
|
| Submits an audio URL. Responds |
|
| Submits an audio file as |
|
| Submits Base64-encoded audio. Responds |
|
| Full job state. |
|
| Transcription and metrics. |
|
| Deletes the job and its audio. |
Submit a URL
curl -X POST http://localhost:8000/v1/transcriptions \
-H "Authorization: Bearer $API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"source": { "type": "url", "url": "https://example.com/audio.m4a" },
"model": "small",
"language": "en"
}'{ "id": "tr_3f9a1c2b8e7d4051", "status": "queued", "model": "small", ... }Submit a file
curl -X POST http://localhost:8000/v1/transcriptions/upload \
-H "Authorization: Bearer $API_TOKEN" \
-F "file=@./audio.m4a" \
-F "model=small" \
-F "language=en"Fetch the result
curl -H "Authorization: Bearer $API_TOKEN" \
http://localhost:8000/v1/transcriptions/tr_3f9a1c2b8e7d4051/result409 while the job is unfinished, 500 if the job failed, 200 once status is
completed.
Possible states are queued, downloading, running, completed and failed.
Job completion webhook
Add callback_url (and callback_headers if needed) to the request. Once the
transcription finishes, successfully or not, the service performs a JSON POST with
the job state. callback_url and callback_headers are excluded from the body sent.
MCP tools
MCP Streamable HTTP server on /mcp, named local-whisper-speech-to-text.
Tool | Purpose |
| Submits an audio URL. |
| Submits Base64-encoded audio. |
| Full job state. |
| Transcription and duration metrics. |
| Deletes a job. |
Every tool takes an optional api_token parameter, to be filled with the same value as
API_TOKEN.
Example MCP client configuration:
{
"mcpServers": {
"local-whisper": {
"type": "http",
"url": "http://localhost:8000/mcp",
"headers": { "Authorization": "Bearer a-long-random-token" }
}
}
}Models
Two models are accepted: small (default) and medium.
They are loaded on demand at first use and kept in memory for the lifetime of the
process. Loading medium needs roughly three times the memory of small; on CPU,
int8 is the right compromise. language: "auto" lets Whisper detect the language.
Supported audio formats
.m4a .mp3 .wav .ogg .opus .webm .flac
The extension is checked on receipt. For an upload, a name without a valid extension is rejected; for a URL or Base64 payload, the extension is derived from the path and, failing that, from the response MIME type.
Docker
The build context is the repository root; the image only contains requirements.txt and
app/ (see .dockerignore).
cp docker/.env.example docker/.env
# Fill in DATA_PATH, MODELS_PATH, CURRENT_UID and CURRENT_GID in docker/.env
docker compose -f docker/docker-compose.service.yml up -ddocker/.env.example documents the expected variables. In particular:
CURRENT_UID/CURRENT_GIDstop the container from writing toDATA_PATHas root.MODELS_PATHacts as a persistent HuggingFace cache, so weights are not re-downloaded on every recreate.
The image is based on python:3.11-slim pinned by digest, with ffmpeg installed. Its
entrypoint is uvicorn app.main:app, so it does not run the duplicate mount described
above.
Tests
The tests are manual and require a running service as well as a local audio file.
cd tests
.\test-speech-service-v2.3.ps1 `
-ApiToken "a-long-random-token" `
-AudioFile "C:\paths\to\my\audio.m4a"The script chains: /health probe, Bearer authentication, REST upload, state polling
until completion, result retrieval, then discovery and invocation of the MCP tools. It
deletes the jobs it created, unless -KeepJobs is passed.
It also automatically creates a virtual environment in tests/.venv and installs the
official MCP SDK there — which is why that directory is ignored by git. Pass
-PythonExe to choose the interpreter, -SkipMcp or -SkipRestUpload to narrow the
scope. Without -AudioFile, the REST and MCP tests are skipped with a warning.
tests/mcp_client.py is a standalone MCP client, usable on its own:
python tests/mcp_client.py \
--endpoint http://localhost:8000/mcp \
--api-token "$API_TOKEN" \
--audio-file ./audio.m4a \
--model small \
--language enSecurity
Authentication
If API_TOKEN is empty, the service requires no authentication and checks none. Do
not expose it to a network in that state. REST uses the Authorization: Bearer header;
MCP tools expect the api_token parameter.
SSRF protection
Download and callback URLs go through the same validator (app/security.py), which
rejects:
any scheme other than
httpandhttps;credentials embedded in the URL (
https://user:pass@host/...);hosts absent from the allow-list, when
ALLOWED_URL_HOSTSis set;localhost;any IP address the host resolves to that is private, loopback, link-local, reserved, multicast or unspecified.
Redirects are followed manually, at most six times, and every target is revalidated
before being requested. Size is checked against the Content-Length header and then
against the body actually received.
To fetch audio from an internal network, set ALLOW_PRIVATE_URLS=true — and prefer a
host allow-list over that relaxation.
Network exposure
The service provides neither TLS nor rate limiting. Put it behind a reverse proxy that terminates TLS, and restrict submissions to trusted clients: transcription is expensive in CPU time.
Project layout
app/
config.py Environment variables and data directories
models.py Pydantic schemas (Job, Source, Result, Statuses)
security.py Bearer token and SSRF validation
source.py URL download, Base64, extension checking
jobs.py JobManager: queue, JSON persistence, locks
whisper.py Model cache and faster-whisper invocation
worker.py Processing thread: download, transcription, webhook
mcp.py MCP server and its five tools
main.py FastAPI application, REST routes, MCP mount
run.py Alternative entrypoint (see Known limitations)
data/
input/ Downloaded or uploaded audio, one directory per job
jobs/ State of each job as JSON
docker/ Dockerfile and compose file for a containerised deployment
tests/ PowerShell test script and Python MCP clientA job is laid out as follows:
data/jobs/tr_3f9a1c2b8e7d4051.json state, metrics, result
data/input/tr_3f9a1c2b8e7d4051/audio.m4a source audio
data/input/tr_3f9a1c2b8e7d4051/source.json original URL, if remote sourceKnown limitations
Single-threaded queue. Only one job is transcribed at a time. The queue lives in memory: it does not survive a restart, and an in-flight job is lost.
Duplicate MCP mount.
app/main.py:549mounts the MCP server on/mcp, thenapp/run.py:5callsmount_mcp(app), which mounts it again. Only the first mount is reachable; the second creates an orphaned session manager. Useuvicorn app.main:app.Dead variables in Docker.
MODELandLANGUAGEare set in the Dockerfile and the compose file, but the application never reads them: model and language are chosen per request. Thewhisper_modelsvolume declared in the compose file is not used either.No rate limiting or quota. An authorised client can saturate the queue.
No automatic expiry. Jobs and their audio occupy disk until explicitly deleted.
No CI. The tests require a live service and a model weight download, which makes them a poor fit for continuous integration without a cache.
License
GNU Affero General Public License v3.0. The full text is in LICENSE.
AGPL v3.0 has a network clause (section 13): if you modify this software and make it available to users who access it over a network, you must offer those users the source code of your modified version. This is the essential difference from the classic GPL, which only covers distribution.
Copyright (C) 2026 Laurent Lemercier.
This server cannot be deployed
Maintenance
Related MCP Connectors
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch.
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
Transcribe any audio or video URL to text, SRT and VTT, with chunking and only-new episodes
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables high-performance audio transcription using Faster Whisper with CUDA acceleration, supporting single and batch audio file processing with multiple output formats (VTT, SRT, JSON).-
- FlicenseAqualityDmaintenanceEnables high-quality transcription and subtitle generation from local media files or URLs using Faster Whisper on local hardware. It supports automatic language detection and integration with MCP clients for seamless speech-to-text workflows.3-
- AlicenseAqualityFmaintenanceProvides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.3792 npm3MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to transcribe audio and video from URLs or local files with high accuracy, speaker diarization, 119 languages, and word-level timestamps, while also supporting transcription management and caption export in SRT, WebVTT, or plain text.1488 npm11MIT