open-transcribe-mcp
Provides speech-to-text transcription through ElevenLabs Scribe, allowing audio from files or HTTPS URLs to be transcribed with configurable routing, diarization, timestamps, and transcript styles.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@open-transcribe-mcptranscribe the audio at https://example.org/meeting.mp3"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
OpenTranscribe MCP
Own the recorder. Choose the intelligence.
OpenTranscribe is an open-source MCP server that routes audio to the speech-to-text model of your choice and returns one provider-independent transcript schema.
It is built around a simple idea: buying a great recorder should not lock you into one transcription subscription. Use Plaud, a phone, an open-source wearable, or any other audio source you are authorized to access, then choose Microsoft MAI, ElevenLabs Scribe, or Groq Whisper without changing the downstream workflow.
OpenTranscribe does not jailbreak hardware or bypass access controls. It works only with audio the operator is authorized to access.
How it works
OpenTranscribe keeps the integration boundary stable: the recorder supplies an audio file or HTTPS URL, the server selects or calls the requested speech-to-text provider, and downstream tools receive the same canonical transcript shape.
Related MCP server: audio-transcriber
What v1 ships
MCP Streamable HTTP at
/mcp, with stateless operation and bearer authenticationtranscribe_audio,list_transcription_models,estimate_transcription_cost,get_transcript_chunk, anddelete_transcriptMicrosoft
MAI-Transcribe-2, ElevenLabsscribe-v2, and Groq Whisper adaptersexplicit capability negotiation, routing policies, retries, and observable fallbacks
HTTPS URL passthrough when supported and a bounded streaming proxy otherwise
SSRF controls, signed-URL redaction, zero content logging, and cost/resource limits
disabled-by-default retention; optional memory or S3-compatible temporary result storage
Docker, production-oriented Scaleway Serverless Containers Terraform, tests, and GitHub Actions CI
On the Linux desktop
An optional Debian package, OpenTranscribe Setup, bundles the Python runtime, the engine, and
a native setup application. Install it, add a provider key, and connect a local MCP client — no
Python, no uv, no configuration files. The engine then runs as one client-owned STDIO process
that opens no network socket.
The hosted and command-line paths below are unaffected by it. See OpenTranscribe Setup on Linux; its release evidence lives in docs/desktop-acceptance.md, where every scenario starts at Not tested.
Ten-minute quickstart
Requirements: Python 3.12 and uv, or Docker.
git clone https://github.com/fbossiere/open-transcribe-mcp.git
cd open-transcribe-mcp
cp .env.example .envEdit .env with one provider credential and a strong random MCP bearer token. For Microsoft:
OT_MICROSOFT__ENDPOINT=https://YOUR-RESOURCE.cognitiveservices.azure.com
OT_MICROSOFT__API_KEY=YOUR-KEY
OT_SECURITY__BEARER_TOKEN=YOUR-RANDOM-TOKENStart the server:
uv sync
uv run open-transcribe-mcpThe MCP endpoint is http://localhost:8000/mcp; probes are available at /healthz and /readyz.
Run the included bilingual, two-voice synthetic recording through the connected provider:
uv run python examples/transcribe.py --token YOUR-RANDOM-TOKENThe fixture and reference transcript are non-sensitive and redistributable. Pass another authorized public HTTPS audio URL as the first argument to use your own source.
Connect a FastMCP client:
import asyncio
from fastmcp import Client
async def main() -> None:
async with Client("http://localhost:8000/mcp", auth="YOUR-RANDOM-TOKEN") as client:
models = await client.call_tool("list_transcription_models", {})
print(models)
result = await client.call_tool(
"transcribe_audio",
{
"request": {
"source": {"type": "url", "url": "https://example.org/authorized-audio.mp3"},
"provider": "auto",
"routing_policy": "quality",
}
},
)
print(result)
asyncio.run(main())Provider choice does not alter the response contract. Set provider and model to switch explicitly, or use auto with default, quality, cost, or latency routing.
diarization, timestamps, and transcript_style are unset above on purpose: an unset capability is not requested, so the request routes to any configured model and the response metadata reports what that model applied. Stating one makes it a requirement — "diarization": True excludes every model that cannot diarize rather than quietly returning a single-speaker transcript. See providers and capabilities.
Docker
docker build -t open-transcribe-mcp:1.1.0 .
docker run --rm -p 8000:8000 \
-e OT_ENVIRONMENT=prod \
-e OT_MICROSOFT__ENDPOINT="https://YOUR-RESOURCE.cognitiveservices.azure.com" \
-e OT_MICROSOFT__API_KEY="YOUR-KEY" \
-e OT_SECURITY__AUTH_MODE=bearer \
-e OT_SECURITY__BEARER_TOKEN="YOUR-RANDOM-TOKEN" \
open-transcribe-mcp:1.1.0The published image is also available as ghcr.io/fbossiere/open-transcribe-mcp:1.1.0.
Configuration
All settings use the OT_ prefix and __ for nesting. See .env.example. Provider credentials are server-side environment variables and are never accepted as MCP tool arguments.
Temporary storage is disabled by default. result_mode=stored requires:
OT_RESULT_STORE__BACKEND=memory # local/test only; use s3 for horizontally scaled production
OT_RESULT_STORE__CURSOR_SECRET=ANOTHER-RANDOM-SECRETFor Scaleway Object Storage, install the s3 extra and configure the S3 bucket/endpoint variables documented in the deployment guide.
The reference Scaleway deployment is codified in infra/scaleway. It provisions a private image registry, a scale-to-zero Serverless Container, health probes, HTTPS-only ingress, and an optional TTL-bound result bucket with a dedicated runtime identity.
Security and privacy defaults
source HTTPS is required;
private, loopback, link-local, multicast, reserved, and metadata destinations are rejected;
every redirect target is resolved and validated;
proxy downloads connect to the validated public IP while preserving HTTPS SNI and the original Host header;
source downloads are streamed to an ephemeral file with byte limits, then deleted;
signed URL queries, audio, transcript text, authorization headers, and phrase hints are not logged;
no project telemetry is emitted;
transcripts are not retained unless a result store is explicitly enabled and used.
ElevenLabs zero-retention requests are enabled by default; disabling them is an explicit operator choice.
Read SECURITY.md, the threat model, and the retention policy before exposing the service publicly.
Known limitations
OpenTranscribe currently targets self-hosted, single-tenant installations. Provider feature parity is deliberately not guaranteed; capability negotiation exposes differences instead of hiding them. URL ingestion is the only remote input type. Synchronous provider limits still apply. OIDC and asynchronous jobs are planned for later releases. The memory store is neither durable nor horizontally scalable. S3 lookups prioritize a simple deployment contract over very-large-bucket indexing; dedicate the result prefix and enforce lifecycle deletion.
Application controls do not replace network policy. Internet-facing operators should still combine exact source-host allow-listing with egress firewall rules.
Provider prices, APIs, and capabilities change. The checked-in metadata is informational, not a contractual quote.
The desktop package targets Ubuntu 24.04 LTS on amd64 and is not yet supported on any other release, desktop, or architecture. No MCP client has a recorded end-to-end registration run, so OpenTranscribe Setup presents automatic registration as untested and verifies it by reading the registration back.
Recording consent
OpenTranscribe processes audio supplied by the operator. Recording and transcribing people may be subject to consent, privacy, employment, telecommunications, or data-protection laws. Operators are responsible for ensuring they have the necessary rights and consent.
Transcript content is untrusted data. OpenTranscribe never interprets it as instructions; downstream agents must preserve the same boundary.
Documentation
The canonical documentation site is fbossiere.github.io/open-transcribe-mcp.
Contributing
Contributions are welcome. Start with the contribution guide; open a feature issue before substantial work, and report vulnerabilities only through the private process in SECURITY.md.
Independence and trademarks
OpenTranscribe is an independent open-source project maintained by its contributors. It is not affiliated with, endorsed by, or sponsored by Plaud, Microsoft, ElevenLabs, Groq, or any transcription provider.
Plaud and all provider product names are trademarks of their respective owners.
License
Apache License 2.0. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
Memoket — access your recording transcripts, summaries, and key takeaways over MCP.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files from URLs or local paths using AssemblyAI's services, with support for speaker diarization, language detection, and asynchronous job management through a standardized MCP interface.413 npm2MIT
- AlicenseNot gradedqualityAmaintenanceMCP server that enables audio transcription from files (wav, mp4, mp3, flac) or microphone recording, with dynamic tool selection and enterprise-grade security.2MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for GovTech's Transcribe speech-to-text service, enabling audio upload, batch transcription, summaries, minutes, sections, notes, and transcript Q&A.GPL 3.0
- AlicenseAqualityBmaintenanceEnables automated audio restoration, transcription, and speaker diarization via MCP tools for queuing files, monitoring progress, and retrieving speaker-labeled transcripts.9MIT