jackai-stt-mcp
OfficialProvides tools for transcribing audio files using OpenAI's speech-to-text models, supporting file paths, URLs, and base64 input, with options for model selection, language, and prompt.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jackai-stt-mcptranscribe meeting.mp3 and tell me who said what"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jackai-stt-mcp
Transcribe audio by asking for it in plain language:
You: transcribe ~/Downloads/voice.opus
Claude: تمام تمام، موضوع اللغة العربية إن شاء الله محلول…
An MCP server that runs on your own machine, so the assistant can open the file directly. Nothing to run, no commands to memorise — you mention a file and the assistant does the rest.
No uploading, no copying, no base64. Your OpenAI key never leaves your computer.
Arabic works well, including dialect.
Install
Add this to your MCP client's config. uvx fetches and runs the package on
first use — nothing to install by hand.
{
"servers": {
"jackai-stt": {
"command": "uvx",
"args": ["jackai-stt-mcp"],
"env": {
"OPENAI_API_KEY": "sk-proj-..."
}
}
}
}Where the config lives:
Client | File |
VS Code |
|
Claude Desktop |
|
Claude Code |
|
Cursor |
|
Restart the client, then ask it to transcribe something.
Get an API key at platform.openai.com/api-keys. The account needs credit; transcription runs about $0.0045 per minute.
Related MCP server: Whisper MCP Server
Usage
Nothing to run, no syntax to learn. Talk to the assistant the way you normally would and mention the file:
transcribe ~/Downloads/voice.opus
what does the voice note on my desktop say?
read ./recordings/call.m4a and summarise what the client is asking for
transcribe meeting.mp3 and tell me who said what
The assistant recognises that it needs this tool and fills in the arguments from what you said. You never call it yourself or write out its parameters.
That last example needs speaker labels, which only one model produces. Mention it in passing and the assistant picks it:
transcribe meeting.mp3 with the diarize model
Same for a language it keeps mishearing ("it's Egyptian Arabic"), or names it should spell correctly ("the speakers are Ahmad and Sara"). Ordinary words — the table below is just what those words map onto.
What the assistant fills in
You don't set these by hand; this is here so you know what it can control.
Argument | Default | Notes |
| — | Path on this machine. Absolute, relative, or |
| — | Public http(s) URL, downloaded then transcribed. |
| — | For short clips. Prefer |
|
| See the table below. |
| auto-detect | ISO-639-1 ( |
| — | Names or jargon likely in the audio, to steer spelling. |
Exactly one audio source per request.
Models
Model | Cost/min | Speaker labels |
| $0.0045 | no |
| $0.003 | no |
| $0.006 | no |
| $0.006 | yes |
| $0.006 | no |
gpt-transcribe is OpenAI's recommended model: cheaper than whisper-1 and
more accurate. There is no reason to pick whisper-1 unless you need it
specifically.
Formats
flac m4a mp3 mp4 mpeg mpga oga ogg wav webm — up to 25 MB.
.opus files work too. OpenAI rejects that extension even though the bytes are
what it accepts as .ogg, so this server renames it in flight. WhatsApp voice
notes are all .opus, which is exactly the case that would otherwise fail.
Why it runs locally
A remote MCP server cannot read a file you attached in chat. There is no mechanism in the protocol that carries attachments to a server — it has been an open issue since early 2025, and the proposals to add one are still drafts. Base64 through a tool argument works in theory but dies on client argument-size limits after a second or two of audio.
Running on your machine avoids all of it. The server is a process you own, reading a file you own, with a key you hold.
Development
git clone https://github.com/jack-ai-net/stt-mcp
cd stt-mcp
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytestTests stub the OpenAI call, so they need no API key and cost nothing.
License
MIT — see LICENSE.
Built by JackAI.
Available Tools
1 tooltranscribe_audioARead-onlyIdempotent
Transcribe ONE audio file and return its text.
Arabic is fully supported: return the Arabic text as transcribed, and do not translate it unless the user asks.
Reads the file directly from this machine, so a path is all that's needed — no uploading, copying, or encoding. Pass exactly one audio source.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | 'gpt-transcribe' (default) is OpenAI's recommended model and the cheapest accurate one. Use 'gpt-4o-transcribe-diarize' to label who is speaking — it is the only model that does. | gpt-transcribe |
| prompt | No | Names or terms likely in the audio, to steer spelling. Ignored by gpt-4o-transcribe-diarize, which rejects prompts. | |
| filename | No | Filename to report for audio_base64. Its extension tells OpenAI the format, so keep it accurate (voice.ogg, note.mp3, clip.wav). | audio.mp3 |
| language | No | ISO-639-1 hint such as 'ar' or 'en'. Leave unset to auto-detect; set it only when detection gets the language wrong. | |
| audio_url | No | Public http(s) URL to download and transcribe. | |
| file_path | No | Path to the audio file on THIS machine. Absolute, relative, or starting with ~. This is the normal way to use the tool — the server runs locally, so no upload is involved. Pass exactly one of file_path, audio_url, audio_base64. | |
| audio_base64 | No | Base64-encoded audio, for short clips. Prefer file_path: base64 passes through the conversation and hits client size limits fast. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already state readOnlyHint, idempotentHint, and destructiveHint false, covering safety. The description adds valuable behavior: it reads files directly from the machine, preserves Arabic text without translation unless requested, and enforces a single audio source constraint. This enriched context goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at four sentences, immediately front-loads the core purpose, and each sentence adds a distinct piece of value: purpose, language handling, local file access, and input constraint. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (not shown but indicated) and strong per-parameter schema descriptions. The description provides a complete picture: what it does, how to invoke it (local path), which inputs to provide, and a notable edge case (Arabic). Combined with annotations, it is fully sufficient for an agent to select and call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema descriptions cover 100% of the 7 parameters, including details about model choices, prompt behavior, base64 limitations, and file_path being the normal approach. The description reinforces the file_path strategy ('no uploading, copying, or encoding') but does not significantly add meaning beyond the schema. Thus, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Transcribe ONE audio file and return its text.' This clearly identifies what the tool does and even distinguishes it from batch transcription via the emphatic 'ONE'. With no sibling tools, no differentiation is required, making this a fully clear purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful usage context: it explains the tool reads directly from the local machine ('no uploading, copying, or encoding'), emphasizes passing exactly one audio source, and mentions Arabic-language handling. It does not explicitly state when not to use the tool, but with no alternatives listed, the exclusionary guidance is less critical; the provided usage guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v0.1.1- First observed
transcribe_audio
TDQS
Only one tool exists, so there is no possibility of confusing it with others. Its purpose is clearly described.
The single tool follows a clear verb_noun pattern (transcribe_audio), which is consistent and predictable.
With only one tool, the server is extremely thin for its apparent scope as an STT service. A typical STT server would include multiple tools for broader coverage.
The server only provides basic single-file transcription, lacking common STT features like remote file handling, language selection, or batch processing. This is a significant gap for a dedicated STT MCP server.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
- mcpOAuthso.transcribe
Transcribe audio and video into speaker-labelled transcripts, subtitles, clips, and cited Q&A.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables advanced audio transcription, text-to-speech generation, and audio processing using OpenAI's Whisper and GPT-4o models with support for multiple audio formats, file management, and parallel processing.858MIT
- AlicenseAqualityFmaintenanceProvides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.31353MIT
- AlicenseAqualityDmaintenanceLocal speech-to-text transcription using Microsoft's VibeVoice-ASR model with speaker diarization, enabling audio transcription directly in AI tools like Claude Code, Cursor, and OpenCode.32MIT
- AlicenseNot gradedqualityAmaintenanceEnables transcription and speaker diarization of audio files, interviews, and YouTube URLs, producing speaker-attributed transcripts with timestamps. Supports multiple backends (local Whisper, OpenAI API) and output formats (txt, vtt, srt, json).Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/jack-ai-net/stt-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server