Skip to main content
Glama
jack-ai-net

jackai-stt-mcp

Official
by jack-ai-net

jackai-stt-mcp

tests PyPI License: MIT

Transcribe audio by asking for it in plain language:

You: transcribe ~/Downloads/voice.opus

Claude: تمام تمام، موضوع اللغة العربية إن شاء الله محلول…

An MCP server that runs on your own machine, so the assistant can open the file directly. Nothing to run, no commands to memorise — you mention a file and the assistant does the rest.

No uploading, no copying, no base64. Your OpenAI key never leaves your computer.

Arabic works well, including dialect.

Install

Add this to your MCP client's config. uvx fetches and runs the package on first use — nothing to install by hand.

{
  "servers": {
    "jackai-stt": {
      "command": "uvx",
      "args": ["jackai-stt-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-proj-..."
      }
    }
  }
}

Where the config lives:

Client

File

VS Code

~/Library/Application Support/Code/User/mcp.json (macOS)

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json

Claude Code

.mcp.json in the project, or claude mcp add

Cursor

~/.cursor/mcp.json

Restart the client, then ask it to transcribe something.

Get an API key at platform.openai.com/api-keys. The account needs credit; transcription runs about $0.0045 per minute.

Related MCP server: Whisper MCP Server

Usage

Nothing to run, no syntax to learn. Talk to the assistant the way you normally would and mention the file:

transcribe ~/Downloads/voice.opus

what does the voice note on my desktop say?

read ./recordings/call.m4a and summarise what the client is asking for

transcribe meeting.mp3 and tell me who said what

The assistant recognises that it needs this tool and fills in the arguments from what you said. You never call it yourself or write out its parameters.

That last example needs speaker labels, which only one model produces. Mention it in passing and the assistant picks it:

transcribe meeting.mp3 with the diarize model

Same for a language it keeps mishearing ("it's Egyptian Arabic"), or names it should spell correctly ("the speakers are Ahmad and Sara"). Ordinary words — the table below is just what those words map onto.

What the assistant fills in

You don't set these by hand; this is here so you know what it can control.

Argument

Default

Notes

file_path

Path on this machine. Absolute, relative, or ~/....

audio_url

Public http(s) URL, downloaded then transcribed.

audio_base64

For short clips. Prefer file_path.

model

gpt-transcribe

See the table below.

language

auto-detect

ISO-639-1 (ar, en). Set only if detection is wrong.

prompt

Names or jargon likely in the audio, to steer spelling.

Exactly one audio source per request.

Models

Model

Cost/min

Speaker labels

gpt-transcribe (default)

$0.0045

no

gpt-4o-mini-transcribe

$0.003

no

gpt-4o-transcribe

$0.006

no

gpt-4o-transcribe-diarize

$0.006

yes

whisper-1

$0.006

no

gpt-transcribe is OpenAI's recommended model: cheaper than whisper-1 and more accurate. There is no reason to pick whisper-1 unless you need it specifically.

Formats

flac m4a mp3 mp4 mpeg mpga oga ogg wav webm — up to 25 MB.

.opus files work too. OpenAI rejects that extension even though the bytes are what it accepts as .ogg, so this server renames it in flight. WhatsApp voice notes are all .opus, which is exactly the case that would otherwise fail.

Why it runs locally

A remote MCP server cannot read a file you attached in chat. There is no mechanism in the protocol that carries attachments to a server — it has been an open issue since early 2025, and the proposals to add one are still drafts. Base64 through a tool argument works in theory but dies on client argument-size limits after a second or two of audio.

Running on your machine avoids all of it. The server is a process you own, reading a file you own, with a key you hold.

Development

git clone https://github.com/jack-ai-net/stt-mcp
cd stt-mcp
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest

Tests stub the OpenAI call, so they need no API key and cost nothing.

License

MIT — see LICENSE.

Built by JackAI.

Available Tools

1 tool
transcribe_audioA
Read-onlyIdempotent

Transcribe ONE audio file and return its text.

Arabic is fully supported: return the Arabic text as transcribed, and do not translate it unless the user asks.

Reads the file directly from this machine, so a path is all that's needed — no uploading, copying, or encoding. Pass exactly one audio source.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo'gpt-transcribe' (default) is OpenAI's recommended model and the cheapest accurate one. Use 'gpt-4o-transcribe-diarize' to label who is speaking — it is the only model that does.gpt-transcribe
promptNoNames or terms likely in the audio, to steer spelling. Ignored by gpt-4o-transcribe-diarize, which rejects prompts.
filenameNoFilename to report for audio_base64. Its extension tells OpenAI the format, so keep it accurate (voice.ogg, note.mp3, clip.wav).audio.mp3
languageNoISO-639-1 hint such as 'ar' or 'en'. Leave unset to auto-detect; set it only when detection gets the language wrong.
audio_urlNoPublic http(s) URL to download and transcribe.
file_pathNoPath to the audio file on THIS machine. Absolute, relative, or starting with ~. This is the normal way to use the tool — the server runs locally, so no upload is involved. Pass exactly one of file_path, audio_url, audio_base64.
audio_base64NoBase64-encoded audio, for short clips. Prefer file_path: base64 passes through the conversation and hits client size limits fast.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint, idempotentHint, and destructiveHint false, covering safety. The description adds valuable behavior: it reads files directly from the machine, preserves Arabic text without translation unless requested, and enforces a single audio source constraint. This enriched context goes beyond the structured annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at four sentences, immediately front-loads the core purpose, and each sentence adds a distinct piece of value: purpose, language handling, local file access, and input constraint. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown but indicated) and strong per-parameter schema descriptions. The description provides a complete picture: what it does, how to invoke it (local path), which inputs to provide, and a notable edge case (Arabic). Combined with annotations, it is fully sufficient for an agent to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema descriptions cover 100% of the 7 parameters, including details about model choices, prompt behavior, base64 limitations, and file_path being the normal approach. The description reinforces the file_path strategy ('no uploading, copying, or encoding') but does not significantly add meaning beyond the schema. Thus, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Transcribe ONE audio file and return its text.' This clearly identifies what the tool does and even distinguishes it from batch transcription via the emphatic 'ONE'. With no sibling tools, no differentiation is required, making this a fully clear purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful usage context: it explains the tool reads directly from the local machine ('no uploading, copying, or encoding'), emphasizes passing exactly one audio source, and mentions Arabic-language handling. It does not explicitly state when not to use the tool, but with no alternatives listed, the exclusionary guidance is less critical; the provided usage guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev0.1.1
    • First observedtranscribe_audio

TDQS

A4.1/5.0
Disambiguation5/5

Only one tool exists, so there is no possibility of confusing it with others. Its purpose is clearly described.

Naming Consistency5/5

The single tool follows a clear verb_noun pattern (transcribe_audio), which is consistent and predictable.

Tool Count2/5

With only one tool, the server is extremely thin for its apparent scope as an STT service. A typical STT server would include multiple tools for broader coverage.

Completeness2/5

The server only provides basic single-file transcription, lacking common STT features like remote file handling, language selection, or batch processing. This is a significant gap for a dedicated STT MCP server.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    F
    maintenance
    Provides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.
    3
    135
    3
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Local speech-to-text transcription using Microsoft's VibeVoice-ASR model with speaker diarization, enabling audio transcription directly in AI tools like Claude Code, Cursor, and OpenCode.
    3
    2
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jack-ai-net/stt-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server