Open Router Audio Transcription MCP
README.md
# Open Router Audio Transcription MCP
An MCP server that transcribes audio files using OpenRouter's audio-capable language models.
## Features
- **Verbatim transcription** — exact word-for-word output including filler words, false starts, and repetitions
- **Cleaned transcription** — lightly edited for readability: removes fillers, adds punctuation, sentence boundaries, and paragraph breaks; omits content not intended for transcription
- **Custom prompt transcription** — direct the transcription with your own prompt for specialized use cases
## Supported Models
| Model | Provider |
|-------|----------|
| `google/gemini-3-flash-preview` (default standard) | Google |
| `google/gemini-3.1-flash-lite-preview` (default budget) | Google |
| `xiaomi/mimo-v2-omni` | Xiaomi |
| `openai/gpt-audio` | OpenAI |
| `openai/gpt-audio-mini` (budget) | OpenAI |
| `mistralai/voxtral-small-24b-2507` | Mistral |
| `openai/gpt-4o-audio-preview` | OpenAI |
## Supported Audio Formats
mp3, wav, ogg, flac, m4a, aac, webm, wma, opus
## Setup
### 1. Get an OpenRouter API key
Sign up at [openrouter.ai](https://openrouter.ai) and create an API key at [openrouter.ai/keys](https://openrouter.ai/keys).
### 2. Add to Claude Code
Run the following command to add the MCP server to Claude Code:
`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- npx -y or-audio-transcription-mcp`
Or add it manually to your Claude Code MCP settings (`~/.claude/settings.json` or project `.mcp.json`):
```json
{
"mcpServers": {
"audio-transcription": {
"command": "npx",
"args": ["-y", "or-audio-transcription-mcp"],
"env": {
"OPENROUTER_API_KEY": "your-api-key-here"
}
}
}
}
```
### Alternative: Install from source
```bash
git clone https://github.com/danielrosehill/OR-Audio-Transcription-MCP.git
cd OR-Audio-Transcription-MCP
npm install
npm run build
```
Then configure with a direct path:
`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- node /path/to/OR-Audio-Transcription-MCP/dist/index.js`
## Tools
### `transcribe_audio`
Transcribe an audio file.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `file_path` | string | Yes | Absolute path to the audio file |
| `mode` | `"verbatim"` \| `"cleaned"` \| `"custom"` | Yes | Transcription mode |
| `custom_prompt` | string | When mode=custom | Custom prompt to direct the transcription |
| `model` | string | No | OpenRouter model ID (defaults to `google/gemini-3-flash-preview`) |
| `budget` | boolean | No | Use budget model (`google/gemini-3.1-flash-lite-preview`). Ignored if `model` is set |
### `list_transcription_models`
Lists all available audio transcription models.
## License
MIT
TDQS
A3.5/5.0
Scored across 2 tools
Disambiguation5/5
The two tools have completely distinct purposes: one discovers available transcription models, the other performs the transcription. There is no plausible way to confuse them.
Naming Consistency5/5
Both names follow a consistent snake_case verb_noun pattern (list_transcription_models, transcribe_audio) and are predictable for the domain.
Tool Count3/5
Two tools is on the thin side for a server, even a focused one; a discovery tool plus a single action tool leaves little room for variation. It is workable but borderline minimal.
Completeness4/5
The core workflow (find a model, then transcribe) is covered end to end. Some useful operations like batch transcription or job status are absent, but they are minor gaps an agent can work around.
Maintenance
ActivityInactive
ResponsivenessNo issues