Open Router Audio Transcription MCP
README.md
# Open Router Audio Transcription MCP
An MCP server that transcribes audio files using OpenRouter's audio-capable language models.
## Features
- **Verbatim transcription** — exact word-for-word output including filler words, false starts, and repetitions
- **Cleaned transcription** — lightly edited for readability: removes fillers, adds punctuation, sentence boundaries, and paragraph breaks; omits content not intended for transcription
- **Custom prompt transcription** — direct the transcription with your own prompt for specialized use cases
## Supported Models
| Model | Provider |
|-------|----------|
| `google/gemini-3-flash-preview` (default standard) | Google |
| `google/gemini-3.1-flash-lite-preview` (default budget) | Google |
| `xiaomi/mimo-v2-omni` | Xiaomi |
| `openai/gpt-audio` | OpenAI |
| `openai/gpt-audio-mini` (budget) | OpenAI |
| `mistralai/voxtral-small-24b-2507` | Mistral |
| `openai/gpt-4o-audio-preview` | OpenAI |
## Supported Audio Formats
mp3, wav, ogg, flac, m4a, aac, webm, wma, opus
## Setup
### 1. Get an OpenRouter API key
Sign up at [openrouter.ai](https://openrouter.ai) and create an API key at [openrouter.ai/keys](https://openrouter.ai/keys).
### 2. Add to Claude Code
Run the following command to add the MCP server to Claude Code:
`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- npx -y or-audio-transcription-mcp`
Or add it manually to your Claude Code MCP settings (`~/.claude/settings.json` or project `.mcp.json`):
```json
{
"mcpServers": {
"audio-transcription": {
"command": "npx",
"args": ["-y", "or-audio-transcription-mcp"],
"env": {
"OPENROUTER_API_KEY": "your-api-key-here"
}
}
}
}
```
### Alternative: Install from source
```bash
git clone https://github.com/danielrosehill/OR-Audio-Transcription-MCP.git
cd OR-Audio-Transcription-MCP
npm install
npm run build
```
Then configure with a direct path:
`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- node /path/to/OR-Audio-Transcription-MCP/dist/index.js`
## Tools
### `transcribe_audio`
Transcribe an audio file.
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `file_path` | string | Yes | Absolute path to the audio file |
| `mode` | `"verbatim"` \| `"cleaned"` \| `"custom"` | Yes | Transcription mode |
| `custom_prompt` | string | When mode=custom | Custom prompt to direct the transcription |
| `model` | string | No | OpenRouter model ID (defaults to `google/gemini-3-flash-preview`) |
| `budget` | boolean | No | Use budget model (`google/gemini-3.1-flash-lite-preview`). Ignored if `model` is set |
### `list_transcription_models`
Lists all available audio transcription models.
## License
MIT
TDQS
A3.5/5.0
Scored across 2 tools
Disambiguation5/5
The two tools have completely distinct purposes: one discovers available transcription models, the other performs the transcription. There is no plausible way to confuse them.
Naming Consistency5/5
Both names follow a consistent snake_case verb_noun pattern (list_transcription_models, transcribe_audio) and are predictable for the domain.
Tool Count3/5
Two tools is on the thin side for a server, even a focused one; a discovery tool plus a single action tool leaves little room for variation. It is workable but borderline minimal.
Completeness4/5
The core workflow (find a model, then transcribe) is covered end to end. Some useful operations like batch transcription or job status are absent, but they are minor gaps an agent can work around.