Skip to main content
Glama
danielrosehill

Open Router Audio Transcription MCP

README.md
# Open Router Audio Transcription MCP

An MCP server that transcribes audio files using OpenRouter's audio-capable language models.

## Features

- **Verbatim transcription** — exact word-for-word output including filler words, false starts, and repetitions
- **Cleaned transcription** — lightly edited for readability: removes fillers, adds punctuation, sentence boundaries, and paragraph breaks; omits content not intended for transcription
- **Custom prompt transcription** — direct the transcription with your own prompt for specialized use cases

## Supported Models

| Model | Provider |
|-------|----------|
| `google/gemini-3-flash-preview` (default standard) | Google |
| `google/gemini-3.1-flash-lite-preview` (default budget) | Google |
| `xiaomi/mimo-v2-omni` | Xiaomi |
| `openai/gpt-audio` | OpenAI |
| `openai/gpt-audio-mini` (budget) | OpenAI |
| `mistralai/voxtral-small-24b-2507` | Mistral |
| `openai/gpt-4o-audio-preview` | OpenAI |

## Supported Audio Formats

mp3, wav, ogg, flac, m4a, aac, webm, wma, opus

## Setup

### 1. Get an OpenRouter API key

Sign up at [openrouter.ai](https://openrouter.ai) and create an API key at [openrouter.ai/keys](https://openrouter.ai/keys).

### 2. Add to Claude Code

Run the following command to add the MCP server to Claude Code:

`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- npx -y or-audio-transcription-mcp`

Or add it manually to your Claude Code MCP settings (`~/.claude/settings.json` or project `.mcp.json`):

```json
{
  "mcpServers": {
    "audio-transcription": {
      "command": "npx",
      "args": ["-y", "or-audio-transcription-mcp"],
      "env": {
        "OPENROUTER_API_KEY": "your-api-key-here"
      }
    }
  }
}
```

### Alternative: Install from source

```bash
git clone https://github.com/danielrosehill/OR-Audio-Transcription-MCP.git
cd OR-Audio-Transcription-MCP
npm install
npm run build
```

Then configure with a direct path:

`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- node /path/to/OR-Audio-Transcription-MCP/dist/index.js`

## Tools

### `transcribe_audio`

Transcribe an audio file.

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `file_path` | string | Yes | Absolute path to the audio file |
| `mode` | `"verbatim"` \| `"cleaned"` \| `"custom"` | Yes | Transcription mode |
| `custom_prompt` | string | When mode=custom | Custom prompt to direct the transcription |
| `model` | string | No | OpenRouter model ID (defaults to `google/gemini-3-flash-preview`) |
| `budget` | boolean | No | Use budget model (`google/gemini-3.1-flash-lite-preview`). Ignored if `model` is set |

### `list_transcription_models`

Lists all available audio transcription models.

## License

MIT

TDQS

A3.5/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: one discovers available transcription models, the other performs the transcription. There is no plausible way to confuse them.

Naming Consistency5/5

Both names follow a consistent snake_case verb_noun pattern (list_transcription_models, transcribe_audio) and are predictable for the domain.

Tool Count3/5

Two tools is on the thin side for a server, even a focused one; a discovery tool plus a single action tool leaves little room for variation. It is workable but borderline minimal.

Completeness4/5

The core workflow (find a model, then transcribe) is covered end to end. Some useful operations like batch transcription or job status are absent, but they are minor gaps an agent can work around.