Skip to main content
Glama
danielrosehill

Open Router Audio Transcription MCP

README.md
# Open Router Audio Transcription MCP

An MCP server that transcribes audio files using OpenRouter's audio-capable language models.

## Features

- **Verbatim transcription** — exact word-for-word output including filler words, false starts, and repetitions
- **Cleaned transcription** — lightly edited for readability: removes fillers, adds punctuation, sentence boundaries, and paragraph breaks; omits content not intended for transcription
- **Custom prompt transcription** — direct the transcription with your own prompt for specialized use cases

## Supported Models

| Model | Provider |
|-------|----------|
| `google/gemini-3-flash-preview` (default standard) | Google |
| `google/gemini-3.1-flash-lite-preview` (default budget) | Google |
| `xiaomi/mimo-v2-omni` | Xiaomi |
| `openai/gpt-audio` | OpenAI |
| `openai/gpt-audio-mini` (budget) | OpenAI |
| `mistralai/voxtral-small-24b-2507` | Mistral |
| `openai/gpt-4o-audio-preview` | OpenAI |

## Supported Audio Formats

mp3, wav, ogg, flac, m4a, aac, webm, wma, opus

## Setup

### 1. Get an OpenRouter API key

Sign up at [openrouter.ai](https://openrouter.ai) and create an API key at [openrouter.ai/keys](https://openrouter.ai/keys).

### 2. Add to Claude Code

Run the following command to add the MCP server to Claude Code:

`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- npx -y or-audio-transcription-mcp`

Or add it manually to your Claude Code MCP settings (`~/.claude/settings.json` or project `.mcp.json`):

```json
{
  "mcpServers": {
    "audio-transcription": {
      "command": "npx",
      "args": ["-y", "or-audio-transcription-mcp"],
      "env": {
        "OPENROUTER_API_KEY": "your-api-key-here"
      }
    }
  }
}
```

### Alternative: Install from source

```bash
git clone https://github.com/danielrosehill/OR-Audio-Transcription-MCP.git
cd OR-Audio-Transcription-MCP
npm install
npm run build
```

Then configure with a direct path:

`claude mcp add audio-transcription -e OPENROUTER_API_KEY=your-api-key-here -- node /path/to/OR-Audio-Transcription-MCP/dist/index.js`

## Tools

### `transcribe_audio`

Transcribe an audio file.

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `file_path` | string | Yes | Absolute path to the audio file |
| `mode` | `"verbatim"` \| `"cleaned"` \| `"custom"` | Yes | Transcription mode |
| `custom_prompt` | string | When mode=custom | Custom prompt to direct the transcription |
| `model` | string | No | OpenRouter model ID (defaults to `google/gemini-3-flash-preview`) |
| `budget` | boolean | No | Use budget model (`google/gemini-3.1-flash-lite-preview`). Ignored if `model` is set |

### `list_transcription_models`

Lists all available audio transcription models.

## License

MIT

TDQS

A3.5/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: one discovers available transcription models, the other performs the transcription. There is no plausible way to confuse them.

Naming Consistency5/5

Both names follow a consistent snake_case verb_noun pattern (list_transcription_models, transcribe_audio) and are predictable for the domain.

Tool Count3/5

Two tools is on the thin side for a server, even a focused one; a discovery tool plus a single action tool leaves little room for variation. It is workable but borderline minimal.

Completeness4/5

The core workflow (find a model, then transcribe) is covered end to end. Some useful operations like batch transcription or job status are absent, but they are minor gaps an agent can work around.

Maintenance

ActivityInactive
ResponsivenessNo issues