Skip to main content
Glama
README.md
# ofaudio-mcp

An **agent-first MCP server** for any audio an AI agent needs — **generate** (music, sound
effects, speech, multi-speaker dialogue) and **transform** (transcribe, caption, isolate/denoise,
voice-convert/dub) — behind one intent-shaped tool surface. It wraps **ElevenLabs** (speech,
dialogue, SFX, voice library/design/cloning, music, transcription, isolation, speech-to-speech)
and **Mureka** (music, lyrics), does the work in-process, and writes files to the local
filesystem. Games, audiobooks, podcasts, video, ads, accessibility, localization — all one MCP.

Built to be driven by an agent, never a human clicking a UI. Every generation call returns
a `job_id` instantly and runs in the background, so an agent can fan out many variants
without blocking, then collect them with `check_jobs`.

## Install

```bash
# from a checkout
pip install -e .
ofaudio-mcp            # or:  python -m ofaudio_mcp
```

Only hard dependency is the `mcp` SDK. Python 3.10+.

## Configure (environment)

Copy [`.env.example`](.env.example) to `.env` and fill in your keys (`.env` is gitignored).
Set at least one provider's key. Keys accept a **singular** or **plural** form; the plural
is comma-separated and becomes a multi-account pool (see [Parallelism](#parallelism)).

| Env var | Default | Purpose |
|---|---|---|
| `ELEVENLABS_API_KEY` / `ELEVENLABS_API_KEYS` | — | ElevenLabs key(s) — voice + SFX. Works with a subscription tier *or* pure pay-as-you-go. |
| `MUREKA_API_KEY` / `MUREKA_API_KEYS` | — | Mureka key(s) — music. Prepaid API credits. |
| `OFAUDIO_OUT_DIR` | `./ofaudio-out` | Where audio files are written. |
| `OFAUDIO_STATE_DIR` | `<OUT_DIR>/.state` | Job records (JSON) + `presets.json`. |
| `OFAUDIO_TOOLS` | all | Comma-separated tool **groups** to mount (`generate,speakers,voices,account,transform,pronunciation`) — fewer tools, less client context. See [Context cost](#context-cost). |
| `OFAUDIO_ELEVENLABS_LANES` | `3` | Concurrent requests **per ElevenLabs key**. |
| `OFAUDIO_MUREKA_LANES` | `1` | Concurrent requests **per Mureka key** (Mureka caps at 1/account). |
| `OFAUDIO_POLL_INTERVAL` | `5` | Seconds between polls for async (Mureka) jobs. |
| `OFAUDIO_JOB_TIMEOUT` | `300` | Max seconds to wait for one generation. |
| `OFAUDIO_MUREKA_MODEL` | `auto` | Pin a Mureka model (e.g. the V9 flagship string) once confirmed on a live key. |
| `OFAUDIO_ELEVENLABS_MODEL` | `eleven_v3` | TTS model; drop to `eleven_multilingual_v2` if v3 isn't on the key's tier. |
| `OFAUDIO_ELEVENLABS_VOICE` | Rachel | Default voice id when a call doesn't specify one. |
| `ELEVENLABS_BASE_URL` / `MUREKA_BASE_URL` | official | Override API base URLs. |

A key entry may be `label:key` to name the lane in logs and job results
(e.g. `MUREKA_API_KEYS="acct-a:ak_...,acct-b:ak_..."`).

## Tools

27 tools in six groups (`generate`, `speakers`, `voices`, `account`, `transform`,
`pronunciation`); mount a subset with `OFAUDIO_TOOLS`. Every parameter is documented in the
tool schema itself, so an agent needs nothing beyond `tools/list`.

| Tool | What it does |
|---|---|
| `generate_music(prompt, duration?, instrumental?, reference?, lyrics?, engine?, format?, count?, composition_plan?, store_for_inpainting?, respect_sections_durations?)` | Themes, BGM, songs, instrumentals. Soft-defaults to Mureka (prompt→song writes its own lyrics; pass `lyrics` for your own; `reference` = local/URL track → vibe-match; `format` defaults to lossless **flac**; `count` 1–3). `engine="elevenlabs"` = Eleven Music (`composition_plan` replaces `prompt`; `store_for_inpainting` → `song_id`) → `{job_id}` |
| `generate_lyrics(prompt)` | Write song lyrics from a theme (Mureka) — **synchronous** `{title, lyrics}`; review/edit before spending song credits |
| `transcribe(audio, language?, diarize?, captions?)` | Speech-to-text (ElevenLabs Scribe). `audio` = local path or a URL (YouTube/TikTok); `diarize` labels speakers; `captions="srt"/"vtt"` writes a subtitle file → `{text, language_code, words, captions_path?}` |
| `forced_alignment(audio, text, captions?)` | Align a known script to existing audio → exact per-word timings (+ optional SRT/VTT) — for captioning recordings you already have the transcript for |
| `isolate_audio(audio)` | Remove background noise / isolate the voice → `{path}` |
| `convert_voice(audio, voice, remove_background_noise?, format?, seed?)` | Speech-to-speech: re-voice a recording onto another `voice`, keeping delivery (dubbing/ADR) → `{path}` |
| `generate_sfx(prompt, duration_seconds?, loop?, prompt_influence?, model?, engine?, format?)` | One-shots, ambience, seamless loops → `{job_id}` |
| `generate_speech(text, voice?, profile?, speaker?, mood?, workspace?, model?, stability?, style?, speed?, similarity_boost?, use_speaker_boost?, seed?, previous_text?, next_text?, language_code?, with_timestamps?, variants?, pronunciation?, engine?, format?)` | Dialogue / character VO. Identity = a `voice` or a saved `speaker` (+ `mood` layered over its base delivery); shape delivery with a `profile` and/or precision knobs; inline v3 tags like `[whispers]` work; `with_timestamps=True` emits a per-character alignment sidecar (lip-sync/captions); `variants=N` (1–8) fans out N takes to audition → `{job_id}` (or `{job_ids}`) |
| `generate_dialogue(lines=[{voice, text}], model?, seed?, language_code?, apply_text_normalization?, engine?, format?)` | Multi-speaker scene in one call (≤10 voices, ~2000 chars) with natural turn-taking → `{job_id}` |
| `save_profile(name, voice?, model?, stability?, style?, speed?, …, tag?, seed?, description?)` | Save a reusable persona/mood preset (a preset over the speech knobs) |
| `save_speaker(name, voice, workspace?, model?, stability?, …, seed?, description?)` | Save a **speaker** (pinned voice + base delivery) under a workspace |
| `list_speakers(workspace?)` | Saved speakers (optionally by workspace), each with its workspace |
| `search_voice_library(query?, gender?, age?, accent?, use_case?, category?, descriptives?, page_size?)` | Search ElevenLabs' shared library of thousands of voices; every result carries a `preview_url` to audition — **no generation, no slots** (keep `query` to a word or two) |
| `add_library_voice(public_owner_id, voice_id, name, speaker?, workspace?, seed?, description?)` | Add a library voice to the account and save it under `name` (uses a voice slot); `speaker=` also saves it as a speaker in one step |
| `design_voice(description, save_as?, text?, model?, guidance_scale?, loudness?, seed?)` | Synthesize a **custom** voice from a text description; returns previews to audition |
| `save_voice(name, generated_voice_id, description?)` | Keep a designed preview as a permanent, reusable voice (uses a voice slot) |
| `clone_voice(name, audio_paths, description?, remove_background_noise?)` | Instant Voice Cloning from your audio file(s) → saved, reusable voice (uses a slot) |
| `delete_voice(voice, confirm=False)` | Delete a custom voice and free its slot (fresh name→id, blocks premade, warns on profile/speaker refs) |
| `account_status()` | **Both providers, one call** — ElevenLabs (tier, char quota, voice slots, cloning, + `quality`/formats the tier allows) and Mureka (credit `balance`, concurrency limit). Check before fan-out |
| `list_models()` | ElevenLabs models available on the account + capability flags |
| `check_jobs(job_ids)` | Batch-poll — returns status + file path per id |
| `list_jobs(status?, limit?)` | Recent jobs, newest first |
| `create_pronunciation_dictionary(name, rules, description?)` | Lock how words/lore-names are said — `rules` = `{word: say_as}` map or EL rule objects. Apply via `generate_speech(pronunciation=)` or `save_speaker(pronunciation=)` |
| `add_pronunciation_rules(dictionary, rules)` | Extend a saved dictionary (mints a new version) |
| `list_voices()` | ElevenLabs voices on the account (`{account, profiles}`) — resolve a voice by **name** or id |
| `get_voice(voice)` | One voice's details (category, labels, description, preview_url) |
| `list_presets()` | Everything saved, one call: voice aliases, music presets, profiles (built-in + custom), speakers, pronunciation dictionaries, workspaces |

Each `generate_*` has a soft-pinned provider (music→Mureka, SFX/speech→ElevenLabs) and an
optional `engine=` override. A result is `{job_id, status, path, duration, format, seed,
provider}` once succeeded.

## Voices, speakers & delivery

**The voice is the speaker's identity — pick or make one, don't dial it out of parameters.** Three
ways to get a voice, all reusable by name afterwards:

1. **Use one of ElevenLabs' many** — `search_voice_library(query="warm narrator", age="middle_aged")`
   → `add_library_voice(public_owner_id, voice_id, "narrator")`. Thousands of ready-made voices.
2. **Design a custom one** — `design_voice("warm, measured audiobook narrator, clear articulation",
   save_as="narrator")`. Describe it, audition the previews, keep the best.
3. **Reuse what's on the account** — `list_voices()` shows the premade voices already there.

Then just `generate_speech("...", voice="narrator")` — the voice stays consistent.

**Make it a first-class Speaker, grouped by workspace.** For a whole cast across several projects,
`save_speaker("narrator", voice, workspace="projectA", ...)` saves the *identity* (pinned voice + base
delivery), and `generate_speech(text, speaker="narrator", mood="calm", workspace="projectA")` speaks
it — the mood layers over the speaker (a speaker is *who*; a profile is *what mood*).
`search_voice_library(query)` returns candidates with a `preview_url` each (no slots), and
`add_library_voice(..., speaker="narrator", workspace="projectA")` adds + saves in one step.
`variants=N` on `generate_speech` fans out N takes so a human picks the best.

**Delivery (mood/tone) is a separate layer on top.** Shape it with a `profile` and/or the
precision knobs. Built-in profiles: `neutral, narrator, calm, whisper, sad, angry, excited,
shout` (v3, emotion via audio tags) and `steady, dynamic` (v2, consistency axis). Save your
own persona/mood presets with `save_profile`.

Two facts worth knowing (both handled for you, but they explain the design):

- **Eleven v3** (default for speech) does emotion via inline **audio tags** (`[whispers]`,
  `[angry]`, `[snorts]`, …) and a 3-mode **stability** (`creative`/`natural`/`robust`) — it
  ignores the numeric `style` slider. Tags only land on an **expressive** voice; a flat
  narration voice mutes them.
- **Eleven v2** uses the numeric sliders, but those control *consistency* (dynamic ↔ steady),
  **not a specific emotion**. Use it via the `steady`/`dynamic` profiles.

## Parallelism

Mureka caps concurrency at **1 request per prepaid account** and it doesn't stack, so the
way to run music generation in parallel is **multiple accounts**. This server treats keys
as a pool: supply N keys and it runs N jobs in parallel, one per key, automatically routing
around a key that hits its rate limit (back-off) or runs out of balance. Supplying one key
is just the N=1 case. (ElevenLabs scales concurrency within a single key by tier, so a lane
count > 1 per key is the norm there.)

## Mount it

**Claude Code**
```bash
claude mcp add ofaudio \
  --env ELEVENLABS_API_KEY=xi_... --env MUREKA_API_KEY=ak_... \
  -- python -m ofaudio_mcp
```

**Cursor** — `.cursor/mcp.json`:
```json
{
  "mcpServers": {
    "ofaudio": {
      "command": "python",
      "args": ["-m", "ofaudio_mcp"],
      "env": { "ELEVENLABS_API_KEY": "xi_...", "MUREKA_API_KEY": "ak_..." }
    }
  }
}
```

**Self-hosted / server** — add the server to your MCP client's config and put the keys in a
`.env` next to it (they stay on the box):
```
ELEVENLABS_API_KEYS=xi_...
MUREKA_API_KEYS=acct-a:ak_...,acct-b:ak_...
OFAUDIO_OUT_DIR=/var/audio/ofaudio-out
```
The intended usage pattern is fan-out: fire many `generate_*` calls, then `check_jobs`.

## Context cost

Most MCP clients load every tool schema into the model's context on every turn, so the size
of `tools/list` is a real cost. Claude Code defers schemas until a tool is used, but truncates
any tool description longer than **2048 characters**. The surface is shaped around both:

- Each tool's docstring is ≤ 700 chars; every parameter is documented in its own schema
  `description` (nothing gets cut, and schema-rendering clients see real docs, not `title`s).
- Tools carry MCP annotations (`readOnlyHint`, `destructiveHint`) so clients can skip approval
  prompts on reads and warn on `delete_voice`.
- The advertised schemas are compacted at startup (pydantic's `title` / `anyOf null` /
  `default: null` noise stripped; validation unchanged).
- `OFAUDIO_TOOLS=generate,account` mounts only the groups a project needs.
- The usage guide is served in-band as the MCP resource `ofaudio://guide` (`text/markdown`)
  for clients that support resources; the same text ships as the `skills/ofaudio` skill.

| Surface | Tools | `tools/list` | ≈ tokens |
|---|---|---|---|
| before this shaping | 33 | 42.3 KB | 12.1k |
| all six groups | 27 | 33.7 KB | 9.6k |
| `OFAUDIO_TOOLS=generate,account` | 9 | 13.7 KB | 3.9k |

Measure it with `python scripts/tool_surface.py` (honours `OFAUDIO_TOOLS`; exits 1 if any
description exceeds 2048 chars).

## Status

Both providers work end-to-end. **ElevenLabs**: speech (+ timestamps, mood profiles,
`variants`), multi-speaker dialogue, SFX, Eleven Music, voice library / design / instant
cloning, voice lifecycle, account introspection, the speaker/workspace layer, and the transform
tools (transcribe, forced alignment, isolate, voice-convert). **Mureka**: instrumental + song,
prompt→song (easy-generate), user-supplied lyrics, standalone `generate_lyrics`, reference-track
vibe-match, and lossless flac-default output.

## Design

See [`CLAUDE.md`](CLAUDE.md) for the architecture, module map, tool surface, and the verified
provider facts that shape the design.

## License

Source-available under **[PolyForm Noncommercial 1.0.0](LICENSE)**, with a licensor carve-out
that draws the line at the **tool**, not the **audio**:

- ✅ **Run it freely for any purpose — including to generate audio you use commercially.** This
  MCP is only a client; it claims no rights in what you generate. Your rights to the generated
  audio are governed by ElevenLabs' and Mureka's terms and your own accounts with them.
- ❌ **Don't commercialize the tool itself** — you may not sell this software, host it as a paid
  service, or turn it (or a derivative) into a commercial product.

In short: *make whatever audio you want, sell it if your provider terms allow — just don't sell
the MCP.* See [`LICENSE`](LICENSE) for the full terms.

TDQS

A3.7/5.0

Scored across 33 tools

Disambiguation2/5

Several voice-related tools have overlapping responsibilities, such as preview_voices vs search_voice_library and add_speaker_from_library vs add_library_voice. These overlaps create ambiguous boundaries, even with detailed descriptions, making it hard for an agent to pick the correct tool confidently.

Naming Consistency4/5

The majority of tools follow a clear verb_noun pattern (generate_*, list_*, save_*, delete_*). Minor deviations include standalone verbs like 'speak' and 'transcribe', plus noun-based names like 'account_status' and 'forced_alignment', but the overall style is recognizable and mostly consistent.

Tool Count2/5

With 33 tools, the server feels overloaded. Many are small variations on voice management (add, save, list, preview, search, clone, delete), inflating the count well beyond the typical 3-15 tool range for a focused server.

Completeness4/5

The domain is broad and well covered: music and speech generation, voice cloning/design/library management, speech customization (profiles, speakers, pronunciation), transcription, forced alignment, noise isolation, voice conversion, and job tracking. Minor gaps include no delete for speakers/profiles, but these can be worked around by overwriting.

Maintenance

ActivityMaintained
ResponsivenessNo issues