Skip to main content
Glama

audio-mcp-mixer

Audio generation with MCP functions, descriptions and file database. Individual sounds are "created" by vector search of the closest audio file from the available file database, based on LAION-CLAP embeddings. Mixes are created by utilizing the EBU ADM Renderer.

Features

  • Semantic Audio Search: Uses CLAP embeddings to find sounds by text description

  • Spatial Audio Mixing: Supports stereo, 5.1 surround, and 7.1.4 immersive formats

  • Mix Analysis: Analyze multichannel mixes to identify sounds, positions, and spatial balance

  • Mix Mimicry: Use embeddings to create new mixes similar to existing ones

  • HuggingFace Integration: Access thousands of sounds from popular audio datasets

  • Orchestration: LLM-powered decision making for intelligent mixing

  • On-Demand Documentation: Agent reads documentation as needed using file tools

Related MCP server: soundgrep

Documentation

  • RENDERER_TOOL_GUIDE.md: API reference for spatial audio rendering and mixing

  • SPATIAL_AUDIO_GUIDE.md: Guide to spatial positioning, speaker layouts, and mixing strategies

Quick Start

1. Install dependencies

pip install -r requirements.txt

This installs the core stack — the EBU EAR renderer (ear), the MCP server/transport (mcp), and the CLAP/FAISS/HuggingFace embedding stack used for audio search.

2. Set up HuggingFace audio database

python database_hf.py

This downloads and indexes popular audio datasets from HuggingFace.

3. Create text corpus index (for mix analysis)

from database_hf import create_text_corpus_index
create_text_corpus_index()

This creates a FAISS index of text descriptions for labeling sounds during mix analysis.

The MCP application of your choice will use available tools to create immersive audio mixes based on language descriptions.

MCP setup

The server speaks MCP over stdio (python mcp_server.py), so any MCP client can drive it. It depends on the mcp package (installed via pip install -r requirements.txt, step 1). Point command at the audio-renderer-agent conda environment's python (an absolute path is safest for desktop clients, which don't inherit an activated shell).

VS Code

A project config is already provided at .vscode/mcp.json:

{
  "servers": {
    "immersive-audio-tools": {
      "type": "stdio",
      "command": "python",
      "args": ["${workspaceFolder}/mcp_server.py"]
    }
  }
}

In Settings, enable chat.tools.automaticallyCreateMcpServers (or use the MCP button in the chat input) to add the project server. Make sure the audio-renderer-agent conda environment is the selected Python interpreter.

Claude Desktop

Add to your claude_desktop_config.json (replace the interpreter and path with your absolute paths):

{
  "mcpServers": {
    "immersive-audio-tools": {
      "command": "<conda_env>/bin/python",
      "args": ["/abs/path/to/audio-renderer-agent/mcp_server.py"]
    }
  }
}

Cursor

Add a project-level .cursor/mcp.json:

{
  "mcpServers": {
    "immersive-audio-tools": {
      "command": "python",
      "args": ["${workspaceFolder}/mcp_server.py"]
    }
  }
}

What the client gets

  • Tools: generate_single_sound, render_spatial_mix, sum_mixes, analyze_mix_spatial, mimic_mix_direct, mark_complete, read_file

  • Resources: SPATIAL_AUDIO_GUIDE.md, RENDERER_TOOL_GUIDE.md (read these before composing positions)

A hermetic smoke test drives the real server over stdio (tool listing, resources, a full stereo render + analysis, and the failure modes):

pytest test_mcp_server.py -v

Licensing

This project is dual-licensed to separate the research content from the functional code:

If you use or adapt this work, please provide attribution to the original authors.

Citation

If you use this research or code in your work, please cite it as follows:

@software{mcp-mixer2026,
  author       = {Hirvonen, Toni},
  title        = {MCP Mixer for Audio},
  month        = sep,
  year         = {2026},
  publisher    = {GitHub},
  version      = {1.0.0},
  url          = {https://github.com/thirv/audio-mcp-mixer},
  doi          = {10.5281/zenodo.23074340}
}

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search of local audio samples by describing sounds, MIDI generation, stem separation, and other music production tools from Claude Desktop.
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides audio inspection, conversion, processing, and generation capabilities via SoX, enabling AI agents to 'hear' and manipulate audio files through structured JSON interfaces.
    MIT