immersive-audio-tools
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@immersive-audio-toolsrender a 7.1.4 mix of a spaceship engine"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
audio-mcp-mixer
Audio generation with MCP functions, descriptions and file database. Individual sounds are "created" by vector search of the closest audio file from the available file database, based on LAION-CLAP embeddings. Mixes are created by utilizing the EBU ADM Renderer.
Features
Semantic Audio Search: Uses CLAP embeddings to find sounds by text description
Spatial Audio Mixing: Supports stereo, 5.1 surround, and 7.1.4 immersive formats
Mix Analysis: Analyze multichannel mixes to identify sounds, positions, and spatial balance
Mix Mimicry: Use embeddings to create new mixes similar to existing ones
HuggingFace Integration: Access thousands of sounds from popular audio datasets
Orchestration: LLM-powered decision making for intelligent mixing
On-Demand Documentation: Agent reads documentation as needed using file tools
Related MCP server: soundgrep
Documentation
RENDERER_TOOL_GUIDE.md: API reference for spatial audio rendering and mixing
SPATIAL_AUDIO_GUIDE.md: Guide to spatial positioning, speaker layouts, and mixing strategies
Quick Start
1. Install dependencies
pip install -r requirements.txtThis installs the core stack — the EBU EAR renderer (ear), the MCP server/transport (mcp), and the CLAP/FAISS/HuggingFace embedding stack used for audio search.
2. Set up HuggingFace audio database
python database_hf.pyThis downloads and indexes popular audio datasets from HuggingFace.
3. Create text corpus index (for mix analysis)
from database_hf import create_text_corpus_index
create_text_corpus_index()This creates a FAISS index of text descriptions for labeling sounds during mix analysis.
The MCP application of your choice will use available tools to create immersive audio mixes based on language descriptions.
MCP setup
The server speaks MCP over stdio (python mcp_server.py), so any MCP client can drive it. It depends on the mcp package (installed via pip install -r requirements.txt, step 1). Point command at the audio-renderer-agent conda environment's python (an absolute path is safest for desktop clients, which don't inherit an activated shell).
VS Code
A project config is already provided at .vscode/mcp.json:
{
"servers": {
"immersive-audio-tools": {
"type": "stdio",
"command": "python",
"args": ["${workspaceFolder}/mcp_server.py"]
}
}
}In Settings, enable chat.tools.automaticallyCreateMcpServers (or use the MCP button in the chat input) to add the project server. Make sure the audio-renderer-agent conda environment is the selected Python interpreter.
Claude Desktop
Add to your claude_desktop_config.json (replace the interpreter and path with your absolute paths):
{
"mcpServers": {
"immersive-audio-tools": {
"command": "<conda_env>/bin/python",
"args": ["/abs/path/to/audio-renderer-agent/mcp_server.py"]
}
}
}Cursor
Add a project-level .cursor/mcp.json:
{
"mcpServers": {
"immersive-audio-tools": {
"command": "python",
"args": ["${workspaceFolder}/mcp_server.py"]
}
}
}What the client gets
Tools:
generate_single_sound,render_spatial_mix,sum_mixes,analyze_mix_spatial,mimic_mix_direct,mark_complete,read_fileResources:
SPATIAL_AUDIO_GUIDE.md,RENDERER_TOOL_GUIDE.md(read these before composing positions)
A hermetic smoke test drives the real server over stdio (tool listing, resources, a full stereo render + analysis, and the failure modes):
pytest test_mcp_server.py -vLicensing
This project is dual-licensed to separate the research content from the functional code:
Code: All software code, scripts, and notebooks (
.py,.ipynb, etc.) are licensed under the MIT License.Research & Documentation: All written research, markdown files (
.md), documentation, and data assets are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
If you use or adapt this work, please provide attribution to the original authors.
Citation
If you use this research or code in your work, please cite it as follows:
@software{mcp-mixer2026,
author = {Hirvonen, Toni},
title = {MCP Mixer for Audio},
month = sep,
year = {2026},
publisher = {GitHub},
version = {1.0.0},
url = {https://github.com/thirv/audio-mcp-mixer},
doi = {10.5281/zenodo.23074340}
}This server cannot be deployed
Maintenance
Related MCP Connectors
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
CC0 sound effects API for AI agents — search, preview, and download via MCP.
Manage ElevenLabs voice agents and generate speech, music, sound effects, images, and video.
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables natural language control of Wwise audio middleware, allowing project auditing, batch editing, event creation, and reference checking through conversational interaction.201MIT
- AlicenseNot gradedqualityDmaintenanceEnables semantic search of local audio samples by describing sounds, MIDI generation, stem separation, and other music production tools from Claude Desktop.2MIT
- AlicenseNot gradedqualityDmaintenanceProvides audio inspection, conversion, processing, and generation capabilities via SoX, enabling AI agents to 'hear' and manipulate audio files through structured JSON interfaces.MIT
- AlicenseNot gradedqualityBmaintenanceEnables automated sound design spotting from Claude Code: analyze video, generate cue sheets, search or generate SFX, and export DAW-synchronized stems.8 npmMIT