immersive-audio-tools
by thirv
README.md
# audio-mcp-mixer
Audio generation with MCP functions, descriptions and file database. Individual sounds are "created" by vector search of the closest audio file from the available file database, based on LAION-CLAP embeddings. Mixes are created by utilizing the [EBU ADM Renderer](https://github.com/ebu/ebu_adm_renderer).
## Features
- **Semantic Audio Search**: Uses CLAP embeddings to find sounds by text description
- **Spatial Audio Mixing**: Supports stereo, 5.1 surround, and 7.1.4 immersive formats
- **Mix Analysis**: Analyze multichannel mixes to identify sounds, positions, and spatial balance
- **Mix Mimicry**: Use embeddings to create new mixes similar to existing ones
- **HuggingFace Integration**: Access thousands of sounds from popular audio datasets
- **Orchestration**: LLM-powered decision making for intelligent mixing
- **On-Demand Documentation**: Agent reads documentation as needed using file tools
## Documentation
- **RENDERER_TOOL_GUIDE.md**: API reference for spatial audio rendering and mixing
- **SPATIAL_AUDIO_GUIDE.md**: Guide to spatial positioning, speaker layouts, and mixing strategies
## Quick Start
### 1. Install dependencies
```bash
pip install -r requirements.txt
```
This installs the core stack — the EBU EAR renderer (`ear`), the MCP server/transport (`mcp`), and the CLAP/FAISS/HuggingFace embedding stack used for audio search.
### 2. Set up HuggingFace audio database
```bash
python database_hf.py
```
This downloads and indexes popular audio datasets from HuggingFace.
### 3. Create text corpus index (for mix analysis)
```python
from database_hf import create_text_corpus_index
create_text_corpus_index()
```
This creates a FAISS index of text descriptions for labeling sounds during mix analysis.
The MCP application of your choice will use available tools to create immersive audio mixes based on language descriptions.
## MCP setup
The server speaks MCP over **stdio** (`python mcp_server.py`), so any MCP client can drive it. It depends on the `mcp` package (installed via `pip install -r requirements.txt`, step 1). Point `command` at the `audio-renderer-agent` conda environment's `python` (an absolute path is safest for desktop clients, which don't inherit an activated shell).
### VS Code
A project config is already provided at `.vscode/mcp.json`:
```json
{
"servers": {
"immersive-audio-tools": {
"type": "stdio",
"command": "python",
"args": ["${workspaceFolder}/mcp_server.py"]
}
}
}
```
In **Settings**, enable `chat.tools.automaticallyCreateMcpServers` (or use the **MCP** button in the chat input) to add the project server. Make sure the `audio-renderer-agent` conda environment is the selected Python interpreter.
### Claude Desktop
Add to your `claude_desktop_config.json` (replace the interpreter and path with your absolute paths):
```json
{
"mcpServers": {
"immersive-audio-tools": {
"command": "<conda_env>/bin/python",
"args": ["/abs/path/to/audio-renderer-agent/mcp_server.py"]
}
}
}
```
### Cursor
Add a project-level `.cursor/mcp.json`:
```json
{
"mcpServers": {
"immersive-audio-tools": {
"command": "python",
"args": ["${workspaceFolder}/mcp_server.py"]
}
}
}
```
### What the client gets
- **Tools:** `generate_single_sound`, `render_spatial_mix`, `sum_mixes`, `analyze_mix_spatial`, `mimic_mix_direct`, `mark_complete`, `read_file`
- **Resources:** `SPATIAL_AUDIO_GUIDE.md`, `RENDERER_TOOL_GUIDE.md` (read these before composing positions)
A hermetic smoke test drives the real server over stdio (tool listing, resources, a full stereo render + analysis, and the failure modes):
```bash
pytest test_mcp_server.py -v
```
## Licensing
This project is dual-licensed to separate the research content from the functional code:
* **Code**: All software code, scripts, and notebooks (`.py`, `.ipynb`, etc.) are licensed under the [MIT License](LICENSE).
* **Research & Documentation**: All written research, markdown files (`.md`), documentation, and data assets are licensed under the [Creative Commons Attribution 4.0 International License (CC BY 4.0)](LICENSE-CONTENT).
If you use or adapt this work, please provide attribution to the original authors.
## Citation
If you use this research or code in your work, please cite it as follows:
```bibtex
@software{mcp-mixer2026,
author = {Hirvonen, Toni},
title = {MCP Mixer for Audio},
month = sep,
year = {2026},
publisher = {GitHub},
version = {1.0.0},
url = {https://github.com/thirv/audio-mcp-mixer},
doi = {10.5281/zenodo.23074340}
}
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues