Audio Sonic MCP
Allows submitting YouTube URLs for asynchronous audio analysis, returning a sonic signature with tempo, key, vibe tags/embeddings, and production profile.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Audio Sonic MCPAnalyze the sonic signature of this track: https://youtu.be/dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ต Audio Sonic MCP
Turn any song into a structured "sonic signature" โ extracting tempo, musical key, a 512-dimension CLAP vibe embedding, human-readable vibe tags, and a production profile โ from a single local call.
Audio Sonic MCP runs entirely on your local machine (requiring no API keys, external servers, or cloud dependencies) and exposes two premium access points to the same underlying high-fidelity audio analysis engine:
Tailored For | Core Interface & Mechanics | |
๐ค MCP Server | LLMs, AI agents, & IDEs (Claude, Cursor, Windsurf, Cline) | Asynchronous, fire-and-forget analysis of YouTube URLs. Avoids blocking client LLMs during heavy audio processing. |
๐๏ธ Local CLI | Musicians, sound producers, & audio engineers | Deep command-line tool targeting local files for full-song multi-window analysis and high-fidelity output. |
๐น Quick Taste: What You Get
1. Musician-Friendly CLI Summary (--summary mode)
๐ต SONIC SIGNATURE โ my_demo.mp3 (3:24)
TEMPO 153.8 BPM (steady)
KEY G Major ยท shifts to G Phrygian @0:30 (confidence 74%)
VIBE aggressive ยท dark ยท driving ยท hip-hop ยท gritty
PRODUCTION
Vocals forward
Punch 0.62 (moderate)
Stereo wide
Low end ~55 Hz dominant
Overall confidence: 88% ยท analyzed in 0:28 (GPU-accelerated)2. Comprehensive JSON (Returned by MCP and CLI by default)
{
"header": {
"job_id": "sig_a3f9b2c1",
"status": "success",
"confidence_score": 0.88,
"source_metadata": {
"title": "Acoustic Vibe Demo",
"duration_sec": 204,
"source_type": "file"
}
},
"sonic_signature": {
"bpm": 153.8,
"bpm_engine": "madmom",
"bpm_variable": false,
"key": "G Major",
"key_variable": true,
"key_map": [
{ "start_sec": 0.0, "end_sec": 30.0, "key": "G Major" },
{ "start_sec": 30.0, "end_sec": 90.0, "key": "G Phrygian" }
],
"mode_confidence": 0.74,
"vibe_vector": [0.012, -0.034, "... 512 float dimensions ..."],
"vibe_tags": ["aggressive", "dark", "driving", "hip-hop", "gritty"],
"production_profile": {
"vocal_presence": "forward",
"transient_punch": 0.62,
"stereo_width": "wide",
"dominant_freq_peaks_hz": {
"harmonic": [55.0, 110.2],
"percussive": [125.0, 250.1]
}
}
},
"telemetry": {
"inference_time_sec": 28.0
}
}Related MCP server: music-perception-mcp
โก Key Features
๐ฅ Tempo & Beat Tracking โ Full BPM computation with variable-tempo drift detection and transient windowing.
๐น Key & Harmonic Mapping โ Computes structural musical key + mode, generating a detailed
key_maptracking section-by-section modulations.๐ Vibe & Style Embeddings โ Compiles a 512-dimensional CLAP embedding and human-readable style tags (covering energy, texture, mood, and genre) using zero-shot music vocab classification.
๐๏ธ Production Analytics โ Measures vocal spatial presence, transient punch coefficients, stereo width, and dominant frequency peaks.
๐ค MCP-Native System โ Fully exposes 4 standardized Model Context Protocol tools for instant integration into AI tools.
๐ชถ Robust Graceful Degradation โ Automatically utilizes a CUDA GPU if present and falls back to CPU; gracefully degrades to HPSS and standard librosa feature arrays if heavy deep learning packages (
[clap]) are omitted.๐ 100% Offline & Private โ All conversion, separation, and inference occur locally.
๐ฆ Installation & Setup
System Prerequisites
Ensure you have Python 3.10+ and FFmpeg installed and accessible on your system PATH.
Installing FFmpeg:
macOS:
brew install ffmpegLinux (Debian/Ubuntu):
sudo apt update && sudo apt install -y ffmpegWindows: Run
winget install Gyan.FFmpegvia PowerShell (Administrator), or download manually from ffmpeg.org and add thebindirectory to your system environment variables.
Step-by-Step Installation
Clone the Repository
git clone https://github.com/ripunjay-kashyap/audio-sonic-mcp.git cd audio-sonic-mcpInitialize Virtual Environment
python -m venv .venv # Activate on macOS/Linux: source .venv/bin/activate # Activate on Windows (PowerShell): .venv\Scripts\activateInstall Dependencies Choose between the lightweight core engine or the full high-fidelity ML suite:
Option A: Full High-Fidelity ML Suite (Recommended) Includes demixing stems (Demucs) and zero-shot vibe vectors (CLAP). Requires ~4 GB disk space.
pip install -e ".[clap]"Option B: Core Lightweight Pipeline Uses standard digital signal processing (HPSS/librosa). Rapid install and minimal footprint.
pip install -e .
The optional[clap] stack installs torch, torchaudio, transformers, and demucs. Without these, the server automatically switches to light fallbacks (HPSS instead of Demucs, standard feature matrices instead of CLAP vectors, and leaves out vibe_tags).
๐ค MCP Client Configuration Guide
Audio Sonic MCP registers itself as a standard package script. This enables you to run it using the global executable name (audio-sonic-mcp) directly from your virtual environment's bin folder, or run the script file manually.
1. Claude Desktop Setup
Open your Claude configuration file:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
Add the server to your mcpServers object:
{
"mcpServers": {
"audio-sonic-mcp": {
"command": "C:\\path\\to\\audio-sonic-mcp\\.venv\\Scripts\\audio-sonic-mcp.exe",
"args": [],
"env": {
"JOBS_ROOT": "C:\\path\\to\\audio-sonic-mcp\\jobs"
}
}
}
}Windows Users: Always use double backslashes (\\) in JSON configuration paths. Point the executable directly to the .exe inside your .venv\Scripts\ directory.
2. Cursor IDE Integration
To integrate Audio Sonic MCP into Cursor's AI pane:
Navigate to Settings โ Features โ MCP.
Click + Add New MCP Server.
Fill in the parameters:
Name:
audio-sonic-mcpType:
commandCommand:
/path/to/audio-sonic-mcp/.venv/bin/audio-sonic-mcp(use.exeextension on Windows)
3. Windsurf Integration
Open your Windsurf MCP configurations file (typically found at ~/.codeium/windsurf/mcp_config.json) and append the configuration:
{
"mcpServers": {
"audio-sonic-mcp": {
"command": "/path/to/audio-sonic-mcp/.venv/bin/python",
"args": ["/path/to/audio-sonic-mcp/server.py"],
"env": {
"JOBS_ROOT": "/path/to/audio-sonic-mcp/jobs"
}
}
}
}4. Cline (VS Code Extension) Setup
Open Cline's MCP setting file (usually located at %APPDATA%\Code\User\globalStorage\saoudrizwan.claude-dev\settings\cline_mcp_settings.json or equivalent platform storage) and add:
{
"mcpServers": {
"audio-sonic-mcp": {
"command": "/path/to/audio-sonic-mcp/.venv/bin/audio-sonic-mcp",
"args": [],
"env": {
"JOBS_ROOT": "/path/to/audio-sonic-mcp/jobs"
}
}
}
}๐ค Interaction Flow for AI Agents & LLMs
LLMs automatically learn how to use this server by reading its exposed tool definitions. Because audio stem separation and CLAP embeddings are computationally demanding, Audio Sonic MCP uses an Asynchronous Fire-and-Forget Job Pattern.
Automated LLM Workflow
[User Prompts LLM]
โ
โผ
1. Submit URL โโโโโโโโโโโโโโโบ [Tool: get_sonic_signature]
โ (Returns Job ID instantly)
โผ
2. Notify User โโโโโโโโโโโโโโ [LLM acknowledges job is queued]
โ
โโโโโบ 3. Wait 10-15s (Or proceed with other tasks)
โ
โผ
4. Check Progress โโโโโโโโโโโบ [Tool: get_job_status]
โ (Checks status: running/success/error)
โผ
5. Present Signature โโโโโโโโ [LLM formats rich output for user]Natural Prompts to Try
"Check the health of my audio-sonic-mcp server to make sure all ML components are ready."
"Submit this YouTube track for sonic analysis:
https://www.youtube.com/watch?v=XXXXXX.""Check the progress of my sonic signature job
sig_a1b2c3d4and summarize the BPM, production width, and vibe once complete."
๐๏ธ CLI Usage (Local Files)
For musicians, engineers, and producers working directly in the terminal, you can analyze a full-length local file directly without running any background servers:
# Get a visual, musician-friendly sonic signature digest (recommended)
python analyze_file.py "my_demo.wav" --summary
# Print full raw JSON directly to the stdout stream
python analyze_file.py "my_demo.wav"
# Dump JSON payload to a file while keeping the stdout clean
python analyze_file.py "my_demo.wav" > signature.jsonCLI Command Options Reference
Option | Shorthand | Description |
| None | Absolute or relative path to the local audio file (Required). |
|
| Print a clean, formatted terminal summary instead of standard JSON. |
| None | Generate JSON signature but omit the heavy 512-dimension vibe float array. |
|
| Output the final JSON signature directly to the specified file. |
|
| Do not delete intermediate WAV files or separated stem files in |
|
| Explicitly define the internal identifier (useful for batch scripts). |
Supported File Formats: wav, mp3, flac, ogg, m4a, aac.
๐ง Environment Variables Reference
Configure environment options by declaring these variables in your active terminal session, container environment, or the env block of your MCP configuration file:
Variable | Default Value | Description / Practical Use |
|
| Workspace directory where audio files, temporary converted WAVs, and stems are processed. |
| Unset | Set to |
|
| Safety ceiling for local file processing duration (YouTube downloads are capped at 60 minutes). |
| Unset | Path to folder containing the |
| Unset | HTTP/SOCKS proxy string passed directly to |
|
| Transport the server listens on: |
|
| Listening port when |
๐ณ Docker / Podman Execution
If you prefer to avoid setting up local Python libraries, running via containers encapsulates FFmpeg, yt-dlp, and the core Python dependencies (CPU-based pipeline):
# Build the container image
docker build -t audio-sonic-mcp .
# Run the MCP server over stdio, mounting local folders for job persistence
docker run -i --rm \
-v "$(pwd)/jobs:/app/jobs" \
-v "$(pwd)/models:/app/models" \
audio-sonic-mcpTo connect Claude Desktop to your Docker container, configure claude_desktop_config.json:
{
"mcpServers": {
"audio-sonic-mcp-docker": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-v", "/absolute/path/to/jobs:/app/jobs",
"-v", "/absolute/path/to/models:/app/models",
"audio-sonic-mcp"
]
}
}
}โ๏ธ How it Works under the Hood
Audio Sonic MCP pipelines are constructed modularly, using transactional checkpoints to ensure reliability.
LLM Agent / Claude Desktop Musician (Terminal)
โ โ
โ MCP (stdio JSON-RPC) โ analyze_file.py
โผ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Modular 6-Stage Analysis Pipeline โ
โ โ
โ Stage 1: Ingestion โ Pre-checks format, scans duration metadata โ
โ Stage 2: Download โ Fetches audio tracks via yt-dlp (URLs only) โ
โ Stage 3: Conversion โ normalizes sample formats to 44.1kHz WAV (FFmpeg)โ
โ Stage 4: Separation โ Splits stems: Vocals, Drums, Bass, Other (Demucs)โ
โ Stage 5: Analysis โ Computes BPM, modulations, key, punch (librosa) โ
โ Stage 6: Embeddings โ Generates 512-dim zero-shot music vibe tags (CLAP)โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โผ
Result Payload: (header ยท sonic_signature ยท telemetry)Stem Demixing: Meta AI's Demucs (
mdx_extra) separates the track into isolation stems (vocals,drums,bass,other). If missing, it gracefully drops back to Harmonic-Percussive Source Separation (HPSS).Analysis engine: librosa extracts rhythmic and tonal structures, matching chord patterns and sub-bass movements against Krumhansl-Schmuckler and Phrygian template engines.
Semantic Vibe Tagging: LAION CLAP (
laion/larger_clap_music_and_speech) runs zero-shot inference against high-coverage aesthetic descriptors (moods, textures, genres), choosing top candidates across stylistic poles.
๐ฉบ Resiliency & Troubleshooting
1. One-Time Setup Download Delays
Upon the very first analysis job utilizing the full ML pipeline, demucs and transformers will download their pre-trained model weights (approximately 400 MB for Demucs, and 200 MB for CLAP).
The server redirects download progress indicators to
stderrso they do not corrupt the JSON-RPC standard stream.During this download,
get_job_statuswill remain inrunning. Allow 1โ3 minutes depending on your network speed. Subsequent startups take under 10 seconds.
2. FastMCP Concurrency Controls
Model inference on multi-staged architectures is highly CPU/VRAM intensive. To protect consumer hardware and virtual environments from crashing (OutOfMemory exceptions), Audio Sonic MCP enforces a strict global serialization lock (CONCURRENCY_LOCK).
If you submit multiple URLs simultaneously, they will be processed sequentially.
Polling
get_job_statusfor subsequent jobs will reportqueuedorrunningwhile they wait in the pipeline queue.
3. Windows Librosa Deadlock Fix
FastMCP thread dispatching under Windows can cause Numba compilation deadlocks inside background worker threads. To prevent this, Audio Sonic MCP incorporates a Pre-warming Routine (_prewarm_librosa() and _prewarm_demucs()) on launch. It forces JIT compile of resampling, HPSS, and mono-mixing functions in the main thread before starting the RPC listener.
4. BPM Accuracy and the bpm_engine Field
Tempo is estimated by madmom's RNN beat tracker. madmom is an optional dependency: it is unmaintained (latest release 0.16.1, classifiers stop at Python 3.7) and requires a Cython build, so it cannot be installed reliably everywhere and is not part of the default install.
When madmom is unavailable the pipeline falls back to librosa. That fallback is good on steady four-on-the-floor material but can lock onto a 2:3 or octave multiple of the true tempo โ on one of our regression fixtures it reports 99.4 BPM against a ground truth of 148.
So the tempo is never reported unqualified. Every payload carries a bpm_engine field naming the engine that actually produced the number:
| Meaning |
| RNN beat tracker โ full accuracy. |
| madmom unavailable; treat BPM as approximate and expect occasional octave/triplet errors. |
check_health reports madmom's status explicitly. To enable the accurate path:
pip install ".[beats]"If the build fails on a recent Python, use 3.10 for the analysis environment โ madmom has no wheels for newer interpreters.
5. Diagnosing with check_health
If the server reports as degraded or tools are missing, call the check_health tool or check CLI warnings. It queries:
Availability of
ffmpegon the execution path.Installation status of Python packages (
librosa,soundfile,mcp, etc.).Presence of the optional
madmombeat tracker, and whichbpm_enginewill be used as a result.Access permissions to the
JOBS_ROOTdirectory.
๐ ๏ธ Development & Testing
Run unit tests inside your virtual environment to verify the mathematical pipelines using synthesized audio waveforms:
# Install development test framework
pip install -e ".[dev]"
# Execute full suite (requires no network or model downloads)
pytest
# Test specifically CLI execution code paths
pytest tests/test_cli.py๐ License
Distributed under the MIT License. See LICENSE for details.
ยฉ 2026 Ripunjay Kashyap. All rights reserved.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceDownloads audio from YouTube, analyzes with Essentia for BPM, mood, energy, spectrograms, and fetches synced lyrics from LRCLIB.6Apache 2.0
- FlicenseNot gradedqualityBmaintenanceAnalyzes audio files to extract exact, reproducible measurements like loudness, tempo, key, spectral balance, and clipping for LLM-based DAW control.
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to analyze audio files, extracting tempo, key, beat drops, volume surges, high tones, loudness, brightness, and structure, and returning structured JSON and visualizations.1MIT
- AlicenseAqualityCmaintenanceProvides local audio analysis tools for LLMs, enabling transcription, conversation dynamics, prosody analysis, and visual inspection without API keys.8MIT
Related MCP Connectors
Privacy-first audio intelligence: BPM, key, waveform. Audio never stored. Pay per second.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ripunjay-kashyap/audio-sonic-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server