A comprehensive audio MCP server that enables AI agents to generate speech, transcribe audio, clone voices, analyze speech quality, design soundscapes, and manage audio assets through a standardized interface.
An MCP server that runs Stability AI's Stable Audio Open 1.0 locally on NVIDIA GPUs, enabling AI agents to generate broadcast-quality 44.1 kHz stereo WAV sound effects from text prompts fully offline with no API costs.
Exposes text-to-audio sound effect generation as an MCP tool, allowing clients like Claude Desktop to generate sound effects locally using a diffusion model, with support for AMD ROCm and Apple Silicon.
An MCP server for local image, speech, music, and SFX generation that preserves character identity across calls and rejects degenerate outputs. It runs with near-zero GPU idle memory and is designed for DeepSeek Harness agents.