A FastAPI server implementing the Model Context Protocol (MCP) for structured tool use, providing utility tools including random number generation, image generation via Azure OpenAI DALL-E, and AI podcast generation.
Enables MCP-capable chat clients to synthesize speech from text, convert it with a local RVC voice model, and render the result in an inline audio player.
MCP server for local speech-to-text using Whisper Large V3 (MLX), enabling audio transcription with text/timestamps/SRT output and LLM-based correction, all running offline on Apple Silicon.
A modular suite of MCP servers for Audiokinetic Wwise, enabling AI agents to browse, edit, audition, profile, and build Wwise projects through the Wwise Authoring API.
Enables transcription, summarization, and action item extraction from audio files on your Mac using MacWhisper and Claude Desktop, all locally without any cloud APIs.
Guaardvark turns your own GPU into a full media studio for any MCP client: generate images, video, music and voice, run a five-role Film Crew, upscale to 8K, and query your indexed documents and code, with nothing leaving your machine. Every tool runs behind a default-deny policy, and every render queues locally and is polled by batch id, so an agent can direct hours of production without a cloud
Enables real-time transcription and heuristic vocal stress analysis of live financial webcasts, providing an LLM with rolling transcripts and a stress score.
Enables LLMs to perform FFmpeg operations like clipping, merging, extracting audio, adding subtitles, and transcoding videos via a set of tools exposed as an MCP server.
Enables MCP clients to search songs, retrieve details and lyrics, and manage personal playlists on NetEase Cloud Music, with optional local player control on macOS, all through a privacy-first, self-hosted service.
MCP server that enables natural language control of internet radio from Claude Code, with access to 30,000+ global stations, auto-playback via mpv, and a real-time status line with audio spectrum visualization.
MCP server that lets AI assistants control Siglent SDG waveform generators over a local network using natural language, supporting signal generation, modulation, sweep, burst, and arbitrary waveforms.
Gemini Audio MCP is a high-performance Model Context Protocol (MCP) server that leverages the power of the Gemini 2.0 Multimodal Live API to generate high-fidelity, environmental soundscapes on-demand.
Enables deterministic room acoustics analysis through MCP tools that compute axial, tangential, and oblique standing wave modes, Bonello and Bolt criteria compliance, Schroeder cutoff frequency, and Sabine/Norris-Eyring RT60 decay across octave bands for rectangular rooms. Also derives optimal speaker and sweet-spot coordinates with SBIR notch predictions and calculates the acoustic treatment area needed to hit target reverb times for podcast, mixing, home theater, or listening use.
A mock-first MCP server for audio generation that creates placeholder WAV files locally, enabling development and testing of audio workflows without APIs or costs.