ElevenLabs Voice-to-Voice Agent
Related Servers
Alternatives to ElevenLabs Voice-to-Voice Agent
No user-submitted related servers found.
Related Servers
- FlicenseNot gradedqualityCmaintenanceAn MCP server that provides text-to-speech, speech-to-text, and voice management via ElevenLabs API.1-
- AlicenseAqualityCmaintenanceA full-featured MCP server for the ElevenLabs API that brings text-to-speech, speech-to-text, voice cloning, sound effects, music, audio isolation, dubbing, and account tools to any MCP client.28MIT
- AlicenseNot gradedqualityDmaintenanceOfficial MCP server that enables interaction with ElevenLabs Text to Speech and audio processing APIs. It allows generating speech, cloning voices, transcribing audio, and creating sound effects through natural language.MIT

ElevenLabs MCP Serverofficial
AlicenseBqualityFmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,536MIT- AlicenseAqualityDmaintenanceAn MCP server that enables LLMs to generate spoken audio from text using OpenAI's Text-to-Speech API, supporting various voices, models, and audio formats.16 npm1MIT
- FlicenseAqualityBmaintenanceAn agent-first MCP server for generating and transforming audio (music, speech, sound effects) via ElevenLabs and Mureka, providing tools for transcription, voice cloning, and multi-speaker dialogue.33-
TDQS
Scored across 7 tools
Each tool targets a distinct operation: listing vs fetching voices, listing models, text-to-speech, speech-to-text, user info, and history. There is no overlap or ambiguity between them.
Most tools follow a clear verb_noun pattern like list_voices, get_voice, get_user_info, and get_history. However, text_to_speech and speech_to_text deviate from this pattern as noun-to-noun phrases, though they are internally consistent and readable.
Seven tools is a well-scoped number for a voice/audio agent, covering core operations without bloat. The count is appropriate for the apparent purpose.
The set covers TTS, STT, voice listing, model listing, user info, and history, but notably lacks a voice-to-voice conversion tool despite the server name. Voice editing/deletion and history item management are also missing, leaving some workflows incomplete.