Skip to main content
Glama
199-mcp
by 199-mcp

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
ELEVENLABS_API_KEYYesYour ElevenLabs API key. Get one at https://elevenlabs.io/app/settings/api-keys
ELEVENLABS_MCP_BASE_PATHNoThe base path the MCP server should look for and output files specified with relative paths.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
text_to_speechC

Converts text to speech (v2/v3/flash models). Returns: audio file path. Use when: single speaker narration. Supports v3 with audio tags.

text_to_speech_v3A

Converts text to speech with v3 model and tags. Returns: audio file path. Use when: single speaker needs emotions, pauses, or sound effects.

speech_to_textC

Transcribes audio to text. Returns: transcript text or file path. Use when: converting speech recordings to text.

text_to_sound_effectsC

Generates sound effects from text. Returns: audio file path. Use when: creating custom sound effects from descriptions.

search_voicesB

Searches available voices. Returns: JSON with voice details. Use when: finding voices by name, gender, or characteristics.

get_voice_id_by_nameA

Resolves voice name to ID. Returns: JSON with voice_id and confidence. Use when: need voice ID from name with fuzzy matching.

list_modelsA

Lists available TTS models. Returns: model list with capabilities. Use when: choosing between v2, v3, or other models.

get_voiceB

Gets voice details. Returns: voice metadata and settings. Use when: need detailed information about a specific voice.

voice_cloneB

Creates voice clone from audio. Returns: new voice ID. Use when: creating custom voice from recordings.

isolate_audioA

Removes background noise from audio. Returns: cleaned audio file path. Use when: extracting voice from noisy recordings.

check_subscriptionA

Checks account subscription. Returns: subscription details and usage. Use when: monitoring API usage and limits.

create_agentB

Creates conversational AI agent. Returns: agent ID and details. Use when: setting up voice-enabled chatbot or assistant.

add_knowledge_base_to_agentB

Adds knowledge to agent. Returns: knowledge base ID. Use when: giving agent access to documents or information.

list_agentsB

Lists all agents. Returns: agent list with IDs. Use when: viewing available conversational AI agents.

get_agentB

Gets agent details. Returns: agent configuration. Use when: viewing specific agent settings and capabilities.

speech_to_speechB

Transforms voice in audio. Returns: audio file with new voice. Use when: changing speaker voice in existing audio.

text_to_voiceA

Creates voice from description. Returns: three voice preview files. Use when: designing custom voice from text prompt.

create_voice_from_previewB

Saves generated voice to library. Returns: permanent voice ID. Use when: keeping voice from text_to_voice previews.

make_outbound_callB

Initiates phone call with agent. Returns: call details. Use when: making automated calls via Twilio integration.

search_voice_libraryB

Searches global voice library. Returns: shared voices list. Use when: finding voices across entire ElevenLabs platform.

list_phone_numbersA

Lists account phone numbers. Returns: phone number list. Use when: viewing available numbers for outbound calls.

play_audioA

Plays audio file locally. Returns: playback confirmation. Use when: previewing generated audio without downloading.

get_conversationB

Gets conversation with transcript. Returns: conversation details and full transcript. Use when: analyzing completed agent conversations.

list_conversationsB

Lists agent conversations. Returns: conversation list with metadata. Use when: browsing conversation history.

get_conversation_transcriptB

Gets conversation transcript in chunks. Returns: transcript chunk with metadata. Use when: retrieving large conversation transcripts.

text_to_dialogueB

Converts multi-speaker text to audio. Returns: dialogue audio file paths. Use when: creating conversations with multiple voices.

enhance_dialogueB

Adds audio tags to dialogue. Returns: enhanced text with v3 tags. Use when: improving dialogue with emotions and effects.

fetch_v3_tagsA

Lists v3 audio tags. Returns: comprehensive tag list. Use when: preparing text for text_to_speech_v3 or text_to_dialogue with emotions and effects.

get_v3_audio_tags_guideA

Gets v3 tag usage guide. Returns: detailed v3 documentation. Use when: learning how to use v3 audio tags effectively.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.3/5.0

Scored across 29 tools

Disambiguation3/5

Most tools have distinct purposes, but there is notable overlap between text_to_speech and text_to_speech_v3, which could cause confusion as both handle single-speaker narration with v3 support mentioned in text_to_speech. Additionally, get_conversation and get_conversation_transcript serve similar functions with minor differences in output format, potentially leading to misselection. However, descriptions help clarify boundaries in most cases.

Naming Consistency4/5

Tool names follow a consistent snake_case pattern throughout, with clear verb_noun structures (e.g., create_agent, list_models, get_voice). Minor deviations exist, such as text_to_speech_v3 including a version suffix, but this is reasonable for differentiation. Overall, the naming is predictable and readable, supporting easy identification.

Tool Count3/5

With 29 tools, the count is borderline high for the ElevenLabs domain, which covers voice generation, agents, and audio processing. While the scope is broad, some tools might be consolidated (e.g., overlapping speech functions), making it feel slightly heavy. However, it's not extreme and remains manageable for the apparent feature set.

Completeness5/5

The tool set provides comprehensive coverage for the ElevenLabs domain, including CRUD operations for agents and voices, audio processing (speech-to-text, isolation), generation (text-to-speech, dialogue, sound effects), and utilities (subscription, playback). There are no obvious gaps; agents can handle full workflows from creation to conversation analysis, and voice management includes cloning, searching, and previewing.

Maintenance

ActivityInactive
ResponsivenessUnresponsive