Hume MCP Server
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Hume MCP ServerSynthesize expressive speech for 'Hello, how are you?' with a friendly tone."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
The Hume MCP Server gives you the the ability to use Octave Text to Speech from within an AI Chat, using an MCP Client Application such as Claude Desktop, Cursor, or Windsurf.
Octave TTS is the first text-to-speech system built on LLM intelligence. Octave is a speech-language model that understands what words mean in context, unlocking a new level of expressiveness and nuance. It performs the source text, it doesn't just pronounce it.
See this video for a demonstration of using the MCP Server to narrate a scene from an audiobook.
Quickstart
Click here to add to Cursor:
Copy the following into your client's MCP configuration (for example, inside the .mcpServers property of claude_desktop_config.json for Claude Desktop, or of the mcp.json for Cursor).
{
...
"hume": {
"command": "npx",
"args": [
"@humeai/mcp-server"
],
"env": {
"HUME_API_KEY": "<your_hume_api_key>",
}
}
}Related MCP server: mcp-ai-voice
Prerequisites
An account and API Key from Hume AI
(optional) A command-line audio player
ffplay from FFMpeg is recommended, but the server will attempt to detect and use any of several common players.
Available Tools
The server exposes the following MCP tools:
tts: Synthesize (and play) speech from text
play_previous_audio: Replay previously generated audio
list_voices: List available voices
save_voice: Save a generated voice to your library
delete_voice: Remove a voice from your library
Command Line Options
Options:
--workdir, -w <path> Set working directory for audio files (default: system temp)
--(no-)embedded-audio-mode Enable/disable embedded audio mode (default: false)
--help, -h Show this help messageEvaluation Framework
The project includes a comprehensive evaluation framework that measures how effectively AI agents can utilize the Hume TTS tools across various real-world scenarios.
Environment Variables
HUME_API_KEY: Your Hume AI API key (required)WORKDIR: Working directory for audio files (default: system temp directory + "/hume-tts")EMBEDDED_AUDIO_MODE: Enable/disable embedded audio mode (default: false, set to 'true' to enable)ANTHROPIC_API_KEY: Required for running evaluations
This server cannot be deployed
Maintenance
Related MCP Connectors
Manage ElevenLabs voice agents and generate speech, music, sound effects, images, and video.
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
- ChamadeOAuthio.chamade
Voice and chat for AI agents — Discord, Teams, Meet, Slack, Zoom, Telegram, WhatsApp, NC Talk, SIP
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to generate and play high-quality text-to-speech audio using the Kokoro model, with support for multiple voices, adjustable speaking speed, and audio caching.-
- AlicenseAqualityDmaintenanceEnables AI agents to synthesize natural speech using either platform system voices or premium OpenAI TTS, with automatic engine selection and graceful fallback.19 npmMIT

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.8 npmMIT- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to generate speech with custom cloned voices via DashScope or ElevenLabs, with an inline audio player and visualizer panel.MIT