voice-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@voice-mcpspeak this in my cloned voice: Hello! I'm ready to help you today."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🎙️ voice-mcp
An MCP (Model Context Protocol) server for AI voice synthesis with an inline audio player. Give your AI assistant a custom cloned voice!
Features
🎤 Custom Voice Cloning — Use MiniMax TTS API with your own cloned voice
🎵 Inline Audio Player — Beautiful WeChat-style player with waveform visualization
📝 Transcript Toggle — Show/hide the spoken text
🌙 Dark Mode Support — Automatic theme adaptation
⚡ Cloudflare Workers — Fast, serverless deployment
Related MCP server: oto
Demo
When you call the speak tool, you get:
A sleek audio player with play/pause button
Animated waveform that follows playback progress
Duration display
Expandable transcript
Quick Start
1. Clone the repository
git clone https://github.com/garan0613/voice-mcp.git
cd voice-mcp2. Install dependencies
npm install3. Configure MiniMax API
You'll need a MiniMax account with voice cloning enabled.
Add your secrets to Cloudflare:
npx wrangler secret put MINIMAX_API_KEY
npx wrangler secret put VOICE_ID
npx wrangler secret put BOT_NAME # Optional, defaults to "AI"Current MiniMax T2A authentication uses the API key only. GroupId is no longer required.
4. Deploy
npx wrangler deploy5. Connect to Claude.ai
Go to Settings → Connectors → Add Connector
Enter your Worker URL:
https://your-worker.workers.dev/mcpDone! The
speaktool is now available.
Configuration
Variable | Required | Description |
| ✅ | Your MiniMax API key |
| ✅ | The cloned voice ID |
| ❌ | Display name (default: "AI") |
API Endpoints
Endpoint | Description |
| MCP server (SSE protocol) |
| Direct audio file |
| Health check |
How to Clone a Voice
Go to MiniMax Console
Navigate to Voice Cloning
Upload 10-30 seconds of clear audio
Wait for processing (usually a few minutes)
Copy the Voice ID
Custom Deployment
Using a Custom Domain
Add your domain to Cloudflare
Create a DNS record pointing to your Worker
Update
wrangler.jsonc:
{
"routes": [
{ "pattern": "voice.yourdomain.com/*", "zone_name": "yourdomain.com" }
]
}Self-Hosting (Node.js)
The core MCP logic can be adapted for other platforms. You'll need to:
Replace
createMcpHandlerwith a standard HTTP/SSE handlerUse
@modelcontextprotocol/sdkdirectlyHandle the SSE transport yourself
Tech Stack
Cloudflare Workers — Serverless runtime
MCP SDK — Model Context Protocol
MiniMax TTS — Voice synthesis
ext-apps — Inline UI rendering
License
MIT © 2026
Credits
Inspired by the need to give AI assistants a voice. Built with ❤️
Related MCP Connectors
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
AI voice generation: text-to-speech and voice cloning from any MCP client.
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
- JellypodOAuthcom.jellypod
Create, import, and publish Jellypod podcast episodes from your AI assistant.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to generate and play high-quality text-to-speech audio using the Kokoro model, with support for multiple voices, adjustable speaking speed, and audio caching.-
- AlicenseNot gradedqualityBmaintenanceEnables text-to-speech conversion using OpenAI's TTS API, with inline audio playback and history within MCP hosts like Claude.7 npmBSD 4-Clause "Original" or "Old"

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.12 npmMIT- FlicenseNot gradedqualityDmaintenanceEnables Claude to speak text with an embedded audio player, supporting 54 voices, voice cloning, and playback controls, all running locally.3-