eleven-v4-mcp-voice-starter
Generates text-to-speech audio through the ElevenLabs Text to Dialogue API using a configured API key and voice ID, producing MP3 files for playback or delivery.
Sends generated voice messages to a fixed Telegram chat via the Telegram Bot API sendVoice method when configured for Telegram delivery.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@eleven-v4-mcp-voice-starterRead this aloud as an MP3: launch is at 3pm today."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Eleven v4 MCP Voice Starter
一个可以自己运行的 ElevenLabs Eleven v4 语音示例:文字生成 MP3、MCP 工具调用、Telegram 发送,以及本机网页播放。代码从零编写,不包含任何人的 API key、voice ID、Bot Token、聊天记录或服务器配置。
先选入口:
想做什么 | 从哪里开始 |
只想生成一个 MP3 | 本页「五分钟试听」 |
想让 AI 通过 MCP 触发语音 | 本页「接入 MCP」 |
想设计自己的声音 | |
想把 v3 改成 v4 | |
想换成自己的前端或 Telegram | |
想在 Codex、ChatGPT Work 或 Chat 播放 | |
让 AI 接手这个仓库 |
五分钟试听
需要 Node.js 22 或更新版本、自己的 ElevenLabs API key,以及一个 voice ID。Voice ID 可以从 ElevenLabs 的 My Voices 页面点声音旁的三个点并复制。没有声音时,先看 Voice Design 指南。
下载或克隆仓库,进入目录,运行 npm ci。
复制 .env.example 为 .env。Windows PowerShell 用 Copy-Item .env.example .env;macOS/Linux 用 cp .env.example .env。
在 .env 中填写 ELEVENLABS_API_KEY 和 ELEVENLABS_VOICE_ID。别把真实 .env 提交到 Git。
运行:
npm run demo -- "你好,这是我的第一段 Eleven v4 语音。"当前目录会生成 voice.mp3。这个脚本只做一件事:请求 ElevenLabs Text to Dialogue API,返回 MP3 并保存。默认模型明确指定为 eleven_v4;如果不写 model_id,接口目前默认 eleven_v3。官方模型 · 官方接口
Related MCP server: MCP TTS Server
在本机网页播放
再在 .env 填入一个随机的 VOICE_INTERNAL_TOKEN,并保持 VOICE_DELIVERY=web。生成随机值的例子:
node -e "console.log(require('node:crypto').randomBytes(32).toString('hex'))"启动服务:
npm start浏览器打开 http://127.0.0.1:8788 。输入文字点「生成试听」就能播放。若 MCP 调用 trigger_voice,语音会保存在本机 data/voices 中并出现在页面下方;页面每五秒检查一次新语音。该目录已被 .gitignore 排除。
本示例固定监听 127.0.0.1,供本机学习与实验。要让互联网上的前端访问,应接入你自己的登录、访问控制、HTTPS、文件保留策略和反向代理;不要直接把这个本机示例当公网服务。
接入 MCP
保持上一节的服务运行。在另一个终端运行 npm run mcp。它是 stdio MCP server:正常启动后会等待客户端连接,不会像网页服务器那样打印一个地址。可以用 MCP Inspector 调试:
npx @modelcontextprotocol/inspector npm run mcp工具有两个:
trigger_voice(text):通过本机语音服务生成一次语音,并交给服务端选定的出口。web 模式进入网页语音列表;telegram 模式发到固定 Telegram 聊天。工具参数只有文字。
preview_voice(text):返回短 MP3 的 MCP audio 内容块;客户端是否显示内嵌播放器取决于客户端自身支持情况。
本地 Codex 可在设置的 MCP servers 页面添加 STDIO server,启动命令指向本仓库的 Node 进程;或按官方 Codex MCP 指南配置。启动命令需从本仓库目录运行 npm run mcp,或使用 Node 的 --env-file 传绝对 .env 路径并传绝对 src/mcp.mjs 路径。MCP 与 gateway 进程都要读取相同的 VOICE_INTERNAL_TOKEN。
调用链: AI → MCP trigger_voice(text) → 本机 gateway 的 /internal/voice → ElevenLabs → server-selected output。模型工具看不到 API key、voice ID、Telegram chat ID,也不能从参数里改收件人。详情见架构图与替换点。
换成 Telegram
在 .env 中改为 VOICE_DELIVERY=telegram,并填写自己的 TELEGRAM_BOT_TOKEN 和 TELEGRAM_CHAT_ID。重启 npm start。此时 trigger_voice 会在服务端调用 Telegram sendVoice;网页的 MCP 语音列表不再新增。网页的「生成试听」仍可本机播放。Telegram Bot API 的 sendVoice 文档说明了语音上传接口。
验证与边界
运行 npm test。测试使用假音频和模拟的供应商/Telegram 响应,不消费 ElevenLabs 额度。真实声音仍需你用自己的 key 和 voice ID 做一次试听。
这是教学示例。实际多用户或生产系统还需要耐久幂等回执、身份与房间解析、队列/重试策略、鉴权的音频存储和清理。本仓库没有复制 Serein Home 的私有服务,也没有假称它的生产可靠性。参考 docs/architecture.md。
资料
MIT License。
This server cannot be deployed
Maintenance
Related MCP Connectors
AI voice generation: text-to-speech and voice cloning from any MCP client.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
Generate Suno AI music (v5.5) from any MCP client. Async; billed only on success.
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.1MIT
- FlicenseNot gradedqualityDmaintenanceProvides text-to-speech conversion through a unified MCP interface, supporting both local Kokoro and cloud OpenAI TTS engines with streaming audio, voice selection, and customization via natural language instructions.7-
- AlicenseNot gradedqualityBmaintenanceEnables text-to-speech conversion using OpenAI's TTS API, with inline audio playback and history within MCP hosts like Claude.7 npmBSD 4-Clause "Original" or "Old"

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.12 npmMIT