edge-tts
# 语音合成服务器
一个集成了Microsoft Edge高质量语音合成能力的MCP服务器,支持多语言语音生成、音频合并和云端存储。
An MCP server integrated with Microsoft Edge's high-quality speech synthesis capabilities, supporting multilingual speech generation, audio merging, and cloud storage.## 工具列表 Tool List
本MCP服务封装下列工具,可让模型通过标准化接口调用以下功能。 本MCP服务封装下列工具,可让模型通过标准化接口调用以下功能。
| 工具 Tool | 描述 Description |
|-------|--------------------|
| generate_speech | Generate speech audio from text using Microsoft Edge TTS. Supports multi-role conversations and audio merging. |
## 检查服务 ## Inspector
工具在线测试: [https://mcp.xiaobenyang.com/inspector/1777316659830787](https://mcp.xiaobenyang.com/inspector/1777316659830787)
Online Tool test [https://mcp.xiaobenyang.com/inspector/1777316659830787](https://mcp.xiaobenyang.com/inspector/1777316659830787)
## 服务配置 MCP Server Config
> #### 如何获取 XBY-APIKEY ? How to get XBY-APIKEY ?
> 访问小笨羊科技网站 [https://xiaobenyang.com](https://xiaobenyang.com),注册用户即可获得APIKEY
> Visit XiaoBenYang website [https://xiaobenyang.com](https://xiaobenyang.com), register and get the APIKEY.
### SSE
```json
{
"mcpServers": {
"语音合成服务器": {
"headers": {
"XBY-APIKEY": "<YOUR_XBY_APIKEY>"
},
"type": "sse",
"url": "https://mcp.xiaobenyang.com/1777316659830787/sse"
}
}
}
```
### STREAMABLE HTTP
```json
{
"mcpServers": {
"语音合成服务器": {
"headers": {
"XBY-APIKEY": "<YOUR_XBY_APIKEY>"
},
"type": "streamable_http",
"url": "https://mcp.xiaobenyang.com/1777316659830787/mcp"
}
}
}
```
### STDIO
```json
{
"mcpServers": {
"语音合成服务器": {
"command": "npx",
"args": [
"-y",
"xiaobenyang-mcp"
],
"env": {
"XBY_APIKEY": "<YOUR_XBY_APIKEY>",
"mcpId": "1777316659830787",
},
"transport": "stdio"
}
}
}
```
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool's purpose is clearly defined and distinct by default.
A single tool inherently has perfect naming consistency. The tool name 'generate_speech' follows a clear verb_noun pattern, and there are no other tools to compare or create inconsistency with.
One tool is too few for most practical purposes, as it severely limits functionality. While it covers the core TTS generation, the server lacks tools for managing voices, configurations, or other related operations, making the scope feel incomplete and thin.
The server is severely incomplete for a TTS domain. It only provides speech generation without tools for listing available voices, adjusting speech parameters, or handling audio playback, leaving significant gaps that agents cannot work around effectively.