Skip to main content
Glama
Megaprompting

MS-Lucidia-Voice-Gateway-MCP

MS-Lucidia-语音网关-MCP

一个模型上下文协议 (MCP) 服务器,使用 Windows 内置语音服务提供文本转语音和语音转文本功能。此服务器通过 PowerShell 命令利用本机 Windows 语音 API (SAPI),从而无需外部 API 或服务。

特征

  • 使用 Windows SAPI 语音的文本转语音 (TTS)

  • 使用 Windows 语音识别进行语音转文本 (STT)

  • 用于测试的简单 Web 界面

  • 无外部 API 依赖

  • 使用原生 Windows 功能

Related MCP server: Edge TTS MCP

先决条件

  • 启用语音识别的 Windows 10/11

  • Node.js 16+

  • PowerShell

安装

  1. 克隆存储库:

git clone https://github.com/ExpressionsBot/MS-Lucidia-Voice-Gateway-MCP.git
cd MS-Lucidia-Voice-Gateway-MCP
  1. 安装依赖项:

npm install
  1. 构建项目:

npm run build

用法

测试接口

  1. 启动测试服务器:

npm run test
  1. 在浏览器中打开http://localhost:3000

  2. 使用 Web 界面测试 TTS 和 STT 功能

可用工具

文本转语音

使用 Windows SAPI 将文本转换为语音。

参数:

  • text (必需):要转换为语音的文本

  • voice (可选):要使用的语音(例如“Microsoft David Desktop”)

  • speed (可选):语速从 0.5 到 2.0(默认值:1.0)

例子:

fetch('http://localhost:3000/tts', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    text: "Hello, this is a test",
    voice: "Microsoft David Desktop",
    speed: 1.0
  })
});

语音转文本

录制音频并使用 Windows 语音识别将其转换为文本。

参数:

  • duration (可选):录制持续时间(秒)(默认值:5,最大值:60)

例子:

fetch('http://localhost:3000/stt', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    duration: 5
  })
}).then(response => response.json())
  .then(data => console.log(data.text));

故障排除

  1. 确保 Windows 语音识别已启用:

    • 打开 Windows 设置

    • 前往“时间和语言”>“语音”

    • 启用语音识别

  2. 检查可用的声音:

    • 打开 PowerShell 并运行:GXP7

  3. 测试语音识别:

    • 在 Windows 设置中打开语音识别

    • 如果尚未完成,请运行安装向导

    • 测试 Windows 是否可以识别你的声音

贡献

  1. 分叉存储库

  2. 创建你的功能分支

  3. 提交你的更改

  4. 推送到分支

  5. 创建新的 Pull 请求

执照

麻省理工学院

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides text-to-speech functionality for AI agents using Microsoft Edge's text-to-speech technology, supporting multiple voices, languages, and voice customization.
    2
    8
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that converts text into lifelike speech using Microsoft Edge's Text-to-Speech service, supporting customizable voice, rate, volume, and pitch.
    4
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    A lightweight MCP server that provides offline text-to-speech via native OS speech engines. It allows AI assistants to speak task summaries instantly with no network calls.
    -