Skip to main content
Glama

mcp-kokoro-tts

本地 Kokoro-82M 文本转语音 MCP 服务器。当你的智能体调用 speak 时,它会在你的机器上合成语音并播放,让你能听到智能体框架说话。

适用于任何 MCP 客户端:Claude Desktop、Claude Code、Cursor、VS Code、opencode、Cline 等。只需一段简短的配置,无需 API 密钥——合成在本地运行,使用 Kokoro-82M。

首次启动时,服务器会准备两项不在 PyPI 上的内容:Kokoro-82M 权重(约 312 MB)会被放入本地缓存,spaCy 的英文模型(en_core_web_sm)会被放入服务器运行所在的同一个 Python 环境。第二次安装是必需的,因为 Kokoro 的 G2P 流水线会加载 spaCy,而 uvx / uv tool 环境不会包含该模型,除非本包将其放入其中。

安装

添加到你的客户端的 MCP 配置中:

{
  "mcpServers": {
    "mcp-kokoro-tts": {
      "command": "uvx",
      "args": ["mcp-kokoro-tts"]
    }
  }
}

需要 Python 3.12 和 uv。首次启动服务器时会自动准备 Kokoro 权重和 spaCy 英文模型。

若要在不启动 MCP 服务器的情况下预下载这两项:

uvx mcp-kokoro-tts-provision

Related MCP server: MCP TTS Server

让智能体调用它

在 AGENTS.md / CLAUDE.md / 系统提示词中添加一行:

When the user wants to hear something spoken aloud, call the `speak` tool with clear, natural text.

工具

speak

合成语音,写入 WAV 文件,并在本地播放。

参数

必填

描述

text

是

要朗读的文本(最多 500 个字符)

voice

否

语音 ID(例如 af_heart)或 .pt 语音文件的绝对路径

speed

否

播放速度倍率(默认 1.0)

list_voices

列出可用的 Kokoro 语音以及当前选定的默认语音。

选择你的语音

解析顺序:

  1. TTS_VOICE 环境变量 — 语音 ID 或 .pt 绝对路径

  2. 包内 voices/ 文件夹中名称以 default 开头的文件

  3. voices/ 中第一个 .pt 文件(按字母顺序)

  4. 模型自带的 af_heart 语音

{
  "mcpServers": {
    "mcp-kokoro-tts": {
      "command": "uvx",
      "args": ["mcp-kokoro-tts"],
      "env": {
        "TTS_VOICE": "af_heart"
      }
    }
  }
}

环境变量

变量

描述

TTS_VOICE

默认语音 ID 或 .pt 绝对路径

TTS_MODEL_DIR

覆盖模型缓存目录

TTS_HF_CACHE_DIR

覆盖 Hugging Face hub 缓存目录

TTS_OUTPUT_DIR

生成的 WAV 文件目录

TTS_PLAY

设置为 0 可仅合成而不在本地播放

HF_TOKEN

可选的 Hugging Face 令牌,用于加速下载

平台

操作系统

合成

播放

macOS

是

afplay

Linux

是

ffplay、paplay 或 aplay

Windows

是

PowerShell MediaPlayer

espeak-ng 是可选的。英语无需它即可工作;安装它可以获得更好的词汇外覆盖,并支持一些非英语语言。

发布

为版本打标签会运行 GitHub Actions 的 publish.yml,它会先上传到 PyPI,再上传到 MCP Registry。

发布到 PyPI 使用仓库密钥 PYPI_TOKEN(一个 PyPI API 令牌)。也可以在 PyPI 项目上配置 GitHub 可信发布;此工作流使用该令牌进行身份验证,因此首次发布不依赖于待处理发布者的匹配。

发布

  1. 在 pyproject.toml 中递增 version(如果尚未打标签,还要在 server.json 中递增)

  2. 提交并打标签:git tag v0.1.2 && git push origin v0.1.2

  3. GitHub Actions 运行 publish.yml:

    • release — 类型检查、测试、构建 wheel/sdist

    • pypi-publish — 使用 PYPI_TOKEN 上传到 PyPI

    • mcp-registry — OIDC → MCP Registry(在 PyPI 成功后)

开发

cd mcps-tts
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pyright
pytest
python -m mcp_kokoro_tts

许可证

Apache-2.0。参见 LICENSE 和 NOTICE。Kokoro-82M 模型权重根据其自身的 Apache-2.0 许可证单独下载。

Available Tools

2 tools
list_voicesList voicesA

List Kokoro voices available to the speak tool and the currently selected default voice. Resolution order: TTS_VOICE env var, voices/default* file, first file in voices/, then the model's default af_heart voice.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does exceptionally well: it reveals the exact resolution order (env var, file, first file, fallback model default), which behaviorally affects what the tool returns. This is precisely the kind of contextual detail an agent needs beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, the first clearly scopes the listing and the second enumerates the resolution order. Every phrase adds value and the default-voice resolution rule is stated in a compact, order-relevant manner.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description does not need to explain return format. It fully explains what this tool does (list voices and default voice) and even how the default is determined, which is the behavioral nuance essential for an agent to pick the right voice via speak.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is already complete (100% coverage) and there is nothing for the description to elaborate on. Baseline 4 applies since there is simply no parameter information needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('Kokoro voices available to the speak tool') and also includes the currently selected default voice. The mention of 'speak tool' clearly distinguishes it from the sibling speak tool, making its function unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool should be used to discover available voice options before using the speak tool, especially because it names the speak tool as the consumer. It clearly states the resolution order for the default voice, giving immediate context, but it does not explicitly say when not to use it or mention alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speakSpeakA

Synthesize speech with Kokoro-82M and play it on the user's machine. Call this when the user wants the harness to talk, read something aloud, or hear a spoken response. Always pass clear, natural text to speak. Default voice resolution: TTS_VOICE env var, voices/default* in the package, first file in voices/, then the bundled af_heart voice from the model.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
speedNo
voiceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it makes a reasonable effort. It discloses the side effect of audio playback ('play it on the user's machine') and thoroughly documents the voice resolution fallback order (TTS_VOICE env var, voices/default*, first file in voices/, then af_heart). However, it doesn't disclose potential blocking behavior, absence of audio hardware, or other side effects that could be expected given the nature of the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, with the core purpose front-loaded first and the resolution detail last. Every sentence earns its place — the usage guidance and voice resolution order are both essential for the agent. The voice resolution sentence is dense but packs critical behavioral information. A slight trim could occur since the resolution order could be seen as implementation detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description captures the core essence—synth and play, when to trigger, how voice is resolved—which is solid for a 3-parameter tool. No annotations or schema descriptions exist to offload this burden to. Missing context includes behavior of the speed parameter, the relationship to the sibling list_voices, and failure modes. Given the total absence of annotations and schema descriptions, the description could do more to guide an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description needed to compensate, and it partially does. It adds value for the text parameter ('clear, natural text to speak') and implicitly for the voice parameter by explaining the resolution order. However, the speed parameter is entirely unexplained (what units, what range), and the voice format is only hinted at. Given the extremely low schema coverage, this partial compensation keeps it at a decent but not exceptional level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Synthesize speech with Kokoro-82M and play it on the user's machine') that clearly identifies what the tool does. It goes further to elaborate intended use cases ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It doesn't explicitly contrast itself with the sibling list_voices, which keeps it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance with concrete examples ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It includes a practical instruction to 'Always pass clear, natural text to speak.' However, it stops short of explicitly naming alternatives (e.g., list_voices) for comparison or stating when not to use it, so it lacks the explicit exclusions of a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.4
    • First observedlist_voices
    • First observedspeak

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one performs speech synthesis ('speak'), the other lists available voices ('list_voices'). No ambiguity between them, as they serve complementary roles in the TTS workflow.

Naming Consistency5/5

Both tool names follow an imperative verb style ('speak', 'list_voices'), which is concise and predictable. 'list_voices' uses a verb_noun pattern, but the naming is consistent in tone and clarity for a small set.

Tool Count5/5

With only 2 tools, the server is minimal yet well-scoped for its purpose: a text-to-speech harness. The count is appropriate; additional tools would likely be redundant.

Completeness4/5

The core operation (speak) is present, and the companion list_voices provides necessary context for voice selection. Minor gaps exist, such as no pause/stop or voice configuration tool, but these are not critical for basic TTS functionality.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers