mcp-kokoro-tts
mcp-kokoro-tts
ローカルKokoro-82Mテキスト読み上げMCPサーバーです。エージェントがspeakを呼び出すと、音声を合成してあなたのマシンで再生するため、ハーネスが話すのを聞くことができます。
あらゆるMCPクライアントで動作します:Claude Desktop、Claude Code、Cursor、VS Code、opencode、Clineなど。短い設定ブロックを1つ追加するだけで、APIキーは不要です。音声合成はKokoro-82Mでローカルに実行されます。
初回起動時に、サーバーはPyPIにない2つのものを用意します: Kokoro-82Mの重みデータ(約312 MB)をローカルキャッシュへ、spaCyの英語モデル(en_core_web_sm)をサーバーが実行しているのと同じPython環境へ。この2つ目のインストールが必要なのは、KokoroのG2PパイプラインがspaCyを読み込むためです。uvx / uv tool環境には、このパッケージが配置しない限りモデルが存在しません。
インストール
クライアントのMCP設定に追加します:
{
"mcpServers": {
"mcp-kokoro-tts": {
"command": "uvx",
"args": ["mcp-kokoro-tts"]
}
}
}Python 3.12とuvが必要です。最初のサーバー起動時に、Kokoroの重みデータとspaCy英語モデルが自動的に用意されます。
MCPサーバーを起動せずに両方を事前にダウンロードするには:
uvx mcp-kokoro-tts-provisionRelated MCP server: MCP TTS Server
エージェントに呼び出させる
AGENTS.md / CLAUDE.md / システムプロンプトに1行追加します:
When the user wants to hear something spoken aloud, call the `speak` tool with clear, natural text.ツール
speak
音声を合成し、WAVファイルを書き出して、ローカルで再生します。
パラメータ | 必須 | 説明 |
| はい | 読み上げるテキスト(最大500文字) |
| いいえ | 音声ID(例: |
| いいえ | 再生速度の倍率(デフォルト |
list_voices
利用可能なKokoroの音声と、現在選択されているデフォルトを一覧表示します。
音声の選択
解決順:
TTS_VOICE環境変数 — 音声IDまたは絶対パスの.ptパスパッケージの
voices/フォルダ内で名前がdefaultで始まるファイルvoices/内の最初の.ptファイル(アルファベット順)モデルに同梱の
af_heart音声
{
"mcpServers": {
"mcp-kokoro-tts": {
"command": "uvx",
"args": ["mcp-kokoro-tts"],
"env": {
"TTS_VOICE": "af_heart"
}
}
}
}環境変数
変数 | 説明 |
| デフォルトの音声IDまたは絶対パスの |
| モデルキャッシュディレクトリを上書き |
| Hugging Faceハブのキャッシュディレクトリを上書き |
| 生成されたWAVファイルのディレクトリ |
|
|
| ダウンロードを高速化するためのオプションのHugging Faceトークン |
プラットフォーム
OS | 合成 | 再生 |
macOS | 対応 |
|
Linux | 対応 |
|
Windows | 対応 | PowerShell |
espeak-ngはオプションです。英語はこれなしで動作します。語彙外(out-of-vocabulary)のカバレッジを向上させ、一部の非英語言語を扱うにはインストールしてください。
公開
バージョンにタグを付けるとGitHub Actionsのpublish.ymlが実行され、PyPIにアップロードした後、MCP Registryにアップロードします。
PyPIへの公開には、リポジトリのシークレットPYPI_TOKEN(PyPI APIトークン)を使用します。GitHubのtrusted publishingはPyPIプロジェクト側でも設定できます。このワークフローはトークンで認証するため、初回リリースはpending-publisherのマッチングに依存しません。
リリース
pyproject.tomlのversionを上げます(まだタグを付けない場合はserver.jsonも)コミットしてタグを付けます:
ギgit tag v0.1.2 && git push origin v0.1.2GitHub Actionsが
publish.ymlを実行します:release— 型チェック、テスト、wheel/sdistのビルドpypi-publish—PYPI_TOKENでPyPIにアップロードmcp-registry— OIDC → MCP Registry(PyPI成功後)
開発
cd mcps-tts
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pyright
pytest
python -m mcp_kokoro_ttsライセンス
Apache-2.0です。LICENSEとNOTICEを参照してください。Kokoro-82Mモデルの重みデータは、Apache-2.0ライセンスのもとで別途ダウンロードされます。
Available Tools
2 toolslist_voicesList voicesA
List Kokoro voices available to the speak tool and the currently selected default voice. Resolution order: TTS_VOICE env var, voices/default* file, first file in voices/, then the model's default af_heart voice.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does exceptionally well: it reveals the exact resolution order (env var, file, first file, fallback model default), which behaviorally affects what the tool returns. This is precisely the kind of contextual detail an agent needs beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, the first clearly scopes the listing and the second enumerates the resolution order. Every phrase adds value and the default-voice resolution rule is stated in a compact, order-relevant manner.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description does not need to explain return format. It fully explains what this tool does (list voices and default voice) and even how the default is determined, which is the behavioral nuance essential for an agent to pick the right voice via speak.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is already complete (100% coverage) and there is nothing for the description to elaborate on. Baseline 4 applies since there is simply no parameter information needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('Kokoro voices available to the speak tool') and also includes the currently selected default voice. The mention of 'speak tool' clearly distinguishes it from the sibling speak tool, making its function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool should be used to discover available voice options before using the speak tool, especially because it names the speak tool as the consumer. It clearly states the resolution order for the default voice, giving immediate context, but it does not explicitly say when not to use it or mention alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speakSpeakA
Synthesize speech with Kokoro-82M and play it on the user's machine. Call this when the user wants the harness to talk, read something aloud, or hear a spoken response. Always pass clear, natural text to speak. Default voice resolution: TTS_VOICE env var, voices/default* in the package, first file in voices/, then the bundled af_heart voice from the model.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| speed | No | ||
| voice | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it makes a reasonable effort. It discloses the side effect of audio playback ('play it on the user's machine') and thoroughly documents the voice resolution fallback order (TTS_VOICE env var, voices/default*, first file in voices/, then af_heart). However, it doesn't disclose potential blocking behavior, absence of audio hardware, or other side effects that could be expected given the nature of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, with the core purpose front-loaded first and the resolution detail last. Every sentence earns its place — the usage guidance and voice resolution order are both essential for the agent. The voice resolution sentence is dense but packs critical behavioral information. A slight trim could occur since the resolution order could be seen as implementation detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the core essence—synth and play, when to trigger, how voice is resolved—which is solid for a 3-parameter tool. No annotations or schema descriptions exist to offload this burden to. Missing context includes behavior of the speed parameter, the relationship to the sibling list_voices, and failure modes. Given the total absence of annotations and schema descriptions, the description could do more to guide an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to compensate, and it partially does. It adds value for the text parameter ('clear, natural text to speak') and implicitly for the voice parameter by explaining the resolution order. However, the speed parameter is entirely unexplained (what units, what range), and the voice format is only hinted at. Given the extremely low schema coverage, this partial compensation keeps it at a decent but not exceptional level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Synthesize speech with Kokoro-82M and play it on the user's machine') that clearly identifies what the tool does. It goes further to elaborate intended use cases ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It doesn't explicitly contrast itself with the sibling list_voices, which keeps it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance with concrete examples ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It includes a practical instruction to 'Always pass clear, natural text to speak.' However, it stops short of explicitly naming alternatives (e.g., list_voices) for comparison or stating when not to use it, so it lacks the explicit exclusions of a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.4- First observed
list_voices - First observed
speak
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one performs speech synthesis ('speak'), the other lists available voices ('list_voices'). No ambiguity between them, as they serve complementary roles in the TTS workflow.
Both tool names follow an imperative verb style ('speak', 'list_voices'), which is concise and predictable. 'list_voices' uses a verb_noun pattern, but the naming is consistent in tone and clarity for a small set.
With only 2 tools, the server is minimal yet well-scoped for its purpose: a text-to-speech harness. The count is appropriate; additional tools would likely be redundant.
The core operation (speak) is present, and the companion list_voices provides necessary context for voice selection. Minor gaps exist, such as no pause/stop or voice configuration tool, but these are not critical for basic TTS functionality.
Maintenance
Related MCP Connectors
MCP server for Text-to-Speech
AI voice generation: text-to-speech and voice cloning from any MCP client.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
MCP server for Speech-to-Text
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.1MIT
- FlicenseNot gradedqualityDmaintenanceProvides text-to-speech conversion through a unified MCP interface, supporting both local Kokoro and cloud OpenAI TTS engines with streaming audio, voice selection, and customization via natural language instructions.7-
- AlicenseAqualityCmaintenanceMCP server for text-to-speech using macOS say command, enabling speech synthesis, audio file generation, and voice management.512 npm1MIT

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.13 npmMIT