mcp-kokoro-tts
mcp-kokoro-tts
로컬 Kokoro-82M 텍스트 음성 변환 MCP 서버입니다. 에이전트가 speak를 호출하면 음성을 합성하여 사용자 컴퓨터에서 재생하므로 하네스가 말하는 소리를 들을 수 있습니다.
Claude Desktop, Claude Code, Cursor, VS Code, opencode, Cline 등 모든 MCP 클라이언트에서 작동합니다. 짧은 구성 블록 하나면 충분하며 API 키가 필요 없습니다. 합성은 Kokoro-82M으로 로컬에서 실행됩니다.
첫 시작 시 서버는 PyPI에 없는 두 가지를 준비합니다: Kokoro-82M 가중치(약 312MB)를 로컬 캐시에, 그리고 spaCy의 영어 모델(en_core_web_sm)을 서버가 실행 중인 동일한 Python 환경에 설치합니다. 두 번째 설치가 필요한 이유는 Kokoro의 G2P 파이프라인이 spaCy를 로드하기 때문이며, uvx/uv tool 환경에는 이 패키지가 설치하지 않는 한 모델이 없기 때문입니다.
설치
클라이언트의 MCP 구성에 추가하세요:
{
"mcpServers": {
"mcp-kokoro-tts": {
"command": "uvx",
"args": ["mcp-kokoro-tts"]
}
}
}Python 3.12와 uv가 필요합니다. 첫 서버 시작 시 Kokoro 가중치와 spaCy 영어 모델이 자동으로 준비됩니다.
MCP 서버를 시작하지 않고 둘 다 미리 다운로드하려면:
uvx mcp-kokoro-tts-provisionRelated MCP server: MCP TTS Server
에이전트가 호출하도록 만들기
AGENTS.md / CLAUDE.md / 시스템 프롬프트에 한 줄을 추가하세요:
When the user wants to hear something spoken aloud, call the `speak` tool with clear, natural text.도구
speak
음성을 합성하고 WAV 파일을 작성한 후 로컬에서 재생합니다.
매개변수 | 필수 | 설명 |
| 예 | 말할 텍스트 (최대 500자) |
| 아니요 | 음성 ID (예: |
| 아니요 | 재생 속도 배수 (기본값 |
list_voices
사용 가능한 Kokoro 음성과 현재 선택된 기본 음성을 나열합니다.
음성 선택
선택 순서:
TTS_VOICE환경 변수 — 음성 ID 또는.pt절대 경로패키지
voices/폴더에서 이름이default로 시작하는 파일voices/의 첫 번째.pt파일 (알파벳순)모델에 포함된
af_heart음성
{
"mcpServers": {
"mcp-kokoro-tts": {
"command": "uvx",
"args": ["mcp-kokoro-tts"],
"env": {
"TTS_VOICE": "af_heart"
}
}
}
}환경 변수
변수 | 설명 |
| 기본 음성 ID 또는 |
| 모델 캐시 디렉터리 재정의 |
| Hugging Face 허브 캐시 디렉터리 재정의 |
| 생성된 WAV 파일용 디렉터리 |
| 로컬 재생 없이 합성하려면 |
| 더 빠른 다운로드를 위한 선택적 Hugging Face 토큰 |
플랫폼
OS | 합성 | 재생 |
macOS | 예 |
|
Linux | 예 |
|
Windows | 예 | PowerShell |
espeak-ng는 선택 사항입니다. 영어는 없어도 작동합니다. 어휘 외 단어 커버리지와 일부 비영어 언어를 위해 설치하세요.
배포
버전 태그를 지정하면 GitHub Actions publish.yml이 실행되어 PyPI에 업로드한 다음 MCP Registry에 업로드합니다.
PyPI 배포는 저장소 시크릿 PYPI_TOKEN(PyPI API 토큰)을 사용합니다. PyPI 프로젝트에 GitHub trusted publishing을 구성할 수도 있습니다. 이 워크플로는 토큰으로 인증하므로 첫 릴리스가 pending-publisher 매칭에 의존하지 않습니다.
릴리스
pyproject.toml의version을 올립니다 (아직 태그를 지정하지 않는 경우server.json도)커밋하고 태그 지정:
git tag v0.1.2 && git push origin v0.1.2GitHub Actions가
publish.yml을 실행합니다:release— 타입체크, 테스트, wheel/sdist 빌드pypi-publish—PYPI_TOKEN으로 PyPI에 업로드mcp-registry— OIDC → MCP Registry (PyPI 성공 후)
개발
cd mcps-tts
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pyright
pytest
python -m mcp_kokoro_tts라이선스
Apache-2.0. LICENSE 및 NOTICE를 참조하세요. Kokoro-82M 모델 가중치는 Apache-2.0 라이선스에 따라 별도로 다운로드됩니다.
Available Tools
2 toolslist_voicesList voicesA
List Kokoro voices available to the speak tool and the currently selected default voice. Resolution order: TTS_VOICE env var, voices/default* file, first file in voices/, then the model's default af_heart voice.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does exceptionally well: it reveals the exact resolution order (env var, file, first file, fallback model default), which behaviorally affects what the tool returns. This is precisely the kind of contextual detail an agent needs beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, the first clearly scopes the listing and the second enumerates the resolution order. Every phrase adds value and the default-voice resolution rule is stated in a compact, order-relevant manner.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description does not need to explain return format. It fully explains what this tool does (list voices and default voice) and even how the default is determined, which is the behavioral nuance essential for an agent to pick the right voice via speak.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is already complete (100% coverage) and there is nothing for the description to elaborate on. Baseline 4 applies since there is simply no parameter information needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('Kokoro voices available to the speak tool') and also includes the currently selected default voice. The mention of 'speak tool' clearly distinguishes it from the sibling speak tool, making its function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool should be used to discover available voice options before using the speak tool, especially because it names the speak tool as the consumer. It clearly states the resolution order for the default voice, giving immediate context, but it does not explicitly say when not to use it or mention alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speakSpeakA
Synthesize speech with Kokoro-82M and play it on the user's machine. Call this when the user wants the harness to talk, read something aloud, or hear a spoken response. Always pass clear, natural text to speak. Default voice resolution: TTS_VOICE env var, voices/default* in the package, first file in voices/, then the bundled af_heart voice from the model.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| speed | No | ||
| voice | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it makes a reasonable effort. It discloses the side effect of audio playback ('play it on the user's machine') and thoroughly documents the voice resolution fallback order (TTS_VOICE env var, voices/default*, first file in voices/, then af_heart). However, it doesn't disclose potential blocking behavior, absence of audio hardware, or other side effects that could be expected given the nature of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, with the core purpose front-loaded first and the resolution detail last. Every sentence earns its place — the usage guidance and voice resolution order are both essential for the agent. The voice resolution sentence is dense but packs critical behavioral information. A slight trim could occur since the resolution order could be seen as implementation detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the core essence—synth and play, when to trigger, how voice is resolved—which is solid for a 3-parameter tool. No annotations or schema descriptions exist to offload this burden to. Missing context includes behavior of the speed parameter, the relationship to the sibling list_voices, and failure modes. Given the total absence of annotations and schema descriptions, the description could do more to guide an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to compensate, and it partially does. It adds value for the text parameter ('clear, natural text to speak') and implicitly for the voice parameter by explaining the resolution order. However, the speed parameter is entirely unexplained (what units, what range), and the voice format is only hinted at. Given the extremely low schema coverage, this partial compensation keeps it at a decent but not exceptional level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Synthesize speech with Kokoro-82M and play it on the user's machine') that clearly identifies what the tool does. It goes further to elaborate intended use cases ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It doesn't explicitly contrast itself with the sibling list_voices, which keeps it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance with concrete examples ('when the user wants the harness to talk, read something aloud, or hear a spoken response'). It includes a practical instruction to 'Always pass clear, natural text to speak.' However, it stops short of explicitly naming alternatives (e.g., list_voices) for comparison or stating when not to use it, so it lacks the explicit exclusions of a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.4- First observed
list_voices - First observed
speak
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one performs speech synthesis ('speak'), the other lists available voices ('list_voices'). No ambiguity between them, as they serve complementary roles in the TTS workflow.
Both tool names follow an imperative verb style ('speak', 'list_voices'), which is concise and predictable. 'list_voices' uses a verb_noun pattern, but the naming is consistent in tone and clarity for a small set.
With only 2 tools, the server is minimal yet well-scoped for its purpose: a text-to-speech harness. The count is appropriate; additional tools would likely be redundant.
The core operation (speak) is present, and the companion list_voices provides necessary context for voice selection. Minor gaps exist, such as no pause/stop or voice configuration tool, but these are not critical for basic TTS functionality.
Maintenance
Related MCP Connectors
MCP server for Text-to-Speech
AI voice generation: text-to-speech and voice cloning from any MCP client.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
MCP server for Speech-to-Text
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.1MIT
- FlicenseNot gradedqualityDmaintenanceProvides text-to-speech conversion through a unified MCP interface, supporting both local Kokoro and cloud OpenAI TTS engines with streaming audio, voice selection, and customization via natural language instructions.7-
- AlicenseAqualityCmaintenanceMCP server for text-to-speech using macOS say command, enabling speech synthesis, audio file generation, and voice management.512 npm1MIT

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.13 npmMIT