Kokoro Text to Speech MCP Server
Related Servers
Alternatives to Kokoro Text to Speech MCP Server
No user-submitted related servers found.
Related Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server that provides text-to-speech capabilities using the Kokoro TTS model, offering multiple voice options and customizable speech parameters.4321MIT
- AlicenseNot gradedqualityDmaintenanceAn OpenAI-compatible text-to-speech API server that enables voice cloning, long-form generation, and can be used as a Claude Code tool for speech synthesis.MIT
- AlicenseNot gradedqualityDmaintenanceText-to-speech synthesis server with multiple AI character voices, supporting streaming playback and Claude Desktop integration.1MIT
- AlicenseNot gradedqualityCmaintenanceHeadless text-to-speech and speech-to-text server with REST and MCP API, supporting Kokoro TTS and Whisper STT.MIT
- AlicenseNot gradedqualityBmaintenanceLocal MCP server for neural text-to-speech using Kokoro ONNX engine on CPU, supporting SSML tags, multiple voice profiles, and zero-GPU operation for low-latency speech synthesis.32MIT
- AlicenseNot gradedqualityDmaintenanceA text-to-speech MCP server with 48 voices across 9 languages, supporting emotion spans, SFX tags, and multi-speaker dialogue. Deployable via a single npx command with built-in guardrails and swappable backends.MIT
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool has a clearly defined purpose of converting text to speech, which cannot be confused with any other functionality.
The single tool name 'text_to_speech' follows a clear verb_noun pattern and uses snake_case consistently. With only one tool, naming consistency is inherently perfect as there are no other tools to compare against.
A single tool is too few for a server's apparent scope of text-to-speech functionality. While the tool itself is well-defined, a complete TTS service would typically include additional tools such as listing available voices, managing audio files, or configuring settings, making this server feel thin and incomplete.
The server is severely incomplete for a text-to-speech domain. It lacks essential operations such as listing available voices, checking service status, managing generated audio files (beyond the single generation tool), or handling configuration. This creates significant gaps that will limit agent workflows and cause failures in more complex tasks.