realtime-tts-mcp
by tsushanth
README.md
# realtime-tts-mcp
An MCP server wrapping [ReadAloud](https://readaloudai.org)'s realtime streaming text-to-speech API (Kokoro-82M) as a single `synthesize_speech` tool. Returns a playable WAV file. Typical warm latency is well under a second.
## Tool
- **`synthesize_speech(text, voice?, speed?)`** — converts text to spoken audio. `voice` defaults to `af_heart` (a Kokoro voice id); `speed` defaults to `1.0`.
## Install
```bash
npx realtime-tts-mcp
```
Or from source:
```bash
git clone https://github.com/tsushanth/realtime-tts-mcp.git
cd realtime-tts-mcp
npm install
npm run build
```
Add to your MCP client config:
```json
{
"mcpServers": {
"realtime-tts": {
"command": "npx",
"args": ["-y", "realtime-tts-mcp"],
"env": {
"REALTIME_TTS_API_KEY": "rtts_your_key_here"
}
}
}
}
```
## Getting an API key
Self-serve signup at **https://readaloudai.org/developers** — sign in with email, generate a key. **10,000 characters free, no card required.** Beyond that, pay-as-you-go at $0.01 per 1,000 characters (no plan to manage). Without a key, calls fail with a clear error pointing you to the signup page.
## How it works
Each call does two steps: authorize (a fast key + free-tier/billing check against the ReadAloud API) then connects directly to the synthesis worker with a short-lived token — no relay hop in between, which is what keeps warm-call latency well under a second. A cold worker (idle for a couple minutes) can take several seconds longer on the first call after a gap.
## Environment variables
| Variable | Default | Purpose |
|---|---|---|
| `REALTIME_TTS_API_KEY` | *(none, required)* | Get one free at https://readaloudai.org/developers |
| `REALTIME_TTS_API_BASE` | `https://api.readaloudai.org` | API base URL |
| `REALTIME_TTS_TIMEOUT_MS` | `30000` | Per-synthesis timeout |
## License
MIT
TDQS
A4.1/5.0
Scored across 1 tool
Disambiguation5/5
With only one tool, there is no possibility of overlap or confusion. synthesize_speech clearly indicates its single purpose of converting text to audio.
Naming Consistency5/5
The tool name follows a clear verb_noun pattern and directly matches the server's TTS purpose. There are no conflicting naming conventions to assess.
Tool Count3/5
A single tool is reasonable for a narrowly scoped TTS server, but the surface feels thin because related capabilities like voice selection or streaming control are absent.
Completeness3/5
The core text-to-speech operation is covered, but the toolset lacks visible voice, format, or streaming options that would be expected in a realtime TTS service. It works for basic synthesis but leaves notable gaps.
Maintenance
ActivityMaintained
ResponsivenessNo issues