Skip to main content
Glama
ahamed-kashif

text-to-speech

README.md
# Text-to-Speech MCP Server

An [MCP](https://modelcontextprotocol.io) server that exposes a **text-to-speech**
tool, powered by the free, no-API-key [`gTTS`](https://pypi.org/project/gTTS/)
library (which wraps Google Translate's internal TTS endpoint).

- **Free** — no API key, no billing, no GCP account.
- **Many languages** — English, Bengali, Spanish, and ~60 more.
- **File output** — saves an MP3 and returns its path.
- **Resilient** — retries transient/rate-limit failures with exponential backoff.

> Note: gTTS uses an *undocumented* Google endpoint. It requires internet access
> and can rate-limit or change without notice — great for personal/internal use,
> less suited to high-volume production. The engine is isolated in `engine.py` so
> it can be swapped for Google Cloud TTS or an offline engine later.

## Requirements

- [uv](https://docs.astral.sh/uv/)
- Python 3.12 (uv will provision it)

## Setup

```bash
uv sync
```

This creates `.venv/` and installs the project (and its deps: `gtts`, `mcp`).

Optional configuration — copy `.env.example` to `.env` and edit:

| Variable | Default | Meaning |
|---|---|---|
| `TTS_OUTPUT_DIR` | `./output` | Where MP3s are written when no `output_path` is given |
| `TTS_MAX_RETRIES` | `3` | Max synthesis attempts before failing |
| `TTS_RETRY_BACKOFF` | `1.0` | Base backoff seconds (1s, 2s, 4s, …) |
| `TTS_DEFAULT_LANG` | `en` | Default language code |
| `TTS_DEFAULT_ACCENT` | `com` | Default accent TLD |

## Run

```bash
uv run tts-server
# or
uv run python -m tts_server.server
```

The server speaks MCP over **stdio**.

## Tools

### `text_to_speech`
| Arg | Type | Default | Description |
|---|---|---|---|
| `text` | string | — (required) | Text to speak |
| `language` | string | `en` | Language code (`en`, `bn`, `es`, …) |
| `slow` | boolean | `false` | Speak more slowly |
| `accent` | string | `com` | Accent TLD: `com` (US), `co.uk` (UK), `com.au` (AU), `co.in` (India) |
| `output_path` | string | auto | Output MP3 path; auto-generated under `TTS_OUTPUT_DIR` if omitted |

Returns: `Saved <N> KB MP3 to <path>`, or `Error: <reason>`.

### `list_languages`
No arguments. Returns supported `code: name` pairs.

## Install in an MCP client

This is a standard **stdio** MCP server, so it works with any MCP-compatible
client (Claude Code, Claude Desktop, Cursor, Windsurf, Continue, etc.). Every
client needs the same three things:

| Field | Value |
|---|---|
| **command** | `uv` |
| **args** | `["--directory", "<ABS_PROJECT_PATH>", "run", "tts-server"]` |
| **env** *(optional)* | e.g. `{ "TTS_OUTPUT_DIR": "<ABS_OUTPUT_PATH>" }` |

Replace `<ABS_PROJECT_PATH>` with the absolute path to this project. Using
`uv --directory <path> run` means the client can launch the server from anywhere
and still resolve the right virtual env and `.env`. Make sure [`uv`](https://docs.astral.sh/uv/)
is on the client's `PATH` (or use the absolute path to the `uv` binary — find it
with `which uv`).

### Claude Code (CLI)

```bash
# user scope = available in every project; use --scope project for just this repo
claude mcp add text-to-speech --scope user -- \
  uv --directory /abs/path/to/text-to-speech-agent run tts-server
```

Verify, inspect, or remove:

```bash
claude mcp list                 # health check
claude mcp get text-to-speech   # show config
claude mcp remove text-to-speech
```

Tools load when a session starts, so open a **new** session after adding.

### Claude Desktop

Edit the config file (Settings → Developer → Edit Config, or directly):

- macOS: `~/Library/Application Support/Claude/claude_desktop_config.json`
- Windows: `%APPDATA%\Claude\claude_desktop_config.json`

Add (see `claude_desktop_config.example.json`):

```json
{
  "mcpServers": {
    "text-to-speech": {
      "command": "uv",
      "args": ["--directory", "/abs/path/to/text-to-speech-agent", "run", "tts-server"],
      "env": { "TTS_OUTPUT_DIR": "/abs/path/to/text-to-speech-agent/output" }
    }
  }
}
```

Then fully restart Claude Desktop.

### Cursor / Windsurf / other JSON-config clients

These use the same `mcpServers` JSON shape as Claude Desktop. Add the block above
to the client's MCP config (e.g. Cursor: Settings → MCP, or `~/.cursor/mcp.json`),
then reload the client.

### Generic stdio client

Any client that can launch a stdio server: configure it to run

```
uv --directory /abs/path/to/text-to-speech-agent run tts-server
```

and (optionally) pass environment variables from the table in
[Setup](#setup).

### Troubleshooting

- **Server fails to start / "uv not found"** — use the absolute path to `uv`
  (`which uv`) as the `command`, since GUI clients often don't inherit your shell `PATH`.
- **Tools don't appear** — restart the client (Claude Code: start a new session).
- **MP3s land in the wrong place** — set `TTS_OUTPUT_DIR` in the client's `env`
  block (it overrides `.env`).
- **Synthesis errors** — gTTS needs internet and may rate-limit; the server
  retries automatically, but check connectivity if it persists.

## Project layout

```
src/tts_server/
  server.py   # MCP tools, validation, response shaping
  engine.py   # gTTS wrapper + retry/backoff (only file importing gTTS)
  config.py   # env-driven settings
  utils.py    # filename gen, language validation, file sizing
tests/        # unit tests (retry paths, mocked gTTS)
output/       # generated MP3s
```

## Tests

```bash
uv run python -m pytest -q
```

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one lists supported languages, and the other converts text to speech. There is no overlap or ambiguity.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern (list_languages, text_to_speech), making the naming predictable and intuitive.

Tool Count3/5

With only two tools, the server is quite thin but still covers the core functionality for a simple text-to-speech service. The count is borderline for the domain.

Completeness4/5

The domain is well-covered by listing languages and converting text to speech with various options. No major gaps are apparent, though additional settings tools could be added.

Maintenance

ActivityInactive
ResponsivenessNo issues