Skip to main content
Glama
README.md
# MCP Coqui TTS Server

A Model Context Protocol (MCP) server that provides text-to-speech synthesis capabilities using Coqui TTS, including voice cloning support.

## Features

- **Text-to-Speech Synthesis**: Convert text to natural-sounding speech
- **Multiple Models**: Support for various TTS models and languages
- **Voice Cloning**: Clone voices from audio samples using XTTS models
- **Long Text Support**: Automatic chunking for longer texts
- **Customizable Output**: Control speed, speaker, and language settings

## Prerequisites

Before using this MCP server, you need to install Coqui TTS:

```bash
pip install TTS
```

For voice cloning and concatenation features, you'll also need ffmpeg:

```bash
# macOS
brew install ffmpeg

# Ubuntu/Debian
sudo apt-get install ffmpeg

# Windows (using chocolatey)
choco install ffmpeg
```

## Installation

### From npm

```bash
npm install -g @s.lfr/mcp-coqui-tts
```

### From Source

```bash
git clone https://github.com/yourusername/mcp-coqui-tts.git
cd mcp-coqui-tts
npm install
npm link
```

## Usage

### With Claude Desktop

Add to your Claude Desktop configuration (`~/Library/Application Support/Claude/claude_desktop_config.json` on macOS):

```json
{
  "mcpServers": {
    "coqui-tts": {
      "command": "npx",
      "args": ["@s.lfr/mcp-coqui-tts"]
    }
  }
}
```

Or if installed from source:

```json
{
  "mcpServers": {
    "coqui-tts": {
      "command": "node",
      "args": ["/path/to/mcp-coqui-tts/index.js"]
    }
  }
}
```

### Available Tools

#### 1. `speak`
Convert text to speech with customizable parameters.

**Parameters:**
- `text` (required): The text to convert to speech
- `model`: TTS model to use (default: "tts_models/en/ljspeech/tacotron2-DDC")
- `output_path`: Where to save the audio file
- `speaker_idx`: Speaker index for multi-speaker models
- `language_idx`: Language index for multi-language models
- `speed`: Speed factor (1.0 is normal speed)

**Example:**
```javascript
{
  "text": "Hello, this is a test of the text to speech system.",
  "model": "tts_models/en/ljspeech/tacotron2-DDC",
  "output_path": "/tmp/output.wav",
  "speed": 1.2
}
```

#### 2. `list_models`
List all available TTS models.

**Parameters:** None

#### 3. `synthesize_long_text`
Synthesize longer texts with automatic chunking and concatenation.

**Parameters:**
- `text` (required): The long text to convert
- `model`: TTS model to use
- `output_path`: Where to save the final audio
- `chunk_size`: Maximum characters per chunk (default: 500)

**Example:**
```javascript
{
  "text": "This is a very long text that will be automatically split into chunks...",
  "output_path": "/tmp/long_speech.wav",
  "chunk_size": 500
}
```

#### 4. `clone_voice`
Clone a voice from an audio sample (requires XTTS model).

**Parameters:**
- `text` (required): Text to speak in the cloned voice
- `reference_audio` (required): Path to reference audio for voice cloning
- `output_path`: Where to save the output
- `language`: Language code (default: "en")

**Example:**
```javascript
{
  "text": "This will be spoken in the cloned voice.",
  "reference_audio": "/path/to/sample.wav",
  "output_path": "/tmp/cloned_voice.wav",
  "language": "en"
}
```

## Popular TTS Models

### English Models
- `tts_models/en/ljspeech/tacotron2-DDC` - High quality English TTS
- `tts_models/en/ljspeech/fast_pitch` - Fast English TTS
- `tts_models/en/vctk/vits` - Multi-speaker English (110 speakers)

### Multilingual Models
- `tts_models/multilingual/multi-dataset/xtts_v2` - Supports voice cloning
- `tts_models/multilingual/multi-dataset/your_tts` - Multilingual with voice cloning

### Other Languages
Run `list_models` to see all available models for different languages.

## Deployment on Smithery

To deploy this MCP server on Smithery:

1. Fork this repository
2. Connect your GitHub account to Smithery
3. Create a new MCP server on Smithery
4. Select this repository
5. Deploy

The server will be automatically available for use with any MCP-compatible client.

## Development

### Running Locally

```bash
npm start
```

### Testing

You can test the server using the MCP inspector:

```bash
npx @modelcontextprotocol/inspector node index.js
```

## Troubleshooting

### Common Issues

1. **"tts: command not found"**
   - Make sure Coqui TTS is installed: `pip install TTS`
   - Ensure Python/pip binaries are in your PATH

2. **"ffmpeg: command not found"**
   - Install ffmpeg for your operating system (see Prerequisites)

3. **Model download fails**
   - First run may take time as models are downloaded
   - Check internet connection
   - Ensure sufficient disk space (~1-5GB per model)

4. **Voice cloning not working**
   - Requires XTTS v2 model
   - Reference audio should be clear, 5-10 seconds long
   - WAV format recommended for reference audio

## Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

## License

MIT

## Acknowledgments

- [Coqui TTS](https://github.com/coqui-ai/TTS) for the excellent TTS library
- [Model Context Protocol](https://modelcontextprotocol.io/) for the MCP SDK

## Support

For issues and questions, please open an issue on GitHub.

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation4/5

speak and synthesize_long_text both perform text-to-speech, but synthesize_long_text is explicitly for longer text with automatic chunking, which sets it apart. clone_voice and list_models are clearly distinct, so overall only minor overlap exists.

Naming Consistency4/5

Most tools use a verb_noun pattern (list_models, synthesize_long_text, clone_voice), but 'speak' is a bare verb without an object, creating a slight inconsistency. The pattern is still mostly predictable away from this.

Tool Count5/5

With four tools, the server is well-scoped for its TTS purpose: listing models, synthesizing short and long text, and cloning voices. Each tool earns its place, and the count feels neither too thin nor too heavy.

Completeness4/5

The tool surface covers the core TTS workflows: model discovery, standard and long-form synthesis, and voice cloning. Minor gaps like audio streaming or detailed model management are not essential for a basic TTS server, so the coverage is solid.

Maintenance

ActivityInactive
ResponsivenessNo issues