Skip to main content
Glama
rocnubie

Voxtral TTS MCP Server

by rocnubie
README.md
# Voxtral TTS MCP Server

> Voxtral TTS Online Free Text to Speech & Voice Clone Speech

[![MCP Badge](https://lobehub.com/badge/mcp/rocnubie-voxtraltts-mcp)](https://lobehub.com/mcp/rocnubie-voxtraltts-mcp)
[![Node](https://img.shields.io/badge/node-%3E%3D18-339933?logo=node.js&logoColor=white)](https://nodejs.org)
[![smithery](https://smithery.ai/badge/voxtraltts)](https://smithery.ai)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](./LICENSE)
[![Read Only](https://img.shields.io/badge/server-read--only-2ea44f)](#tools)
[![Zero Config](https://img.shields.io/badge/setup-zero--config-7c3aed)](#installation)
[![MCP](https://img.shields.io/badge/MCP-1.0-blue)](https://modelcontextprotocol.io)

<p align="center"><a href="https://voxtraltts.online"><img src="./assets/hero.svg" alt="Voxtral TTS" width="720" /></a></p>

A Model Context Protocol server that exposes the canonical Voxtral TTS knowledge surface — voice and TTS workflows, FAQ, official links — to MCP-compatible AI clients such as Claude Desktop, Cursor, Windsurf, and Continue. Read-only, no API keys, no quota, ~50 ms cold start.

Official website: https://voxtraltts.online

## 🎙️ About Voxtral TTS

Voxtral TTS is Mistral AI's text-to-speech platform, accessible at voxtraltts.online, that converts written text into realistic, emotionally expressive speech across nine languages. Built on a 3.4-billion-parameter transformer decoder backbone paired with a flow-matching acoustic transformer and a neural audio codec, the system is designed for production environments where voice quality, low latency, and speaker consistency across languages all matter. The service is available as a hosted API at $0.016 per 1,000 characters, and the underlying model weights are also published on Hugging Face for teams that prefer self-hosted deployment. Whether you need a single narrator voice that holds up across French and Arabic, or a real-time voice agent that responds in under 100 milliseconds, Voxtral TTS is built to cover that range.

## Key Features

- **Low-latency inference**: The model delivers approximately 70 ms latency with a real-time factor of around 9.7x, making it suitable for interactive voice applications and live agents rather than just batch narration jobs.
- **Zero-shot voice cloning**: A 3-to-25-second reference audio clip is enough to adapt the model to a custom speaker identity, with no fine-tuning pipeline required.
- **Cross-lingual speaker consistency**: The same cloned voice can be reused across all nine supported languages — English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic — without losing the speaker's characteristic sound.
- **Emotional style control**: Outputs can be guided toward distinct expressive styles including neutral, happy, and sarcastic, giving teams control over tone without recording separate voice assets.
- **Built-in voice library**: A set of ready-to-use named voices (Margaret, Paul, Marie, Oliver, and others) lets teams get started immediately without providing reference audio.
- **Flexible deployment**: The API integrates directly through Mistral Studio and standard REST endpoints; the open-source weights option allows air-gapped or on-premise deployments for compliance-sensitive environments.

## Use Cases

- **Customer support automation**: Generate voice responses for IVR systems, automated routing menus, and support bots that need to sound consistent across long sessions and multiple languages.
- **Product onboarding and explainer narration**: Turn documentation, tooltips, or walkthrough scripts into spoken audio without recording studio sessions, and update the audio as fast as you update the text.
- **Multilingual marketing and localization**: Produce regional campaign audio from a single script using the same speaker voice across language variants, keeping brand voice coherent in every market.
- **Real-time voice agents**: Power conversational AI agents, virtual assistants, or phone bots where the round-trip from text to audible speech needs to stay well under a second.
- **Regulated-industry workflows**: Use the self-hosted model weights to run speech synthesis entirely within a private infrastructure, meeting data residency requirements in financial services, healthcare, or manufacturing contexts.

## Who Is It For

Voxtral TTS is aimed at product teams, engineers, and growth operators who are building voice features into applications rather than looking for a one-off audio tool. The API-first design and per-character pricing model suit developers who want to integrate TTS into a larger pipeline — whether that is a customer-facing chatbot, a localization workflow, or an internal voice agent. The zero-shot cloning capability and cross-lingual consistency make it especially useful for teams serving multilingual audiences who cannot afford to maintain separate voice recordings per language. Teams in regulated industries benefit from the open-source weight option, which lets them run inference entirely on their own infrastructure.

## Tools

### `list_voices`
Return the canonical voice and TTS configuration exposed on the site. (Voxtral TTS)

_Input:_ no parameters. _Returns:_ text/markdown.

### `get_official_links`
Return the canonical list of official links for Voxtral TTS (website, support, docs when available).

_Input:_ no parameters. _Returns:_ text/markdown.

## Resources

- `site://voxtraltts/voices` — Supported voices, languages, and TTS modes.
- `site://voxtraltts/faq` — Short FAQ generated from public site metadata.
- `site://voxtraltts/links` — Canonical URLs to share with users.

## Prompts

### `tell_me_about_voxtraltts`
Summarize what the site is, who it's for, and how it works. — Voxtral TTS

### `read_aloud_demo_voxtraltts`
Plan a read-aloud workflow with the site's voices. — Voxtral TTS

## Installation

### Install via Smithery

```bash
npx -y @smithery/cli install voxtraltts-mcp --client claude
```

(Replace `claude` with `cursor`, `windsurf`, or `continue` for those clients.)

### Install from source

```bash
git clone https://github.com/rocnubie/voxtraltts-mcp.git
cd voxtraltts-mcp
pnpm install
```

Then add to your MCP client config (`claude_desktop_config.json` for Claude Desktop, `mcp.json` for Cursor / Windsurf / Continue):

```json
{
  "mcpServers": {
    "voxtraltts-mcp": {
      "command": "node",
      "args": [
        "/absolute/path/to/voxtraltts-mcp/src/index.mjs"
      ]
    }
  }
}
```

### Debug with MCP Inspector

```bash
npx @modelcontextprotocol/inspector node src/index.mjs
```

## Official Links

- Website: https://voxtraltts.online
- Support: support@voxtraltts.online

## Development

```bash
pnpm install
pnpm start                 # run the server over stdio
```

## License

MIT

TDQS

A3.8/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have entirely different purposes: one returns voice and TTS configuration, the other returns official links. No ambiguity.

Naming Consistency5/5

Both tools follow the verb_noun pattern (list_voices, get_official_links), providing consistent and predictable naming.

Tool Count2/5

Only 2 tools for a TTS server is too few; core functionality like text-to-speech synthesis is missing, making the set feel incomplete.

Completeness1/5

The server claims to be a TTS MCP but lacks any actual speech synthesis tool. Essential operations for the domain are absent.