Skip to main content
Glama
README.md
# Rime MCP 

[![rime](rime-logo.png)](https://www.rime.ai)

A Model Context Protocol (MCP) server that provides text-to-speech capabilities using the Rime API. This server downloads audio and plays it using the system's native audio player.

## Features

- Exposes a `speak` tool that converts text to speech and plays it through system audio
- Uses Rime's high-quality voice synthesis API

## Requirements

- Node.js 16.x or higher
- A working audio output device
- macOS: Uses `afplay`

There's sample code from Claude for the following that is not tested 🤙✨
  - Windows: Built-in Media.SoundPlayer (PowerShell)
  - Linux: mpg123, mplayer, aplay, or ffplay

## MCP Configuration

```
"ref": {
  "command": "npx",
  "args": ["rime-mcp"],
  "env": {
      RIME_API_KEY=your_api_key_here

      # Optional configuration
      RIME_GUIDANCE="<guide how the agent speaks>"
      RIME_WHO_TO_ADDRESS="<your name>"
      RIME_WHEN_TO_SPEAK="<tell the agent when to speak>"
      RIME_VOICE="cove" 
  }
}
```

All of the optional env vars are part of the tool definition and are prompts to 

All voice options are [listed here](https://users.rime.ai/data/voices/all-v2.json).

You can get your API key from the [Rime Dashboard](https://rime.ai/dashboard/tokens).

The following environment variables can be used to customize the behavior:

- `RIME_GUIDANCE`: The main description of when and how to use the speak tool
- `RIME_WHO_TO_ADDRESS`: Who the speech should address (default: "user")
- `RIME_WHEN_TO_SPEAK`: When the tool should be used (default: "when asked to speak or when finishing a command")
- `RIME_VOICE`: The default voice to use (default: "cove")

## Example use cases

[![Demo of Rime MCP in Cursor](https://img.youtube.com/vi/tYqTACgijxk/0.jpg)](https://www.youtube.com/watch?v=tYqTACgijxk)


### Example 1: Coding agent announcements

```
"RIME_WHEN_TO_SPEAK": "Always conclude your answers by speaking.",
"RIME_GUIDANCE": "Give a brief overview of the answer. If any files were edited, list them."
```

### Example 2: Learn how the kids talk these days

```
RIME_GUIDANCE="Use phrases and slang common among Gen Alpha."
RIME_WHO_TO_ADDRESS="Matt"
RIME_WHEN_TO_SPEAK="when asked to speak"
```

### Example 3: Different languages based on context

```
RIME_VOICE="use 'cove' when talking about Typescript and 'antoine' when talking about Python"
```


## Development

1. Install dependencies:
```bash
npm install
```

2. Build the server:
```bash
npm run build
```

3. Run in development mode with hot reload:
```bash
npm run dev
```


## License

MIT

## Badges

<a href="https://glama.ai/mcp/servers/@MatthewDailey/rime-mcp">
  <img width="380" height="200" src="https://glama.ai/mcp/servers/@MatthewDailey/rime-mcp/badge" alt="Rime MCP server" />
</a>
<a href="https://smithery.ai/server/@MatthewDailey/rime-mcp"><img alt="Smithery Badge" src="https://smithery.ai/badge/@MatthewDailey/rime-mcp"></a>

### Installing via Smithery

To install Rime Text-to-Speech Server for Claude Desktop automatically via [Smithery](https://smithery.ai/server/@MatthewDailey/rime-mcp):

```bash
npx -y @smithery/cli install @MatthewDailey/rime-mcp --client claude
```

TDQS

A3.7/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The speak tool has a single, clearly defined purpose for text-to-speech conversion.

Naming Consistency5/5

A single tool inherently has perfect naming consistency, as there are no other tools to compare it against. The name 'speak' follows a clear verb pattern appropriate for its function.

Tool Count2/5

A single tool is too few for most MCP server purposes, even for a text-to-speech service. This feels thin and limited, lacking complementary tools like volume control, voice selection, or speech status checks that would enhance functionality.

Completeness2/5

The tool surface is severely incomplete for a text-to-speech domain. While the speak tool covers the core output function, there are obvious gaps such as no tools for managing voices, adjusting speech parameters, stopping speech, or checking speech status, which limits agent capabilities.

Maintenance

ActivityInactive
ResponsivenessNo issues