Rime MCP
The Rime MCP server provides text-to-speech capabilities using Rime's API with the following features:
Converts text to speech and plays it through the system's native audio player
Supports customization of voice (speakers), speech speed, and latency optimization
Can be configured via environment variables for guidance, addressing, timing, and default voice settings
Offers cross-platform support for macOS, Windows, and Linux
Integrates with MCP workflows for contextual speech generation
Supports various Linux audio players (mpg123, mplayer, aplay, ffplay) for playing synthesized speech
Uses macOS native 'afplay' audio player to output synthesized speech
Supports context-based voice selection when discussing Python, using specific voices like 'antoine'
Provides text-to-speech capabilities using Rime's API, allowing AI agents to convert text to speech and play it through the system's audio output
Supports context-based voice selection when discussing TypeScript, using specific voices like 'cove'
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Rime MCPread my latest commit message out loud"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Rime MCP

A Model Context Protocol (MCP) server that provides text-to-speech capabilities using the Rime API. This server downloads audio and plays it using the system's native audio player.
Features
Exposes a
speaktool that converts text to speech and plays it through system audioUses Rime's high-quality voice synthesis API
Related MCP server: AivisSpeech MCP Server
Requirements
Node.js 16.x or higher
A working audio output device
macOS: Uses
afplay
There's sample code from Claude for the following that is not tested 🤙✨
Windows: Built-in Media.SoundPlayer (PowerShell)
Linux: mpg123, mplayer, aplay, or ffplay
MCP Configuration
"ref": {
"command": "npx",
"args": ["rime-mcp"],
"env": {
RIME_API_KEY=your_api_key_here
# Optional configuration
RIME_GUIDANCE="<guide how the agent speaks>"
RIME_WHO_TO_ADDRESS="<your name>"
RIME_WHEN_TO_SPEAK="<tell the agent when to speak>"
RIME_VOICE="cove"
}
}All of the optional env vars are part of the tool definition and are prompts to
All voice options are listed here.
You can get your API key from the Rime Dashboard.
The following environment variables can be used to customize the behavior:
RIME_GUIDANCE: The main description of when and how to use the speak toolRIME_WHO_TO_ADDRESS: Who the speech should address (default: "user")RIME_WHEN_TO_SPEAK: When the tool should be used (default: "when asked to speak or when finishing a command")RIME_VOICE: The default voice to use (default: "cove")
Example use cases

Example 1: Coding agent announcements
"RIME_WHEN_TO_SPEAK": "Always conclude your answers by speaking.",
"RIME_GUIDANCE": "Give a brief overview of the answer. If any files were edited, list them."Example 2: Learn how the kids talk these days
RIME_GUIDANCE="Use phrases and slang common among Gen Alpha."
RIME_WHO_TO_ADDRESS="Matt"
RIME_WHEN_TO_SPEAK="when asked to speak"Example 3: Different languages based on context
RIME_VOICE="use 'cove' when talking about Typescript and 'antoine' when talking about Python"Development
Install dependencies:
npm installBuild the server:
npm run buildRun in development mode with hot reload:
npm run devLicense
MIT
Badges
Installing via Smithery
To install Rime Text-to-Speech Server for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @MatthewDailey/rime-mcp --client claudeAvailable Tools
1 toolspeakA
Speak text aloud using Rime's text-to-speech API. Should be used when user asks you to speak or to announce and explain when you finish a command
User configuration:
WHO_TO_ADDRESS: user
WHEN_TO_SPEAK: when asked to speak or when finishing a command
VOICE: cove
GUIDANCE: Use the speak tool to convert text to speech when the user requests audio output or when providing verbal responses
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to speak aloud | |
| speaker | No | The voice to use (defaults to 'cove') | |
| speedAlpha | No | Speech speed multiplier (default: 1.0) | |
| reduceLatency | No | Whether to optimize for lower latency (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the API ('Rime's text-to-speech API') and usage context, but lacks details on behavioral traits such as rate limits, authentication requirements, error handling, or output format. The description does not contradict annotations, but it provides only basic operational context without deeper behavioral insights.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat verbose and includes redundant sections like 'User configuration:' with WHO_TO_ADDRESS, WHEN_TO_SPEAK, VOICE, and GUIDANCE, which could be integrated more efficiently. While it provides useful information, the structure is not optimally front-loaded, and some sentences (e.g., the configuration headers) do not add significant value beyond the core description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (4 parameters, no output schema, no annotations), the description is fairly complete. It covers purpose, usage guidelines, and basic context, but lacks details on behavioral aspects like performance or errors. Without annotations or output schema, it does enough to guide usage but could be more comprehensive for full transparency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description does not add any parameter-specific information beyond what the schema provides (e.g., it mentions 'VOICE: cove' but the schema already describes the 'speaker' parameter with a default). Baseline score of 3 is appropriate as the schema handles parameter documentation effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Speak text aloud using Rime's text-to-speech API.' It specifies the verb ('speak') and resource ('text'), and distinguishes it from potential alternatives by mentioning the specific API. However, since there are no sibling tools, the differentiation aspect is not applicable, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: 'Should be used when user asks you to speak or to announce and explain when you finish a command' and 'Use the speak tool to convert text to speech when the user requests audio output or when providing verbal responses.' It clearly defines when to use the tool, including specific scenarios, making it highly actionable for an AI agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool, there is no possibility of ambiguity or overlap between tools. The speak tool has a single, clearly defined purpose for text-to-speech conversion.
A single tool inherently has perfect naming consistency, as there are no other tools to compare it against. The name 'speak' follows a clear verb pattern appropriate for its function.
A single tool is too few for most MCP server purposes, even for a text-to-speech service. This feels thin and limited, lacking complementary tools like volume control, voice selection, or speech status checks that would enhance functionality.
The tool surface is severely incomplete for a text-to-speech domain. While the speak tool covers the core output function, there are obvious gaps such as no tools for managing voices, adjusting speech parameters, stopping speech, or checking speech status, which limits agent capabilities.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
AI voice generation: text-to-speech and voice cloning from any MCP client.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
MCP server for Text-to-Speech
Related MCP Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server that provides text-to-speech capabilities using the Kokoro TTS model, offering multiple voice options and customizable speech parameters.4251MIT
- FlicenseDqualityCmaintenanceA Model Context Protocol server that enables AI assistants to utilize AivisSpeech Engine's high-quality voice synthesis capabilities through a standardized API interface.11
- AlicenseNot gradedqualityCmaintenanceA Model Context Protocol server that integrates high-quality text-to-speech capabilities with Claude Desktop and other MCP-compatible clients, supporting multiple voice options and audio formats.171MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that provides text-to-speech functionality for AI agents using Microsoft Edge's text-to-speech technology, supporting multiple voices, languages, and voice customization.28MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/MatthewDailey/rime-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server