cursor-tts-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cursor-tts-mcpSpeak progress updates as you work, and stop when I ask."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cursor-tts MCP
Progress-triggered TTS for Cursor Agent (speak / stop). Voice follows the system UI language (edge-tts; Windows-first).
Marketplace / Plugin install
Requires uv (uvx).
Way | Steps |
Cursor Marketplace | After listing: search |
Local Plugin | Clone repo → Cursor: add local plugin / copy into |
Manual MCP | Point MCP at |
Plugin layout:
.cursor-plugin/plugin.jsonmcp.json(uvx --from ${PLUGIN_ROOT})rules/cursor-tts-speak.mdcassets/logo.svg
Submit checklist: docs/guide/Marketplace提交清单.md · Form: https://cursor.com/marketplace/publish
Related MCP server: Kokoro TTS MCP Server
Docs
See docs/README.md.
Quick start (dev)
cd E:\WorkSpace\mcp
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
pytestDefault Marketplace mode is embedded (MCP process speaks). Optional P1: tts_service on 127.0.0.1:18765 + thin HTTP client.
# optional standalone service
python -m tts_serviceLocal Cursor wiring:
.cursor/mcp.json→ MCPcursor-tts.cursor/rules/cursor-tts-speak.mdc→ speak-on-progress rule
If tools missing: reload MCP / restart Cursor. See docs/guide/Cursor接入与配置说明.md.
Layout
.cursor-plugin/ # Marketplace plugin manifest
mcp.json # Portable MCP launch for Plugin
rules/ # Agent rule shipped with Plugin
src/cursor_tts_mcp/ # MCP + orchestrator + engines + playback
config/default.yaml # defaults
docs/ # design & contracts
tests/Available Tools
2 toolsspeakA
Speak a short status line for the user. Use for real progress, task conclusions, errors/blocks, and when the user must decide or act. Always call this when delivering a conclusion — do not only type. Do not read code or long text. Do not wait until the whole task finishes for the first call. Speak in the system UI language (en): use short English phrases (≤80 characters). Do not force Chinese if the system language is not Chinese.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| category | No | progress | |
| priority | No | normal | |
| interrupt | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to carry safety or behavioral signals, the description does most of the work itself. It discloses that the tool should not be used to read code or long text, that it should be called even when a conclusion is typed, and that spoken output must be short English (≤80 characters). It does not explain interrupt or priority behavior, but the core behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and usage, and each instruction serves a distinct function. The final 'Do not force Chinese' clause is somewhat redundant after the 'system UI language (en)' note, but it is a cheap and useful reminder.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple output tool, the description covers required inputs, category mapping, output style, language, and call timing. The main gap is the undocumented 'priority' and 'interrupt' parameters, but these are optional and defaulted, so the description remains largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does add meaning for 'text' (short, English, ≤80 characters) and 'category' (progress, errors, conclusions, decisions), corresponding to the enum values. However, 'priority' and 'interrupt' are never explained, leaving their semantics to names and defaults alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-plus-resource statement ('Speak a short status line for the user') and then enumerates the exact event types it is for: progress, task conclusions, errors/blocks, and decisions. This makes the tool's role obvious and distinguishes it from the sibling tool 'stop'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to call it ('Use for real progress, task conclusions, errors/blocks, and when the user must decide or act') and when not to ('Do not read code or long text'). It also gives operational guidance about calling early instead of waiting for full-task completion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stopA
Immediately stop current TTS playback. Default clear_queue=true; if false, only stop current and keep the queue.
| Name | Required | Description | Default |
|---|---|---|---|
| clear_queue | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the immediate stop action, the default clear_queue=true, and the conditional behavior when false ('only stop current and keep the queue'). This is meaningful behavioral detail beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action, and no filler. Every word adds value, covering the action and the parameter behavior efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool, this is complete: it states the action, the parameter semantics, and the default. Since an output schema exists, return-value documentation is not required from the description, and nothing essential for calling the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully explains the single parameter clear_queue: its default value and the effect when false. This completely compensates for the schema's lack of semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Immediately stop current TTS playback.' This clearly distinguishes it from the sibling tool speak, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool, and the only sibling (speak) makes the use case obvious. While it does not explicitly name an alternative, the parameter guidance about clearing vs. keeping the queue adds practical usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
speak - First observed
stop
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: speak outputs audio, stop halts playback. No overlap or ambiguity exists between them.
Both tool names are single-word imperative verbs (speak, stop) that directly describe their actions. The naming pattern is perfectly consistent and predictable.
With only two tools, the set is minimal but appropriate for a narrow TTS status-line server. It feels slightly thin, but each tool earns its place.
For the stated purpose of speaking short status lines and stopping playback, the tools cover the full lifecycle with no obvious gaps. Stop even offers queue-clearing behavior.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
Turn any article, document, or chat reply into a one-word-at-a-time RSVP reading session.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables coding agents to speak aloud using text-to-speech functionality. Works with agents running inside devcontainers and provides configurable voice settings for creating chatty AI companions.4-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to generate and play high-quality text-to-speech audio using the Kokoro model, with support for multiple voices, adjustable speaking speed, and audio caching.-
- AlicenseNot gradedqualityDmaintenanceProvides text-to-speech generation using the Kokoro-82M model, enabling AI assistants to generate voiceovers and audio content directly within Claude Desktop and Cursor.14Apache 2.0
- AlicenseNot gradedqualityDmaintenanceAdds text-to-speech capabilities to Cursor IDE, allowing your AI assistant to speak responses, summaries, and explanations out loud using OpenAI or ElevenLabs.1Apache 2.0