Skip to main content
Glama

cursor-tts MCP

Progress-triggered TTS for Cursor Agent (speak / stop). Voice follows the system UI language (edge-tts; Windows-first).

Marketplace / Plugin install

Requires uv (uvx).

Way

Steps

Cursor Marketplace

After listing: search cursor-tts in Cursor

Local Plugin

Clone repo → Cursor: add local plugin / copy into ~/.cursor/plugins/local/cursor-tts

Manual MCP

Point MCP at uvx --from <this-repo> cursor-tts-mcp (see root mcp.json)

Plugin layout:

  • .cursor-plugin/plugin.json

  • mcp.json (uvx --from ${PLUGIN_ROOT})

  • rules/cursor-tts-speak.mdc

  • assets/logo.svg

Submit checklist: docs/guide/Marketplace提交清单.md · Form: https://cursor.com/marketplace/publish

Related MCP server: Kokoro TTS MCP Server

Docs

See docs/README.md.

Quick start (dev)

cd E:\WorkSpace\mcp
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
pytest

Default Marketplace mode is embedded (MCP process speaks). Optional P1: tts_service on 127.0.0.1:18765 + thin HTTP client.

# optional standalone service
python -m tts_service

Local Cursor wiring:

  • .cursor/mcp.json → MCP cursor-tts

  • .cursor/rules/cursor-tts-speak.mdc → speak-on-progress rule

If tools missing: reload MCP / restart Cursor. See docs/guide/Cursor接入与配置说明.md.

Layout

.cursor-plugin/       # Marketplace plugin manifest
mcp.json              # Portable MCP launch for Plugin
rules/                # Agent rule shipped with Plugin
src/cursor_tts_mcp/   # MCP + orchestrator + engines + playback
config/default.yaml   # defaults
docs/                 # design & contracts
tests/

Available Tools

2 tools
speakA

Speak a short status line for the user. Use for real progress, task conclusions, errors/blocks, and when the user must decide or act. Always call this when delivering a conclusion — do not only type. Do not read code or long text. Do not wait until the whole task finishes for the first call. Speak in the system UI language (en): use short English phrases (≤80 characters). Do not force Chinese if the system language is not Chinese.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
categoryNoprogress
priorityNonormal
interruptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to carry safety or behavioral signals, the description does most of the work itself. It discloses that the tool should not be used to read code or long text, that it should be called even when a conclusion is typed, and that spoken output must be short English (≤80 characters). It does not explain interrupt or priority behavior, but the core behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and usage, and each instruction serves a distinct function. The final 'Do not force Chinese' clause is somewhat redundant after the 'system UI language (en)' note, but it is a cheap and useful reminder.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple output tool, the description covers required inputs, category mapping, output style, language, and call timing. The main gap is the undocumented 'priority' and 'interrupt' parameters, but these are optional and defaulted, so the description remains largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does add meaning for 'text' (short, English, ≤80 characters) and 'category' (progress, errors, conclusions, decisions), corresponding to the enum values. However, 'priority' and 'interrupt' are never explained, leaving their semantics to names and defaults alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb-plus-resource statement ('Speak a short status line for the user') and then enumerates the exact event types it is for: progress, task conclusions, errors/blocks, and decisions. This makes the tool's role obvious and distinguishes it from the sibling tool 'stop'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to call it ('Use for real progress, task conclusions, errors/blocks, and when the user must decide or act') and when not to ('Do not read code or long text'). It also gives operational guidance about calling early instead of waiting for full-task completion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stopA

Immediately stop current TTS playback. Default clear_queue=true; if false, only stop current and keep the queue.

ParametersJSON Schema
NameRequiredDescriptionDefault
clear_queueNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the immediate stop action, the default clear_queue=true, and the conditional behavior when false ('only stop current and keep the queue'). This is meaningful behavioral detail beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action, and no filler. Every word adds value, covering the action and the parameter behavior efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, this is complete: it states the action, the parameter semantics, and the default. Since an output schema exists, return-value documentation is not required from the description, and nothing essential for calling the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully explains the single parameter clear_queue: its default value and the effect when false. This completely compensates for the schema's lack of semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Immediately stop current TTS playback.' This clearly distinguishes it from the sibling tool speak, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool, and the only sibling (speak) makes the use case obvious. While it does not explicitly name an alternative, the parameter guidance about clearing vs. keeping the queue adds practical usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedspeak
    • First observedstop

TDQS

A4.5/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: speak outputs audio, stop halts playback. No overlap or ambiguity exists between them.

Naming Consistency5/5

Both tool names are single-word imperative verbs (speak, stop) that directly describe their actions. The naming pattern is perfectly consistent and predictable.

Tool Count4/5

With only two tools, the set is minimal but appropriate for a narrow TTS status-line server. It feels slightly thin, but each tool earns its place.

Completeness5/5

For the stated purpose of speaking short status lines and stopping playback, the tools cover the full lifecycle with no obvious gaps. Stop even offers queue-clearing behavior.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables coding agents to speak aloud using text-to-speech functionality. Works with agents running inside devcontainers and provides configurable voice settings for creating chatty AI companions.
    4
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Adds text-to-speech capabilities to Cursor IDE, allowing your AI assistant to speak responses, summaries, and explanations out loud using OpenAI or ElevenLabs.
    1
    Apache 2.0