stts-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@stts-mcplisten to me, then read the README aloud"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
stts
stts is a voice window for coding agents: an MCP server (stts-mcp, tools stt and tts) talks to a small local daemon on 127.0.0.1:15986, which drives a Chrome app window that listens with on-device speech recognition and speaks with Piper. It ships as the Claude Code plugin stts from the stts-marketplace and also runs over stdio behind an MCP gateway. Data lives in %LOCALAPPDATA%\cc-gc-stts\. MIT licence, see LICENSE.
Install
/plugin marketplace add markkennethbadilla/stts/plugin install stts@stts-marketplaceRun
/sttsto start a voice conversation.
Gateway (MCPJungle, every agent): register a stdio server with command node <repo>/dist/mcp.js (after npm run build) and env STTS_WHO=agent, so a helper agent cannot close the session's window. The entry lives in mkb-agentops (spec 010).
For development: npm ci, then npm run check and npm run build. npm run e2e builds, starts the daemon on STTS_TEST_PORT (default 15990) with its own data dir under test-results/, and drives the page in headless Chromium with a fake speech recogniser, fake media, page.clock and a fake Piper server. Locally it uses the shared Playwright headless shell from mkb-agentops/versions.json.
CI (.github/workflows/ci.yml): one job on ubuntu-latest, pull requests and pushes to main that touch code only, cancel-in-progress, 15-minute cap. A run takes about 3 minutes; at about 40 runs a month that is about 120 of the free 2,000 minutes.
Related MCP server: VoxMesh
Specs
This server cannot be deployed
Maintenance
Related MCP Connectors
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
The Listenetic MCP server is a remote, cloud-hosted server that enables AI assistants like ChatGPT and Claude to convert articles, documents, websites, and videos into high-quality AI-generated audio. It provides multi-format support for text and binary files, natural-sounding text-to-audio conversion using AI, and specialized processing for SSML, markup, markdown, and various media formats through three core tools: listentic_supported_mimetypes, listentic_add_content_text, and listentic_add_content_binary.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.16Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables voice-first interactions with AI agents and MCP tools, supporting speech input/output, STT/TTS, and a provider-independent agent core.1MIT
- AlicenseNot gradedqualityAmaintenanceAdds voice conversation capabilities to AI agents via MCP, enabling local speech recognition and synthesis with tools like speak, listen, and ask_by_voice for interactive voice interactions.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to access local voice synthesis, zero-shot cloning, and voice catalog tools through native MCP tool calls for applications like Claude and Cursor.174 npmAGPL 3.0