Skip to main content
Glama

stts

stts is a voice window for coding agents: an MCP server (stts-mcp, tools stt and tts) talks to a small local daemon on 127.0.0.1:15986, which drives a Chrome app window that listens with on-device speech recognition and speaks with Piper. It ships as the Claude Code plugin stts from the stts-marketplace and also runs over stdio behind an MCP gateway. Data lives in %LOCALAPPDATA%\cc-gc-stts\. MIT licence, see LICENSE.

Install

  1. /plugin marketplace add markkennethbadilla/stts

  2. /plugin install stts@stts-marketplace

  3. Run /stts to start a voice conversation.

Gateway (MCPJungle, every agent): register a stdio server with command node <repo>/dist/mcp.js (after npm run build) and env STTS_WHO=agent, so a helper agent cannot close the session's window. The entry lives in mkb-agentops (spec 010).

For development: npm ci, then npm run check and npm run build. npm run e2e builds, starts the daemon on STTS_TEST_PORT (default 15990) with its own data dir under test-results/, and drives the page in headless Chromium with a fake speech recogniser, fake media, page.clock and a fake Piper server. Locally it uses the shared Playwright headless shell from mkb-agentops/versions.json.

CI (.github/workflows/ci.yml): one job on ubuntu-latest, pull requests and pushes to main that touch code only, cancel-in-progress, 15-minute cap. A run takes about 3 minutes; at about 40 runs a month that is about 120 of the free 2,000 minutes.

Related MCP server: VoxMesh

Specs

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.
    16
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Adds voice conversation capabilities to AI agents via MCP, enabling local speech recognition and synthesis with tools like speak, listen, and ask_by_voice for interactive voice interactions.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to access local voice synthesis, zero-shot cloning, and voice catalog tools through native MCP tool calls for applications like Claude and Cursor.
    174 npm
    AGPL 3.0