Skip to main content
Glama

ThotStream Lite 🎙️

License: MIT Python: 3.8+ Multi-OS: Linux | macOS | Windows Dependencies: Zero

Minimalist, zero-dependency, cross-compatible AI-to-audio speech engine and MCP server.
Stream internal thoughts, tool actions, and responses in real-time with 1 local voice.


⚡ Core Highlights

  • Zero Mandatory Dependencies: Runs entirely on Python 3.8+ standard library (subprocess, threading, queue, shutil).

  • Verified Cross-Platform:

    • macOS: Built-in say CLI and afplay CoreAudio playback (0 MB install).

    • Linux: Intelligent auto-cascade across spd-say (speech-dispatcher), espeak-ng, espeak, plus ALSA (aplay), PulseAudio (paplay), and PipeWire (pw-play).

    • Windows: Built-in System.Speech via PowerShell (0 MB install).

    • Any OS (Neural Upgrade): Automatically detects local Piper TTS ONNX models for neural voice output.

  • Universal Multi-Adapter:

    • MCP Server: Stdio JSON-RPC 2.0 compliant with Claude Desktop, Antigravity IDE, Cursor, and Continue.

    • CLI Pipe & Wrapper: Transparent stdout interception (thotstream-wrap) for zero token overhead in terminal harnesses (Freebuff, Gemini CLI, Claude Code CLI).

    • Skill Card: Standard SKILL.md instruction specification for prompt-driven agents.

  • Non-Blocking Threaded Architecture: Speech synthesis and audio playback occur on an asynchronous worker thread, ensuring LLM text generation is never blocked.


Related MCP server: mcp-ai-voice

🚀 Quickstart (60 Seconds)

1. Install (Editable / Zero Dependencies)

cd packages/thotstream-lite
pip install -e .

2. Verify Your System Audio Driver

thotstream-lite --status

Example outputs:

  • macOS: Audio Driver: say

  • Linux: Audio Driver: spd-say (or espeak-ng)

  • Windows: Audio Driver: sapi5

  • With Piper: Audio Driver: piper


🔌 Integration Modes

Mode A: Claude Desktop & Antigravity IDE (MCP)

Add to claude_desktop_config.json or .gemini/settings.json:

{
  "mcpServers": {
    "thotstream": {
      "command": "python",
      "args": ["-m", "thotstream_lite.mcp_server"]
    }
  }
}

The agent receives three dedicated audio tools:

  • speak_thought(text): Narrate internal reasoning or hypotheses.

  • speak_action(text): Announce tool execution intent before running.

  • speak_response(text): Speak the final answer aloud.

See docs/INTEGRATIONS.md for full setup screenshots.


Mode B: CLI Streaming Interception (Freebuff / Terminal CLIs)

Zero token overhead. The LLM generates text normally with XML tags; ThotStream Lite intercepts and speaks them out-of-band:

# 1. Pipe streaming stdout from any agent
freebuff --mode agent | thotstream-lite --listen

# 2. Or wrap the CLI command directly
thotstream-lite freebuff --mode agent

Mode C: Python SDK

from thotstream_lite import ThotStreamLite

engine = ThotStreamLite()

# Asynchronous, non-blocking queue calls
engine.speak_thought("Formulating system architecture hypothesis.")
engine.speak_action("Querying vector database for matching nodes.")
engine.speak_response("Operation completed successfully.")

# Drain and shutdown
engine.drain()
engine.shutdown()

📂 Documentation & Examples


THOTSTREAM LITE IS DISTRIBUTED UNDER THE MIT LICENSE ON AN "AS IS" AND "AS AVAILABLE" BASIS, WITHOUT WARRANTIES OR GUARANTEES OF ANY KIND.

  • User Responsibility: The user/operator assumes 100% legal and operational responsibility for any text ingested, commands executed, and audio synthesized through this software.

  • Voice Rights & Regulations: Maintainers do not bundle or license third-party voice models. Users are solely responsible for compliance with voice likeness, copyright, and AI synthesis regulations.

Read the full legal notice in DISCLAIMER.md and LICENSE.

Related MCP Connectors

Related MCP Servers