Skip to main content
Glama

livechat-mcp

Ein Model Context Protocol (MCP) Server, der es dir ermöglicht, eine kontinuierliche Sprachkonversation mit deinem KI-Coding-Assistenten zu führen. Du sprichst, deine Sprache wird lokal mit Whisper transkribiert und jede Äußerung wird an den Assistenten übermittelt, als hättest du sie getippt. Kein Tab-Wechsel, kein Kopieren/Einfügen, keine Batch-Aufnahme.

Funktioniert mit jedem MCP-Host. Erstklassige Unterstützung für:

  • Claude Code

  • Codex CLI

  • Gemini CLI

Anforderungen

  • macOS, Linux oder Windows (nativ via PowerShell oder unter WSL2 / Git Bash).

  • Python 3.10+

  • Ein installierter MCP-Host (Claude Code, Codex, Gemini, etc.)

  • Ein funktionierendes Mikrofon

  • ~500 MB Speicherplatz für den Whisper-Modell-Cache + Abhängigkeiten

  • uv für das Projektmanagement (empfohlen)

Related MCP server: Voice MCP

Schnellinstallation (empfohlen)

Von einem Klon des Repositories aus:

# macOS / Linux / Git Bash on Windows
./install.sh
# Native Windows PowerShell
.\install.ps1

Der Bootstrap erkennt das Betriebssystem, installiert bei Bedarf Portaudio (brew / apt / dnf / pacman / zypper — Windows-Wheels enthalten es bereits), installiert uv, falls es fehlt, führt uv sync aus, legt den Wizard in ~/.local/bin ab und startet den interaktiven Einrichtungsassistenten.

Windows: Die native Sperrung verwendet msvcrt und die Übernahmesignalisierung ist dateibasiert, daher gibt es keine fcntl-Abhängigkeit. Der interaktive Wizard ist ein Bash-Skript — install.ps1 ruft es über Git Bash auf, dessen Installation via winget angeboten wird, falls es fehlt.

Manuelle Einrichtung

Wenn du die Installation lieber Schritt für Schritt durchführen möchtest, hier ist, was install.sh tut:

1. Portaudio installieren

sounddevice benötigt Portaudio.

  • macOS: brew install portaudio

  • Debian/Ubuntu: sudo apt-get install libportaudio2 portaudio19-dev

  • Fedora/RHEL: sudo dnf install portaudio portaudio-devel

  • Arch: sudo pacman -S portaudio

2. uv installieren, falls nicht vorhanden

curl -LsSf https://astral.sh/uv/install.sh | sh

3. Klonen und Abhängigkeiten installieren

cd livechat-mcp
uv sync

Dies erstellt .venv/ und installiert mcp, faster-whisper, sounddevice, silero-vad, torch usw.

4. Einrichtungsassistent ausführen

install -m 0755 bin/livechat-mcp ~/.local/bin/livechat-mcp
livechat-mcp setup

Der Assistent wird:

  1. Fragen, für welche Assistenten die Installation erfolgen soll (Claude Code / Codex / Gemini, jede Kombination).

  2. Die Slash-Befehle /livechat und /endlivechat in Hosts kopieren, die benutzerdefinierte Slash-Befehle unterstützen. Für Codex installiert er sowohl Legacy-Prompt-Dateien als auch einen livechat-Skill, da aktuelle Codex CLI-Releases keine benutzerdefinierten Prompts als /livechat bereitstellen.

  3. Den MCP-Server in der Konfigurationsdatei jedes Hosts registrieren.

  4. Dich durch die anpassbaren Umgebungsvariablen führen (Stille-Schwellenwert, Whisper-Modell usw.) — drücke Enter, um die Standardwerte beizubehalten.

Stelle sicher, dass ~/.local/bin in deinem PATH enthalten ist (dies ist bereits der Fall, wenn du den offiziellen uv-Installer verwendet hast).

Wenn du die Dinge lieber manuell konfigurieren möchtest, findest du die manuellen Schritte für jeden Host unten.

5. Mikrofonberechtigung erteilen

  • macOS: Wenn der Server zum ersten Mal versucht, Audio aufzunehmen, wird macOS deine Terminal-App (Terminal, iTerm, Ghostty, Warp usw.) nach dem Mikrofonzugriff fragen. Wenn du die Aufforderung verpasst, aktiviere sie manuell:

    Systemeinstellungen → Datenschutz & Sicherheit → Mikrofon → für dein Terminal aktivieren

    Wenn du dies überspringst, liefert die Audioaufnahme lautlos Stille und nichts wird transkribiert.

  • Windows: Einstellungen → Datenschutz & Sicherheit → Mikrofon → Desktop-Apps den Zugriff auf das Mikrofon erlauben (und sicherstellen, dass dein Terminal berechtigt ist).

  • Linux: Normalerweise keine Aufforderung — stelle nur sicher, dass dein Benutzer den richtigen ALSA- / PulseAudio- / Pipewire-Zugriff hat (normalerweise die audio-Gruppe).

6. Whisper-Modell vorab herunterladen (optional)

Der erste Durchlauf lädt base.en (~150 MB) herunter. Du kannst es vorab laden:

uv run python -c "from faster_whisper import WhisperModel; WhisperModel('base.en', device='cpu', compute_type='int8')"

Manuelle Installation (überspringen, wenn du livechat-mcp setup verwendet hast)

Claude Code

Kopiere die Slash-Befehle:

mkdir -p ~/.claude/commands
cp commands/livechat.md ~/.claude/commands/
cp commands/endlivechat.md ~/.claude/commands/

Registriere den MCP-Server:

claude mcp add livechat -- uv --directory "$(pwd)" run livechat-mcp

Oder bearbeite ~/.claude.json direkt:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
    }
  }
}

Codex CLI

Installiere den Codex-Skill und die Legacy-Prompt-Dateien:

mkdir -p ~/.codex/skills/livechat
cp skills/livechat/SKILL.md ~/.codex/skills/livechat/
mkdir -p ~/.codex/prompts
cp commands/livechat.md ~/.codex/prompts/
cp commands/endlivechat.md ~/.codex/prompts/

Registriere den MCP-Server in ~/.codex/config.toml:

[mcp_servers.livechat]
command = "uv"
args = ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]

Gemini CLI

Gemini verwendet TOML für benutzerdefinierte Befehle. Der Wizard generiert diese für dich; um es manuell zu tun, siehe commands/gemini/livechat.toml.template (erstellt durch einmaliges Ausführen von livechat-mcp setup).

Registriere den MCP-Server in ~/.gemini/settings.json:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
    }
  }
}

Verwendung

Öffne das CLI deines Assistenten in einem beliebigen Terminal:

claude    # or: codex    or: gemini

Dann im Prompt des Assistenten:

/livechat            # Claude Code, Gemini CLI
use livechat         # Codex CLI

Codex-Neustart erforderlich. Codex lädt Skills und MCP-Server nur beim Start. Wenn du den Wizard bei geöffnetem Codex ausgeführt hast, beende ihn und starte ihn neu, bevor du use livechat verwendest.

Codex 0.128.0 unterstützt keine benutzerdefinierten /livechat Slash-Befehle; / ist derzeit für die eingebauten Befehle von Codex reserviert. Die Einrichtung installiert stattdessen einen auffindbaren livechat-Skill, sodass du use livechat eingeben oder /skills öffnen und livechat auswählen kannst.

Der Assistent ruft get_voice_input auf und beginnt zuzuhören. Sprich normal. Wenn du für ca. 1,5 Sekunden pausierst, wird deine Äußerung abgeschlossen, transkribiert und als Prompt gesendet. Der Assistent antwortet und hört dann sofort auf die nächste Äußerung.

Während der Assistent eine Antwort generiert, ist das Mikrofon weiterhin aktiv — alles, was du während dieser Zeit sagst, wird in die Warteschlange gestellt und beim nächsten get_voice_input-Aufruf auf einmal übermittelt.

Eine Sitzung beenden

Drei Möglichkeiten:

  1. /endlivechat — am saubersten, wird vom Prompt des Assistenten aus ausgeführt. (Du musst zuerst die aktuelle Antwort unterbrechen, falls sie gerade generiert wird.)

  2. Wake-Phrase — sage terminate voice session now. Die Transkription löst das Herunterfahren aus. Die Phrase ist absichtlich umständlich gewählt, um Kollisionen mit echtem Review-Inhalt zu vermeiden. Konfigurierbar über LIVECHAT_END_PHRASE.

  3. Strg+C — beendet den MCP-Server. Der Assistent sieht beim nächsten Aufruf einen Tool-Fehler und beendet die Schleife.

Konfiguration

Alle anpassbaren Werte befinden sich in livechat_mcp/config.py und können über Umgebungsvariablen überschrieben werden:

Variable

Standardwert

Hinweise

LIVECHAT_WHISPER_MODEL

base.en

Nur Englisch: tiny.en, base.en, small.en, medium.en. Mehrsprachig (ohne .en): tiny, base, small, medium

LIVECHAT_WHISPER_LANGUAGE

en

Sprachcode (en, pt, es, …) oder auto zur Erkennung pro Äußerung. auto erfordert ein mehrsprachiges Modell

LIVECHAT_WHISPER_DEVICE

auto

cpu, cuda, auto

LIVECHAT_WHISPER_COMPUTE

int8

int8 (CPU), float16 (GPU)

LIVECHAT_SILENCE_SEC

1.5

Stille nach dem Sprechen, um eine Äußerung zu beenden

LIVECHAT_VAD_THRESHOLD

0.5

Silero VAD Sprachwahrscheinlichkeitsschwellenwert

LIVECHAT_MIN_UTTERANCE_SEC

0.4

Minimale Äußerungslänge (filtert Husten)

LIVECHAT_MAX_UTTERANCE_SEC

120

Erzwingt das Abschneiden zu langer Äußerungen

LIVECHAT_LONG_POLL_SEC

300

Wie lange get_voice_input blockiert, bevor __NO_INPUT__ erfolgt

LIVECHAT_END_PHRASE

terminate voice session now

Gesprochene Phrase zum Beenden der Sitzung

LIVECHAT_DEBUG

nicht gesetzt

Auf 1 setzen für VAD/Segmentierungs-Debug-Logs nach stderr

Der einfache Weg, diese zu setzen, ist livechat-mcp set KEY VALUE — dies bearbeitet den env-Block in jeder gefundenen Host-Konfiguration (Claude / Codex / Gemini).

livechat-mcp show           # print current env block(s)
livechat-mcp set LIVECHAT_SILENCE_SEC 1.5
livechat-mcp unset LIVECHAT_DEBUG

Starte dein Assistenten-CLI nach jeder Änderung neu — MCP-Umgebungsvariablen werden vom Server beim Start gelesen.

Um es manuell zu tun, bearbeite das env-Feld des Livechat-MCP-Eintrags in der Konfiguration jedes Hosts. Beispiel für Claude Code:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/abs/path", "run", "livechat-mcp"],
      "env": {
        "LIVECHAT_WHISPER_MODEL": "small.en",
        "LIVECHAT_DEBUG": "1"
      }
    }
  }
}

Fehlerbehebung

Nichts passiert, wenn ich spreche. Überprüfe (in dieser Reihenfolge): Mikrofonberechtigung für deine Terminal-App, Mikrofon-Eingangspegel (Systemeinstellungen → Ton), setze LIVECHAT_DEBUG=1 und beobachte stderr auf VAD-Ereignisse, senke LIVECHAT_VAD_THRESHOLD auf 0.3.

Transkriptionen sind ungenau. Modell aktualisieren: LIVECHAT_WHISPER_MODEL=small.en oder medium.en. medium.en ist auf der CPU spürbar langsamer (immer noch nahezu in Echtzeit), aber viel besser für technisches Vokabular.

Äußerung endet zu schnell / zu langsam. Optimiere LIVECHAT_SILENCE_SEC (oder führe livechat-mcp set LIVECHAT_SILENCE_SEC 1.5 aus). 1.0–4.5 ist der nützliche Bereich — niedriger fühlt sich reaktionsschneller an, riskiert aber das Abschneiden bei Pausen mitten im Gedanken.

uv nicht gefunden. Installiere entweder uv (empfohlen) oder ändere den MCP-Konfigurationsbefehl command in einen direkten Aufruf von python -m livechat_mcp.server aus einem aktivierten venv heraus.

Der Server startet, aber der Assistent ruft das Tool nie auf. Stelle sicher, dass /livechat aufgerufen wurde. Ohne den Slash-Befehl hat der Assistent keine Anweisung, in die Schleife einzutreten.

Server-Logs landen als Müll in der UI des Assistenten / unterbrechen das Protokoll. Dies sollte nicht passieren — alle Server-Logs gehen an stderr. Wenn du es siehst, melde einen Bug. Stelle sicher, dass du keine print(...)-Anweisungen ohne file=sys.stderr hinzugefügt hast.

portaudio-Fehler beim Start. Installiere es: brew install portaudio. Wenn es installiert ist und immer noch fehlschlägt, versuche brew reinstall portaudio und installiere sounddevice neu: uv sync --reinstall.

Funktionsweise (Kurzfassung)

[mic] → [Silero VAD] → [Whisper] → [queue] ← [get_voice_input tool] ← [Assistant]
   ↑________background thread, always running________↑

Die Audio-Pipeline ist vom MCP-Tool entkoppelt, sodass das Mikrofon immer aktiv ist, während der Server läuft. Äußerungen, die gesprochen werden, während der Assistent eine Antwort generiert, werden in die Warteschlange gestellt und beim nächsten Tool-Aufruf übermittelt.

Lizenz

MIT.

Available Tools

4 tools
end_voice_sessionA

Cleanly end the current voice session. After calling this, any further get_voice_input calls will return 'END_SESSION'. Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the key side effect: subsequent get_voice_input calls return '__END_SESSION__'. It doesn't mention idempotency or error conditions, but for a zero-parameter tool this is sufficient disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the main action, then the effect and trigger. No wasted words, perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description covers purpose, usage, and behavioral consequence. An agent has everything needed to decide when to call and what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description correctly adds no parameter info, and none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Cleanly end the current voice session.' It uses a specific verb (end) and resource (voice session), and the mention of the sentinel return on get_voice_input distinguishes it from siblings like take_over_voice_session or reset_voice_session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit usage guidance: 'Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.' This clearly tells the agent when to invoke this tool versus the alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_inputA

Returns the next voice utterance from the user as text. Used in a loop during voice review sessions. If multiple utterances are queued, they are joined with ' / '. Returns the literal string 'END_SESSION' when the user has ended the session (Ctrl+C, /endlivechat, or wake phrase) — stop calling this tool when you see that. Returns 'NO_INPUT' if the long-poll timed out with no speech; in that case, call this tool again. Returns 'ALREADY_RUNNING:' if another livechat MCP process (e.g. another Claude Code window) currently holds the session lock — ask the user to confirm a takeover, then call take_over_voice_session if they agree.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the behavioral disclosure burden. It reveals long-poll timeout behavior, queue joining with ' / ', sentinel values, the session lock with PID, and the exact follow-up action for each sentinel. This is exceptionally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than most, but every sentence earns its place: purpose, loop context, queue joining, and each sentinel case. It is well-structured, front-loaded with the core purpose, and groups related behavioral details logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description documents every possible return value and the appropriate agent reaction, including edge cases like concurrent process lock and takeover. It also names the relevant sibling tool. No meaningful gaps remain for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters)Skip, so the schema is trivially complete (100% coverage). The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Returns the next voice utterance from the user as text.' It clearly frames the tool as a polling operation used in a loop during voice review sessionsasha. The sentinel behaviors and the mention of take_over_voice_session distinguish it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to call the tool (in a loop during voice review) and how to react to each possible return value: retry on __NO_INPUT__, stop on __END_SESSION__, and ask the user then call take_over_voice_session on __ALREADY_RUNNING__. This is clear, actionable usage guidance with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reset_voice_sessionA

Clear stale shutdown state from a previous /endlivechat in this same MCP server process so a new voice session can start cleanly. Call this once at the very beginning of a /livechat session, after the announcement and before the first get_voice_input. Safe to call mid-session: if a session is already running healthily this is a no-op and no in-flight utterances are dropped.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It discloses that it clears stale state, is a no-op on a healthy session, and explicitly guarantees 'no in-flight utterances are dropped'. These behavioral details are specific and transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose and state, exact placement in the session flow, and safety guarantee. No redundancy, front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter utility tool with no output schema, the description covers purpose, timing, safety, and idempotency. It provides all necessary context for an agent to decide when and how to call it, and it clearly differentiates from sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is empty and description coverage is trivially 100%. The baseline for zero-parameter tools is 4; the description appropriately says nothing about parameters since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Clear'), a resource ('stale shutdown state'), and the purpose ('so a new voice session can start cleanly'). It clearly differentiates from siblings by placing the call before get_voice_input, making its role obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to call: 'at the very beginning of a /livechat session, after the announcement and before the first get_voice_input'. It also notes it is safe mid-session and is a no-op if the session is healthy, giving clear usage conditions without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

take_over_voice_sessionA

Forcibly take the cross-process session lock from another livechat MCP instance. Signals the holder to release, waits briefly, and starts a new session here. Only call this after the user explicitly confirms taking over from the other window. Returns 'OK' on success or an error string on failure.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral disclosure burden. It fully discloses the mechanism: the tool signals the current holder to release, waits briefly, starts a new session, and returns 'OK' or an error string. This gives the agent enough detail to anticipate side effects and outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences with no filler. The first sentence names the action and target, the second explains the process, and the third adds the safety condition and expected return value. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description is complete: it explains what happens, when to call it, and what the response will be. Nothing essential is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline is 4. The description correctly adds no parameter details because none exist, and it still clarifies the operational scope and return behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('take') and resource ('cross-process session lock from another livechat MCP instance'), making the action unmistakable. It also differentiates this tool from siblings like get_voice_input or end_voice_session by focusing on the takeover of a lock rather than reading or ending a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when the tool is appropriate: 'Only call this after the user explicitly confirms taking over from the other window.' This is a clear, actionable condition that prevents premature or accidental invocation, and it implies the alternative context (continue using the current instance) without needing to name a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedend_voice_session
    • First observedget_voice_input
    • First observedreset_voice_session
    • First observedtake_over_voice_session

TDQS

A4.9/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: retrieving input, ending the session, taking over a lock, and resetting state. No overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern (e.g., get_voice_input, end_voice_session). The pattern is uniform and predictable.

Tool Count5/5

Four tools is well-scoped for a voice session lifecycle. Each tool serves a necessary function with no redundancy, fitting the server's narrow purpose.

Completeness5/5

The tool surface covers the full voice session lifecycle: starting clean, retrieving input, ending, and handling cross-process takeover. No obvious dead ends or missing operations for the intended use case.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to speak aloud using text-to-speech functionality. Works with agents running inside devcontainers and provides configurable voice settings for creating chatty AI companions.
    6
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables natural voice interaction with Claude Code through speech-to-text, supporting wake word activation and multiple backends like Whisper and Google. It allows users to execute commands and control their coding environment hands-free via their microphone.
    2
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.
    16
    Apache 2.0