livechat-mcp
livechat-mcp
AIコーディングアシスタントと継続的な音声会話ができるModel Context Protocol (MCP) サーバーです。あなたが話すと、その音声がWhisperでローカルに文字起こしされ、各発話がまるでタイピングしたかのようにアシスタントに送信されます。タブの切り替え、コピー&ペースト、バッチ録音は不要です。
あらゆるMCPホストで動作します。以下のホストを公式サポートしています:
Claude Code
Codex CLI
Gemini CLI
要件
macOS、Linux、またはWindows (PowerShell経由、またはWSL2 / Git Bash環境)。
Python 3.10以上
インストール済みのMCPホスト (Claude Code, Codex, Geminiなど)
動作するマイク
Whisperモデルキャッシュと依存関係用に約500MBのディスク容量
プロジェクト管理用に
uv(推奨)
Related MCP server: Voice MCP
クイックインストール (推奨)
リポジトリをクローンした後:
# macOS / Linux / Git Bash on Windows
./install.sh# Native Windows PowerShell
.\install.ps1ブートストラップがOSを検出し、必要に応じてportaudioをインストールし (brew / apt / dnf / pacman / zypper — Windowsのwheelには同梱されています)、uvがなければインストールし、uv syncを実行し、ウィザードを ~/.local/bin に配置して、対話型セットアップウィザードを起動します。
Windows: ネイティブのロックには
msvcrtを使用し、テイクオーバーのシグナル伝達はファイルベースで行われるため、fcntlへの依存はありません。対話型ウィザードはbashスクリプトです —install.ps1はGit Bashを通じてこれを呼び出します。Git Bashがない場合はwingetでインストールを提案します。
手動セットアップ
ステップバイステップでインストールしたい場合は、install.sh が行う内容は以下の通りです:
1. portaudioのインストール
sounddevice には portaudio が必要です。
macOS:
brew install portaudioDebian/Ubuntu:
sudo apt-get install libportaudio2 portaudio19-devFedora/RHEL:
sudo dnf install portaudio portaudio-develArch:
sudo pacman -S portaudio
2. uv がない場合はインストール
curl -LsSf https://astral.sh/uv/install.sh | sh3. クローンと依存関係のインストール
cd livechat-mcp
uv syncこれにより .venv/ が作成され、mcp、faster-whisper、sounddevice、silero-vad、torch などがインストールされます。
4. セットアップウィザードの実行
install -m 0755 bin/livechat-mcp ~/.local/bin/livechat-mcp
livechat-mcp setupウィザードは以下のことを行います:
どのインストール先のアシスタント(Claude Code / Codex / Gemini、またはその組み合わせ)かを確認します。
カスタムスラッシュコマンドをサポートするホストに
/livechatおよび/endlivechatスラッシュコマンドをコピーします。Codexの場合、現在のCodex CLIリリースではカスタムプロンプトが/livechatとして公開されていないため、レガシープロンプトファイルとlivechatスキルの両方をインストールします。各ホストの設定ファイルにMCPサーバーを登録します。
調整可能な環境変数(無音しきい値、Whisperモデルなど)を案内します — デフォルトのままにする場合はEnterを押してください。
~/.local/bin が PATH に含まれていることを確認してください(公式の uv インストーラーを使用した場合は既に含まれています)。
手動で設定したい場合は、各ホストの手順を以下に示します。
5. マイク権限の付与
macOS: サーバーが初めてオーディオをキャプチャしようとすると、macOSはターミナルアプリ(Terminal, iTerm, Ghostty, Warpなど)に対してマイクへのアクセスを求めます。プロンプトを見逃した場合は、手動で有効にしてください:
システム設定 → プライバシーとセキュリティ → マイク → ターミナルアプリを有効にする
これをスキップすると、オーディオキャプチャは無音を返し、何も文字起こしされません。
Windows: 設定 → プライバシーとセキュリティ → マイク → デスクトップアプリによるマイクへのアクセスを許可する(ターミナルに許可が与えられていることを確認してください)。
Linux: 通常プロンプトは表示されません — ユーザーが適切なALSA / PulseAudio / Pipewireアクセス権(通常は
audioグループ)を持っていることを確認してください。
6. Whisperモデルの事前ダウンロード (オプション)
初回実行時に base.en (約150MB) がダウンロードされます。事前にダウンロードしておくことも可能です:
uv run python -c "from faster_whisper import WhisperModel; WhisperModel('base.en', device='cpu', compute_type='int8')"手動インストール (livechat-mcp setup を使用した場合はスキップ)
Claude Code
スラッシュコマンドをコピーします:
mkdir -p ~/.claude/commands
cp commands/livechat.md ~/.claude/commands/
cp commands/endlivechat.md ~/.claude/commands/MCPサーバーを登録します:
claude mcp add livechat -- uv --directory "$(pwd)" run livechat-mcpまたは ~/.claude.json を直接編集します:
{
"mcpServers": {
"livechat": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
}
}
}Codex CLI
Codexスキルとレガシープロンプトファイルをインストールします:
mkdir -p ~/.codex/skills/livechat
cp skills/livechat/SKILL.md ~/.codex/skills/livechat/
mkdir -p ~/.codex/prompts
cp commands/livechat.md ~/.codex/prompts/
cp commands/endlivechat.md ~/.codex/prompts/~/.codex/config.toml にMCPサーバーを登録します:
[mcp_servers.livechat]
command = "uv"
args = ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]Gemini CLI
GeminiはカスタムコマンドにTOMLを使用します。ウィザードがこれらを生成しますが、手動で行う場合は commands/gemini/livechat.toml.template を参照してください(livechat-mcp setup を一度実行すると作成されます)。
~/.gemini/settings.json にMCPサーバーを登録します:
{
"mcpServers": {
"livechat": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
}
}
}使用方法
任意のターミナルでアシスタントのCLIを開きます:
claude # or: codex or: gemini次に、アシスタントのプロンプトで以下を入力します:
/livechat # Claude Code, Gemini CLI
use livechat # Codex CLICodexの再起動が必要です。 Codexは起動時にのみスキルとMCPサーバーを読み込みます。Codexが開いている間にウィザードを実行した場合は、
use livechatを使用する前に終了して再起動してください。
Codex 0.128.0はユーザー定義の /livechat スラッシュコマンドをサポートしていません。/ は現在Codexの組み込みコマンド用に予約されています。セットアップでは代わりに検出可能な livechat スキルがインストールされるため、use livechat と入力するか、/skills を開いて livechat を選択できます。
アシスタントが get_voice_input を呼び出し、聞き取りを開始します。通常通り話してください。 約1.5秒間停止すると、発話が確定し、文字起こしされてプロンプトとして送信されます。アシスタントが応答し、その後すぐに次の発話を聞き取ります。
アシスタントが応答を生成している間もマイクはアクティブです。その間に話した内容はキューに入れられ、次の get_voice_input 呼び出し時にまとめて配信されます。
セッションの終了
3つの方法があります:
/endlivechat— 最もクリーンな方法で、アシスタントのプロンプトから実行します。(応答の途中である場合は、まず現在のターンを中断する必要があります。)ウェイクフレーズ —
terminate voice session nowと言います。文字起こしがシャットダウンをトリガーします。このフレーズは、実際のレビュー内容との衝突を避けるために意図的に不自然なものにしています。LIVECHAT_END_PHRASEで設定可能です。Ctrl+C — MCPサーバーを強制終了します。アシスタントは次の呼び出しでツールエラーを確認し、ループを停止します。
設定
すべての調整可能な項目は livechat_mcp/config.py にあり、環境変数で上書きできます:
変数 | デフォルト | 備考 |
|
| 英語のみ: |
|
| 言語コード ( |
|
|
|
|
|
|
|
| 発話を終了するための無音時間 |
|
| Silero VAD音声確率しきい値 |
|
| 最小発話長(咳などをフィルタリング) |
|
| 長すぎる発話を強制終了 |
|
|
|
|
| セッションを終了するための音声フレーズ |
| 未設定 |
|
設定を簡単に行うには livechat-mcp set KEY VALUE を使用します。これは見つかったすべてのホスト設定(Claude / Codex / Gemini)の env ブロックを編集します。
livechat-mcp show # print current env block(s)
livechat-mcp set LIVECHAT_SILENCE_SEC 1.5
livechat-mcp unset LIVECHAT_DEBUG変更後はアシスタントのCLIを再起動してください。MCPの環境変数はサーバー起動時に読み込まれます。
手動で行う場合は、各ホストの設定にあるlivechat MCPエントリの env フィールドを編集してください。Claude Codeの例:
{
"mcpServers": {
"livechat": {
"command": "uv",
"args": ["--directory", "/abs/path", "run", "livechat-mcp"],
"env": {
"LIVECHAT_WHISPER_MODEL": "small.en",
"LIVECHAT_DEBUG": "1"
}
}
}
}トラブルシューティング
話しても何も起こらない。
以下の順序で確認してください:ターミナルアプリのマイク権限、マイク入力レベル(システム設定 → サウンド)、LIVECHAT_DEBUG=1 を設定してstderrでVADイベントを確認、LIVECHAT_VAD_THRESHOLD を 0.3 に下げる。
文字起こしが不正確。
モデルをアップグレードしてください:LIVECHAT_WHISPER_MODEL=small.en または medium.en。medium.en はCPUでは顕著に遅くなりますが(それでもリアルタイムに近い)、技術用語にははるかに適しています。
発話がすぐに終了する / 終了するのが遅すぎる。
LIVECHAT_SILENCE_SEC を調整してください(または livechat-mcp set LIVECHAT_SILENCE_SEC 1.5 を実行)。1.0〜4.5が有効な範囲です。低くすると反応が速く感じられますが、思考中の休止で途切れるリスクがあります。
uv が見つからない。
uvをインストールする(推奨)か、MCP設定の command を、アクティブ化されたvenv内からの python -m livechat_mcp.server の直接呼び出しに変更してください。
サーバーは起動するが、アシスタントがツールを呼び出さない。
/livechat が呼び出されたことを確認してください。スラッシュコマンドがないと、アシスタントはループに入る指示を受け取りません。
サーバーログがアシスタントのUIにゴミとして表示される / プロトコルを壊す。
これは発生しないはずです。すべてのサーバーログはstderrに送信されます。もし表示される場合はバグを報告してください。file=sys.stderr を指定せずに print(...) 文を追加していないことを確認してください。
起動時に portaudio エラーが発生する。
インストールしてください:brew install portaudio。インストール済みでも失敗する場合は brew reinstall portaudio を試し、sounddeviceを再インストールしてください:uv sync --reinstall。
仕組み(要約)
[mic] → [Silero VAD] → [Whisper] → [queue] ← [get_voice_input tool] ← [Assistant]
↑________background thread, always running________↑オーディオパイプラインはMCPツールから切り離されているため、サーバーが起動している間はマイクが常にアクティブです。アシスタントが応答を生成している間に話された発話はキューに入れられ、次のツール呼び出し時に配信されます。
ライセンス
MIT。
Available Tools
4 toolsend_voice_sessionA
Cleanly end the current voice session. After calling this, any further get_voice_input calls will return 'END_SESSION'. Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the key side effect: subsequent get_voice_input calls return '__END_SESSION__'. It doesn't mention idempotency or error conditions, but for a zero-parameter tool this is sufficient disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main action, then the effect and trigger. No wasted words, perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description covers purpose, usage, and behavioral consequence. An agent has everything needed to decide when to call and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description correctly adds no parameter info, and none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Cleanly end the current voice session.' It uses a specific verb (end) and resource (voice session), and the mention of the sentinel return on get_voice_input distinguishes it from siblings like take_over_voice_session or reset_voice_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit usage guidance: 'Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.' This clearly tells the agent when to invoke this tool versus the alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_voice_inputA
Returns the next voice utterance from the user as text. Used in a loop during voice review sessions. If multiple utterances are queued, they are joined with ' / '. Returns the literal string 'END_SESSION' when the user has ended the session (Ctrl+C, /endlivechat, or wake phrase) — stop calling this tool when you see that. Returns 'NO_INPUT' if the long-poll timed out with no speech; in that case, call this tool again. Returns 'ALREADY_RUNNING:' if another livechat MCP process (e.g. another Claude Code window) currently holds the session lock — ask the user to confirm a takeover, then call take_over_voice_session if they agree.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the behavioral disclosure burden. It reveals long-poll timeout behavior, queue joining with ' / ', sentinel values, the session lock with PID, and the exact follow-up action for each sentinel. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than most, but every sentence earns its place: purpose, loop context, queue joining, and each sentinel case. It is well-structured, front-loaded with the core purpose, and groups related behavioral details logically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description documents every possible return value and the appropriate agent reaction, including edge cases like concurrent process lock and takeover. It also names the relevant sibling tool. No meaningful gaps remain for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters)Skip, so the schema is trivially complete (100% coverage). The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Returns the next voice utterance from the user as text.' It clearly frames the tool as a polling operation used in a loop during voice review sessionsasha. The sentinel behaviors and the mention of take_over_voice_session distinguish it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to call the tool (in a loop during voice review) and how to react to each possible return value: retry on __NO_INPUT__, stop on __END_SESSION__, and ask the user then call take_over_voice_session on __ALREADY_RUNNING__. This is clear, actionable usage guidance with no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_voice_sessionA
Clear stale shutdown state from a previous /endlivechat in this same MCP server process so a new voice session can start cleanly. Call this once at the very beginning of a /livechat session, after the announcement and before the first get_voice_input. Safe to call mid-session: if a session is already running healthily this is a no-op and no in-flight utterances are dropped.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses that it clears stale state, is a no-op on a healthy session, and explicitly guarantees 'no in-flight utterances are dropped'. These behavioral details are specific and transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose and state, exact placement in the session flow, and safety guarantee. No redundancy, front-loaded with the primary action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter utility tool with no output schema, the description covers purpose, timing, safety, and idempotency. It provides all necessary context for an agent to decide when and how to call it, and it clearly differentiates from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is empty and description coverage is trivially 100%. The baseline for zero-parameter tools is 4; the description appropriately says nothing about parameters since none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Clear'), a resource ('stale shutdown state'), and the purpose ('so a new voice session can start cleanly'). It clearly differentiates from siblings by placing the call before get_voice_input, making its role obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to call: 'at the very beginning of a /livechat session, after the announcement and before the first get_voice_input'. It also notes it is safe mid-session and is a no-op if the session is healthy, giving clear usage conditions without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_over_voice_sessionA
Forcibly take the cross-process session lock from another livechat MCP instance. Signals the holder to release, waits briefly, and starts a new session here. Only call this after the user explicitly confirms taking over from the other window. Returns 'OK' on success or an error string on failure.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral disclosure burden. It fully discloses the mechanism: the tool signals the current holder to release, waits briefly, starts a new session, and returns 'OK' or an error string. This gives the agent enough detail to anticipate side effects and outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tight sentences with no filler. The first sentence names the action and target, the second explains the process, and the third adds the safety condition and expected return value. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description is complete: it explains what happens, when to call it, and what the response will be. Nothing essential is missing for an agent to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the baseline is 4. The description correctly adds no parameter details because none exist, and it still clarifies the operational scope and return behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('take') and resource ('cross-process session lock from another livechat MCP instance'), making the action unmistakable. It also differentiates this tool from siblings like get_voice_input or end_voice_session by focusing on the takeover of a lock rather than reading or ending a session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool is appropriate: 'Only call this after the user explicitly confirms taking over from the other window.' This is a clear, actionable condition that prevents premature or accidental invocation, and it implies the alternative context (continue using the current instance) without needing to name a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
end_voice_session - First observed
get_voice_input - First observed
reset_voice_session - First observed
take_over_voice_session
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: retrieving input, ending the session, taking over a lock, and resetting state. No overlap or ambiguity between them.
All tool names follow a consistent verb_noun snake_case pattern (e.g., get_voice_input, end_voice_session). The pattern is uniform and predictable.
Four tools is well-scoped for a voice session lifecycle. Each tool serves a necessary function with no redundancy, fitting the server's narrow purpose.
The tool surface covers the full voice session lifecycle: starting clean, retrieving input, ending, and handling cross-process takeover. No obvious dead ends or missing operations for the intended use case.
Maintenance
Related MCP Connectors
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
Command your AI agents by voice: PTT rooms, channels, direct messages, agent email, memory (mRAG).
- mcpOAuthcom.attendmeet
Bring meeting decisions, tasks and user stories into your AI editor (Claude, Cursor, Copilot).
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables coding agents to speak aloud using text-to-speech functionality. Works with agents running inside devcontainers and provides configurable voice settings for creating chatty AI companions.6-
- FlicenseNot gradedqualityDmaintenanceEnables voice interaction with Claude Code through local speech-to-text (Whisper) and text-to-speech (Supertonic), allowing verbal input/output without external API calls.1-
- FlicenseNot gradedqualityCmaintenanceEnables natural voice interaction with Claude Code through speech-to-text, supporting wake word activation and multiple backends like Whisper and Google. It allows users to execute commands and control their coding environment hands-free via their microphone.2-
- AlicenseNot gradedqualityDmaintenanceEnables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.16Apache 2.0