Skip to main content
Glama

livechat-mcp

一个模型上下文协议 (MCP) 服务器,让您可以与 AI 编程助手进行连续的语音对话。您说话,语音会在本地通过 Whisper 转录,每段话语都会像您输入一样发送给助手。无需切换标签页,无需复制/粘贴,无需批量录音。

适用于任何 MCP 主机。对以下工具提供一流支持:

  • Claude Code

  • Codex CLI

  • Gemini CLI

要求

  • macOS、Linux 或 Windows(通过 PowerShell 原生运行,或在 WSL2 / Git Bash 下运行)。

  • Python 3.10+

  • 已安装 MCP 主机(Claude Code、Codex、Gemini 等)

  • 可用的麦克风

  • ~500 MB 磁盘空间用于 Whisper 模型缓存 + 依赖项

  • uv 用于项目管理(推荐)

Related MCP server: Voice MCP

快速安装(推荐)

从仓库克隆后:

# macOS / Linux / Git Bash on Windows
./install.sh
# Native Windows PowerShell
.\install.ps1

引导程序会自动检测操作系统,在需要时安装 portaudio(brew / apt / dnf / pacman / zypper — Windows wheel 已捆绑),如果缺少 uv 则安装它,运行 uv sync,将向导放入 ~/.local/bin,并启动交互式设置向导。

Windows:原生锁定使用 msvcrt,接管信号基于文件,因此没有 fcntl 依赖。交互式向导是一个 bash 脚本 — install.ps1 通过 Git Bash 调用它,如果缺少 Git Bash,它会建议通过 winget 安装。

手动设置

如果您更喜欢一步步安装,以下是 install.sh 所做的工作:

1. 安装 portaudio

sounddevice 需要 portaudio。

  • macOS: brew install portaudio

  • Debian/Ubuntu: sudo apt-get install libportaudio2 portaudio19-dev

  • Fedora/RHEL: sudo dnf install portaudio portaudio-devel

  • Arch: sudo pacman -S portaudio

2. 如果没有 uv,请安装它

curl -LsSf https://astral.sh/uv/install.sh | sh

3. 克隆并安装依赖项

cd livechat-mcp
uv sync

这将创建 .venv/ 并安装 mcp、faster-whisper、sounddevice、silero-vad、torch 等。

4. 运行设置向导

install -m 0755 bin/livechat-mcp ~/.local/bin/livechat-mcp
livechat-mcp setup

向导将:

  1. 询问要为哪些助手安装(Claude Code / Codex / Gemini,任意组合)。

  2. 将 /livechat 和 /endlivechat 斜杠命令复制到支持自定义斜杠命令的主机。对于 Codex,它会同时安装旧版提示文件和 livechat 技能,因为当前的 Codex CLI 版本不会将自定义提示公开为 /livechat。

  3. 在每个主机的配置文件中注册 MCP 服务器。

  4. 指导您完成可调环境变量(静音阈值、Whisper 模型等)的设置 — 按回车键保留默认值。

确保 ~/.local/bin 在您的 PATH 中(如果您使用了官方的 uv 安装程序,它已经在其中了)。

如果您更喜欢手动连接,每个主机的具体步骤如下。

5. 授予麦克风权限

  • macOS: 服务器第一次尝试捕获音频时,macOS 会提示您的 终端应用(Terminal、iTerm、Ghostty、Warp 等)获取麦克风访问权限。如果您错过了提示,请手动启用:

    系统设置 → 隐私与安全性 → 麦克风 → 为您的终端启用

    如果您跳过此步骤,音频捕获将静默返回静音,且永远不会进行转录。

  • Windows: 设置 → 隐私与安全性 → 麦克风 → 允许桌面应用访问麦克风(并确保您的终端已获得许可)。

  • Linux: 通常没有提示 — 只需确保您的用户拥有正确的 ALSA / PulseAudio / Pipewire 访问权限(通常是 audio 组)。

6. 预下载 Whisper 模型(可选)

首次运行时会下载 base.en(约 150 MB)。您可以预先加载它:

uv run python -c "from faster_whisper import WhisperModel; WhisperModel('base.en', device='cpu', compute_type='int8')"

手动安装(如果您使用了 livechat-mcp setup,请跳过此部分)

Claude Code

复制斜杠命令:

mkdir -p ~/.claude/commands
cp commands/livechat.md ~/.claude/commands/
cp commands/endlivechat.md ~/.claude/commands/

注册 MCP 服务器:

claude mcp add livechat -- uv --directory "$(pwd)" run livechat-mcp

或者直接编辑 ~/.claude.json:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
    }
  }
}

Codex CLI

安装 Codex 技能和旧版提示文件:

mkdir -p ~/.codex/skills/livechat
cp skills/livechat/SKILL.md ~/.codex/skills/livechat/
mkdir -p ~/.codex/prompts
cp commands/livechat.md ~/.codex/prompts/
cp commands/endlivechat.md ~/.codex/prompts/

在 ~/.codex/config.toml 中注册 MCP 服务器:

[mcp_servers.livechat]
command = "uv"
args = ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]

Gemini CLI

Gemini 使用 TOML 进行自定义命令。向导会为您生成这些文件;若要手动操作,请参阅 commands/gemini/livechat.toml.template(通过运行 livechat-mcp setup 一次创建)。

在 ~/.gemini/settings.json 中注册 MCP 服务器:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/livechat-mcp", "run", "livechat-mcp"]
    }
  }
}

使用方法

在任何终端中打开助手的 CLI:

claude    # or: codex    or: gemini

然后在助手提示符中:

/livechat            # Claude Code, Gemini CLI
use livechat         # Codex CLI

需要重启 Codex。 Codex 仅在启动时加载技能和 MCP 服务器。如果您在 Codex 打开时运行了向导,请在执行 use livechat 之前退出并重新启动。

Codex 0.128.0 不支持用户定义的 /livechat 斜杠命令;/ 目前保留给 Codex 的内置命令。设置程序会安装一个可发现的 livechat 技能,因此您可以输入 use livechat 或打开 /skills 并选择 livechat。

助手将调用 get_voice_input 并开始监听。正常说话即可。 当您暂停约 1.5 秒时,您的语音会被定稿、转录并作为提示发送。助手会做出响应,然后立即监听下一段话语。

当助手正在生成响应时,麦克风仍然处于开启状态 — 在此期间您说的任何内容都会排队,并在下一次 get_voice_input 调用时一次性发送。

结束会话

有三种方法:

  1. /endlivechat — 最干净的方法,从助手提示符运行。(如果正在响应中,您需要先中断当前轮次。)

  2. 唤醒短语 — 说 terminate voice session now。转录会触发关闭。该短语特意设置得比较生硬,以避免与实际评论内容冲突。可通过 LIVECHAT_END_PHRASE 配置。

  3. Ctrl+C — 终止 MCP 服务器。助手将在下一次调用时看到工具错误并停止循环。

配置

所有可调参数位于 livechat_mcp/config.py 中,可以通过环境变量覆盖:

变量

默认值

说明

LIVECHAT_WHISPER_MODEL

base.en

仅限英语:tiny.en, base.en, small.en, medium.en。多语言(去掉 .en):tiny, base, small, medium

LIVECHAT_WHISPER_LANGUAGE

en

语言代码(en, pt, es 等)或 auto 以按话语检测。auto 需要多语言模型

LIVECHAT_WHISPER_DEVICE

auto

cpu, cuda, auto

LIVECHAT_WHISPER_COMPUTE

int8

int8 (CPU), float16 (GPU)

LIVECHAT_SILENCE_SEC

1.5

语音后的静音时间,用于结束一段话语

LIVECHAT_VAD_THRESHOLD

0.5

Silero VAD 语音概率阈值

LIVECHAT_MIN_UTTERANCE_SEC

0.4

最小话语长度(过滤咳嗽声)

LIVECHAT_MAX_UTTERANCE_SEC

120

强制截断过长的话语

LIVECHAT_LONG_POLL_SEC

300

get_voice_input 在返回 __NO_INPUT__ 前阻塞的时间

LIVECHAT_END_PHRASE

terminate voice session now

用于结束会话的口语短语

LIVECHAT_DEBUG

未设置

设置为 1 可将 VAD/分段调试日志输出到 stderr

设置这些参数的简单方法是 livechat-mcp set KEY VALUE — 它会编辑它找到的每个主机配置(Claude / Codex / Gemini)中的 env 块。

livechat-mcp show           # print current env block(s)
livechat-mcp set LIVECHAT_SILENCE_SEC 1.5
livechat-mcp unset LIVECHAT_DEBUG

更改后请重启您的助手 CLI — MCP 环境变量在服务器启动时读取。

若要手动操作,请编辑每个主机配置中 livechat MCP 条目的 env 字段。以 Claude Code 为例:

{
  "mcpServers": {
    "livechat": {
      "command": "uv",
      "args": ["--directory", "/abs/path", "run", "livechat-mcp"],
      "env": {
        "LIVECHAT_WHISPER_MODEL": "small.en",
        "LIVECHAT_DEBUG": "1"
      }
    }
  }
}

故障排除

说话时没有反应。 检查(按顺序):终端应用的麦克风权限、麦克风输入电平(系统设置 → 声音)、设置 LIVECHAT_DEBUG=1 并观察 stderr 中的 VAD 事件、将 LIVECHAT_VAD_THRESHOLD 降低到 0.3。

转录不准确。 升级模型:LIVECHAT_WHISPER_MODEL=small.en 或 medium.en。medium.en 在 CPU 上明显较慢(但仍接近实时),但对技术词汇的处理效果好得多。

话语结束得太快 / 太慢。 调整 LIVECHAT_SILENCE_SEC(或运行 livechat-mcp set LIVECHAT_SILENCE_SEC 1.5)。1.0–4.5 是有效范围 — 较低的值感觉更灵敏,但有中断思考停顿的风险。

找不到 uv。 要么安装 uv(推荐),要么将 MCP 配置中的 command 更改为在激活的 venv 中直接调用 python -m livechat_mcp.server。

服务器启动了,但助手从不调用该工具。 确保已调用 /livechat。如果没有斜杠命令,助手就没有进入循环的指令。

服务器日志作为乱码进入助手的 UI / 破坏了协议。 这不应该发生 — 所有服务器日志都发送到 stderr。如果您看到了,请提交 bug。确保您没有添加任何没有 file=sys.stderr 的 print(...) 语句。

启动时出现 portaudio 错误。 安装它:brew install portaudio。如果已安装但仍然失败,请尝试 brew reinstall portaudio 并重新安装 sounddevice:uv sync --reinstall。

工作原理(简述)

[mic] → [Silero VAD] → [Whisper] → [queue] ← [get_voice_input tool] ← [Assistant]
   ↑________background thread, always running________↑

音频管道与 MCP 工具解耦,因此只要服务器运行,麦克风就始终处于开启状态。在助手生成响应时所说的话语会被排队,并在下一次工具调用时发送。

Available Tools

4 tools
end_voice_sessionA

Cleanly end the current voice session. After calling this, any further get_voice_input calls will return 'END_SESSION'. Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the key side effect: subsequent get_voice_input calls return '__END_SESSION__'. It doesn't mention idempotency or error conditions, but for a zero-parameter tool this is sufficient disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the main action, then the effect and trigger. No wasted words, perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description covers purpose, usage, and behavioral consequence. An agent has everything needed to decide when to call and what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description correctly adds no parameter info, and none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Cleanly end the current voice session.' It uses a specific verb (end) and resource (voice session), and the mention of the sentinel return on get_voice_input distinguishes it from siblings like take_over_voice_session or reset_voice_session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit usage guidance: 'Use this when the user invokes /endlivechat or otherwise asks to stop voice mode.' This clearly tells the agent when to invoke this tool versus the alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_inputA

Returns the next voice utterance from the user as text. Used in a loop during voice review sessions. If multiple utterances are queued, they are joined with ' / '. Returns the literal string 'END_SESSION' when the user has ended the session (Ctrl+C, /endlivechat, or wake phrase) — stop calling this tool when you see that. Returns 'NO_INPUT' if the long-poll timed out with no speech; in that case, call this tool again. Returns 'ALREADY_RUNNING:' if another livechat MCP process (e.g. another Claude Code window) currently holds the session lock — ask the user to confirm a takeover, then call take_over_voice_session if they agree.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the behavioral disclosure burden. It reveals long-poll timeout behavior, queue joining with ' / ', sentinel values, the session lock with PID, and the exact follow-up action for each sentinel. This is exceptionally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than most, but every sentence earns its place: purpose, loop context, queue joining, and each sentinel case. It is well-structured, front-loaded with the core purpose, and groups related behavioral details logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description documents every possible return value and the appropriate agent reaction, including edge cases like concurrent process lock and takeover. It also names the relevant sibling tool. No meaningful gaps remain for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters)Skip, so the schema is trivially complete (100% coverage). The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Returns the next voice utterance from the user as text.' It clearly frames the tool as a polling operation used in a loop during voice review sessionsasha. The sentinel behaviors and the mention of take_over_voice_session distinguish it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to call the tool (in a loop during voice review) and how to react to each possible return value: retry on __NO_INPUT__, stop on __END_SESSION__, and ask the user then call take_over_voice_session on __ALREADY_RUNNING__. This is clear, actionable usage guidance with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reset_voice_sessionA

Clear stale shutdown state from a previous /endlivechat in this same MCP server process so a new voice session can start cleanly. Call this once at the very beginning of a /livechat session, after the announcement and before the first get_voice_input. Safe to call mid-session: if a session is already running healthily this is a no-op and no in-flight utterances are dropped.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It discloses that it clears stale state, is a no-op on a healthy session, and explicitly guarantees 'no in-flight utterances are dropped'. These behavioral details are specific and transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose and state, exact placement in the session flow, and safety guarantee. No redundancy, front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter utility tool with no output schema, the description covers purpose, timing, safety, and idempotency. It provides all necessary context for an agent to decide when and how to call it, and it clearly differentiates from sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is empty and description coverage is trivially 100%. The baseline for zero-parameter tools is 4; the description appropriately says nothing about parameters since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Clear'), a resource ('stale shutdown state'), and the purpose ('so a new voice session can start cleanly'). It clearly differentiates from siblings by placing the call before get_voice_input, making its role obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to call: 'at the very beginning of a /livechat session, after the announcement and before the first get_voice_input'. It also notes it is safe mid-session and is a no-op if the session is healthy, giving clear usage conditions without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

take_over_voice_sessionA

Forcibly take the cross-process session lock from another livechat MCP instance. Signals the holder to release, waits briefly, and starts a new session here. Only call this after the user explicitly confirms taking over from the other window. Returns 'OK' on success or an error string on failure.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral disclosure burden. It fully discloses the mechanism: the tool signals the current holder to release, waits briefly, starts a new session, and returns 'OK' or an error string. This gives the agent enough detail to anticipate side effects and outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences with no filler. The first sentence names the action and target, the second explains the process, and the third adds the safety condition and expected return value. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description is complete: it explains what happens, when to call it, and what the response will be. Nothing essential is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline is 4. The description correctly adds no parameter details because none exist, and it still clarifies the operational scope and return behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('take') and resource ('cross-process session lock from another livechat MCP instance'), making the action unmistakable. It also differentiates this tool from siblings like get_voice_input or end_voice_session by focusing on the takeover of a lock rather than reading or ending a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when the tool is appropriate: 'Only call this after the user explicitly confirms taking over from the other window.' This is a clear, actionable condition that prevents premature or accidental invocation, and it implies the alternative context (continue using the current instance) without needing to name a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedend_voice_session
    • First observedget_voice_input
    • First observedreset_voice_session
    • First observedtake_over_voice_session

TDQS

A4.9/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: retrieving input, ending the session, taking over a lock, and resetting state. No overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern (e.g., get_voice_input, end_voice_session). The pattern is uniform and predictable.

Tool Count5/5

Four tools is well-scoped for a voice session lifecycle. Each tool serves a necessary function with no redundancy, fitting the server's narrow purpose.

Completeness5/5

The tool surface covers the full voice session lifecycle: starting clean, retrieving input, ending, and handling cross-process takeover. No obvious dead ends or missing operations for the intended use case.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to speak aloud using text-to-speech functionality. Works with agents running inside devcontainers and provides configurable voice settings for creating chatty AI companions.
    6
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables natural voice interaction with Claude Code through speech-to-text, supporting wake word activation and multiple backends like Whisper and Google. It allows users to execute commands and control their coding environment hands-free via their microphone.
    2
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.
    16
    Apache 2.0