Skip to main content
Glama

Claude says something through your Mac's speakers, listens to your answer the way a person would, and gets back what you said as text. Speech recognition runs on your Mac, so no audio leaves your computer.

Claude ──speak_and_listen("Tests pass. Open the PR?")──▶  🔊 "Tests pass. Open the PR?"
                                                         🎙 you: "yes, and tag Priya"
Claude ◀──────────────── "Yes, and tag Priya." ─────────  whisper.cpp on your Mac

It's built for Apple Silicon Macs (M1–M4). Linux works too, with SoX and espeak-ng installed.

🤖 Vibe-coded

This project was designed and written with Claude, in conversation. A human (me) steered it, tried it on a real Mac and checked the test suite, but most of the code was written by AI. It's a proof of concept: it works and has tests, but expect rough edges, and read the code before relying on it for anything important. Issues and pull requests are welcome.

It's an independent project, not made or endorsed by Anthropic. It works with any MCP client, including Claude Desktop, Claude Code and Cursor.

How it works

The server has two tools and two prompts:

What it does

speak_and_listen

Speaks a short message, listens for one conversational turn and returns the transcript.

voice_setup

Checks what's installed. After you agree, it installs only what's missing.

/mcp__voice-mcp__setup

A guided setup: it checks, asks you, installs, then runs a spoken test.

/mcp__voice-mcp__voice_mode

A hands-free session where Claude checks in by voice at natural points.

Listening works like a conversation. A soft chime plays when the mic opens. The server waits for you to start talking and hands back to Claude about a second after you stop. It adjusts to background noise, doesn't cut you off at pauses mid-sentence, and ignores coughs and clicks. If you say nothing for 15 seconds, Claude gets "no speech", which it is told never to treat as a yes.

Replies come back fast. whisper.cpp's server keeps the speech model loaded between turns, and the model starts loading while Claude is still talking. You don't wait for a model load on each reply. After 15 idle minutes the server shuts down to free memory. It is stopped automatically even if the MCP server crashes.

Claude sends text meant to be heard. The tool description gives Claude rules for writing short spoken sentences. If code, paths, links or markdown still get through, the server rewrites them before speaking and tells Claude, so the next message is cleaner. See below.

Related MCP server: Voice MCP

Install

You need Node.js 22 or newer. Whichever way you install, run setup once afterwards (see Then, below).

One click

From a marketplace

  • Claude Code plugin marketplace. This repo is its own marketplace:

    /plugin marketplace add jeet0007/mac-voice-mcp
    /plugin install mac-voice-mcp@mac-voice-mcp

    The plugin adds /mac-voice-mcp:setup and /mac-voice-mcp:talk, a skill that teaches Claude how to use voice well and fix common problems, and a hook that keeps a voice conversation in voice (see Voice mode). Each plugin version runs the matching npm release.

  • The official MCP Registry. It's listed as io.github.jeet0007/mac-voice-mcp. Apps and directories that read the registry pick it up from there. In VS Code, open the Extensions view (⇧⌘X), search @mcp mac-voice, and click Install. Smithery, Glama, PulseMCP and mcp.so copy the registry, so it shows up there too.

By hand

Claude Code

claude mcp add voice-mcp -s user -- npx -y mac-voice-mcp@latest

Claude Desktop. Add this to ~/Library/Application Support/Claude/claude_desktop_config.json, then quit (⌘Q) and reopen the app:

{
  "mcpServers": {
    "voice-mcp": {
      "command": "npx",
      "args": ["-y", "mac-voice-mcp@latest"]
    }
  }
}

@latest makes npx check for a new release each time the app starts. Without it, npx keeps running whichever version it cached first.

If you get spawn npx ENOENT, use the full path from which npx, e.g. "command": "/opt/homebrew/bin/npx".

Cursor. Add the same mcpServers block to ~/.cursor/mcp.json.

VS Code. Run MCP: Add Server from the Command Palette, or:

code --add-mcp '{"name":"voice-mcp","command":"npx","args":["-y","mac-voice-mcp@latest"]}'

Any other MCP client. Run npx -y mac-voice-mcp@latest as a stdio server.

Then: set up and allow the mic

Run setup once. Ask Claude to "set up voice". In Claude Code you can also run /mac-voice-mcp:setup (plugin) or /mcp__voice-mcp__setup (added by hand), or from a terminal run npx -y mac-voice-mcp@latest setup.

Setup checks what's already there before it changes anything:

Needed

Provided by

If it's missing

Voice

macOS say, using the most natural voice installed

Nothing to do. For a far better voice, add a free Premium one (see below).

Microphone capture

SoX (rec). ffmpeg works as a fallback, but setup recommends SoX

brew install sox

Speech-to-text

whisper.cpp (whisper-cli and whisper-server, Metal-accelerated)

brew install whisper-cpp

Speech model

base.en, ~140 MB

Downloaded once to ~/.cache/mac-voice-mcp/models/

  • Nothing is redone. Tools already on your PATH are used as they are. If the model is already somewhere on disk (a whisper.cpp checkout, Homebrew's share folder, another tool's cache, or anything Spotlight can find), it's symlinked, not downloaded again. brew install runs only for the missing formulae.

  • Nothing happens without your OK. Claude calls voice_setup to check first, shows you the checklist, and asks before calling it with install=true.

  • Slow installs don't time out. If brew install whisper-cpp takes a while, setup reports INSTALLING. The install carries on in the background, and the next check picks up the result.

Allow the microphone. The first time Claude listens, macOS asks whether Claude (or Cursor, or your terminal) can use the microphone. Click Allow.

Get a better voice (recommended). macOS includes free Premium voices that sound far more natural than the default. Open System Settings → Accessibility → Spoken Content → System Voice → Manage Voices…, and download one, for example English → Ava (Premium) or Zoe (Premium). The next voice turn uses it automatically. To choose a specific voice, or keep the system voice, see VOICE_MCP_VOICE under Configuration.

Allow voice turns without prompts

By default, Claude Code asks for approval every time Claude wants to speak, which breaks the flow of a conversation. To allow voice turns, add the tool to the permissions.allow list in ~/.claude/settings.json. Use the name that matches how you installed it:

{
  "permissions": {
    "allow": [
      "mcp__plugin_mac-voice-mcp_voice-mcp__speak_and_listen",
      "mcp__voice-mcp__speak_and_listen"
    ]
  }
}

The first name is for the plugin, the second for claude mcp add voice-mcp …. Leave voice_setup out, so installs still ask you first. While /mac-voice-mcp:talk runs, voice turns are already allowed.

Installing from a clone

One command does everything above: it builds the project, runs setup (asking before installing anything), adds voice-mcp to Claude Desktop (backing up your config first) and to Claude Code, and offers a spoken test. It's safe to re-run, because each step checks first and skips anything already done.

git clone https://github.com/jeet0007/mac-voice-mcp && bash mac-voice-mcp/install.sh

Upgrading

  • Claude Code plugin: run /plugin marketplace update mac-voice-mcp, then open /plugin, choose mac-voice-mcp under your installed plugins, and update it. Restart Claude Code. If there's no update option, uninstall and reinstall it.

  • Everything installed with mac-voice-mcp@latest (Claude Desktop, Cursor, VS Code, claude mcp add): restart the app. npx fetches the new release when the server starts.

  • Configs without @latest: change mac-voice-mcp to mac-voice-mcp@latest in the config, then restart the app. Otherwise npx keeps running the version it cached first.

Check which version you'd get with npx -y mac-voice-mcp@latest --version, and see what changed in the changelog. Your model, voice and settings carry over.

Using it

  • "Work on X and check in with me by voice when you need a decision." Claude works quietly and only speaks at decision points.

  • /mac-voice-mcp:talk fix the flaky login test (plugin), or /mcp__voice-mcp__voice_mode fix the flaky login test (added by hand). Claude reads its plan back to you, then checks in at each checkpoint. Say "stop voice mode" or "I'm back" to end it.

  • "Read me a 20-second summary of this PR and ask if I should approve it." Use this for one-off briefings.

  • Just talk after the chime. You don't need to hurry or fill silence. If you're still talking at the 30-second safety cap (listen_seconds), Claude is told your reply may be cut off and asks you to continue.

Voice mode

Once you answer out loud, you're in a voice conversation. Claude replies by voice, not in text, until one of these happens:

  • You type something. You're back at the keyboard.

  • You say you're done. Claude says a short goodbye without opening the mic (speak_and_listen with listen: false).

  • You don't answer. After 15 seconds of silence Claude asks once more. If you still don't answer, it pauses and summarizes on screen, and the mic stays off. Type anything, or run /mac-voice-mcp:talk, to pick up again.

In Claude Code, the plugin enforces this with a hook. If Claude tries to answer in text mid-conversation, the hook sends it back once to answer by voice. If it stops again, the hook lets it. To turn the hook off, add "VOICE_MCP_STAY_IN_VOICE": "0" to the env block in ~/.claude/settings.json. Other apps rely on the instructions alone.

Several sessions, one mic. Every mac-voice-mcp on your Mac takes turns: Claude Code windows, Claude Desktop and Cursor. While one is speaking or listening, the others wait for that turn to finish (shown as "waiting for another voice session"). They give up after 2 minutes with a message. If a session crashes, the next one takes the mic over.

Getting Claude to sound natural

Guidance reaches Claude through several channels, because each client shows different ones:

Channel

Who sees it

Tool description (rules plus a good and a bad example)

Every client

Server instructions (when to use voice, how to handle replies, setup)

Claude Code (it reads up to 2 KB)

The voice_mode and setup prompts

Claude Code (as slash commands), Claude Desktop, Cursor

Server-side rewrite plus a voice-mcp note back to Claude

Always on

The plugin's /mac-voice-mcp:talk, voice-help skill and stay-in-voice hook

Claude Code, with the plugin

The rules Claude is given:

- 1–3 short sentences, under ~40 words; lead with the outcome, then one question.
- Plain words only: no markdown, bullets, emoji, code, file paths, URLs, stack traces or tables.
- Describe code instead of reading it, say file names not paths, round numbers, spell out symbols.
- Ask one question at a time, answerable in a few words.
- Put the details (diffs, logs, links) in the on-screen reply, and say so out loud.

The safety net turns ## Results\n- \npm test` ✅ 42/42\n- see /Users/x/repo/src/index.ts:120` into "Results. npm test 42 of 42. see index.ts."

Claude Desktop ignores server instructions. To make the rules stick there, paste examples/CLAUDE.md into Settings → Profile → personal preferences or into a Project's instructions. For Claude Code, add it to CLAUDE.md. For Cursor, copy examples/voice-mcp.mdc to .cursor/rules/.

Configuration

Everything is optional. Set these in your client config's "env": { … } block, or with -e NAME=value in claude mcp add.

Speaking

Variable

Default

VOICE_MCP_VOICE

most natural installed

Unset: the best Premium or Enhanced voice installed for the language, else the system voice. Set a name, e.g. Ava (Premium), Daniel, Kanya (list them with say -v '?'), or default to always use the system voice.

VOICE_MCP_RATE

system rate

Words per minute, e.g. 200.

VOICE_MCP_MAX_SPEAK_WORDS

120

Longer text is cut at a sentence boundary ("the rest is on screen").

VOICE_MCP_CHIME

1

Set to 0 to turn off the mic open/close sounds.

VOICE_MCP_LOCK_WAIT_SECONDS

120

How long a turn waits while another session on this Mac is using the mic.

Listening

Variable

Default

VOICE_MCP_END_SILENCE_MS

1200

How long a pause ends your turn. Use 1800 if it cuts you off while you think, 800 for snappier replies.

VOICE_MCP_START_TIMEOUT_SECONDS

15

How long to wait for you to start talking.

VOICE_MCP_SPEECH_MARGIN_DB

12

How much louder than room noise counts as speech. Raise it in noisy rooms.

VOICE_MCP_MIN_SPEECH_DB

-48

The quietest level that ever counts as speech (dBFS).

VOICE_MCP_RECORDER

auto

sox or ffmpeg (ffmpeg is macOS only). auto uses SoX, falling back to ffmpeg.

VOICE_MCP_FFMPEG_DEVICE

:default

Which input ffmpeg records from. :default follows System Settings; :1 picks device 1 (list them with ffmpeg -f avfoundation -list_devices true -i "").

Speech-to-text

Variable

Default

VOICE_MCP_WHISPER_MODEL

base.en

Which model to use (see the table below).

VOICE_MCP_LANGUAGE

en for *.en models, otherwise auto

en, th, ja, de, …

VOICE_MCP_WHISPER_PROMPT

—

Words to bias toward: names, product terms, jargon.

VOICE_MCP_WHISPER_MODEL_PATH

—

Use this exact ggml-*.bin file.

VOICE_MCP_MODEL_SEARCH_PATHS

—

Extra folders to check for an existing model (:-separated).

VOICE_MCP_WHISPER_SERVER

1

Set to 0 to always use whisper-cli, with no warm server.

VOICE_MCP_SERVER_IDLE_MINUTES

15

How long the warm server stays up without use.

VOICE_MCP_THREADS

min(8, cores)

Number of whisper.cpp threads.

VOICE_MCP_CACHE_DIR

~/.cache/mac-voice-mcp

Where models are downloaded or symlinked.

VOICE_MCP_DEBUG

0

Verbose logs with per-turn timings, written to stderr.

Claude Code plugin hook. Set this in the env block of ~/.claude/settings.json, not in the server's config:

Variable

Default

VOICE_MCP_STAY_IN_VOICE

1

Set to 0 to stop the plugin's hook from sending Claude back to answer by voice.

Models (whisper.cpp names):

Model

Size

Good for

tiny.en

75 MB

Yes/no answers, the lowest latency

base.en

142 MB

Default. English conversation.

small.en

466 MB

Noticeably more accurate English

large-v3-turbo-q5_0

547 MB

Other languages, e.g. Thai with VOICE_MCP_LANGUAGE=th

Troubleshooting

Symptom

Fix

"voice-mcp is not set up yet"

Ask Claude to set up voice, or run npx -y mac-voice-mcp@latest setup.

"microphone returned pure digital silence"

macOS is blocking the mic for the host app. Go to System Settings → Privacy & Security → Microphone, enable Claude / Cursor / your terminal, then restart that app. If the message says it's recording with ffmpeg, the input device is the likelier cause: run brew install sox.

No permission prompt ever appears

Run tccutil reset Microphone <bundle id> and restart the app. Running test in Terminal only gives permission to Terminal, not to Claude Desktop.

It cuts me off while I'm thinking

Set VOICE_MCP_END_SILENCE_MS=1800 (or up to 2500).

It never stops listening

The room is too noisy for the defaults. Set VOICE_MCP_SPEECH_MARGIN_DB=18, or use a headset.

It hears its own voice

Use headphones, or turn the speaker volume down. It only listens after it finishes speaking, but echo can linger.

"Homebrew: not installed"

Install it from brew.sh. It needs your password, so it can't run from Claude. Then run setup again.

It garbles names or jargon

Set VOICE_MCP_WHISPER_PROMPT="Priya, Postgres, Kubernetes", or switch to small.en.

The voice sounds robotic

Download a Premium voice (see Get a better voice). It's used automatically.

It asks for approval every turn

Add the tool to Claude Code's allow list: see Allow voice turns without prompts.

"Another voice session on this Mac…"

Another Claude window, Claude Desktop or Cursor held the speaker and mic for over 2 minutes, which means one very long turn. End that conversation, then try again.

Claude keeps answering by voice after I'm done

Type anything, or say "stop voice mode". To switch the plugin's hook off entirely, see Voice mode.

Turns feel slow

Each result ends with a timing line, e.g. spoke 3.1 s · listened 4.0 s · transcribed 0.3 s. Ask Claude what it says. Transcribing should take well under a second.

Known limitations

  • You can't interrupt it. It finishes speaking, then listens. Barge-in would mean listening while the speakers play, which needs headphones or echo cancellation.

  • It's macOS-first. Linux works with SoX and espeak-ng. Windows is untested.

  • Turn-taking is based on loudness, not a speech model. It adapts to background noise, but very noisy rooms, music or TV can confuse it. A headset helps, and so do the listening settings above.

  • One conversation at a time. There's one speaker and one microphone, so turns from every session on the Mac are queued.

Privacy and safety

  • Audio stays on your machine. Recordings go to a temporary file that's deleted after each turn. The only network use is the one-time model download from Hugging Face.

  • The warm whisper server is local only. It listens on 127.0.0.1 on a random port, and stops when idle or when this server exits.

  • Setup can only install known packages. Its install list is fixed in the code (sox, whisper-cpp), so nothing Claude says can make it install anything else. It never uninstalls or modifies other software.

Security

  • Secrets: every push and pull request is scanned for leaked secrets with TruffleHog, and the whole history is scanned before the first push. GitHub secret scanning with push protection is also on.

  • Dependencies: Dependabot opens weekly update pull requests. CI fails on high-severity advisories (npm audit), and dependency review blocks pull requests that add vulnerable packages.

  • Code: CodeQL runs with the security-extended queries.

  • Releases: releases publish through npm trusted publishing, with no long-lived npm token and a signed provenance attestation for every version.

To report a vulnerability, see SECURITY.md.

Development

npm install
npm test               # build + unit tests, end-to-end tests over MCP with stub binaries, and release-metadata checks
npm run audit          # known-vulnerability and signature checks on dependencies
npm run setup          # check what's installed; offers to install what's missing
npm run test:voice     # one real speak → listen → transcribe turn
npm run inspect        # MCP Inspector

Module

Responsibility

config.ts

Environment settings, logging, the PATH fix-up for GUI apps

speech-text.ts

Rewriting screen text for speech, cleaning up transcripts (pure, unit-tested)

endpointer.ts

Turn-taking voice-activity detection (pure, unit-tested)

audio.ts

Text-to-speech, chimes, streaming mic capture

model.ts

Finding, symlinking or downloading the model

stt.ts

The warm whisper-server with orphan guard, and the whisper-cli fallback

setup.ts

Requirement checks and consent-based background installs

lock.ts

One voice turn at a time across every session on the Mac

voice.ts, server.ts, index.ts

The round trip, the MCP tools and prompts, and the CLI

skills/, hooks/

The Claude Code plugin's /mac-voice-mcp:talk and :setup commands, the voice-help skill, and the stay-in-voice hook

The package installs two commands: mac-voice-mcp (the one npx -y mac-voice-mcp runs), and voice-mcp.

Testing the plugin: don't start Claude Code inside this repo. In this folder, npx mac-voice-mcp@<this version> finds the checkout itself instead of downloading the package, can't run it, and the server fails with CONNECTION_CLOSED. Start claude in any other folder. To try unreleased plugin files (skills, hooks), swap the marketplace to your checkout: run /plugin marketplace remove mac-voice-mcp, then /plugin marketplace add /path/to/checkout, and install at user scope. The server still comes from npm, so test server changes with npm test and npm run test:voice.

Releasing

  • First release: bash publish.sh. It asks before each public step and uses your own GitHub and npm logins. It creates the GitHub repo, publishes to npm, and lists the server in the official MCP Registry, which Smithery, Glama, PulseMCP and mcp.so pick up from.

  • Later releases:

    1. Add a section to CHANGELOG.md for the new version, in Keep a Changelog format, dated today.

    2. Run npm version patch (bug fixes), npm version minor (new features) or npm version major (breaking changes), per semver. It updates package.json, package-lock.json, server.json and the plugin (including the npm version the plugin runs), commits, and tags v<version>.

    3. Run git push --follow-tags.

    The Publish workflow then runs the tests, which also check that every version and the changelog entry match. It publishes to npm with provenance, lists the release in the MCP Registry, and creates a GitHub Release from the changelog section. It needs no secrets: npm and the registry both use GitHub's OIDC identity, once you've set the package's Trusted Publisher on npmjs.com (publish.sh prints the steps). If a step fails, fix the cause and use Re-run failed jobs. Steps that already finished are skipped.

License

MIT

Available Tools

2 tools
speak_and_listenSpeak and listenA

Say something out loud to the user and hear their spoken reply. Speaks text_to_speak through the computer's speakers, then listens like a conversation turn — it waits for the user to start talking and stops when they finish — and returns an on-device transcript of what they said.

Write text_to_speak for the ear, not the screen:

  • 1–3 short sentences, under ~40 words; lead with the outcome, then one question.

  • Plain words only: no markdown, bullets, emoji, code, file paths, URLs, stack traces or tables.

  • Describe code instead of reading it ("I added a retry to the upload function"), say file names not paths ("in index.ts"), round numbers ("about two hundred ms"), spell out symbols.

  • Ask one question at a time, answerable in a few words ("Should I deploy it — yes or no?").

  • Put the details (diffs, logs, links) in your normal on-screen reply, and say so out loud.

  • Tool output (Bash stdout, file writes) is NOT the on-screen reply. Only your own assistant message text renders.

  • A spoken turn with no assistant message text shows the user nothing — never skip the on-screen reply.

Good: "The build passed and all tests are green. Want me to open the pull request?" Bad: "## Results\n- npm test ✅ 42/42\n- see /Users/x/repo/src/index.ts:120"

Once the user is talking with you by voice, keep the conversation in voice: answer each transcript with another speak_and_listen call (not a text reply) until they say stop or start typing. To end voice mode, or for a one-way announcement, pass listen: false: it speaks without opening the mic. A reply of "(No speech detected …)" means the user did not answer — never treat it as consent. Ask once more; if still nothing, stop and say on screen that voice mode is paused and they can type to carry on (not speak: the mic is off). The "voice-mcp timing" line at the end is diagnostics: ignore it unless the user asks why things feel slow. If it reports that voice-mcp is not set up, call voice_setup.

ParametersJSON Schema
NameRequiredDescriptionDefault
listenNoDefault true. false: only speak, without opening the microphone — for a one-way announcement, or to say goodbye when the user ends voice mode.
text_to_speakYesThe summary or message to read aloud to the user. Plain, conversational text.
listen_secondsNoUpper limit on how long to listen, in seconds (default 30, max 120). Listening already stops when the user finishes talking, so you rarely need this.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare only the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description goes well beyond that: it explains the blocking turn semantics, that listening self-terminates when the user finishes, the on-device transcript return, default 30s / max 120s listen window, and how a no-speech result should be interpreted. It also flags the 'voice-mcp timing' line as ignorable diagnostics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded in the first sentence and the remainder is organized as scannable rules with a good/bad example. It is longer than a 3-parameter tool strictly needs — the 'Tool output is NOT the on-screen reply' and timing-diagnostics notes drift toward general assistant behavior — but nearly every line changes how the tool is called, so it earns most of its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description covers the return value (spoken transcript, plus the '(No speech detected …)' case), the blocking/stopping behavior, the optional non-blocking mode, and the diagnostic footer. An agent has everything needed to call it correctly and to recover from the failure modes it will actually hit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters are documented in structured form and the baseline is 3. The description adds meaning the schema does not: how to author text_to_speak for the ear (40-word ceiling, no markdown/emoji/paths, spell out symbols, one question at a time) and concrete scenarios for listen: false. It doesn't restate listen_seconds beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete verb pair and resource: speaks text_to_speak through the speakers, then listens for the user's reply and returns an on-device transcript. The interaction shape (a conversational turn that ends when the user stops talking) is stated, so an agent can distinguish it from voice_setup, which is named explicitly at the end.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when/when-not coverage: keep replying via speak_and_listen while in voice mode until the user stops or types; pass listen: false to end voice mode or make a one-way announcement; call voice_setup if the diagnostics say it is not configured. It also prescribes exact behavior for the edge case of a '(No speech detected …)' reply — never treat as consent, ask once more, then stop.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voice_setupVoice setupA
Idempotent

Check whether this computer has everything speak_and_listen needs — text-to-speech, a microphone recorder (SoX), whisper.cpp speech-to-text and the speech model — and optionally install what's missing. Anything already installed is reused (an existing model elsewhere on disk is symlinked, not re-downloaded). Call it with install=false (the default) first and tell the user what is missing. Only call it with install=true after the user agrees: it runs brew install for the missing packages and downloads the speech model once. It never uninstalls or changes anything else. If it reports INSTALLING, the install continues in the background — wait a minute and call it again to check.

ParametersJSON Schema
NameRequiredDescriptionDefault
installNofalse (default): only check. true: brew install missing packages and download the model. Ask the user first.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate mutability and idempotency, but the description adds valuable specifics: existing installs are reused, an existing model is symlinked rather than re-downloaded, it uses brew install, downloads only once, never uninstalls or changes other things, and installation may continue in the background. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but every sentence carries necessary operational information: what is checked, what installation does, reuse behavior, user-consent requirement, safety constraints, and background-install handling. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers prerequisites, side effects, safety boundaries, invocation sequence, and the background-install scenario. The only missing detail is the full set of possible return statuses, but the mention of the INSTALLING report is enough to guide the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description enriches the single install parameter substantially: false is the default and means check-only, true means install missing packages and download the model, and true requires prior user consent. This goes well beyond the schema's brief parameter description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific purpose: check for and optionally install the dependencies that speak_and_listen needs, naming the exact components (text-to-speech, SoX, whisper.cpp, model). It clearly distinguishes itself from the sibling speak_and_listen by being the setup tool rather than the communication tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit sequencing: call with install=false first and report missing items; only call with install=true after user agreement. It also explains the background-install behavior and tells the agent to wait and re-check, which is strong operational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.3.1
    • Changedspeak_and_listen1 field changed
      • addedInput schema / properties / listen
        Added value: +{
        +  "description": "Default true. false: only speak, without opening the microphone — for a one-way announcement, or to say goodbye when the user ends voice mode.",
        +  "type": "boolean"
        +}
  2. 2 tool updatesv0.1.0
    • First observedspeak_and_listen
    • First observedvoice_setup

TDQS

A4.7/5.0

Scored across 2 tools

Disambiguation5/5

The two tools serve clearly separate phases: 'speak_and_listen' handles the conversational voice turn, while 'voice_setup' checks/installs dependencies. Descriptions explicitly link them (e.g., call voice_setup if not set up), so an agent can easily distinguish them.

Naming Consistency4/5

Both names use consistent snake_case, but the semantic structure differs: 'speak_and_listen' is a verb phrase while 'voice_setup' is a noun phrase. This is a minor deviation from a uniform verb_noun pattern, yet still readable and predictable.

Tool Count4/5

Two tools are slightly below the typical 3-15 range, but the server's narrow scope (voice interaction plus setup) makes each tool rich and necessary. The count is minimal but well-suited to the focused purpose.

Completeness4/5

The core voice lifecycle is covered: setup, a combined speak-and-listen turn (with one-way mode via listen=false), and ending voice mode. Minor gaps exist for voice/rate configuration or explicit interruption, but these are not critical to the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables hands-free voice conversations with Claude using real-time speech recognition and text-to-speech on macOS. Creates a self-sustaining conversation loop where Claude can autonomously listen, respond, and continue the interaction without keyboard input.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.
    16
    Apache 2.0
  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    A voice-enabled interface for Claude Desktop that supports speech-to-text input and text-to-speech output via ElevenLabs, turning Claude into a voice assistant.
    1
    -