vox
Provides high-quality multilingual text-to-speech via ElevenLabs API, with automatic fallback to macOS system voice.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@voxlisten"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
vox
A native macOS MCP server that gives Claude a voice.
Built in Swift. No Node.js. No Python. No Electron. Just a single binary.
{
"mcpServers": {
"vox": {
"command": "/path/to/vox"
}
}
}Then just say: listen — and Claude hears you.
Tools
Tool | What it does |
| Activates the mic, waits for you to speak, returns transcript + detected language when silence is detected |
| Speaks text aloud — ElevenLabs TTS with automatic fallback to macOS system voice |
Related MCP server: whisper-windows-mcp
Why Swift + macOS only
AVFoundation — native mic capture, zero overhead
SFSpeechRecognizer — Apple on-device speech recognition, works offline, 50+ languages
NLLanguageRecognizer — automatic language detection per utterance
AVSpeechSynthesizer — built-in TTS fallback, no API key needed
Single ~200KB binary. Zero npm install. Zero Python venv.
Requirements
macOS 13+
Microphone set as default in System Settings → Sound → Input
Microphone + Speech Recognition permission for your terminal app
Optional:
ELEVENLABS_API_KEYin~/.claude/.envfor high-quality multilingual TTS
Install
Option 1: Prompt for Claude Code
Copy and paste this into Claude Code:
Install vox (native macOS voice input/output MCP server):
1. Download the code-signed binary to ~/vox:
curl -L https://github.com/boska/vox/releases/download/v1.1.0/vox-1.1.0-darwin-arm64 -o ~/vox && chmod +x ~/vox
2. Add to ~/.claude.json in the mcpServers section:
{
"mcpServers": {
"vox": {
"command": "/Users/$(whoami)/vox"
}
}
}
3. Restart Claude Code
4. Test it by saying: listenClaude will handle the download, configuration, and restart.
Option 2: Manual Install (Pre-built Binary)
curl -L https://github.com/boska/vox/releases/download/v1.1.0/vox-1.1.0-darwin-arm64 -o ~/vox && chmod +x ~/voxThen add to ~/.claude.json:
{
"mcpServers": {
"vox": {
"command": "/Users/$(whoami)/vox"
}
}
}Restart Claude Code.
Option 3: Build from Source
git clone https://github.com/boska/vox
cd vox
swift build -c release --product voxAdd to ~/.claude.json:
{
"mcpServers": {
"vox": {
"command": "/path/to/vox/.build/release/vox"
}
}
}Restart Claude Code.
Voice
Default TTS voice is Hana via ElevenLabs (eleven_multilingual_v2). Supports Chinese, English, Czech, Vietnamese, and 20+ languages in the same voice. Without an ElevenLabs key, falls back to macOS system voice automatically.
How it works
Claude Code
↓ JSON-RPC 2.0 over stdio
vox binary
↓ ↑
AVAudioEngine AVSpeechSynthesizer
SFSpeechRecognizer ElevenLabs API
NLLanguageRecognizerNo ports. No sockets. No daemon. Just stdin/stdout.
Troubleshooting
"No speech detected" — speak within ~1s of calling listen, VAD cuts off after 0.8s silence.
No default input device — Mac mini has no built-in mic. Connect USB/Bluetooth mic or iPhone via Continuity Camera, set default in System Settings → Sound → Input.
ElevenLabs silent — add ELEVENLABS_API_KEY=sk-... to ~/.claude/.env, or leave it out to use system voice.
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP server for Speech-to-Text
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for Text-to-Speech
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Related MCP Servers
- AlicenseAqualityDmaintenanceA cross-platform MCP server that enables Claude to speak using Microsoft Edge TTS with support for over 300 voices across 50+ languages. It requires no API keys and allows for customization of speech rate, volume, and pitch.318 PyPI2MIT
- FlicenseAqualityAmaintenanceA Windows-native MCP server that lets Claude Desktop transcribe audio files locally using whisper.cpp, with no internet connection required.1319 npm1-
- AlicenseNot gradedqualityCmaintenanceText-to-speech MCP server using the Kokoro-82M model accelerated with MLX on Apple Silicon, enabling local Claude and Codex clients to speak text aloud and convert text to audio.4MIT
- AlicenseNot gradedqualityDmaintenanceLocal MCP server that enables voice input to Claude by transcribing microphone audio locally using faster-whisper, avoiding cloud services.MIT