voix
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@voixread me the build results out loud"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
voix
Local voice output for AI agents. One executable, one MCP tool, no cloud.
Agent ──MCP──▶ voix ──▶ Kokoro (local ONNX) ──▶ your speakersAn agent gathers whatever it needs with its other tools, writes a spoken summary, and calls speak. voix synthesizes it with Kokoro-82M on the CPU and plays it. First audio starts in about a second; nothing leaves the machine.
MVP scope: macOS on Apple Silicon. See docs/SPEC.md for every design decision.
Install
One line (installs to ~/.local/bin/voix; set VOIX_INSTALL_DIR to change that):
curl -fsSL https://github.com/osmelmora/voix-mcp/releases/latest/download/install.sh | shOr download the executable from the latest release (direct link) and put it on your PATH:
curl -fsSL -o voix https://github.com/osmelmora/voix-mcp/releases/latest/download/voix-darwin-arm64
chmod +x voix && mv voix ~/.local/bin/voixOr build it yourself (needs Bun 1.4.2 (pinned in mise.toml)):
git clone https://github.com/osmelmora/voix-mcp && cd voix-mcp
bun install
bun run build # → dist/voix-darwin-arm64 (≈120 MB)
cp dist/voix-darwin-arm64 ~/.local/bin/voixThe binary is not code-signed. curl downloads run as-is; if you downloaded it with a browser run xattr -d com.apple.quarantine ~/.local/bin/voix once.
Related MCP server: mcp-ai-voice
First run
voix say "Hello"On first use voix downloads the Kokoro model (326 MB, fp32, Apache-2.0) into ~/.cache/voix and verifies its SHA-256. After that everything is offline. To do the download ahead of time:
voix setupConnect an agent
Add voix to your MCP client configuration (Claude Code, Claude Desktop, Cursor, ...):
{
"mcpServers": {
"voix": { "command": "voix", "args": ["mcp"] }
}
}If voix is not on the client's PATH, use the absolute path, for example "/Users/you/.local/bin/voix".
The agent gets two tools:
tool | what it does |
| Speak text. Returns as soon as the first sentence starts playing (or |
| Stop immediately and drop the queue. |
Errors come back as typed tool errors (EmptyText, TextTooLong, InvalidVoice, InvalidSpeed, ModelDownloading, DownloadFailed, ChecksumMismatch, PlayerNotFound, PlaybackFailed, SynthFailed). The server also publishes short usage instructions and the resource skill://voix/SKILL.md, which is the same text as skills/voix/SKILL.md. Copy that folder into your agent's skills directory if it supports skills.
Then try:
Give me my daily update out loud.
CLI
voix say "Build finished." # speak
echo "Deploy done." | voix say # from stdin
voix say --voice bm_george --speed 1.1 "Good evening."
voix say --out hello.wav "Hello" # write a WAV instead of playing
voix voices # 28 English voices, * marks the default (af_heart)
voix status # model, cache, player, runtime
voix setup # download the model now
voix mcp # MCP server over stdioConfiguration
There is no config file. Defaults: voice af_heart, speed 1.0.
variable | effect |
| where models are stored (default |
| disable audio output (tests, CI) |
How it works
Runtime: TypeScript on Bun, compiled with
bun build --compile. Effect 4 for services, typed errors, the queue, interruption, and the MCP and CLI layers.Inference:
onnxruntime-nodeon the CPU with about 150 lines of Kokoro glue and the eSpeak NG phonemizer in WebAssembly. No Python, no transformers.js.Pipelining: text is split into sentences; the first plays while the rest synthesize, and each later playback chunk is everything that finished in the meantime.
Playback: a temp WAV played by
/usr/bin/afplayas a scoped child process, killed onstop, on request cancellation, and on exit.What is inside the binary: Bun runtime, the ONNX Runtime library, 28 voice files, the tokenizer, the code. Only the model is downloaded.
Development
Use mise to install and activate the Bun, Node.js, hk, Pkl, and cocogitto versions pinned in mise.toml. Node runs the lint and format tooling. mise also puts node_modules/.bin on PATH. The release workflow uses the same mise configuration.
Linting and formatting use Ultracite with Oxlint and Oxfmt. Run bun run check to check both, or bun run fix to apply automatic fixes. The CI workflow checks commit messages, linting, formatting, types, and tests on every push to main and every pull request; the release workflow repeats those checks before building.
Git hooks run through hk, configured in hk.pkl. The pre-commit hook runs Oxlint, Oxfmt, and tsc on staged files and stages automatic fixes. The commit-msg hook requires Conventional Commits subjects such as feat(mcp): add stop tool. The pre-push hook runs the same checks and the tests. Run hk check --all to check everything at once, or hk fix --all to apply fixes. Set HK=0 to skip the hooks once.
Releases use cocogitto, configured in cog.toml. cog bump --auto picks the next version from the commits since the last tag, sets it in package.json, prepends the release to CHANGELOG.md, then commits and tags it. Pushing the tag runs the release workflow, which uses that version's changelog entry as the GitHub release notes. Run cog changelog to preview unreleased changes.
mise install
bun install
hk install --mise # install the git hooks
hk check --all # lint, format, and type check
VOIX_PLAYER=none bun test # synthesis and MCP tests run only if the model is cached
bun run src/main.ts say "dev mode"
bun run build && VOIX_BIN=./dist/voix-darwin-arm64 VOIX_PLAYER=none bun test tests/mcp.test.tsKnown limitations
darwin-arm64 only. Linux and Windows cross-compile but are untested; Intel Macs lack a prebuilt ONNX Runtime in this version.
English voices only. Kokoro's other languages need a different grapheme-to-phoneme stack.
The embedded eSpeak NG build is GPL-3; see NOTICE.
License
MIT for voix. Third-party components are listed in NOTICE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
Command your AI agents by voice: PTT rooms, channels, direct messages, agent email, memory (mRAG).
Text to speech for your AI. Your AI can send text to Doc Player to read it aloud. You will see a reader window with the text and you can control the playback sentence by sentence. Find an example here: https://documentplayer.com/connect-ai/
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to speak and listen in real-time with interruption handling, using local ML models and hot-swappable adapters.37 npmMIT
- AlicenseAqualityDmaintenanceEnables AI agents to synthesize natural speech using either platform system voices or premium OpenAI TTS, with automatic engine selection and graceful fallback.118 npmMIT
- AlicenseAqualityAmaintenanceEnables AI agents to speak using MacOS native text-to-speech, with support for blocking and non-blocking speech and a sequential queue.215MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to speak aloud by generating and playing audio through the system output. Supports multiple TTS providers, playback queue management, and configurable voice profiles.37 npm1MIT