Skip to main content
Glama
  • Transcribe locally25 languages, up to ~19x faster than Whisper on Apple Silicon, ~2.5x on CPU

  • Speak back — text-to-speech in 9 languages

  • Plug into agents — ship voice workflows as CLI commands, an MCP server, an OpenClaw skill, or a Hermes agent

  • Small Rust engine — single ~65MB binary, no ffmpeg, no Python, no native Node addons

Quick Start

Runtime: Bun >= 1.3.0.

# 1. Install Bun (skip if you have it)
curl -fsSL https://bun.sh/install | bash        # macOS/Linux — or: brew install oven-sh/bun/bun
powershell -c "irm bun.sh/install.ps1 | iex"    # Windows

# 2. Install Kesha
bun add -g @drakulavich/kesha-voice-kit
kesha --version                                 # confirms `kesha` resolved on PATH

# 3. Download the engine and models — pick one path
kesha init                                      # guided: TTS languages and optional VAD / diarization
kesha install --plan && kesha install           # manual: preview the sizes, then download

# 4. Transcribe
kesha audio.ogg                                 # transcript to stdout

kesha install pulls ~2.5 GB on Linux/Windows and ~0.6 GB on Apple Silicon, whose CoreML engine reads a smaller model set. It is always explicit — nothing downloads behind your back — and reports download progress on stderr. If bun --version fails right after step 1, reload your PATH: exec $SHELL -l.

Prefer Homebrew or Docker? See Other install methods. Air-gapped or behind a corporate mirror? See docs/model-mirror.md.

Platform support

All three targets transcribe, detect the spoken language, run VAD, and speak. The macOS-only rows need Apple frameworks — they are not a missing port. Windows is a tested path rather than a published binary nobody ran: CI does a cold kesha install on windows-latest, transcribes a fixture, and round-trips a synthesis (#216, #667).

macOS arm64

Linux x64

Windows x64

Transcribe · audio language ID · VAD

CoreML / ANE

ONNX CPU

ONNX CPU

TTS — en ru es fr it pt

TTS — hi ja zh and macOS system voices

Mic capture and live dictation (kesha record)

Speaker diarization (--speakers)

Word-level timestamps (words in --json)

Voice auto-routing from the text's language

pass --lang

pass --lang

Intel Macs get no published engine binary. Full matrix with maturity labels: docs/product-positioning.md.

Related MCP server: STT2TTS MCP

Speech-to-text

kesha audio.ogg                            # transcribe (plain text)
kesha --format transcript audio.ogg        # text + language/confidence
kesha --format json audio.ogg              # full JSON with lang fields
kesha --json --timestamps audio.ogg        # JSON with timestamped segments
kesha --itn audio.ogg                      # spelled-out numbers -> digits
kesha --toon audio.ogg                     # compact LLM-friendly TOON
kesha status                               # show installed backend info
kesha status --disk                        # + recursive cache disk usage
kesha status --json                        # machine-readable, for scripts

Multiple files get head-style headers; stdout is the transcript, stderr is errors — pipe-friendly:

$ kesha freedom.ogg tahiti.ogg
=== freedom.ogg ===
Свободу попугаям! Свободу!

=== tahiti.ogg ===
Таити, Таити! Не были мы ни в какой Таити! Нас и тут неплохо кормят.
  • Record from the mic (macOS): kesha record --out hello.wav writes microphone audio to a WAV file (kesha hello.wav transcribes it). macOS prompts for microphone access on first use — grant it under System Settings → Privacy & Security → Microphone if it was denied. On Linux/Windows or headless boxes, pass any existing audio file straight to kesha instead.

  • Dictate straight to text (darwin-arm64): kesha record --live transcribes the mic as it captures and prints the transcript to stdout — no WAV in between, so it pipes (kesha record --live | pbcopy). To end after trailing silence, explicitly install VAD then opt in: kesha install --vad && kesha record --live --auto-stop. The defaults are 1,000 ms of silence after 250 ms of speech; tune them with --auto-stop-silence-ms, --auto-stop-min-speech-ms, and --auto-stop-threshold. Progress goes to stderr. Linux and Windows do not capture the microphone; pass an existing audio file to kesha to transcribe it. An interruption is recoverable: Ctrl-C (or SIGTERM) stops the session, still prints what you dictated, and exits 130/143, and the audio is spilled to a recovery WAV under ~/.cache/kesha/recordings/ — named on stderr when the session starts, deleted once the transcript has actually been delivered, kept if anything — a signal, a crash, a closed terminal, a dead pipe — got in the way first (#962).

  • Long / silence-heavy audio: install VAD (kesha install --vad); Kesha auto-uses it past 120 s. Without VAD, long audio falls back to fixed ASR chunks. See docs/vad.md.

  • Speaker diarization (darwin-arm64): kesha install --diarize (which installs VAD too), then kesha --json --speakers meeting.m4a stamps each segment with a speaker id. --speakers engages VAD windowing itself at any duration, so it cannot be combined with --no-vad. Linux/Windows return a clear "darwin-arm64 only" error (#199).

  • Word-level timestamps (every platform): kesha --json --timestamps audio.ogg adds a words array to each segment — { "word": "email", "start": 0.72, "end": 1.12 } — on the same file-relative clock as the segment, so a word always lies inside the segment carrying it. Read them off the decoder's own frame grid, so: times are quantised to 0.08 s, consecutive spans may overlap (each end is a per-word duration prediction, not the next word's start), end >= start rather than strictly greater, and punctuation stays attached to its word. The key is simply absent where a segment has none — any segment --itn rewrote, for one — so check the transcribe.words capability rather than expecting an empty array (#720).

  • Text-language detection: JSON and TOON results include textLanguage with a language code, confidence, and its source. On macOS Kesha uses Apple NLLanguageRecognizer; elsewhere it uses the bundled tinyld fallback, whose confidence scale is different. This is separate from audioLanguage, which identifies the spoken audio when available.

  • Written-form numbers: --itn rewrites what the model spells out — "two hundred thirty two""232", "five dollars and fifty cents""$5.50". Opt-in, every platform, timestamps untouched. English-only in practice; Russian and the rest pass through unchanged. Spoken punctuation names stay words ("dot", "comma", "the period of growth") because Kesha transcribes speech rather than dictation — so "example dot com" keeps its words too (#822). A sentence "and" survives the number that follows it ("cats and three dogs""cats and 3 dogs"), while an "and" the number owns still joins it ("three hundred and five""305") (#1000) — and no longer splits the number around it ("two hundred and thirty two""232", not "230 2") (#1006). A hyphenated number reads the same as the spaced form ("twenty-five apples""25 apples"), while a hyphen between ordinary words is left alone ("well-known", "state-of-the-art", "twenty-something") (#1004).

Text-to-speech

Kesha speaks back in 9 languages. Kokoro runs natively through FluidAudio CoreML/ANE on Apple Silicon and through ONNX on Linux and Windows; Russian uses Vosk-TTS, while macos-* system voices need no model download. On macOS Kesha picks the voice from the text's own language; on Linux and Windows, state the language with --lang <code> (or the voice with --voice <id>) — otherwise the engine default speaks.

kesha install --tts                              # English voices; sizes differ per platform — preview: kesha install --plan
kesha install --tts en ru                        # + Russian (+~890 MB, Vosk)
kesha say "Hello, world" > hello.wav
kesha say "Привет, мир" > privet.wav             # auto-routes by language (macOS)
kesha say --lang ru "Привет, мир" > privet.wav   # explicit — the Linux/Windows path
kesha say --voice ru-vosk-m02 "Голос в текст." > ru.wav

Output formats (--format, or inferred from the --out extension):

kesha say "Hello" --out hi.wav                    # WAV (default, uncompressed)
kesha say "Hello" --format ogg-opus --out hi.ogg  # OGG/Opus — messenger voice notes
kesha say "Hello" --format flac --out hi.flac     # FLAC — lossless, plays in every browser incl. Safari/iOS

kesha say --list-voices lists what's installed. Voices, the full catalogue, macOS system voices, SSML, speaking rate (--rate, <prosody>), Russian word stress, and Russian/English abbreviation handling are all in docs/tts.md.

Languages

Speech-to-text spans 25 languages and text-to-speech 9 — full tables with codes, flags, and per-platform availability in docs/languages.md. Audio language detection identifies 107 languages.

Performance

Up to ~19x faster than Whisper on Apple Silicon (M2), ~2.5x faster on CPU

Compared against Whisper large-v3-turbo, all engines auto-detecting language:

Benchmark: openai-whisper vs faster-whisper vs Kesha Voice Kit

Full per-file breakdown (Russian + English): BENCHMARK.md. The CPU figure is the ONNX engine on an M2's CPU cores; no x86 numbers are published yet.

Other install methods

All of these install the Bun CLI wrapper; engine + models still download explicitly via kesha install. (Nix is the exception — it currently builds only the engine from source; see below.)

  • Homebrewbrew install drakulavich/tap/kesha-voice-kit · docs/homebrew.md

  • Linux packages (.deb/.rpm, x64) — published on CLI releases, see docs/linux-packages.md

  • Docker (GHCR image) — docs/docker.md

  • Nix (aarch64-darwin / x86_64-linux) — builds the engine from source (nix build github:drakulavich/kesha-voice-kit#kesha-engine). The full kesha CLI via nix run / nix profile install is not yet available — it needs a maintainer with Nix to populate a build hash (#946). · docs/nix-install.md

  • Shell completions + manpagekesha completions bash|zsh|fish and kesha manpage print the packaged files to install wherever your shell expects them.

Integrations

  • MCP serverkesha mcp exposes transcribe/synthesize/list tools to any MCP client (Claude, Cursor, Codex, Gemini). Setup: docs/mcp.md.

  • OpenClaw — give your LLM agent ears. Install & config: docs/openclaw.md.

  • Hermes Agent — local STT/TTS through Hermes command providers. Setup: docs/hermes.md.

  • Raycast (macOS) — offline microphone dictation from the launcher: Dictate to Clipboard records with a live signal meter, auto-stops on silence, transcribes locally, and copies the text. Install from the Raycast Store · source: raycast/.

  • Programmatic API@drakulavich/kesha-voice-kit/core for use inside a Bun program. See docs/api.md.

More

  • Architecture — runtime data flow, the models that ship, the CLI ↔ Rust engine boundary, model pinning, and where tests live.

  • Use cases — copy-paste recipes (transcribe a meeting, speak from OpenClaw, run offline, move the cache).

  • Product positioning — supported workflows, non-goals, maturity labels, platform matrix.

  • Changelog — every release, with the behaviour changes spelled out.

  • Diagnostics: kesha doctor, kesha support-bundle (redacted .tar.gz for issues), and kesha logs produce local, content-free diagnostics — see docs/diagnostic-logs.md. Every failure prints a stable error [CODE]: … line and a documented process exit code.

  • Scripting & CI: --json (or --toon) for machine-readable output, --include-errors (with either) to get per-file failures on stdout alongside the results, --quiet/-q to silence progress, and --no-color (or NO_COLOR=1) for plain logs. Colors switch off automatically when CI=true.

  • Privacy / Local Stats: Stats are off by default and fully local. Opt in with kesha stats enable to record content-free operational metrics in a local SQLite database — never networked, never storing audio, transcripts, text, or paths. Full commands & lifecycle: docs/local-stats.md.

Contributing

See CONTRIBUTING.md, the Roadmap (Now / Next / Later), and the Decision log (why platform/model choices were made — and reversed). Dev setup: just dev-setup (Bun, Rust, nextest, platform libs).

License

Made with 💛🩵 and 🥤 energy under MIT License

A
license - permissive license
Not graded
quality - not tested
A
maintenance

Maintenance

Maintainers
13hResponse time
1dRelease cycle
90Releases (12mo)
Commit activity
Issues opened vs closed

Related MCP Servers

  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    A local voice interface providing high-performance speech recognition and natural text-to-speech with voice cloning capabilities. It enables AI assistants to speak, listen, and engage in character-based voice conversations through integrated MCP tools.
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local-first speech-to-text and text-to-speech MCP server. Hot-swappable engines via config.yaml — no code changes, no API keys required.
    2
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Local, private audio transcription MCP server enabling AI agents to transcribe audio files entirely on-device without uploading data.
    3
    MIT

View all related MCP servers

Related MCP Connectors

  • LibreTranslate MCP — open-source machine translation (BYO endpoint)

  • MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/drakulavich/kesha-voice-kit'

If you have feedback or need assistance with the MCP directory API, please join our Discord server