Talkies
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Talkiestranscribe /data/audio/meeting.wav using whisper-large-v3-turbo"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
talkies
Self-hosted speech services in one Docker image: OpenAI-compatible file transcription and text-to-speech, Talkies live ASR over WebSocket, file staging, model lifecycle controls, and an MCP endpoint for ASR workflows.
Contents
Related MCP server: kokoro-tts
Start here
Restrict the first boot to the models you need; otherwise the entrypoint downloads every model in the bundled registry.
docker run --rm -it --name talkies \
-p 127.0.0.1:8000:8000 \
-v "$PWD/talkies-data:/data" \
-e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo,kokoro-82m \
psyb0t/talkies:latest
curl -s http://127.0.0.1:8000/healthz
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F "file=@/path/to/clip.wav" \
-F "model=whisper-large-v3-turbo"For CUDA-only models — Parakeet-TDT, the larger Canary models, Qwen3 TTS and
Chatterbox Turbo — use psyb0t/talkies:latest-cuda with --gpus all. The
loopback port mapping keeps the service local; see
Getting started for first boot and authentication.
What it provides
Surface | Purpose | Reference |
| File transcription and subtitles | |
| Live 16 kHz PCM ASR | |
| Speech synthesis in six formats | |
| Enabled slugs and their modality | |
| Per-model voice catalog with origin tags | |
| Server-side file staging | |
| Model inspection and eviction | |
| Streamable HTTP MCP with ASR/file tools | |
| Liveness probe; the only unauthenticated route |
The HTTP transcription and speech routes use the corresponding OpenAI wire shapes where those contracts overlap. Streaming ASR, files, lifecycle controls, and MCP are Talkies extensions.
Models at a glance
CPU: two Whisper models, Canary-180M-Flash, Nemotron ASR via parakeet.cpp, four English Sherpa-ONNX Zipformer choices, Vosk small English, two phoneme recognizers, and two Kokoro TTS backends.
CUDA: the CPU set plus Parakeet-TDT, Canary 1B/Qwen ASR, five Qwen3 TTS variants, and Chatterbox Turbo.
Live ASR: bundled Nemotron, Sherpa-ONNX, and Vosk are native; bundled Whisper is a bounded rolling decoder. Sherpa and Vosk also work through the OpenAI-compatible file-transcription endpoint.
Phoneme recognition:
wav2vec2-xlsr-53-espeakandzipa-ipareturn the IPA phones that were spoken, not words, with no language model correcting them toward the nearest dictionary entry. Same transcription endpoint and timestamp options as the other ASR models; see Phoneme recognition.Per-model concurrency limits cover WebSocket, HTTP, MCP, ASR, and TTS; the bundled Nemotron CPU and CUDA entries admit two requests.
Streaming TTS: Qwen3 returns incremental raw PCM for
response_format="pcm"; other TTS formats and Kokoro are buffered.Expressive TTS: Chatterbox Turbo (English) takes 19 inline tags such as
[sigh],[whispering]and[laugh]directly in the input text. Its output carries a neural watermark by default; setTALKIES_CHATTERBOX_WATERMARKto false to emit unmarked audio.Voice cloning: drop a
.wavinto/data/custom-voicesand it appears onGET /v1/audio/voices. Qwen3 pairs it with an optional sibling.txttranscript; Chatterbox needs only the clip, longer than five seconds.
Exact slugs, executors, tag list, and registry format: Models and registries.
Reading phonemes
wav2vec2-xlsr-53-espeak and zipa-ipa use the same transcription call as
every other ASR slug; only the model changes. text comes back as a
space-separated IPA phone stream rather than words, and no language model
corrects a mispronunciation toward a real word.
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F "file=@/path/to/clip.wav" \
-F "model=zipa-ipa"
# {"text": "a ɪ m k ə n f j u z ...", ...}Add -F "response_format=verbose_json" (or timestamp_granularities[]=word)
to get each phone as a words entry with start and end in seconds.
Prompting Chatterbox with emotion
Tags go inline in input, in square brackets, lowercase. They are real tokens
in the model's tokenizer, so only these 19 do anything — any other bracketed
word is spoken as literal text:
[angry] [fear] [surprised] [whispering] [advertisement] [dramatic] [narration]
[crying] [happy] [sarcastic] [clear throat] [sigh] [shush] [cough] [groan]
[sniff] [gasp] [chuckle] [laugh]curl -s http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "chatterbox-turbo",
"voice": "builtin",
"input": "Oh, that is hilarious. [chuckle] Anyway [sigh] back to work.",
"response_format": "mp3"
}' --output out.mp3Swap "voice" for the name of any .wav you dropped in /data/custom-voices
(extension stripped) to speak the same line in a cloned voice.
Documentation
Guide | Contents |
Run CPU/CUDA, persist data, authenticate, verify | |
Bundled slugs, image availability, custom registries | |
Request flow, backend selection, on-disk layout | |
Requests, responses, files, lifecycle, MCP | |
Live ASR protocol, streaming backends, PCM TTS | |
Supported environment variables and limits | |
Exposure, model memory, data retention, logs | |
Make targets, test suites, image builds |
Agent integrations
The Talkies skill teaches agents to use the HTTP,
WebSocket, and MCP surfaces. Install it through the shared psyb0t marketplace
or let Codex discover it directly from this checkout.
Claude Code
claude plugin marketplace add psyb0t/agents
claude plugin install talkies@psyb0tClaude Code prompts for the Talkies URL and, when enabled, the bearer token; the sensitive token is stored through the client's protected configuration.
Codex
codex plugin marketplace add psyb0t/agents
codex plugin add talkies@psyb0tA marketplace install invokes the skill as $talkies:talkies. Codex also
discovers .agents/skills/talkies directly in this repository, where it is
invoked as $talkies without installation.
OpenClaw
The skill and MCP bridge are published through ClawHub:
openclaw skills install @psyb0t/talkies
openclaw plugins install clawhub:@psyb0t/talkiesThe bridge connects local stdio MCP clients to a running Talkies /v1/mcp
endpoint. Set TALKIES_URL and, when authentication is enabled,
TALKIES_AUTH_TOKEN.
Security in one minute
TALKIES_AUTH_TOKEN enables a shared bearer token for every HTTP and WebSocket
route except /healthz. It is unset by default. Keep the port loopback-only or
put Talkies behind TLS, authentication, and rate limiting. If untrusted callers
can supply remote file_path URLs, set TALKIES_BLOCK_PRIVATE_DOWNLOADS=true.
See Operations and security for the complete posture.
Development
make check # lint + unit tests in the dev image
make lint # flake8 + mypy only
make test-unit # fast offline unit tests
make run # run the CPU image locally
make test-streaming # real CPU native WebSocket ASR test
make test-streaming-custom # real CPU Sherpa/Vosk WebSocket + HTTP tests
make test-streaming-custom-cuda # real CUDA Sherpa WebSocket + HTTP test
make compile-heavy # regenerate the hash-locked ML requirements
make build-all # CPU and CUDA production imagesmake help lists every target.
Talkies is released under the WTFPL. Model weights are downloaded at runtime and have their own terms; image component notices are in THIRD_PARTY.md. Release notes are in CHANGELOG.md.
This server cannot be deployed
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
Give AI agents real phone numbers, messages, and voice calls via MCP.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables MCP-compatible clients to leverage OpenAI's multimodal capabilities (vision, image generation, speech-to-text, text-to-speech) through file-oriented tools with a security-first architecture.101MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to generate high-quality speech with 54+ voices in multiple languages via MCP tools.19Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables voice-first interactions with AI agents and MCP tools, supporting speech input/output, STT/TTS, and a provider-independent agent core.1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI clients like Claude and ChatGPT to generate images and videos, animate images, create lip-synced videos, list TTS voices, and manage media via remote MCP tools.-