loop-tts
by Mnemoclaw
README.md
# loop-tts — self-verifying TTS for agents and pipelines
Your TTS engine does not tell you what it actually said. **loop-tts closes the
loop**: every utterance is generated, transcribed by an *independent* ASR
(Whisper), measured, and re-generated with escalating instructions until it
passes quality gates. The result is a WAV file **and** a structured QC report.
```
┌──────────────────────────────────────────────────────┐
text ──────▶ │ TTS backend (any HTTP endpoint: POST text → WAV) │
└──────────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ independent ASR (faster-whisper) transcribes it │
└──────────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ quality gates on what was ACTUALLY said: │
│ • real WPM ≥ target (words transcribed / seconds)│
│ • integrity ≥ min (no swallowed words) │
│ • max gap ≤ limit (no mid-sentence dead air) │
└──────────────────────────┬───────────────────────────┘
pass? no ──▶ escalate (see below)
│ yes and retry (keep best)
▼
audio + QCReport
```
## Why
TTS models fail in ways you cannot see by reading the text: they talk too
slowly, swallow a word, or freeze for two seconds mid-sentence. The classic
"generate → listen → maybe retry" workflow is manual and unreliable. Verifying
with an independent ASR makes the quality check *objective* and *automatic* —
and the QC report is machine-readable, which is exactly what agents need.
Verifying TTS output with ASR has been described recently for audiobook
pipelines and in closed-loop TTS research. What loop-tts adds:
1. **A complete, dependency-light implementation** — one package, any backend.
2. **Pluggable escalation** (below) instead of a fixed retry strategy.
3. **A structured QC report designed for agents** — exposed as MCP tools and
a REST API, so any agent can *read* the failure and fix the *text* itself.
4. **`verify` mode** — the quality gates run on audio from *any* TTS engine,
turning the loop into a generic pre-publication QA gate.
## Escalation strategies
When a chunk fails, the loop retries with a stronger "boost". Which knob it
turns depends on your engine's control surface:
| strategy | what changes per retry | for engines that… |
|-------------|----------------------------------------------------------|--------------------------------|
| `instruct` | the *instruct* string climbs a ladder (default: live on air → live urgent → breaking-news pace) | accept style instructions (e.g. Qwen3-TTS and other instruct-aware models) |
| `parameter` | a payload parameter climbs a ladder, e.g. `[{"rate": "+8%"}, {"rate": "+18%"}]` | expose a speed/rate parameter |
| `none` | nothing — plain retry (pure quality gate) | everything else |
The **verification loop is engine-agnostic**: WPM, integrity, gap detection
and keep-best work identically with any backend. Only the escalation ladder is
swapped.
## Languages
The loop is **language-agnostic**: it works for any language covered by
*both* your TTS engine and the Whisper verification model (the reference
Qwen3-TTS server covers 10 languages; Whisper `small` covers ~99). Numbers
are spelled out with num2words (20+ languages) before comparing the expected
text with the transcription. For unspaced scripts (Chinese, Japanese),
integrity is computed **per character** and the rate gate counts characters
per minute — set `wpm_min` accordingly (Mandarin broadcast is roughly
200–260 chars/min).
The default escalation ladder is written in English (it is a style
instruction read by the model); translate `config.boosts` for best results
in other languages.
## Quality gates
- **Real WPM** — words in the *transcription* divided by the real audio
duration (not the requested text). Catches slow, "mushy" speech.
- **Word integrity** — multiset recall of the expected words in the
transcription, with numbers spelled out in the target language
(`1500` ↔ `mille cinq cents`) so ASR phrasing differences do not false-alarm.
Catches swallowed words.
- **Max internal gap** — longest silence strictly *inside* the audio
(chunk edges are excluded), detected by RMS. Catches mid-sentence dead air.
- **Keep best** — all attempts are scored
(`2·ΔWPM + 100·(1−integrity) + 50·Δgap`); the best one is kept even if
none fully passed.
## Install
```bash
pip install loop-tts
```
You need any TTS backend that answers `POST /tts` with a WAV body (see
[docs/tts-server.md](docs/tts-server.md) for a minimal example and notes on
open instruction-aware models). The Whisper model (~460 MB for `small`) is
downloaded on first use.
## Quickstart (Python)
```python
from loop_tts import TTSLoop, LoopTTSConfig
cfg = LoopTTSConfig(
tts_url="http://127.0.0.1:8092/tts", # your TTS backend
language="fr", # fr / en / de / ...
wpm_min=145, # minimum real speaking rate
escalation="instruct", # or "parameter" / "none"
)
loop = TTSLoop(cfg)
path, report = loop.speak(
"Welcome to the show. Today we talk about sleep, training and one "
"very opinionated kettlebell.",
instruct="Warm, conversational tone.",
)
print(report.ok, report.wpm_real, report.integrity, report.attempts)
# True 168.3 1.0 2
```
## CLI
```bash
# single utterance (QC report on stdout, exit code 0 = passed)
loop-tts speak "Bonjour, tout le monde." --instruct "Dynamique" --out hello.wav
# multi-voice piece from a JSON file
loop-tts batch segments.json --out episode.wav
# quality-gate ANY existing audio (any TTS engine)
loop-tts verify hello.wav --text "Bonjour, tout le monde."
# HTTP API (localhost by default) — docs at /docs
loop-tts serve --port 8300
# MCP server (stdio)
loop-tts mcp
```
## HTTP API
| route | body | returns |
|-------|------|---------|
| `POST /v1/tts` | `{text, speaker?, instruct?, min_wpm?}` | `{ok, audio_url, report}` |
| `POST /v1/tts/batch` | `{segments: [{text, speaker?, instruct?}], ...}` | `{ok, audio_url, reports[]}` |
| `POST /v1/verify` | `{audio_path, text}` | `{ok, report}` |
| `GET /v1/health` | — | backend + STT status, active gates |
| `GET /v1/audio/{name}` | — | the WAV file |
The server binds to `127.0.0.1` by default. If you expose it, set
`LOOP_TTS_API_KEY`: every route then requires the `X-API-Key` header.
## For agents (MCP)
`loop-tts mcp` exposes four tools:
- `tts_speak(text, speaker?, instruct?, min_wpm?)` → audio path + QC report
- `tts_batch(segments)` → assembled audio + one report per segment
- `tts_verify(audio_path, text)` → QC report on any existing audio
- `tts_status()` → health
The interesting part is the **report**. When `report.ok` is `false`, the
report tells the agent *what* failed (`failed_checks`), *what was actually
said* (`transcribed`), and per-attempt detail. The agent can then do what a
human would do but faster: **rewrite the offending text** (paraphrase a
swallowed proper noun, split a long sentence, remove an ambiguous opening
word) and call `tts_speak` again. The loop escalates the *delivery*; the agent
can escalate the *content*.
## Configuration
All knobs are plain values in `LoopTTSConfig` (or env vars):
| setting | default | meaning |
|---------|---------|---------|
| `tts_url` | `http://127.0.0.1:8092/tts` | TTS backend endpoint (`LOOP_TTS_URL`) |
| `language` | `en` | any code your TTS + Whisper share: `fr`, `de`, `ru`, `ja`… (`LOOP_TTS_LANGUAGE`) |
| `speaker` | `None` | opaque pass-through to the backend (`LOOP_TTS_SPEAKER`) |
| `wpm_min` | `145` | real WPM gate (`LOOP_TTS_MIN_WPM`) |
| `integrity_min` | `0.80` | word-integrity gate |
| `max_internal_gap` | `1.6` | silence gate (seconds) |
| `max_retries` | `4` | total attempts = 1 + max_retries |
| `escalation` | `instruct` | `instruct` / `parameter` / `none` |
| `boosts` | 3-level pace ladder | instruct ladder (level 0 = baseline) |
| `param_boosts` | `[]` | payload ladder for `parameter` mode |
| `chunk_max_words` | `45` | split long texts on sentences (0 = off) |
| `guard_enabled` / `guard_words` / `guard_prefix` | off | opt-in guard for FR/EN mixes: chunks starting with an ambiguous word (e.g. `Salut`) get a native filler prepended so the engine doesn't switch accent mid-show; customize for your own language pair |
| `stt_model` / `stt_device` / `stt_compute` | `small` / `cpu` / `int8` | Whisper verification model |
| `out_dir` | `out/` | where WAV files are written (`LOOP_TTS_OUT`) |
| `segment_gap_s` | `0.2` | silence between segments in `batch()` |
## Security
- No secrets in code or config defaults; the TTS URL is a plain setting.
- HTTP API binds to localhost by default; optional `LOOP_TTS_API_KEY` shared
key; audio is only served from the configured out dir (no path traversal).
- Generated audio and QC reports are local files — nothing is uploaded.
## Project layout
```
loop_tts/
├── config.py # every knob, one place
├── metrics.py # WPM, word integrity, internal gap detection
├── stt.py # lazy faster-whisper wrapper
├── core.py # the loop: chunk → generate → transcribe → measure → escalate
├── api.py # FastAPI app
├── mcp.py # MCP server
└── cli.py # loop-tts speak|batch|verify|serve|mcp
```
## Disclaimer
loop-tts is an independent project. Qwen3-TTS is referenced only as an
example of an open instruction-aware TTS model; loop-tts works with any
backend exposing a simple HTTP endpoint.
## License
MIT — see [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues