io.github.attenlabs/hotato
OfficialAllows placing calls on Twilio to recapture a recording for a pinned check against the current agent.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@io.github.attenlabs/hotatoscore ./call.wav and tell me what broke in the booking"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
hotato
Find what broke in your agent calls. Pin it and CI stays red until you fix it.
pip install hotato
hotato check --demo # a bundled failing call: what broke, and when
hotato check ./call.wav # then your own recording
hotato vapi health # or your last 100 Vapi callsZero config. Vapi, Retell, Bland, Synthflow, Millis, or local audio. Works with any recording: mono, dual-channel, or a transcript. Timing math, not a judge. Offline. Free. MIT.
What it finds
Say-do gaps: the caller interrupts to cancel (a barge-in), the agent says "canceled", the booking tool fires anyway. hotato takes the turn timing from the audio and the tool call from your OTel trace.
Latency spikes: the pause before a reply going from 800 ms to over 2 s.
Dead air: that pause reaching 5 s, or the line going quiet.
Talk-over: the agent starts a fresh utterance over the caller.
Related MCP server: voice-analysis-mcp
Quickstart
Vapi
pip install hotato
export VAPI_API_KEY=...
hotato vapi health --last 7d --output report.htmlOpen report.html: every critical incident, timestamped, and your Voice Stability Score.
Retell
export RETELL_API_KEY=...
hotato retell health --call-id CALL_ID--call-id is required and repeatable: you name the Retell calls to pull.
hotato bland health, hotato synthflow health, and hotato millis health
follow the Vapi shape. Those stacks mix both voices onto one channel, so
they measure silence timing, dead air and latency gaps, each finding with
its measured confidence; barge-in and talk-over need two channels.
Local audio
hotato autopsy ./call.wavWrites a self-contained HTML report to hotato-output/, plus the JSON
pin reads. Open the HTML in your browser.
From finding a bug to gating on it
autopsy turns one recording into timestamped incidents. scan reads a
folder and tracks the trend. When you are ready, move a finding into CI:
hotato pin turns one incident into a portable failure check, and hotato prove re-runs every stored check and fails the build rather than pass on
evidence it cannot re-read. Every verdict carries its own evidence across
five dimensions: outcome, policy, conversation, speech, reliability.
A pinned check re-measures its own stored recording, so it holds the build
red on that bug and catches an edited bundle, a loosened policy, or an
engine upgrade that scores the same bytes differently. Clearing it takes a
fresh recording of the same moment against the current agent: hotato drive <bundle> places that call on Vapi or Twilio, and
docs/RECAPTURE.md is the walkthrough for every other
stack.
For continuous use: run hotato vapi health on a schedule, and open
hotato console --production-db evidence.db to watch calls land live.
Wire it into CI
The exit code is the verdict: 0 pass, 1 fail, 2 refuse (could not tell).
# .github/workflows/voice-qa.yml
on: [pull_request]
jobs:
hotato:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: attenlabs/hotato@v1.20.0
with:
contracts: contracts/
hotato-version: 1.20.0contracts/ holds what pin wrote. Full workflow: docs/CI.md.
Point your agent at it
Point Claude Code, Cursor, or any coding agent at this repo: it reads
AGENTS.md and runs the loop end to end offline, no key. Over local
stdio the MCP server adds the scorer plus read/verify/propose tools:
uvx --from "hotato[mcp]" hotato-mcp (docs/MCP.md). Deploying is yours.
Nothing leaves your machine
hotato runs offline, on the machine that invokes it. The core is stdlib-only Python: no account, no key, no network call of its own. Your traces, prompts, and audio stay on your disk. Scoring with a language model you host yourself is a separate opt-in add-on, outside that core.
Go deeper
The whole loop, command by command: docs/LIFECYCLE.md.
First touch to a CI gate: docs/GETTING-STARTED.md.
Feed it what you already have: docs/CONNECT.md ·
docs/TRACE.md · docs/SIMULATE.md.
What every verdict stands on: docs/EVIDENCE-CONTRACT.md.
Next to the hosted alternatives: docs/COMPARE.md.
The deep toolkit -- capture, simulation, load, benchmarking, the fix ladder,
the fleet control plane -- lives under hotato lab (hotato lab --help).
The public commands are durable. hotato lab moves faster, and every
command name that worked before 1.17 still runs unchanged.
Specifications
Property | Value |
Footprint | ~10 MiB installed, 0 runtime dependencies (stdlib-only) |
Reproducibility | byte-for-byte: the same recording, the same report |
Exit codes |
|
Release integrity | OIDC Trusted Publishing + build-provenance attested |
Runtime | offline, off the production data path |
PYTHONPATH=src python3 -m hotato.benchmark \
--scenarios corpus/real/scenarios --audio corpus/real/audioOn 13 recorded AMI Meeting Corpus clips, the median error between measured caller-onset and the human word-alignment label is 20 ms. Provenance: corpus/real/README.md · method: METHODOLOGY.md.
Timing is measurable only when the two voices arrive on separate channels; a mono or mixed export is marked NOT SCORABLE and refused (hotato trust --stereo call.wav). The full four-tier evidence policy (what each verdict stands on, per input) is docs/EVIDENCE-CONTRACT.md.
Contribute
Issues and PRs welcome: CONTRIBUTING.md · SECURITY.md · CHANGELOG · docs/
License
MIT (LICENSE)
mcp-name: io.github.attenlabs/hotato
This server cannot be deployed
Maintenance
Related MCP Connectors
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
- OkareoOAuthcom.okareo
Simulation, evaluation and monitoring for voice agents.
Run in-product voice interviews with AI agents and analyze source-linked evidence.
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides a stdio MCP bridge for coding agents to query and record engineering knowledge locally, preserving debugging history, failed attempts, and verified solutions.2 npmMIT
- AlicenseAqualityCmaintenanceProvides local audio analysis tools for LLMs, enabling transcription, conversation dynamics, prosody analysis, and visual inspection without API keys.8MIT
- AlicenseNot gradedqualityAmaintenanceServes offline document validation and evaluation operations of the Judgment Pack Specification to MCP clients over stdio, enabling agents to validate and evaluate JPS documents as tool calls. Supports conformance validation, experimental evaluation with disposition and error classes, and corpus testing.2Apache 2.0
- AlicenseAqualityBmaintenanceEnables an AI agent to run local tools on the user's own machine via stdio, including command execution, workspace file read/write, and system status checks.52MIT