voxpipe
Provides audio and video transcription through the ChatGPT (Codex) backend, converting speech to text with support for language and prompt hints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@voxpipeTranscribe this audio file and save the transcript to a .txt file"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
voxpipe
Turn video/audio into text from the command line. Audio is sent to the ChatGPT (Codex) transcription backend, long files are split on silences so nothing is lost, and a pluggable offline backend lets you use local tools such as whisper.cpp.
⚠️ 免责声明(重要)
本项目使用的是未公开的接口(
https://chatgpt.com/backend-api/transcribe),并非 OpenAI 官方产品,与 OpenAI 没有任何关联。
如果你的权益受到侵犯,请联系 d4n.for.sec@gmail.com,我们会在 7 天内处理。
该接口可能随时变更、限流或要求设备校验,使用风险(包括账号被限流/封禁、数据被上传到第三方)由使用者自行承担。
请仅用于你自己有权处理的内容,并遵守当地法律与相关服务条款。
本项目按“原样”提供,不提供任何担保。
Features
Any ffmpeg-readable input, text out (
txt; a directory when segmented).Silence-aware chunking: target ~4 min, hard ceiling 10 min, prefer a cut at a silence.
Fallback chunking: when no silence is available, blind cuts with 20s overlap; output is a directory with per-segment files,
manifest.json, and a best-effortmerged.txt.Read-only auth: never refreshes and never writes
~/.codex/auth.json.Resumable: segmented runs checkpoint progress under
.voxpipe/.Pluggable backends: built-in
chatgpt, plus an offlinecommandbackend.Runs on Bun (required: curl/Python/Node TLS fingerprints are rejected with HTTP 403 by the endpoint).
Related MCP server: Whisper MCP Server
Install
The CLI ships as a small Node launcher plus per-platform prebuilt binaries
(declared as optionalDependencies, so npm installs only the one matching your
OS/CPU):
npm i -g @d4n-sec/voxpipe # global install
npx @d4n-sec/voxpipe --help # one-off, no global installAt runtime the launcher prefers a local Bun runtime
(authoritative source, always in sync with the installed version) and only falls
back to the prebuilt @d4n-sec/voxpipe-<platform>-<arch> binary when bun is not
on PATH.
From source:
git clone https://github.com/d4n-sec/voxpipe.git && cd voxpipe
bun install
bun run bin/voxpipe.ts --helpEither way you need ffmpeg and ffprobe on PATH.
Usage
voxpipe video.mp4 # text to stdout
voxpipe audio.m4a -l zh -o out.txt # language hint, write a file
voxpipe a.mp4 b.wav -o outdir/ # multiple inputs
voxpipe long.mp4 --dry-run # preview the plan, no network
voxpipe long.mp4 --json # newline-delimited JSON events on stdoutFlags
Flag | Description | Default |
| Output file, or directory when segmented | |
| Language hint, e.g. | auto |
| Prompt hint | |
| Custom vocabulary, one term per line (joined into the prompt) | |
| Model name |
|
|
|
|
| Command for the | |
| Chunking strategy: |
|
| Target chunk length |
|
| Hard chunk ceiling |
|
| Overlap when cutting blind |
|
| Silence threshold in dB |
|
| Minimum silence length in seconds |
|
| Total attempts, 1–3 (1 = no retry) |
|
| Newline-delimited JSON events on stdout (progress → stderr otherwise) | |
| Print the segmentation plan (reflects the backend and | |
| Keep intermediate audio | |
| Keep resume state after a successful run | |
| Config file |
|
| Help |
JSON output
With --json, stdout is newline-delimited JSON: every line is exactly one
JSON object and nothing else is written to stdout (human-readable progress goes
to stderr when --json is absent). Progress events carry a type; the final
line for each input is an input event.
| Emitted when | Key fields |
| the input has been decoded and measured |
|
| the segmentation plan is ready |
|
| a segment upload begins |
|
| a segment finishes |
|
| a failed segment is retried |
|
|
|
|
| a run finishes (progress event) |
|
| final result for each input |
|
voxpipe talk.mp3 --json | while IFS= read -r line; do
printf '%s\n' "$line" | node -e 'JSON.parse(require("fs").readFileSync(0,"utf8"))'
doneCommands
voxpipe clean [--dir <path>]— remove.voxpipe/state directories (default: cwd).voxpipe mcp --out-dir <path> [--http] [--host 127.0.0.1] [--port 8765]— MCP server (stdio by default, Streamable HTTP with--http; refuses to start without--out-dir).voxpipe serve --out-dir <path> [--host 127.0.0.1] [--port 8787]— HTTP API for transcription (refuses to start without--out-dir).
MCP server
Runs an MCP server on top of the official @modelcontextprotocol/sdk, exposing two tools and streaming progress notifications while a transcription runs.
--out-dir is required at startup (same rule as serve); without it the process prints the usage and exits with code 2, so segmented output can never land in the MCP client's working directory.
voxpipe mcp --out-dir /var/voxpipe/out # stdio transport (for MCP clients that spawn a process)
voxpipe mcp --out-dir /var/voxpipe/out --http # Streamable HTTP at http://127.0.0.1:8765/mcp
voxpipe mcp --out-dir /var/voxpipe/out --http --host 0.0.0.0 --port 9000Tools:
transcribe— args{ path, language?, prompt?, model?, backend?, command?, chunking?, outDir? }. Calls the coretranscribe(); returns the transcript for single/joined runs, or{ mode: "segmented", outDir, segments, files, manifest?, merged? }for segmented runs. WhenoutDiris omitted it uses the server's--out-dir; a per-calloutDiroverrides it.transcribe_plan— args{ path, targetSeconds?, maxSeconds?, overlapSeconds?, minSegmentSeconds?, silenceWindowFraction?, silenceDb?, silenceDur? }. Returns the segmentation plan frompreviewInput()and never contacts the API.
While transcribe runs, the server emits notifications/progress mapped from the core ProgressEvents (with total = segment count once known) whenever the client requested progress. Errors are returned as MCP tool errors; an auth failure tells you to run codex login.
Example client entry (stdio):
{ "mcpServers": { "voxpipe": { "command": "voxpipe", "args": ["mcp"] } } }HTTP API
voxpipe serve --out-dir /var/voxpipe/out [--host 127.0.0.1] [--port 8787]--out-dir is required at startup; without it the process prints the usage and exits with code 2, so segmented output can never land in the server's cwd. A request may still pass outDir to override the server default for that request.
Never exposes tokens, binds to 127.0.0.1 by default, and shuts down cleanly on SIGINT/SIGTERM.
GET /healthz→{ "ok": true, "version": "0.1.0" }.POST /transcribe— accepts eithermultipart/form-datawith afilepart (optionallanguage,prompt,model,backend,command,chunking,outDir) or JSON{ path, language?, prompt?, model?, backend?, command?, chunking?, outDir? }.chunkingis one ofauto,none, orsilence; any other value is rejected with HTTP 400. Uploads are written to a temp file and cleaned up; the request body is capped at 200 MB.
Response format is controlled by ?format=:
format=json(default) →{ "mode": "single"|"joined", "text": "..." }, or{ "mode": "segmented", "outDir", "segments", "files", ... }.format=text→ plaintext/plainfor single/joined runs.
Send Accept: text/event-stream (or ?progress=1) to stream server-sent events: one data: <ProgressEvent JSON> per progress event, followed by a final event: result with the outcome (or event: error).
# multipart upload, JSON result
curl -s http://127.0.0.1:8787/transcribe \
-F file=@media/audio.m4a -F language=zh -F backend=command \
-F 'command=sh -c "echo hi"' -F outDir=/tmp/voxpipe-out
# JSON with a server-side path, plain text out
curl -s 'http://127.0.0.1:8787/transcribe?format=text' \
-H 'content-type: application/json' \
-d '{"path":"/tmp/audio.m4a","backend":"command","command":"sh -c \"echo hi\"","outDir":"/tmp/voxpipe-out"}'
# SSE progress stream
curl -N http://127.0.0.1:8787/transcribe \
-H 'Accept: text/event-stream' \
-F file=@media/audio.m4a -F backend=command \
-F 'command=sh -c "echo hi"' -F outDir=/tmp/voxpipe-outChunking and output
Clean silence cuts throughout → a single text (segments joined; CJK-aware, no spaces for zh/ja/ko/yue).
Any blind/overlapped cut → an output directory containing:
seg_0001_000000-000240.txtper segment,manifest.json(time ranges, overlap flag),merged.txtonly when every boundary merges cleanly (longest-suffix overlap removal).
Progress is always rendered to stderr; stdout stays clean for pipes.
Configuration
~/.config/voxpipe/config.toml (flat key = value):
language = "zh"
model = "gpt-4o-transcribe"
backend = "chatgpt"
chunking = "auto"
target_seconds = 240
max_seconds = 600
overlap_seconds = 20
min_segment_seconds = 15
silence_window_fraction = 0.7
silence_db = -35
silence_dur = 0.35
retries = 1
# command = "whisper-cli -m ggml-base.bin -f {file} -l {language}"Precedence: CLI > env (VOXPIPE_*) > config file > defaults. Environment keys mirror the file: VOXPIPE_LANGUAGE, VOXPIPE_MODEL, VOXPIPE_BACKEND, VOXPIPE_COMMAND, VOXPIPE_CHUNKING, VOXPIPE_TARGET_SECONDS, VOXPIPE_MAX_SECONDS, VOXPIPE_OVERLAP_SECONDS, VOXPIPE_MIN_SEGMENT_SECONDS, VOXPIPE_SILENCE_WINDOW_FRACTION, VOXPIPE_SILENCE_DB, VOXPIPE_SILENCE_DUR, VOXPIPE_RETRIES, VOXPIPE_CONFIG.
Backends / plugins
chatgpt(default) — POSTs audio tohttps://chatgpt.com/backend-api/transcribeusing your existing login.command— runs a local command; stdout (trimmed) is the transcript. This is how you plug inwhisper.cpp,faster-whisper, etc. This project bundles no models.
voxpipe talk.mp3 --backend command \
--command "whisper-cli -m models/ggml-base.bin -f {file} -l {language}"Non-zero exit from the command is an error. Placeholders are replaced per argument (quotes are respected).
Chunking policy
The core owns the chunking mechanism (ffmpeg extract/slice/silence/overlap/merge/resume); each backend only declares a policy — a chunking mode plus optional input limits. Users override the policy with --chunking (or VOXPIPE_CHUNKING / chunking in config.toml); an explicit value always wins over the backend default.
Backend | Default chunking | Declared limits |
|
|
|
|
| none |
backend that declares nothing |
| none |
auto— use the backend's declared policy (the default).none— send the whole file in one request; still splits only if the backend declares a byte limit the input exceeds.silence— prefer a cut at a silence near--chunk-seconds, never past--max-seconds(clamped to the backend'smaxInputSeconds), falling back to blind cuts with overlap.
Auth (read-only)
voxpipe never manages your credentials:
Reads
~/.codex/auth.json(or$CODEX_HOME/auth.json) and usestokens.access_token/tokens.account_idas-is.Never refreshes, never writes, never modifies any file.
Missing login, expired JWT, or a rejected session produce an actionable error telling you to run
codex login.Bypass files entirely with
CODEX_STT_TOKEN/CODEX_STT_ACCOUNT_ID.
Known limits / risks
Undocumented, unofficial endpoint; it may change, rate-limit, or require device attestation at any time.
A single request is silently truncated for long audio (observed around 14–15 minutes), which is why files are chunked.
Transcription content passes through a third-party service; mind your privacy.
Development
bun install
bun run typecheck
bun test
bun run build # compiles bin/voxpipe.ts per platform into dist/
bun run pack:platforms # builds + assembles dist/npm/@d4n-sec/voxpipe-<key>/
bun run scripts/pack-platforms.ts bun-linux-x64 # single targetLicense
MIT © d4n-sec. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
Hosted MCP tools for FFmpeg-style video and audio processing through FFMPEG API.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server for Speech-to-Text
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables high-quality transcription and subtitle generation from local media files or URLs using Faster Whisper on local hardware. It supports automatic language detection and integration with MCP clients for seamless speech-to-text workflows.3-
- AlicenseAqualityFmaintenanceProvides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.3792 npm3MIT
- FlicenseNot gradedqualityCmaintenanceMCP server for whisper-based transcription and translation, supporting local stdio and remote HTTP transports with file workflow safety.-
- AlicenseAqualityBmaintenanceEnables automated audio restoration, transcription, and speaker diarization via MCP tools for queuing files, monitoring progress, and retrieving speaker-labeled transcripts.9MIT