voice-ingest
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@voice-ingestTranscribe the file meeting.m4a and show me the markdown transcript."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Voice Ingest
Long-form audio transcription for developers and AI agents.
Upload once. Submit a durable job. Retrieve structured transcripts from your terminal, Python code, or MCP client.
English · 简体中文
Quickstart · CLI · Python SDK · MCP · Deployment · Architecture
Voice Ingest handles the work around cloud ASR: resumable uploads, asynchronous jobs, restart recovery, and consistent exports. It is built for personal use and trusted teams sharing an API-key-protected workspace.
Why Voice Ingest?
Long recordings, bounded memory. Upload in 16 MiB parts with four concurrent parts per file; validate media with ffprobe before recognition.
Jobs survive client disconnects. PostgreSQL persists progress and provider task IDs. Workers use leases and execution generations to recover work safely.
One workflow, four interfaces. HTTP, CLI, an async Python SDK, and FastMCP 4.0.2 with MCP Python SDK v2 address the same jobs.
Results you can reuse. Keep raw provider output and normalized JSON; export TXT, Markdown, SRT, and VTT. Missing timestamps stay missing.
Explicit retry and billing behavior. Idempotency keys prevent duplicate requests. Uncertain provider submissions require attention instead of automatic resubmission.
Develop without cloud credentials. The mock provider exercises the workflow without recognition charges or provider network calls.
Related MCP server: Talky Talky
Quickstart
From a local checkout, use Docker Compose for the backend and Python 3.12 + uv for CLI/SDK development. No local GPU is required.
1. Start the backend
cp .env.example .env
# Edit .env: replace the API key, database password, and S3 credentials.
docker compose --env-file .env -f deploy/compose.yaml up -d --build
curl --fail http://127.0.0.1:18080/health/readyThis starts a separate API, worker, PostgreSQL, and MinIO stack. The default local ports are 18080 (API) and 19000 (S3); 80/443 are unused.
The default provider is mock. Its output is labeled [MOCK] and does not transcribe the contents of your recording. The health request should return HTTP 200 once initialization completes.
2. Submit a recording
uv sync --all-extras --frozen
export VOICE_URL=http://localhost:18080
export VOICE_API_KEY='the-service-key-from-your-dotenv'
uv run voice-ingest transcribe meeting.m4a --wait --format markdownReplace meeting.m4a with your audio file. --wait polls the job and prints Markdown when it succeeds. Omit it to return the job ID immediately; interrupting a wait does not cancel server-side work.
Validation status: Local tests and a real Aliyun recording test have passed. A complete Compose startup has not yet been validated; the last image build was interrupted by dependency download timeouts. See the acceptance record.
Providers
Provider / model | Available behavior | Verification |
Mock | Complete upload/job/export workflow; synthetic text | Offline and PostgreSQL/MinIO integration tests |
Aliyun | Default whole-file asynchronous ASR | Real 87-minute recording completed |
Aliyun | Explicit model selection | Adapter contract tests; no live acceptance yet |
The current model checks allow files up to 12 hours / 2 GB. Speaker diarization is rejected above two hours. Language hints, diarization, and context support depend on the model; inspect voice-ingest models or /v1/models for capabilities. Files are sent whole, without automatic compression or VAD splitting.
To enable Aliyun, edit .env and recreate the API and worker:
VOICE_PROVIDER=aliyun
VOICE_ALIYUN_REGION=beijing
VOICE_ALIYUN_API_KEY=YOUR_REGIONAL_DASHSCOPE_KEY
VOICE_S3_PUBLIC_ENDPOINT=https://files.example.comIn the default signed_url mode, the file endpoint must actually route to your S3 service and be reachable by clients and Aliyun. localhost cannot serve cloud recognition. HTTP and HTTPS origins are supported; use HTTPS for public deployments. Internal storage access and public signing endpoints are configured separately.
For local real-ASR evaluation without public S3, set VOICE_ALIYUN_SOURCE_MODE=temporary_upload and restart the API and worker. The worker uses Aliyun’s official temporary file service; this mode is for local evaluation only. Production keeps the default signed_url mode. See local browser setup.
Billing: This adapter uses regular DashScope ASR. Token Plan / Coding Plan is not integrated, and there is no automatic fallback between billing channels. VOICE_API_KEY protects your backend; VOICE_ALIYUN_API_KEY authenticates the backend to Aliyun. See deployment and credential configuration.
Web workspace
The optional React frontend provides a transcription workspace with separate upload and recognition steps, status filters, transcript search and five export formats. English is the default; Chinese is available in the sidebar.
cd web
npm ci
npm run devOpen http://127.0.0.1:5174 and connect your workspace with its service key, or explicitly explore a sample transcript without credentials. Uploading stops at a review step; Start transcription submits the job. See the web guide for a customer walkthrough, storage CORS, deployment and tests.
CLI
uv run voice-ingest transcribe meeting.m4a
uv run voice-ingest batch ./recordings --recursive --resume
uv run voice-ingest --json jobs list
uv run voice-ingest jobs get JOB_ID
uv run voice-ingest jobs cancel JOB_ID
uv run voice-ingest export JOB_ID --format srt --output meeting.srtUse --model fun-asr on transcribe or batch to select the other Aliyun model.
Single-file transcription resumes by default and reuses the local job record for the same file and options. Use --no-resume for a fresh submission. Batch resume is explicit; one file's failure does not stop the remaining files.
--json is a global option and goes before the subcommand. Progress goes to stderr. Resume state lives under ~/.local/state/voice-ingest/; credentials are not stored there.
Exit code | Meaning |
| Success |
| Network or service error |
| Invalid arguments or local file |
| Job failure or partial batch failure |
| User interruption |
Python SDK
For SDK-only use, install from the checkout with uv pip install . in an activated environment. The default package depends on HTTPX and Pydantic, without database, HTTP server, or MCP dependencies. Optional extras are cli, server, and mcp.
import asyncio
import os
from voice_ingest import AsyncVoiceClient, TranscriptionOptions
async def main():
async with AsyncVoiceClient("http://localhost:18080", os.environ["VOICE_API_KEY"]) as client:
asset = await client.upload("meeting.m4a")
job = await client.submit(
asset.id,
options=TranscriptionOptions(language_hints=["zh"], diarization=True),
idempotency_key="meeting-001-v1",
)
job = await client.wait(job.id)
if job.state == "succeeded":
transcript = await client.result(job.id)
print(transcript.text)
else:
print(job.model_dump_json())
asyncio.run(main())Reuse an idempotency key when retrying the same submission; use a new key for changed parameters. A job in needs_attention may already have been accepted by the provider. A cancellation with remote_may_run=true means the provider may still execute and charge for recognition.
MCP
Recipe: cloud backend, local recording
Transcribe the meeting on my laptop. Prepare Markdown and SRT exports, and keep the job ID so I can return later.
Run a lightweight local MCP bridge to upload files to your cloud backend, or upload in the web workspace and let a remote agent continue with the existing job. The backend owns the long-running work.
Follow the recipe → — client configuration, upload flow, example tool calls, and authenticated downloads. Includes the boundaries for cloud-hosted chat attachments.
Connect a remote client
Use http://localhost:18080/mcp/ with Authorization: Bearer YOUR_VOICE_INGEST_API_KEY. Configure the URL and header using your client's HTTP MCP settings.
Task | Tools |
Discover models |
|
Manage jobs |
|
Read and export |
|
Submissions return durable business job IDs immediately. MCP Tasks support and a persistent MCP connection are not required. Transcript reads support pagination and time ranges to keep long recordings out of a single tool response.
submit_transcription only requires asset_id. MCP automatically deduplicates the same asset and
normalized options across reconnections, including completed jobs. An optional explicit
idempotency_key supports intentional new recognition; reuse it for retries of that request.
Use retry_transcription to retry a failed job. HTTP/SDK idempotency contracts are unchanged.
Upload local files from an agent
After uv sync --all-extras --frozen, configure a local stdio bridge in a client that supports mcpServers configuration:
{
"mcpServers": {
"voice-ingest": {
"command": "uv",
"args": [
"run", "--directory", "/absolute/path/to/voice-ingest",
"voice-ingest-mcp", "--url", "http://localhost:18080",
"--allow-dir", "/absolute/path/to/recordings"
],
"env": {"VOICE_API_KEY": "YOUR_VOICE_INGEST_API_KEY"}
}
}
}Replace both absolute paths and the service key; uv must be on the client's PATH. The backend must already be running. Other clients may use a different configuration schema.
The bridge adds upload_local_audio and checks resolved paths against explicitly allowed directories, including symlink boundaries. It uploads the file to the backend and returns an asset ID for a subsequent transcription submission. Remote MCP does not accept paths on your computer.
HTTP API
All business routes use /v1 and Bearer authentication. Creation of a transcription requires Idempotency-Key and returns 202 Accepted.
Resource | Purpose |
| Create, inspect, sign parts, complete, or abort an upload |
| Inspect or delete source audio |
| Query model capabilities |
| Submit jobs and list them with cursor pagination |
| Inspect a job or delete its results |
| Cancel or retry explicitly |
| Read normalized results and exports |
Download the OpenAPI schema for exact methods and request bodies:
curl -H "Authorization: Bearer $VOICE_API_KEY" \
"$VOICE_URL/openapi.json" -o openapi.json/docs also requires authentication. /health/live and /health/ready are public process/readiness checks; /metrics requires the service key.
How it works
flowchart LR
CLI[CLI / Python SDK] --> API[HTTP API]
Agent[Agent] --> MCP[Remote MCP]
Agent --> Bridge[Local MCP bridge]
Bridge --> API
MCP --> Service[Transcription service]
API --> Service
CLI -->|Presigned upload| S3[(Private S3)]
Service --> PG[(PostgreSQL jobs)]
PG --> Worker[Worker]
Worker --> Provider[Aliyun / Mock]
Worker --> S3
Provider -->|Signed audio URL| S3Uploading creates a stable asset_id; transcription creates a separate job_id. PostgreSQL owns job state, while private S3 stores audio, raw results, and exports. HTTP and remote MCP share business use cases; CLI and the local MCP bridge reuse the SDK.
Workers claim due jobs with SKIP LOCKED, run network operations outside transactions, and fence writes by lease ownership and generation. After a restart, a saved provider task ID is polled again. A lost submission response becomes needs_attention. Success requires checking file-level status, downloading the result, and normalizing it.
Development
uv sync --all-extras --frozen
make check
uv buildmake check runs Ruff, format checks, Pyright, and offline tests. Native backend development also requires ffprobe, PostgreSQL, and S3; run migrations before starting the API and worker. See the deployment guide.
For real PostgreSQL/MinIO tests, configure dedicated test resources and run make integration. These tests use the mock ASR provider. Cloud recognition is a separate, explicitly billable acceptance step.
Read AGENTS.md before coding. Modules are organized by capability (transcription, media, providers, jobs, exports) with thin interfaces and shared runtime wiring. Keep changes within these boundaries, add behavior tests for changed contracts or recovery semantics, and keep both READMEs in sync.
Scope and documentation
The current release focuses on offline ASR for a shared, trusted workspace. Realtime recognition, TTS, multi-tenancy, and automatic knowledge-base ingestion are outside the current implementation. TTS is a future direction, without a committed release date.
Document | Contents |
Configuration, private storage, HTTP/HTTPS, credentials, operations | |
Module boundaries, persistence, recovery, and tradeoffs | |
Durable jobs and capability-oriented organization | |
Tested behavior and remaining validation work (Chinese) | |
Repository rules for contributors and coding agents |
License
MIT. You may use, modify, distribute, and use this project commercially, provided you retain the copyright and license notice. Third-party dependencies and cloud services remain subject to their own licenses and terms.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
No tool schema history has been recorded yet.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
- mcpOAuthso.transcribe
Transcribe audio and video into speaker-labelled transcripts, subtitles, clips, and cited Q&A.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files from URLs or local paths using AssemblyAI's services, with support for speaker diarization, language detection, and asynchronous job management through a standardized MCP interface.4252MIT
- AlicenseNot gradedqualityDmaintenanceA comprehensive audio MCP server that enables AI agents to generate speech, transcribe audio, clone voices, analyze speech quality, design soundscapes, and manage audio assets through a standardized interface.2MIT
- AlicenseAqualityCmaintenanceProvides voice transcription control and polling for MCP-compatible agents, enabling start/stop/pause/resume and retrieval of new text via tools.8MIT
- AlicenseAqualityBmaintenanceEnables automated audio restoration, transcription, and speaker diarization via MCP tools for queuing files, monitoring progress, and retrieving speaker-labeled transcripts.9MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/yuehong136/voice-ingest'
If you have feedback or need assistance with the MCP directory API, please join our Discord server