Local-only transcription server for MCP agents, powered by CrispASR. Transcribes audio/video files without cloud uploads, supporting English and Chinese.
Enables AI agents and applications to send audio notifications with text-to-speech, message streaming, agent-to-agent conversations, and web push notifications through a persistent message store and MCP integration.
Enables AI coding assistants to verbally announce task progress and events through local text-to-speech, with cloud synthesis via Volcano Engine or Edge TTS, scene-based voice settings, and a Windows SAPI fallback so announcements continue even when the cloud fails.
Enables MCP clients to run project-scoped video editing workflows: propose and approve editing strategies, apply validated plans, review immutable versions, and export final renders via FFmpeg, with durable persistence and approval gates.
Provides a catalog of paid micro-work tools for text processing, speech, and image generation with fixed USDC pricing via x402. Enables agents to discover capabilities, get quotes, and prepare calls without handling wallet keys.
Provides local vision and audio perception for MCP-compatible agents, enabling them to read images, transcribe text from visual media, analyze videos, and convert speech to text entirely on-device. It is privacy-focused with no cloud upload or API keys by default, using Ollama for inference.
Enables automated video learning workflows by ingesting video URLs, managing remote GPU ASR transcription, pulling transcripts, generating digest summaries, and searching local notes via MCP tools.
An MCP server that enables voice-to-voice AI conversations using ElevenLabs for speech synthesis and recognition, with tools for voice management, text-to-speech, and speech-to-text.
Transcribe local audio/video to text + SRT/VTT with hosted Whisper large-v3, plus summaries, study notes, flashcards, document Q&A, and PDF compression. No API key required — runs via npx -y vocce-transcribe-mcp.
A self-hosted MCP server that connects AI clients to local media-generation backends like ComfyUI and Blender, and provides built-in utilities for video/audio analysis, editing, subtitles, speech, and scene detection via a persistent file cache.
A multi-agent human-computer interaction system that enables natural interaction through integrated visual recognition, speech recognition, and speech synthesis capabilities.
A FastAPI backend that integrates with ElevenLabs MCP to create an AI agent capable of making outbound calls and delivering friendly, jargon-free tech news updates.
An MCP server providing tools for speech-to-text, translation, language detection, question answering, and text-to-speech using Sarvam AI models, enabling multilingual voice agents.
Let your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.
Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
Extract structured knowledge from voice recordings. Transcribes audio using Mistral's Voxtral model and lets your LLM agent handle post-processing within your existing setup.