Speech Processing
Voice interaction and speech processing capabilities. Enables converting speech to text, audio commands, and voice generation.
MCP ServersBrowse all →
AlicenseAqualityDmaintenanceLets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.Last updated361AGPL 3.0
Anam MCP Serverofficial
AlicenseBqualityFmaintenanceEnables managing AI personas, avatars, voices, and sessions from any MCP client, for integration with Anam AI.Last updated5429MIT
@vocea.app/mcp-serverofficial
AlicenseAqualityCmaintenanceEnables AI agents to generate speech, transcribe audio, and manage voices via the Vocea API.Last updated6MIT
Speak AI MCP Serverofficial
AlicenseAqualityBmaintenanceConnects Speak AI transcription and insight data to Claude and ChatGPT, enabling natural language queries for summaries, action items, and quotes from recordings.Last updated100348MIT
SeaMeet MCPofficial
AlicenseAqualityCmaintenanceSeaMeet MCP connects Claude, Cursor, Codex, and other AI agents to SeaMeet meeting recordings, transcripts, AI summaries, screenshots, action items, webhooks, and desktop recording controls. Use it to search meeting memory, read synced cloud recordings, and automate meeting notes through the Model Context Protocol.Last updated10249MIT- Apache 2.0

mocoVoice MCP Serverofficial
AlicenseAqualityBmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.Last updated63MIT
TypeWhisper MCPofficial
AlicenseAqualityCmaintenanceConnects to the TypeWhisper macOS app to let coding agents transcribe local files, inspect model status, search history, and manage dictionary terms and corrections.Last updated1078GPL 3.0
ElevenLabs MCP Serverofficial
AlicenseAqualityBmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.Last updated261,487MIT- AlicenseAqualityBmaintenanceGive your AI agent a voice with x402 pay-per-call speech synthesis, offering 20 voices, 10 personas, 31 languages, and granular controls.Last updated4622MIT

supertone-mcpofficial
AlicenseAqualityBmaintenanceMCP server for the Supertone TTS API. Generate natural speech, browse and preview the voice catalog, predict synthesis cost, and create cloned voices — directly from Claude Desktop, Cursor, or any MCP-compatible client. Supports Korean, English, Japanese, and 20+ other languages, with speed, pitch, and emotion-style control.Last updated144MIT- AlicenseAqualityCmaintenanceMCP server that brings ElevenLabs to Claude Code — text-to-speech, sound effects, music generation, voice cloning, speech-to-speech, transcription, and voice isolation. 8 tools for industry-leading AI audio.Last updated8MIT
- AlicenseAqualityAmaintenanceLet your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.Last updated316Apache 2.0
- AlicenseAqualityDmaintenanceA cross-platform MCP server that enables Claude to speak using Microsoft Edge TTS with support for over 300 voices across 50+ languages. It requires no API keys and allows for customization of speech rate, volume, and pitch.Last updated32MIT
- AlicenseAqualityBmaintenanceProvides text-to-speech synthesis using Microsoft Edge's free TTS engine, supporting multiple voices, languages, and audio output options (base64 or file).Last updated3MIT
- AlicenseAqualityCmaintenanceBitcoin-powered AI tools via Lightning Network micropayments (L402). Image generation, text generation, video, music, speech, 3D models, file conversion, and SMS — no signup or API keys required.Last updated491971MIT
- AlicenseBqualityDmaintenanceEnables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.Last updated26ISC
- AlicenseBqualityDmaintenanceProvides intelligent transcript processing capabilities for Claude, featuring natural formatting, contextual repair, and smart summarization powered by Deep Thinking LLMs.Last updated419MIT
- AlicenseAqualityCmaintenanceProvides local audio analysis tools for LLMs, enabling transcription, conversation dynamics, prosody analysis, and visual inspection without API keys.Last updated8MIT
- AlicenseAqualityDmaintenanceA text-to-speech MCP server that enables AI assistants to speak using the VOICEVOX engine with support for multi-character conversations. It features queue management, low-latency streaming via FFplay, and cross-platform playback across Windows, macOS, and Linux.Last updated714916ISC
- AlicenseAqualityBmaintenanceReal-time English ↔ Mandarin Chinese speech translation for Claude. Transcribes audio locally with Whisper, translates via Claude API, and synthesises speech locally with Piper TTS. Pass a WAV file path. Claude handles the rest.Last updated31673MIT
- AlicenseBqualityCmaintenanceLocal MCP server for agents to search, structure, and export authorized WhatsApp Web conversations. Uses Playwright for DOM interaction, supports message search, export, media transcription, and controlled message sending.Last updated18MIT
- AlicenseAqualityCmaintenanceProvides speech recognition (STT) and synthesis (TTS) tools via the Sber SaluteSpeech API, enabling audio transcription and voice generation through natural language.Last updated5861MIT
- AlicenseAqualityAmaintenanceA Windows-native MCP server that lets Claude Desktop transcribe audio files locally using whisper.cpp, with no internet connection required.Last updated12871Sleepycat
- AlicenseAqualityCmaintenanceManage voice AI agents from Claude Code, Cursor, VS Code, or any MCP-compatible assistant.Last updated353MIT
- AlicenseAqualityAmaintenanceGive your AI assistant eyes and ears — analyze any video, audio, or image, entirely on your machine.Last updated21241Apache 2.0
- AlicenseAqualityFmaintenanceProvides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.Last updated31032MIT
- AlicenseAqualityCmaintenanceAI-powered speech tools by Brainiall: pronunciation assessment with phoneme-level feedback, speech-to-text with language detection, and text-to-speech with multiple voices.Last updated41MIT
- AlicenseAqualityDmaintenanceExtracts and formats Bilibili video content into structured text for LLM processing and analysis.Last updated14MIT
- AlicenseBqualityFmaintenanceA server that enables Claude 3.7 and other AI agents to access VOICEVOX-compatible speech synthesis engines (AivisSpeech, VOICEVOX, COEIROINK) through the Model Context Protocol.Last updated111MIT
MCP ConnectorsBrowse all →
AI phone secretary: place calls, read transcripts, list calls, agents, and stats.
Voice-powered bug reporting with 13 MCP tools. Record bugs by talking; let AI find and fix them.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
Voice-to-PDF daily construction reports. Court-ready documentation for subcontractors
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
Create, inspect, and manage Wubble music, speech, voice, and sound-effect requests through MCP.
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
AI voice agents: assistants, calls, campaigns, leads, knowledge bases, WhatsApp, SMS & SIP trunks.
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
AI voice agents on SMB websites — fully autonomous build in 2–3 min. 23 MCP tools. EU, GDPR.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
YouTube video search with transcript extraction as first-class output.
Manage Speko voice-AI agents, sessions, calls, phone numbers, knowledge bases, evals, and docs.
AI voice interviewer: create roles, screen CVs, schedule interviews, read scored reports.
Search recordings, summarize meetings, create clips, and automate workflows from your AI assistant.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
User research workspace to transcribe interviews and turn conversations into insights.
Voice-interview your ideas into LinkedIn posts and X threads via a Digital Brain memory.
Create voice-agent scenarios, pull session analytics, place SIP calls, schedule meeting bots.