Speech Processing
Voice interaction and speech processing capabilities. Enables converting speech to text, audio commands, and voice generation.
MCP ServersBrowse all →
AlicenseAqualityAmaintenanceVoice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.245MIT
@vocea.app/mcp-serverofficial
AlicenseAqualityCmaintenanceEnables AI agents to generate speech, transcribe audio, and manage voices via the Vocea API.6MIT
mocoVoice MCP Serverofficial
AlicenseAqualityBmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.63MIT- AlicenseAqualityAmaintenanceMCP server that converts video links into AI-generated Markdown notes, with tools for task management, transcription engines, and LLM providers.22226MIT
- AlicenseAqualityAmaintenanceGive your AI agent a voice with x402 pay-per-call speech synthesis, offering 20 voices, 10 personas, 31 languages, and granular controls.4626MIT

Anam MCP Serverofficial
AlicenseBqualityCmaintenanceEnables managing AI personas, avatars, voices, and sessions from any MCP client, for integration with Anam AI.5430MIT
ElevenLabs MCP Serverofficial
AlicenseAqualityAmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,530MIT
Speak AI MCP Serverofficial
AlicenseAqualityBmaintenanceConnects Speak AI transcription and insight data to Claude and ChatGPT, enabling natural language queries for summaries, action items, and quotes from recordings.100494MIT
Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.224MIT
ContextPulseofficial
AlicenseAqualityBmaintenanceLets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.361AGPL 3.0
@speechweave/mcpofficial
AlicenseAqualityBmaintenanceMCP server for SpeechWeave transcription, enabling AI assistants to transcribe local files and URLs via wait-first or async tools.6564MIT
SeaMeet MCPofficial
AlicenseAqualityBmaintenanceSeaMeet MCP connects Claude, Cursor, Codex, and other AI agents to SeaMeet meeting recordings, transcripts, AI summaries, screenshots, action items, webhooks, and desktop recording controls. Use it to search meeting memory, read synced cloud recordings, and automate meeting notes through the Model Context Protocol.1034MIT- AlicenseAqualityAmaintenanceText to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.4MIT

supertone-mcpofficial
AlicenseAqualityFmaintenanceMCP server for the Supertone TTS API. Generate natural speech, browse and preview the voice catalog, predict synthesis cost, and create cloned voices — directly from Claude Desktop, Cursor, or any MCP-compatible client. Supports Korean, English, Japanese, and 20+ other languages, with speed, pitch, and emotion-style control.145MIT
jackai-stt-mcpofficial
AlicenseAqualityCmaintenanceTranscribes audio files by referencing them in chat, using OpenAI's speech-to-text models locally without uploading audio, and supports speaker diarization.1MIT
Neuratel MCP Serverofficial
AlicenseAqualityBmaintenanceControl your voice AI platform through natural language from any MCP-compatible assistant.469MIT
TypeWhisper MCPofficial
AlicenseAqualityCmaintenanceConnects to the TypeWhisper macOS app to let coding agents transcribe local files, inspect model status, search history, and manage dictionary terms and corrections.10131GPL 3.0- AlicenseAqualityDmaintenanceManage voice AI agents from Claude Code, Cursor, VS Code, or any MCP-compatible assistant.3103MIT
- AlicenseBqualityCmaintenanceProvides VOICEVOX text-to-speech as an MCP tool. Requires a running VOICEVOX engine on localhost.142572Apache 2.0
- AlicenseBqualityDmaintenanceEnables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.213ISC
- AlicenseAqualityCmaintenanceMCP server for local speech-to-text using Whisper Large V3 (MLX), enabling audio transcription with text/timestamps/SRT output and LLM-based correction, all running offline on Apple Silicon.2MIT
- AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files from URLs or local paths using AssemblyAI's services, with support for speaker diarization, language detection, and asynchronous job management through a standardized MCP interface.4132MIT
- AlicenseAqualityDmaintenanceMCP server that synthesizes Claude Code responses into Japanese speech using VOICEVOX, enabling audible feedback during development.31MIT
- AlicenseAqualityDmaintenanceAn MCP server that enables AI agents to analyze videos locally by extracting transcripts, detecting scene changes, and returning key frames.56MIT
- AlicenseBqualityDmaintenanceProvides intelligent transcript processing capabilities for Claude, featuring natural formatting, contextual repair, and smart summarization powered by Deep Thinking LLMs.420MIT
- AlicenseAqualityDmaintenanceMCP Server for automated conversational phone calls using Asterisk with Speech-to-Speech capabilities, allowing users to make phone conversations as easily as writing a prompt.9126MIT
- AlicenseAqualityCmaintenanceEnables AI agents to transcribe audio and video with speaker labels, timestamps, and captions via Pepys API.954MIT
- AlicenseAqualityBmaintenanceLet your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.317Apache 2.0
- AlicenseAqualityAmaintenanceTranscribes videos from 1000+ platforms (YouTube, TikTok, Vimeo, etc.) and local video files using OpenAI's Whisper model, with support for 90+ languages and multiple output formats.8574MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that provides text-to-speech functionality for AI agents using Microsoft Edge's text-to-speech technology, supporting multiple voices, languages, and voice customization.28MIT
MCP ConnectorsBrowse all →
Generate highly realistic Text to Speech voiceovers.
Voice-powered bug reporting with 13 MCP tools. Record bugs by talking; let AI find and fix them.
ElevenLabs in natural language: generate speech in any language, create and manage voices, compose m
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Any video URL to LLM-ready transcript. ASR built in, no captions needed. TikTok, X, TED and more.
Your AI rings your iPhone, speaks its question, and gets your spoken answer back as text.
Construction daily-log generation, jurisdiction compliance requirements, construction FAQs.
Free IELTS prep: band-scored student essays and interactive Listening/Reading drills graded in-chat.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Kurdish (Sorani & Kurmanji) text-to-speech & speech-to-text — 664 AI voices. API key required.
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
Create, inspect, and manage Wubble music, speech, voice, and sound-effect requests through MCP.
AI voice agents: assistants, calls, campaigns, leads, knowledge bases, WhatsApp, SMS & SIP trunks.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
AI voice agents on SMB websites — fully autonomous build in 2–3 min. 23 MCP tools. EU, GDPR.
YouTube video search with transcript extraction as first-class output.
AI phone secretary: place calls, read transcripts, list calls, agents, and stats.