Speech Processing
Voice interaction and speech processing capabilities. Enables converting speech to text, audio commands, and voice generation.
MCP ServersBrowse all →
AlicenseAqualityCmaintenanceEnables AI agents to generate speech, transcribe audio, and manage voices via the Vocea API.6MIT- AlicenseAqualityBmaintenanceEnables MCP clients such as Claude Code, Codex and Cursor to retrieve timestamped transcripts of public YouTube videos, either from cache, from native captions, or via asynchronous audio-generation jobs that are polled by job ID. It returns the source URL, selected language and millisecond-offset segments without summarizing the video itself.356 npm1MIT

Neuratel MCP Serverofficial
AlicenseAqualityDmaintenanceControl your voice AI platform through natural language from any MCP-compatible assistant.469MIT
CyberWareX MCP Serversofficial
AlicenseAqualityBmaintenancePay-per-call x402 APIs for AI agents: DeFi token safety (honeypot/tax simulation, A-F grade), EVM chain data, web access (markdown/screenshot/PDF), speech-to-text. USDC on Base, no account, no API key3MIT
SeaMeet MCPofficial
AlicenseAqualityDmaintenanceSeaMeet MCP connects Claude, Cursor, Codex, and other AI agents to SeaMeet meeting recordings, transcripts, AI summaries, screenshots, action items, webhooks, and desktop recording controls. Use it to search meeting memory, read synced cloud recordings, and automate meeting notes through the Model Context Protocol.10196 npm1MIT
VoiceStudio MCPofficial
AlicenseBqualityCmaintenanceEnables access to a backend's native MCP tools and broader HTTP API for speech generation, voice cloning/design, transcription, dubbing, translation, subtitles, audiobooks, and related batch/long-form workflows over Streamable HTTP or local stdio.21AGPL 3.0
dtelecom-sttofficial
AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files using dTelecom's real-time speech-to-text with pay-per-use USDC micropayments, no API keys required.310 npm1MIT
jackai-stt-mcpofficial
AlicenseAqualityCmaintenanceTranscribes audio files by referencing them in chat, using OpenAI's speech-to-text models locally without uploading audio, and supports speaker diarization.1MIT- AlicenseAqualityAmaintenanceText to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.468 PyPIMIT
- AlicenseAqualityDmaintenanceOfficial MCP server for Theta EdgeCloud's On-Demand Model APIs, providing access to 20+ AI models including image generation, audio transcription, and LLMs, directly from MCP-compatible clients.418 npm1MIT

movoice-mcpofficial
AlicenseAqualityCmaintenanceCreate and manage voice AI agents using natural language from Claude Desktop or Cursor. Enables agent creation, listing, updating, deletion, and viewing call logs.78 npmMIT- Apache 2.0

Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.224MIT
noisy-codingofficial
AlicenseAqualityAmaintenanceVoice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.412MIT- AlicenseAqualityCmaintenanceEnables AI agents to manage phone numbers, send/receive SMS, and place voice calls through natural language, connecting to the phone network via the AgentPhone API.28446 npm124MIT

mocoVoice MCP Serverofficial
AlicenseAqualityCmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.63MIT
Speak AI MCP Serverofficial
AlicenseAqualityBmaintenanceConnects Speak AI transcription and insight data to Claude and ChatGPT, enabling natural language queries for summaries, action items, and quotes from recordings.1002,226 npmMIT
Equalang Translationofficial
AlicenseAqualityBmaintenanceTranslate PDF, Word (DOCX), Excel (XLSX) and PowerPoint (PPTX) files, images, subtitles and text. Turn audio and video into transcripts or translated subtitles. Preserve layout where supported. Requires an Equalang API key and credits.861 npm1Apache 2.0- AlicenseAqualityFmaintenance提取抖音无水印视频链接,视频文案,douyin-mcp-server,mcp,claude skill,支持龙虾123269 PyPI1,276Apache 2.0

ElevenLabs MCP Serverofficial
AlicenseBqualityFmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,534MIT
Anam MCP Serverofficial
AlicenseBqualityDmaintenanceEnables managing AI personas, avatars, voices, and sessions from any MCP client, for integration with Anam AI.5430MIT
@speechweave/mcpofficial
AlicenseAqualityBmaintenanceMCP server for SpeechWeave transcription, enabling AI assistants to transcribe local files and URLs via wait-first or async tools.665 npmMIT
TypeWhisper MCPofficial
AlicenseAqualityDmaintenanceConnects to the TypeWhisper macOS app to let coding agents transcribe local files, inspect model status, search history, and manage dictionary terms and corrections.1010 npm2GPL 3.0
ContextPulseofficial
AlicenseAqualityBmaintenanceLets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.36AGPL 3.0
@vidofy/mcpofficial
AlicenseAqualityBmaintenanceEnables generating images, video, audio, and speech from MCP clients using your own Vidofy account, with access to hundreds of models for text-to-video, image-to-video, image editing, lipsync, text-to-speech, and voice cloning.982 npm2MIT
supertone-mcpofficial
AlicenseAqualityFmaintenanceMCP server for the Supertone TTS API. Generate natural speech, browse and preview the voice catalog, predict synthesis cost, and create cloned voices — directly from Claude Desktop, Cursor, or any MCP-compatible client. Supports Korean, English, Japanese, and 20+ other languages, with speed, pitch, and emotion-style control.1434 PyPI4MIT- AlicenseBqualityDmaintenanceEnables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.12211 npmISC
- AlicenseAqualityDmaintenanceAn MCP server that enables AI agents to analyze videos locally by extracting transcripts, detecting scene changes, and returning key frames.57MIT
- AlicenseAqualityDmaintenanceMCP server integrating VolcEngine for automatic speech recognition (ASR) and text-to-speech (TTS), converting audio to text or text to audio files.22MIT
- AlicenseAqualityCmaintenanceEnables local audio and video transcription via the audiototext engine, returning detected language, full text, and timestamped segments.1MIT
MCP ConnectorsBrowse all →
Transcripts, summaries, chapters and timestamped answers for video links, for AI agents.
Search, browse and read your Contextli voice notes and transcriptions from any MCP client.
Get live AI suggestions for what to say in job interviews, sales calls, and meetings, tailored to your CV, job description, client brief, and notes.
Start with classify_media_route to route long or oversized audio/video by size or duration.
Multilingual AI workspace: 1,000+ prompts, Compare Mode, 35 languages, BYOM. Free trial.
Podcast post-production for an AI agent: transcribe a recording, cut fillers and pauses, clean up the voice, add licensed music that ducks under speech, set chapters and export a finished episode.
Search, read and reply to your Telegram Business chats, transcribed voice included.
Give your AI assistant a real phone line. Place and end real phone calls with AI voice agents, read call transcripts, run a live two-way interpreter between two people who share no language (31 languages, browser link or phone), create and edit voice agents, and search, buy and bind phone numbers in 21 countries. OAuth 2.1 with PKCE — the model never sees your API key; a read-only scope is available. Pay as you go from $0.10/min, $5 free credit for new accounts.
Meeting transcripts for AI agents: search calls, read who said what, transcribe files and links.
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
Spaced-repetition flashcards your AI writes, quizzes you on by voice, and schedules with FSRS.
Summarize and transcribe videos, audio, documents and web pages; subtitles; search your library.
AI phone secretary: place calls, read transcripts, list calls, agents, and stats.
AI voice generation: text-to-speech and voice cloning from any MCP client.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
Podcast transcripts as clean Markdown with real speaker names — via the Spoken API.
One key, 100+ models: chat with any LLM, generate images and speech. Live model list. Free trial.
Pressure-test a startup idea in a live voice interview with an AI mentor; get a GO/NO-GO verdict.
Text to speech in 149 languages: MP3 links from any assistant. Free without an account.
Turn a recording you own into a timestamped transcript with SRT and VTT captions and clips.