MCP server exposing NaN API media tools (image generation/editing, text-to-speech, speech-to-text, embeddings, reranking) for any MCP-compatible client.
Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.
Local MCP voice coach that provides English pronunciation, grammar, and fluency feedback from microphone recordings. It supports free-form voice conversation, focused practice drills, phoneme-level feedback, prosody hints, and learner-profile guidance.
Integrates OpenAI APIs into MCP-compatible clients, providing tools for text generation, chat completions, model discovery, image creation/editing, audio transcription, speech synthesis, embeddings, and content moderation.
MCP server that lets Claude place outbound telephone calls via ElevenLabs Agents over a SIP trunk and report the outcome including status, success evaluation, summary, and transcript.
Enables AI agents to download, transcribe, and inspect video or audio URLs from YouTube, TikTok, X, and 1000+ other sites using server-side yt-dlp, residential proxies, and speech-to-text.
Enables AI agents to interact with the ParrotScribe transcription service on macOS, providing tools to start/stop transcription, retrieve real-time and historical transcripts, and search across sessions.
Enables transcription, summarization, and action item extraction from audio files on your Mac using MacWhisper and Claude Desktop, all locally without any cloud APIs.
Bitcoin-powered AI tools via Lightning Network micropayments (L402). Image generation, text generation, video, music, speech, 3D models, file conversion, and SMS — no signup or API keys required.
Local speech-to-text transcription using Microsoft's VibeVoice-ASR model with speaker diarization, enabling audio transcription directly in AI tools like Claude Code, Cursor, and OpenCode.
Enables AI assistants to transcribe audio and video from URLs or local files with high accuracy, speaker diarization, 119 languages, and word-level timestamps, while also supporting transcription management and caption export in SRT, WebVTT, or plain text.
Official MCP server for Theta EdgeCloud's On-Demand Model APIs, providing access to 20+ AI models including image generation, audio transcription, and LLMs, directly from MCP-compatible clients.
Provides local audio transcription using whisper.cpp, supporting multiple models and audio formats. Enables transcription of audio files via MCP tools with optional timestamps.