A zero-cost, China-accessible MCP server providing four free AI capabilities: text chat, image generation, speech-to-text, and text-to-speech, ready to use after cloning.
Enables MCP clients to leverage the Intel Arrow Lake NPU for local speech transcription, screenshot OCR, private semantic search, and hardware diagnostics, all processed locally.
MCP server that wraps MiniMax platform APIs (speech, video, image, music, and file management) as tools over stdio, enabling natural language interaction with MiniMax's AI services.
Profile-driven MCP server for Google Cloud Text-to-Speech that exposes tools to synthesize text to speech, run diagnostics, and stop playback, with voice and settings locked per profile.
MCP server providing tools to fetch YouTube video transcripts with metadata, supporting direct YouTube transcripts and audio transcription via multiple backends (whisper, AssemblyAI, OpenAI, Gemini).
Official MCP server for Theta EdgeCloud's On-Demand Model APIs, providing access to 20+ AI models including image generation, audio transcription, and LLMs, directly from MCP-compatible clients.
A low-latency text-to-speech MCP server that uses local Kokoro GGUF inference via TTS.cpp, providing say, get_voices, and get_status tools for AI agents to synthesize speech and manage playback queues.
An MCP server providing tools for speech-to-text, translation, language detection, question answering, and text-to-speech using Sarvam AI models, enabling multilingual voice agents.
Provides various AI capabilities through DeepInfra's OpenAI-compatible API including image generation, text processing, embeddings, speech recognition, object detection, and classification tasks. Enables users to access multiple AI models for diverse tasks like generating images from prompts, transcribing audio, analyzing text sentiment, and performing computer vision operations.
Enables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.
Enables text-to-speech and speech-to-text through MCP tools, using Groq's free hosted endpoints when configured and falling back to fully local keyless models otherwise. Also provides voice listing and provider health checks.
Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
Let your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.
Enables MCP clients such as Claude Code, Codex and Cursor to retrieve timestamped transcripts of public YouTube videos, either from cache, from native captions, or via asynchronous audio-generation jobs that are polled by job ID. It returns the source URL, selected language and millisecond-offset segments without summarizing the video itself.