Enables MCP clients to leverage the Intel Arrow Lake NPU for local speech transcription, screenshot OCR, private semantic search, and hardware diagnostics, all processed locally.
An MCP server providing tools for speech-to-text, translation, language detection, question answering, and text-to-speech using Sarvam AI models, enabling multilingual voice agents.
Provides accurate meeting transcription with speaker diarization and multilingual support, allowing users to submit audio URLs, poll transcription status, get transcripts, and summarize via MCP tools in their IDE.
Provides various AI capabilities through DeepInfra's OpenAI-compatible API including image generation, text processing, embeddings, speech recognition, object detection, and classification tasks. Enables users to access multiple AI models for diverse tasks like generating images from prompts, transcribing audio, analyzing text sentiment, and performing computer vision operations.
Enables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.
Enables AI agents to provision phone numbers, send SMS, place AI voice calls, and react to inbound events via the Dial communication stack, all through MCP tools.
Enables supported local AI clients to control a visible desktop presence through bounded presentation tools such as animations, window control, and short local speech over a loopback endpoint, without granting decision or execution authority.
Self-hosted WhatsApp management over Model Context Protocol, exposing a streamable HTTP MCP endpoint with 30 tools for session/QR pairing, messaging, media storage, and optional on-CPU voice note transcription via whisper.cpp.
Enables AI agents to interact with a personal Telegram account via MCP, offering tools for reading and searching chats, summarizing unread messages, sending and editing messages, handling files and voice transcripts, and managing groups and channels, with read-only mode by default.
Provides translation and language detection tools to AI agents, processing text, audio, and Google Meet recordings with emotional voice style preservation via Google's Gemini Live API.
Gives your AI a live, speaker-labeled transcript of the meeting or call happening right now, plus the ability to push advice into the meeting window and speak out loud on the Mac. Requires the VoxAI macOS app — the server reads and writes that app's local files, so tools only return real data on macOS.
Lets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.