Read-only MCP server for Telegram that enables reading messages and transcribing voice, audio, and video notes via the Telegram API for use with Codex and Claude Code.
Turns a YouTube video or allowlisted local video into a timestamped transcript, chronological timeline, and retrievable image resources for transparent media preprocessing.
Self-hosted WhatsApp management over Model Context Protocol, exposing a streamable HTTP MCP endpoint with 30 tools for session/QR pairing, messaging, media storage, and optional on-CPU voice note transcription via whisper.cpp.
Enables watching videos with an AI coding agent that transcribes, prepares concepts and visuals, and answers questions in context of the video. Provides MCP tools for transcript reading, concept mapping, artifact building, and interactive Q&A synchronized with playback.
This service provides fast and reliable transcriptions for audio/video files and voice memos. It allows LLMs to interact with the text content of audio/video file.
Local-first meeting capture and transcription for Claude Code. Records audio from meeting apps, transcribes locally with whisper.cpp, and produces structured notes via Claude.
Turns Craig Discord recordings and Foundry VTT chat logs into speaker-labelled transcripts, player recaps, combat reports, GM notes, and party snapshots for a D&D 5e campaign repository.
Enables seamless integration with ElevenLabs Conversational AI to manage agents, tools, and knowledge base sources. It supports RAG indexing, webhook integration, and document management for building advanced voice-enabled AI agents.
Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
Let your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.
Enables AI video dubbing by letting users submit video URLs for dubbing into Russian, English, or Spanish while preserving original speakers' voices through voice cloning, plus job status and cost estimation tools.
Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
MCP server exposing NaN API media tools (image generation/editing, text-to-speech, speech-to-text, embeddings, reranking) for any MCP-compatible client.
Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.
MCP server for local speech-to-text using Whisper Large V3 (MLX), enabling audio transcription with text/timestamps/SRT output and LLM-based correction, all running offline on Apple Silicon.
Local MCP server that adds multimodal capabilities to text-only models like Codex/DeepSeek, offering tools for image description, audio transcription, video analysis, image/video generation, and speech synthesis.