Analyze an hour-long YouTube video sentence by sentence in three minutes, and extract the information you need from the screen. Open-source local MCP server for AI clients.
Turns a YouTube video or allowlisted local video into a timestamped transcript, chronological timeline, and retrievable image resources for transparent media preprocessing.
Enables AI agents on macOS to securely read and search the local Messages database, catch up on missed messages via a persistent inbox, and send texts or files to allowlisted chats, with optional voice note transcription and text-to-speech.
Provides voice recognition and text extraction capabilities with support for both stdio and MCP modes, processing audio files or base64 encoded data and returning structured results with language, emotion, and speaker information.
Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
Lets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.
Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.
Enables text-to-speech and speech-to-text through MCP tools, using Groq's free hosted endpoints when configured and falling back to fully local keyless models otherwise. Also provides voice listing and provider health checks.
Let your AI agent call your phone and talk to you — MCP servers for live, interruptible voice calls + tiered alerts, using free self-hosted pieces (pjsua2 + whisper.cpp + Linphone). No paid telephony, no extra API key.
Turns a YouTube link into a readable, timestamped Markdown transcript by preferring existing subtitles and falling back to local Whisper transcription, with optional Traditional-to-Simplified Chinese conversion and Obsidian-friendly output.
Enables AI video dubbing by letting users submit video URLs for dubbing into Russian, English, or Spanish while preserving original speakers' voices through voice cloning, plus job status and cost estimation tools.
Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.