Enables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.
Fetches YouTube transcripts and metadata (title, channel, duration) for URLs, using subtitles or on-device Whisper STT when no subtitles are available, enabling chat-based YouTube video analysis.
Enables AI assistants to summarize, take notes on, and answer questions about YouTube, Bilibili, and Xiaohongshu videos by providing subtitles and local speech transcription with timestamps via MCP.