Enables AI agents to download, convert, and visually analyze videos by providing timestamped frames and locally transcribed speech, all without API keys.
Enables coding agents to turn videos from social platforms or local files into a small set of distinct frames plus a manifest, so they can answer visual questions by reading image paths. Also provides metadata lookup for captions and authors without downloading the video.
Turns a YouTube video or allowlisted local video into a timestamped transcript, chronological timeline, and retrievable image resources for transparent media preprocessing.
Enables local video inspection and offline transcription with timestamped JSON and Markdown output, enforced filesystem boundaries, and reproducible content-addressed caching.
Let AI agents watch videos: local transcripts, speaker labels, scenes, chapters and exact-moment search from any video URL or file. Fully local, no API keys.