Enables agents to analyze long videos by downloading them, extracting transcripts and storyboards, and zooming into specific moments with high-resolution frames and OCR.
Enables AI agents to extract, analyze, and manipulate images, PDFs, video, and audio while conserving context through downscaling, truncation, and frame/page caps.
Local MCP server that compresses images and videos, extracts video frames, and prepares media for AI vision agents by returning file paths instead of inline base64.