Enables analysis of YouTube videos using the Gemini API to generate summaries and answer specific questions via direct URLs. It supports standard videos and shorts, allowing users to interact with video content without requiring manual downloads.
Enables an agent to inspect an audio folder and report what is inside each file — duration, sample rate, channels, codec and bitrate — and to group byte-identical files, identical decoded audio stored in different containers, and near-duplicate candidates, with suggested keeps and explicit differences, all read-only.
Provides Claude with detailed image inspection capabilities, including metadata extraction, histogram analysis, tonal and color analysis, sharpness detection, and more, supporting both standard and RAW formats.
Enables AI models to analyze audio files through numerical fingerprints, pitch tracking, and visual spectrograms without requiring direct audio playback. It provides tools for comparing audio iterations and detecting patterns using token-efficient analysis operations.
Plays built-in Windows system sounds or WAV files to notify users when tasks finish or need attention, using the winsound module. Provides a single play_sound tool with presets and customizable file paths, repeats, and intervals.
An MCP server that parses Douyin share links and performs intelligent content analysis using the Doubao video understanding model. It provides structured outputs including video summaries, categorized outlines, and step-by-step tutorial information.
Enables MCP-compatible agents to access KaiRouter's video generation API, providing tools to list available video models, generate videos (text-to-video or image-to-video) asynchronously, check job status, and list recent jobs.
Enables AI assistants to locally process images with tools for cropping, zooming, enhancement, edge detection, segmentation, and text region extraction, all without external API keys. It uses PIL, OpenCV, and scikit-image for robust image analysis.
Enables AI clients to watch local video files or YouTube/Bilibili and other supported URLs, receiving timestamped transcripts, subtitles, searchable text and keyframe contact sheets. Runs fully offline with local speech recognition and bundled ffmpeg, so no API key or cloud upload is needed.
An MCP server that reads and builds CapCut projects locally, enabling natural language queries about project contents, missing media, and creation of new edits including beat-synced cuts.
An AI-powered automation bridge for Adobe Premiere Pro that enables controlling video edits with natural language and automating workflows through Claude or other AI agents.
Enables AI agents to download, transcribe, and inspect video or audio URLs from YouTube, TikTok, X, and 1000+ other sites using server-side yt-dlp, residential proxies, and speech-to-text.