Enables LLMs to generate images via MCP by calling AI models from providers like SiliconFlow, OpenAI, or custom APIs, with tools for image generation and model listing.
Provides a audio/video creation toolbox via MCP protocol, enabling natural language-based video editing tasks such as image-to-video, video merging, subtitle extraction, and more.
Exposes text-to-audio sound effect generation as an MCP tool, allowing clients like Claude Desktop to generate sound effects locally using a diffusion model, with support for AMD ROCm and Apple Silicon.