A
licenseA
qualityC
maintenanceMulti-provider AI video, speech, music, and transcription MCP server enabling video generation, image-to-video, TTS, music creation, and speech-to-text via a unified interface.
3
MIT
No user-submitted related servers found.
Scored across 6 tools
Each tool targets a distinct media type and action (edit_image, generate_audio, etc.), with no overlap in purpose. An agent can easily distinguish them.
All tool names follow a consistent verb_noun pattern (e.g., generate_image, list_providers), making the API predictable and easy to navigate.
Six tools cover image, audio, video, and transcription domains. This is a well-scoped set for a multimodal server, not too few or too many.
Core generation and transcription tasks are covered, and editing is available for images. Missing audio/video editing and image-to-video, but overall it's a reasonable surface for common multimodal workflows.