Skip to main content
Glama
457,697 tools. Updated 2026-08-14 14:35

"Using Hugging Face for Text-to-Audio, Image, and Video Generation" matching MCP tools:

Matching MCP Servers

Matching MCP Connectors

  • Retrieve the complete JSON schema for video generation to understand all configuration options, constraints, and examples before creating valid video configurations.
    MIT
  • Generate a JSON configuration for video generation with support for clips, layers, transitions, and audio options, ready for validation or API submission.
    MIT
  • Explore video generation models with detailed specs: modes, aspect ratios, durations, resolutions, media inputs, audio support, strengths, and pricing. Use to select the right model or verify parameters before generation.
    Apache 2.0
  • Diagnose SuperCMO setup by checking which media-generation keys are set and which capabilities (image, video, audio) are ready. Use first when setting up or after a 'no_provider_configured' error.
    Apache 2.0
  • Create a lipsync video from an audio track or script paired with a still image or video. Use a voice ID to make the visual say the provided text.
    MIT
  • Open the Sync upload widget to choose a ChatGPT file or upload a local image/audio, converting it to a durable assetId for later processing.
    MIT
  • Uploads user-provided image, video, or audio to Sync asset storage and returns a durable assetId for use in subsequent lipsync generation.
    MIT
  • Upload a local image, video, or audio file to obtain a public HTTPS URL for use in video generation tasks, avoiding base64 truncation issues.
    MIT
  • Generate a video from a text prompt or a start image URL. Returns request ID and cost instantly; video renders in 1-10 minutes.
    MIT
  • Generate AI images or videos asynchronously by submitting a task and polling for the result. Supports text-to-image, image-to-video, reference-to-video, and video editing models.
    MIT
  • List all element types a Zvid project supports, including visuals (IMAGE, VIDEO, GIF, SVG, TEXT), AUDIO, SUBTITLE, and SCENE, with summaries and required fields.
    MIT