Skip to main content
Glama
538,021 tools. Updated 2026-09-09 01:40

"Tools for speech generation, video translation, and voice conversion" matching MCP tools:

  • Retrieve the complete JSON schema for video generation to understand all configuration options, constraints, and examples before creating valid video configurations.
    MIT
  • Generate AI subtitles for a hosted video by providing its video ID. Speech is transcribed into a WebVTT track with automatic language detection or a specified ISO code, and attached to the video when complete.
    MIT
  • Turn a script into a spoken MP3 URL by selecting a voice engine and preset, producing standalone voiceover clips up to 900 characters.
    MIT
  • Convert text to speech with selectable voice and language. Get an mp3 URL for standalone use or as voiceover in video composition.
    MIT
  • Automatically lower background music during speech and restore it in pauses. Accepts audio or video as voice input and returns a mixed output.
    MIT
  • Create or update an AI Target for evaluation simulations. Supports generation models, custom endpoints, and voice agents with templates, tools, and HTTP settings.
    Apache 2.0

Matching MCP Servers

Matching MCP Connectors

  • Convert text to speech audio files using specified voices and models, saving results to your chosen directory for accessibility or content creation.
    MIT
  • Convert audio files to different voices while preserving speech content. Transform MP3 or WAV files using voice IDs with adjustable similarity and optional background noise removal.
    MIT
  • Turn JSON data into formatted documents by filling Carbone templates. Supports format conversion, translations, currency conversion, and batch generation.
    Apache 2.0
  • Create custom voice profiles from audio samples for text-to-speech and speech-to-speech applications. Analyze MP3 or WAV files to generate voice replicas that mimic original audio characteristics.
    MIT
  • Find specific video segments using semantic search across speech, on-screen text, and visuals. Returns timestamps and metadata for precise moment retrieval.
    MIT
  • List all available voices, personas, languages, pricing tiers, character limits, and speed/quality controls to plan your speech generation call.
    MIT
  • Swap a finished video's narrator voice while keeping the original performance, lip-sync, and background audio. Provide a video URL and voice preset to receive the served URL.
    MIT
  • Translate text with AI localization using translation memory, style guides, and brand voice. Specify target language, context, glossary, and formality for accurate results.
    MIT
  • Create a new AI voice agent by configuring the language model, text-to-speech, and speech-to-text settings. Define the system prompt and voice to deploy a custom virtual assistant.
    MIT
  • Create a custom voice clone from an audio/video sample via URL or asset ID, and receive a voice ID to use in text-to-speech requests.
    MIT