"How to configure Azure Speech" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Cut audio from YouTube, TikTok or Instagram links to MP3/WAV, and fetch video transcripts.
Cut audio from YouTube, TikTok or Instagram links to MP3/WAV, and fetch video transcripts.
Podcast post-production for an AI agent: transcribe a recording, cut fillers and pauses, clean up the voice, add licensed music that ducks under speech, set chapters and export a finished episode.
Write original songs with SoundBreak AI artists. Built with artists, not on them. Co-write in this chat, play a preview (not a release), and save to your SoundBreak account. Lyrics you wrote stay yours. After you save, you can submit to artists, SoundBreak Radio, or distribute (Spotify, Apple Music, and more) from SoundBreak. Terms: https://app.soundbreak.ai/terms Privacy: https://app.soundbreak.ai/privacy-policy
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
I do everything related to music and lyrics
AI voice generation: text-to-speech and voice cloning from any MCP client.
Text to speech in 149 languages: MP3 links from any assistant. Free without an account.
Read existing YouTube and podcast transcripts with summaries and chapters. No sign-in.
Generate images, video, speech and music, and run LLM chat, across 10,000+ models on ModelsLab.
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
AI audio tools for music producers — stem splitting, vocal removal, BPM & key detection, audio-to-MIDI, format conversion, trimming, video-to-audio extraction and AI song generation.
Generate text, images, speech, music, and video with any AI model, from one credit balance.
Human-made production music for sync — search by brief or reference, preview, score to picture.
One API for 100+ AI video, image, music and speech models.
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
Pronunciation assessment, phoneme scoring, speaker voice ID, audio transcription, speech synthesis.
Remote MCP for AI video, image, music and speech generation.
ElevenLabs in natural language: generate speech in any language, create and manage voices, compose m
Decode an SSTV audio recording and anchor its fingerprint and image hash to the Knox event chain.