"Using OpenAI to Generate Images" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
SEC-grounded stock facts phrased as single spoken sentences for a voice assistant to read.
Sorani & Kurmanji TTS+STT: Kurdish speech most APIs lack. 885 voices, free tier, no key to browse.
Pay-per-call AI gateway: models, speech, web search, on-chain reads and NLP tools, via x402.
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Bambara AI over MCP: text-to-speech, transcription and translation (Bamanankan + more).
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
Any video URL to LLM-ready transcript. ASR built in, no captions needed. TikTok, X, TED and more.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
ElevenLabs in natural language: generate speech in any language, create and manage voices, compose m
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
Generate highly realistic Text to Speech voiceovers.
Give ears to Claude/Openclaw/Hermes/Codex/Grok Bot. Voibe turns recordings into text your AI agent can work with. Ask your agent to transcribe a meeting, interview, call, lecture, podcast episode or voice memo. The raw transcript arrives in the chat with speaker labels, timestamps and a summary. Attach the file in the chat, or point at a file or folder in Claude Code, where a whole folder of recordings works in
Transcribe any audio or video URL to text, SRT and VTT with timestamped segments
Audio AI tools: text-to-speech, voice cloning, music generation, stem separation, transcription.
Generate video, images, audio and speech with Vidofy — Veo 3.1, Kling 3.0, Flux 2 and 570+ models.
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Audio & video to text, Russian-first: diarization, timestamps, summary, action items, subtitles.
- ElevenLabsOAuth unavailableio.elevenlabs
Manage ElevenLabs voice agents and generate speech, music, sound effects, images, and video.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)