"Exploring text-to-image generation techniques" matching MCP connectors:
GET /v1/connectors — MCP directory API referenceMatching Connector Tools:
Generate highly realistic Text to Speech voiceovers.
Bambara AI over MCP: text-to-speech, transcription and translation (Bamanankan + more).
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Remote MCP for AI video, image, music and speech generation.
Any video URL to LLM-ready transcript. ASR built in, no captions needed. TikTok, X, TED and more.
Your AI rings your iPhone, speaks its question, and gets your spoken answer back as text.
Construction daily-log generation, jurisdiction compliance requirements, construction FAQs.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
Kurdish (Sorani & Kurmanji) text-to-speech & speech-to-text — 664 AI voices. API key required.
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Give ears to Claude/Openclaw/Hermes/Codex/Grok Bot. Voibe turns recordings into text your AI agent can work with. Ask your agent to transcribe a meeting, interview, call, lecture, podcast episode or voice memo. The raw transcript arrives in the chat with speaker labels, timestamps and a summary. Attach the file in the chat, or point at a file or folder in Claude Code, where a whole folder of recordings works in
Transcribe public videos & audio (YouTube, TikTok, IG) into accurate, timestamped text via API.
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
AI speech-to-text for public TikTok videos: SRT, VTT, word timings, speaker labels, 90+ languages.
AI music studio: song generation with vocals, covers, stems, voice conversion, mastering, editing.