Skip to main content
Glama
510,032 tools. Updated 2026-09-03 17:32

"How to record desktop audio and transcribe it to a VTT file" matching MCP tools:

  • Transcribe local audio files to text or subtitle formats (SRT, VTT) with speaker diarization. Supports JSON and verbose output, optional context prompts, and language selection.
    MIT
  • Check if the Diagrammo desktop app is installed on macOS to decide how to deliver diagrams: open the source file in the app when available, otherwise fall back to a shareable online URL.
    MIT
  • Burn styled, karaoke captions into a video using word timestamps from a transcript. Use transcribe output to add social-style captions locally with ffmpeg.
    Apache 2.0
  • Generate word-level timestamped transcripts from audio or video to caption videos, read voiceovers, or analyze competitor ad scripts.
    Apache 2.0
  • Open the Sync upload widget to choose a ChatGPT file or upload a local image/audio, converting it to a durable assetId for later processing.
    MIT
  • Convert text into spoken audio and save it as a file. Choose AI voice, language, speed, and format to generate speech on demand.
    MIT

Matching MCP Servers

Matching MCP Connectors

  • Transform any blog post or article URL into ready-to-post social media content for Twitter/X threads, LinkedIn posts, Instagram captions, Facebook posts, and email newsletters. Pay-per-event: $0.07 for all 5 platforms, $0.03 for single platform.

  • Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.

  • Transcribe audio to text in 32+ languages via Smallest AI's Pulse STT. Provide a file path or URL and specify language; optional diarization, PII redaction, timestamps, and emotion detection.
    MIT
  • Convert an audio file URL into accurate, punctuated text with timestamps, language detection, and optional speaker labels. Upload up to 15 minutes of audio or split longer recordings for transcription.
    MIT
  • Identify the active connection mode (DESKTOP or CLOUD) and obtain instructions to enable the other mode.
    MIT
  • Transcribe audio files into text with automatic language detection and optional speaker diarization. Save transcripts to file or return them directly to the client.
    MIT
  • Transcribe audio or video into word-timestamped captions for your video project. Runs Whisper locally—no API key or upload required—and accepts a URL or local file path.
    Apache 2.0
  • Convert text to audio with customizable voice, speed, and emotion, saving the file to a specified directory. Integrates with MiniMax API for high-quality speech synthesis.
    MIT
  • Convert any video or audio URL into formatted Markdown notes: downloads, transcribes, and saves to Desktop. Also writes saved note content to a file.
    MIT
  • Transcribe an audio file into text directly from your machine—no upload needed. Supports Arabic and optional speaker labels.
    MIT
  • Convert subtitles between SRT and VTT, and shift cue timings by milliseconds. Runs locally in your browser—no server needed.
    MIT
  • Transcribe audio from a URL or file path using Deepgram Nova-2. Get transcript, confidence score, word-level timestamps, and a sealed ProofLink receipt for verification.
    MIT
  • Transcribe local audio or video files by providing their absolute path to the TypeWhisper app. Supports language hints, task type selection, engine overrides, and dictionary corrections.
    GPL 3.0
  • Transcribe base64-encoded audio to text. Supports multiple formats (mp3, wav, ogg, webm, flac) up to 10MB and various languages via BCP-47 codes.
    MIT