Skip to main content
Glama

Sats4AI - Bitcoin-Powered AI Tools

transcribe_audio

Transcribe audio to text with WORD-LEVEL timestamps (timestamps:'word' returns per-word start/end times — subtitle alignment, karaoke captions, cutting video to speech) or segment timestamps. Uses Mistral Transcription — high-accuracy speech recognition that handles accents, background noise, and overlapping speakers. 13 languages: en, zh, hi, es, ar, fr, pt, ru, de, ja, ko, it, nl. Up to 500 MB / 60 minutes per file. Async — returns requestId, poll with check_job_status(jobType='transcription'), then get_job_result. 10 sats/min. Privacy: audio and transcripts are ephemeral — processed, returned, and discarded. Never persisted. Pay per request with Bitcoin Lightning — no API key or signup needed. Requires create_payment with toolName='transcribe_audio'.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
diarizeNoIdentify different speakers (default false). Forces segment granularity upstream. If the speaker-label request fails, you still get the text, and result.degraded includes 'diarization'.
languageNoLanguage code (e.g., 'en', 'es')
paymentIdYesValid payment ID (must be paid)
timestampsNoTimestamp granularity in result.segments. 'word' returns per-word start/end times (subtitle alignment, karaoke captions, cutting video to speech). Default 'segment'. If the timestamp request fails, you still get the text, and result.degraded includes 'timestamps' (no segments, srt or vtt).
audioBase64YesBase64 encoded audio file
callback_idNoOptional correlation string echoed back in the webhook body. Max 128 chars.
callback_urlNoOptional HTTPS webhook we POST when the job finishes (HMAC-signed). Polling still works.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • changedInput schema / properties / diarize / description
      Previous value: -"Identify different speakers (default false). Forces segment granularity upstream."New value: +"Identify different speakers (default false). Forces segment granularity upstream. If the speaker-label request fails, you still get the text, and result.degraded includes 'diarization'."
    • changedInput schema / properties / timestamps / description
      Previous value: -"Timestamp granularity in result.segments. 'word' returns per-word start/end times (subtitle alignment, karaoke captions, cutting video to speech). Default 'segment'."New value: +"Timestamp granularity in result.segments. 'word' returns per-word start/end times (subtitle alignment, karaoke captions, cutting video to speech). Default 'segment'. If the timestamp request fails, you still get the text, and result.degraded includes 'timestamps' (no segments, srt or vtt)."
  2. Changed4 schema fields changed
    • addedInput schema / properties / callback_id
      Added value: +{
      +  "description": "Optional correlation string echoed back in the webhook body. Max 128 chars.",
      +  "type": "string"
      +}
    • addedInput schema / properties / callback_url
      Added value: +{
      +  "description": "Optional HTTPS webhook we POST when the job finishes (HMAC-signed). Polling still works.",
      +  "type": "string"
      +}
    • addedInput schema / properties / diarize
      Added value: +{
      +  "description": "Identify different speakers (default false). Forces segment granularity upstream.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / timestamps
      Added value: +{
      +  "description": "Timestamp granularity in result.segments. 'word' returns per-word start/end times (subtitle alignment, karaoke captions, cutting video to speech). Default 'segment'.",
      +  "enum": [
      +    "none",
      +    "segment",
      +    "word"
      +  ],
      +  "type": "string"
      +}
  3. First observed

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses async behavior returning a requestId, the polling lifecycle, the exact pricing (10 sats/min), the ephemerality/privacy guarantee ('never persisted'), the payment precondition, and the degraded-result behavior on partial failure. Nearly everything an agent needs to run this safely is spelled out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but tightly packed with a clear front-loaded purpose statement followed by timestamps, engine, languages, limits, workflow, cost, and privacy. A few clauses (word-timestamp use cases, language enumeration) repeat the schema, but there is little true filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining the return shape (requestId) and that the agent must poll and fetch results. For a paid, asynchronous, mutation-style operation with 7 parameters, the coverage of cost, prerequisites, and lifecycle is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the timestamps:'word' use cases and the 13-language list, but this largely duplicates the enum and per-parameter descriptions already in the schema rather than adding new syntax or format guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: transcribe audio to text with word- or segment-level timestamps, plus the underlying engine, language support, and file limits. It distinguishes the tool by what it produces (timestamps, not translation), though it never names transcribe_translate as the sibling an agent should choose instead when translation is needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete operational workflow: create_payment with toolName='transcribe_audio' first, then poll via check_job_status(jobType='transcription'), then get_job_result. This is clear context for invoking it. What's missing is an explicit exclusion rule versus transcribe_translate or other audio siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.