Skip to main content
Glama

Transcript Kit

Transcribe an audio or video file

transcribe_media

Transcribe an audio or video file. Use this when the user wants the words of a recording they own or may use, for example "transcribe this podcast episode", "make subtitles for my video", "what is said in this file". Pass the uploaded file as file, or a direct https link to an audio or video file as url (not a page of YouTube, TikTok, Douyin or another platform). Set confirm_rights true only after the user says they own the recording or have the rights to transcribe it; otherwise ask. Limits: 15 MB and 15 minutes per file (mp3, m4a, mp4, wav, ogg, flac, webm), and a daily allowance of minutes per user and in total. Returns the language, timestamped segments, SRT and VTT captions (trim with include), the transcript_id, when it expires and the minutes left today. The transcript is kept 24 hours for search and clips; the media file is never kept. Transcribed with Cloudflare Workers AI Whisper.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoDirect https link to an audio or video file the user owns or may use, such as https://example.com/talk.mp3. Pages of video or social platforms are not accepted. Give this or file.
fileNoAn audio or video file the user uploaded to the chat.
titleNoName to show for the transcript (default: the file name).
includeNoWhat to return besides the summary (default all three): timestamped segments, SRT captions, VTT captions.
languageNoSpoken language as an ISO 639-1 code such as en or fr. Leave out to detect it.
library_idNoThe library_id that transcribe_media returned. Only needed when transcribe_media returned one; leave it out otherwise.
confirm_rightsYesSet true only after the user has said they own this recording or have the rights to transcribe it.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
srtNo
vttNo
as_ofYes
modelYes
notesYes
titleYes
usageYes
sourceYes
statusYes
languageYes
segmentsNo
expires_atYes
library_idNoPresent when the caller has no user id; pass it to the other tools to reach this transcript.
word_countNo
transcript_idYes
duration_secondsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false) leave the safety profile underspecified, and the description fills it in: media file never retained, transcript kept 24 hours, size/duration caps (15 MB, 15 minutes), format list, per-user and global daily minute allowances, and the underlying model. This is behavioral context an agent could not infer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then usage triggers, then input routing, then the rights gate, then limits and outputs. Every sentence carries information, though the single dense paragraph could be broken up; nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with a nested file object, an output schema, and a legal/rights gate, the description covers inputs, constraints, retention, and return expectations (language, segments, SRT/VTT, transcript_id, expiry, minutes left). Nothing an agent needs before calling is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description earns more: it clarifies file-vs-url as alternatives ('Give this or file'), explains include as trimming what is returned beyond the summary, and spells out the confirm_rights policy that the schema only tersely states. It stops short of adding syntax detail the schema lacks value over.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Transcribe') and resource ('audio or video file') and immediately scopes it with examples like 'transcribe this podcast episode' and 'make subtitles for my video'. Sibling tools are unrelated (search_transcripts, translate_subtitles, suggest_clips), and the description makes clear this is the ingestion step, not retrieval or translation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers, the two input routes (uploaded file vs direct https link), an explicit when-NOT (platform pages of YouTube/TikTok/Douyin are rejected), and a gated workflow for confirm_rights: only set true after the user asserts ownership, otherwise ask.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources