Skip to main content
Glama

US Company Intelligence for AI Agents (SEC EDGAR, x402)

transcribe

Transcribe the actual audio of a YouTube, TikTok or Instagram video with Whisper large-v3. Not caption scraping: the audio is downloaded and run through speech recognition, so it works on videos with no subtitles, in any language, and on TikTok and Instagram where no caption track exists at all. Extraction runs from a real residential IP, reaching sources that refuse datacenter ranges. Returns full text plus sentence-level timestamps. — $0.020000/call, paid per request via x402 (USDC). Use when asked: "transcribe this youtube video", "get the transcript of this tiktok", "speech to text from a video URL".

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden, and it delivers: it discloses that audio is download-and-run through speech recognition, works without subtitles, works on platforms with no caption tracks, runs from a residential IP to bypass source restrictions, returns full text plus sentence-level timestamps, and includes pricing and payment method. This is unusually transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence adds distinct value: model and mechanism, contrast with alternatives, network behavior, return payload, cost/payment, and trigger examples. The description is front-loaded with the core purpose and remains appropriately sized for the amount of unique behavioral information it conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no annotations and no output schema, the description is mostly complete: it covers input, method, supported sources, output, and commercial terms. Minor gaps remain, such as failure behavior for private/protected videos or video length limits, but nothing an agent needs to make an initial correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage)Skip, the description must compensate for the single url parameter. It explains that the URL must be for a YouTube, TikTok, or Instagram video and implies a direct video URL rather than a channel or playlist. It could specify public/accessibility requirements, but it gives enough context for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Transcribe the actual audio of a YouTube, TikTok or Instagram video with Whisper large-v3.' It clearly distinguishes itself from caption scraping and names the video platforms it supports, so an agent can tell it apart from siblings like extract or understand immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger phrases: 'Use when asked: "transcribe this youtube video", "get the transcript of this tiktok", "speech to text from a video URL".' It also explicitly states what it is not for by contrasting with caption scraping, leaving no ambiguity about when to select it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources