Skip to main content
Glama

Transcribe clip

transcribe_clip
Read-only

Return the cached transcript of a clip's audio. Generated client-side in the Morpha editor (transformers.js Whisper) and cached in R2 next to the clip; if absent, it is produced when the clip is opened in the editor. Independent of the clip's video codec — runs on the audio track only, so it works on HEVC/AV1 clips that the OCR pipeline can't decode. Returns { ok: true, status: 'ready' | 'not-ready', data: { text, word_count, words: [{ word, start, end }], vtt? } }.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
clipYesClip filename (video.<id>.clip).
projectIdYesProject the clip belongs to.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYesWhether the call succeeded.
dataNoThe payload, shaped by the tool.
noteNoWhat to do next when not ready.
errorNoWhy it failed.
statusNoFor cache-backed readers: whether the answer was ready.
editorUrlNoOpens this project in the editor.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • removedInput schema / properties / anonymousToken
      Removed value: -{
      -  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "projectId",
      -  "clip",
      -  "anonymousToken"
      -]New value: +[
      +  "projectId",
      +  "clip"
      +]
  2. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "The result envelope every Morpha tool returns.",
      +  "properties": {
      +    "data": {
      +      "description": "The payload, shaped by the tool.",
      +      "type": [
      +        "object",
      +        "array",
      +        "string",
      +        "number",
      +        "boolean",
      +        "null"
      +      ]
      +    },
      +    "editorUrl": {
      +      "description": "Opens this project in the editor.",
      +      "type": "string"
      +    },
      +    "error": {
      +      "description": "Why it failed.",
      +      "type": "string"
      +    },
      +    "note": {
      +      "description": "What to do next when not ready.",
      +      "type": "string"
      +    },
      +    "ok": {
      +      "description": "Whether the call succeeded.",
      +      "type": "boolean"
      +    },
      +    "status": {
      +      "description": "For cache-backed readers: whether the answer was ready.",
      +      "enum": [
      +        "ready",
      +        "not-ready"
      +      ],
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "ok"
      +  ],
      +  "type": "object"
      +}
  3. Changed2 schema fields changed
    • addedInput schema / properties / anonymousToken
      Added value: +{
      +  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "projectId",
      -  "clip"
      -]New value: +[
      +  "projectId",
      +  "clip",
      +  "anonymousToken"
      +]
  4. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint and destructiveHint annotations already establish it as a safe read operation. The description adds valuable behavior beyond that: the transcript is cached in R2, generated client-side in the editor when absent, and may return a 'not-ready' status, giving the agent a clear model of what to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, generation/caching behavior, codec independence, and return shape. The key purpose is front-loaded first, and the rest is relevant detail without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the important non-obvious behaviors: caching, client-side generation, codec independence, and the ready/not-ready status. An output schema exists so return-value documentation is not a gap; one minor omission is explicit guidance on what to do when status is 'not-ready', though it is inferable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes both parameters with 100% coverage, so the baseline of 3 applies. The description does not add new parameter-level meaning beyond what the schema already states; its extra detail is about return behavior, not parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource: it returns the cached transcript of a clip's audio. It also distinguishes itself from the OCR/video pipeline by noting it runs on the audio track and works on HEVC/AV1 clips that OCR can't decode, which clearly separates it from siblings like describe_video.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when the goal is an audio transcript rather than video/OCR analysis, especially for codecs the OCR pipeline can't handle. It does not explicitly name alternatives or state 'use this instead of X', so it stops short of full explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources