Skip to main content
Glama

Detect text regions

detect_text_regions
Read-only

Return the OCR text regions detected in a clip's video frames OR a still image asset. Pass either clip (a video..clip filename) or image (an image layer's filename) — not both. Video clips are sampled every 0.3s at upload time; images are OCR'd once. Results are cached in R2 next to the source. Returns { status: 'ready' | 'not-ready', frames: [{ frame, time, words: [{ text, x0, y0, x1, y1, confidence }] }], videoWidth, videoHeight }. For an image there's a single frame (frame 0); videoWidth/videoHeight are the image's pixel dimensions. Each words entry is one detected text line — text may hold several words, the box is axis-aligned, and confidence is 0–100. Coordinates are in the SOURCE pixel space (not canvas space). Use to know where titles / subtitles / lower-thirds are baked into the video or image so the Morpha title + intro graphics can be positioned to not collide.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
clipNoVideo clip filename (the same value as `video.<id>.clip`). Pass this OR `image`.
imageNoImage layer filename (the same value as `image.<id>.filename`). Pass this OR `clip`.
projectIdYesProject the clip/image belongs to.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYesWhether the call succeeded.
dataNoThe payload, shaped by the tool.
noteNoWhat to do next when not ready.
errorNoWhy it failed.
statusNoFor cache-backed readers: whether the answer was ready.
editorUrlNoOpens this project in the editor.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • removedInput schema / properties / anonymousToken
      Removed value: -{
      -  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "projectId",
      -  "anonymousToken"
      -]New value: +[
      +  "projectId"
      +]
  2. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "The result envelope every Morpha tool returns.",
      +  "properties": {
      +    "data": {
      +      "description": "The payload, shaped by the tool.",
      +      "type": [
      +        "object",
      +        "array",
      +        "string",
      +        "number",
      +        "boolean",
      +        "null"
      +      ]
      +    },
      +    "editorUrl": {
      +      "description": "Opens this project in the editor.",
      +      "type": "string"
      +    },
      +    "error": {
      +      "description": "Why it failed.",
      +      "type": "string"
      +    },
      +    "note": {
      +      "description": "What to do next when not ready.",
      +      "type": "string"
      +    },
      +    "ok": {
      +      "description": "Whether the call succeeded.",
      +      "type": "boolean"
      +    },
      +    "status": {
      +      "description": "For cache-backed readers: whether the answer was ready.",
      +      "enum": [
      +        "ready",
      +        "not-ready"
      +      ],
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "ok"
      +  ],
      +  "type": "object"
      +}
  3. Changed2 schema fields changed
    • addedInput schema / properties / anonymousToken
      Added value: +{
      +  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "projectId"
      -]New value: +[
      +  "projectId",
      +  "anonymousToken"
      +]
  4. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false)Skip; the description adds substantial behavioral context: 0.3s sampling at upload, one-time OCR for images, R2 caching, status values, source-pixel coordinate space, and confidence range. This goes well beyond what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries distinct information: scope, input selection, processing cadence, caching, return shape, coordinate space, and an explicit use case. It is front-loaded with purpose and avoids filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analysis tool with a clear output shape, the description covers all call-relevant concerns: input selection, coordinate system, confidence semantics, and a practical application. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents clip, image, and projectId. The description enriches this by clarifying the mutual-exclusion rule, what 'clip' and 'image' filenames refer to, how each source is processed, and that image results have a single frame. That is meaningful semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Return the OCR text regions detected in a clip's video frames OR a still image asset.' It clearly distinguishes itself from transcribe_clip and other media-analysis tools by focusing on text-region detection for graphics collision avoidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage context: 'Use to know where titles / subtitles / lower-thirds are baked into the video or image so the Morpha title + intro graphics can be positioned to not collide.' It also explains the clip-vs-image choice with 'not both.' It does not explicitly name alternatives or state when not to use the tool, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources