Skip to main content
Glama

Get clip details

get_clip
Read-only

Read one clip: its elements (positions/sizes in canvas pixels), voiceover (text, voice, duration, voiceover_volume), background and transition. Pass render to also get a PNG of the frame.

ASK FOR WHAT YOU NEED. A full read is large — on a dense clip the per-word voiceover array and the element type_data blobs dominate it, and repeated full reads are the main way a long session runs out of context. select returns exactly the parts you name:

select: ['elements.x','elements.y','elements.width','elements.height'] → geometry only, to fix a layout select: ['elements.name','elements.start_time','elements.end_time'] → a timing pass select: ['voiceover_words'] → word timings only, to sync visuals to narration select: ['elements.textdata','voiceover_words'] → rewrite copy against the VO select: ['elements'] → whole element rows, no words select: ['groups'] → group rows only, to get a group_id for update_groups select: [] → no JSON at all (pair with render for the PNG alone — smallest read) (omit select) → everything; fine for a first look, expensive to repeat

render is the other output, and it is separate from select: select shapes the JSON, render produces a PNG.

render: {} → the frame at t=0 render: { timestamps: 2.5 } → the frame 2.5s into the clip render: { timestamps: [0.5, 2, 4] } → those three moments as ONE labelled grid render: { timestamps: [...], layout:'separate'} → the same moments as full-size images (~4x the tokens) render: { save: true } → also uploads the frame and returns presigned_url render: { max_width: 1280 } → a sharper frame when you must read small print select: [], render: {} → the PNG alone, no JSON select: ['elements'], render: {} → element rows AND the frame

A single frame renders 960px wide by default — legible for this design system and about half the tokens of a 1280 frame. A GRID defaults to 1280, because that is the width of the whole grid and a third of 960 would leave each cell unreadable. Either way max_width caps the image you get back; raise it only to read genuinely small print.

Omitting render renders nothing. timestamps, layout and save live inside it because they only mean anything for a render — there is no way to ask for them without asking for the image.

element_ids is the other axis: it picks WHICH element rows come back, independently of select. Combine them for the leanest read — e.g. element_ids: ['el_9'], select: ['elements.x','elements.y'].

Element shape: universal wrapper fields (id, element_type, name, x, y, width, height, start_time, end_time, rotation) plus type-specific data (textdata/shapedata/imagedata/videodata/zoomdata) plus an optional keyframes array when animated. Keyframes come back in the same flat wire shape add_elements takes — { timestamp, positionX?, positionY?, width?, height?, interpolation? } in canvas pixels — so you can round-trip read → edit → update_elements without reshaping.

Clip-level fields include transition (the current transition object — sibling of the update_clips transition arg; null if none) and voiceover_words (per-word timestamps, {word, start, end, punctuated_word} with start/end in SECONDS; null on clips with no transcription).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
renderNoRender the clip as an image. Presence of this object IS the request to render — omit it and nothing is rendered. `{}` renders at t=0. Independent of `select`, which only shapes the JSON: pair `select: []` with `render: {}` for the PNG alone (smallest read). CHECKING YOUR WORK: pass `timestamps` with SEVERAL mid-clip moments, not the t=0 default — text and image elements have entry animations (a ~0.4s slide/fade by default), so at t=0 they have not arrived yet and a correct edit renders as an empty frame, while shapes have no entry animation and do show at t=0. That mix is what makes a single t=0 render actively misleading: some elements appear and others don't. A list comes back as one grid for about a quarter of the tokens of the same frames separately, so checking several moments is the cheap option, not the expensive one. `animations: false` draws everything settled if you would rather not pick moments at all.
selectNoAsk for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none. Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes), 'busy' (generations still writing to this clip, as `[{entity_path, job_type}]` — EMPTY means nothing is pending, which is how you know a voiceover or AI image has landed; it is the same lock that would refuse your write, so a non-empty list also tells you what not to touch yet). Each section returns the key it is named after. Rendering is `render`, not a value here. Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id, dropShadows. y_top is text-only and present ONLY when alignment is 'center' — the unambiguous TOP edge, which is exactly the case where `y` is NOT the top but the vertical CENTRE. To put that position back, send it as `y` with y_anchor:'top'. Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO. Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this.
clip_indexYesZero-based clip index
project_idYesThe project ID
element_idsNoWHICH element rows to return — all others are dropped. Independent of `select`, which chooses the sections/keys. Use it to re-inspect just what you added or updated; most add_elements/update_elements already echo the element's resolved layout, so often you don't need this at all.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • changedInput schema / properties / select / description
      Previous value: -"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes), 'busy' (generations still writing to this clip, as `[{entity_path, job_type}]` — EMPTY means nothing is pending, which is how you know a voiceover or AI image has landed; it is the same lock that would refuse your write, so a non-empty list also tells you what not to touch yet). Each section returns the key it is named after. Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id. y_top is text-only and present ONLY when alignment is 'center' — the unambiguous TOP edge, which is exactly the case where `y` is NOT the top but the vertical CENTRE. To put that position back, send it as `y` with y_anchor:'top'.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."New value: +"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes), 'busy' (generations still writing to this clip, as `[{entity_path, job_type}]` — EMPTY means nothing is pending, which is how you know a voiceover or AI image has landed; it is the same lock that would refuse your write, so a non-empty list also tells you what not to touch yet). Each section returns the key it is named after. Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id, dropShadows. y_top is text-only and present ONLY when alignment is 'center' — the unambiguous TOP edge, which is exactly the case where `y` is NOT the top but the vertical CENTRE. To put that position back, send it as `y` with y_anchor:'top'.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."
    • changedInput schema / properties / select / items / enum
      Previous value: -[
      -  "elements",
      -  "voiceover_words",
      -  "groups",
      -  "busy",
      -  "elements.name",
      -  "elements.element_type",
      -  "elements.x",
      -  "elements.y",
      -  "elements.width",
      -  "elements.height",
      -  "elements.start_time",
      -  "elements.end_time",
      -  "elements.rotation",
      -  "elements.keyframes",
      -  "elements.textdata",
      -  "elements.shapedata",
      -  "elements.imagedata",
      -  "elements.videodata",
      -  "elements.zoomdata",
      -  "elements.codedata",
      -  "elements.parent_id"
      -]New value: +[
      +  "elements",
      +  "voiceover_words",
      +  "groups",
      +  "busy",
      +  "elements.name",
      +  "elements.element_type",
      +  "elements.x",
      +  "elements.y",
      +  "elements.width",
      +  "elements.height",
      +  "elements.start_time",
      +  "elements.end_time",
      +  "elements.rotation",
      +  "elements.keyframes",
      +  "elements.textdata",
      +  "elements.shapedata",
      +  "elements.imagedata",
      +  "elements.videodata",
      +  "elements.zoomdata",
      +  "elements.codedata",
      +  "elements.parent_id",
      +  "elements.dropShadows"
      +]
  2. Changed4 schema fields changed
    • removedInput schema / properties / context
      Removed value: -{
      -  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      -  "type": "string"
      -}
    • removedInput schema / properties / conversation_id
      Removed value: -{
      -  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      -  "type": "string"
      -}
    • removedInput schema / properties / llm_model
      Removed value: -{
      -  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "project_id",
      -  "clip_index",
      -  "context",
      -  "llm_model"
      -]New value: +[
      +  "project_id",
      +  "clip_index"
      +]
  3. Changed4 schema fields changed
    • addedInput schema / properties / context
      Added value: +{
      +  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      +  "type": "string"
      +}
    • addedInput schema / properties / conversation_id
      Added value: +{
      +  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      +  "type": "string"
      +}
    • addedInput schema / properties / llm_model
      Added value: +{
      +  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "project_id",
      -  "clip_index"
      -]New value: +[
      +  "project_id",
      +  "clip_index",
      +  "context",
      +  "llm_model"
      +]
  4. Changed5 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • removedInput schema / additionalProperties
      Removed value: -false
    • addedInput schema / properties / clip_index / maximum
      Added value: +9007199254740991
    • changedInput schema / properties / select / description
      Previous value: -"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes). Each section returns the key it is named after. Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."New value: +"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes), 'busy' (generations still writing to this clip, as `[{entity_path, job_type}]` — EMPTY means nothing is pending, which is how you know a voiceover or AI image has landed; it is the same lock that would refuse your write, so a non-empty list also tells you what not to touch yet). Each section returns the key it is named after. Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id. y_top is text-only and present ONLY when alignment is 'center' — the unambiguous TOP edge, which is exactly the case where `y` is NOT the top but the vertical CENTRE. To put that position back, send it as `y` with y_anchor:'top'.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."
    • changedInput schema / properties / select / items / enum
      Previous value: -[
      -  "elements",
      -  "voiceover_words",
      -  "groups",
      -  "elements.name",
      -  "elements.element_type",
      -  "elements.x",
      -  "elements.y",
      -  "elements.width",
      -  "elements.height",
      -  "elements.start_time",
      -  "elements.end_time",
      -  "elements.rotation",
      -  "elements.keyframes",
      -  "elements.textdata",
      -  "elements.shapedata",
      -  "elements.imagedata",
      -  "elements.videodata",
      -  "elements.zoomdata",
      -  "elements.codedata",
      -  "elements.parent_id"
      -]New value: +[
      +  "elements",
      +  "voiceover_words",
      +  "groups",
      +  "busy",
      +  "elements.name",
      +  "elements.element_type",
      +  "elements.x",
      +  "elements.y",
      +  "elements.width",
      +  "elements.height",
      +  "elements.start_time",
      +  "elements.end_time",
      +  "elements.rotation",
      +  "elements.keyframes",
      +  "elements.textdata",
      +  "elements.shapedata",
      +  "elements.imagedata",
      +  "elements.videodata",
      +  "elements.zoomdata",
      +  "elements.codedata",
      +  "elements.parent_id"
      +]
  5. Changed9 schema fields changed
    • changedInput schema / properties / render / description
      Previous value: -"Render a PNG of the frame. Presence of this object IS the request to render — omit it and nothing is rendered. `{}` renders at t=0. Independent of `select`, which only shapes the JSON: pair `select: []` with `render: {}` for the PNG alone (smallest read). CHECKING YOUR WORK: pass a mid-clip `timestamp`, not the t=0 default — text and image elements have entry animations (a ~0.4s slide/fade by default), so at t=0 they have not arrived yet and a correct edit renders as an empty frame. Shapes have no entry animation and do show at t=0, which makes a t=0 render especially misleading: some elements appear and others don't."New value: +"Render the clip as an image. Presence of this object IS the request to render — omit it and nothing is rendered. `{}` renders at t=0. Independent of `select`, which only shapes the JSON: pair `select: []` with `render: {}` for the PNG alone (smallest read). CHECKING YOUR WORK: pass `timestamps` with SEVERAL mid-clip moments, not the t=0 default — text and image elements have entry animations (a ~0.4s slide/fade by default), so at t=0 they have not arrived yet and a correct edit renders as an empty frame, while shapes have no entry animation and do show at t=0. That mix is what makes a single t=0 render actively misleading: some elements appear and others don't. A list comes back as one grid for about a quarter of the tokens of the same frames separately, so checking several moments is the cheap option, not the expensive one. `animations: false` draws everything settled if you would rather not pick moments at all."
    • addedInput schema / properties / render / properties / animations
      Added value: +{
      +  "description": "Whether entry/exit animations play at this timestamp. Default true — the frame as the video will show it, which is what you want to judge motion. Pass false to draw every element settled at its final position and full opacity: use that to check COMPOSITION, because with staggered starts and short windows there may be no single timestamp at which nothing is mid-flight, and an element part-way through a slide reads as mis-positioned when it is fine.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / render / properties / layout
      Added value: +{
      +  "description": "Only meaningful when `timestamps` holds several. 'sheet' (default) tiles them into one image; 'separate' returns them full-size, which costs ~4x more — use it only when a grid cell is too small to judge.",
      +  "enum": [
      +    "sheet",
      +    "separate"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / render / properties / max_width
      Added value: +{
      +  "description": "Width in px of the IMAGE you get back — for a grid that is the whole grid, not each cell, so a cell is roughly a third of it. Default 960 for a single frame, 1280 for a grid. An image costs roughly (w×h)/750 tokens, so 1280 is ~1229 and 960 ~691; above 1568 nothing is gained because the model downscales to that anyway. The default is legible for this design system's type scale — a 40px canvas caption lands at ~20px. Raise it only to read genuinely small print, such as UI text inside a pasted screenshot.",
      +  "maximum": 1568,
      +  "minimum": 320,
      +  "type": "integer"
      +}
    • changedInput schema / properties / render / properties / save / description
      Previous value: -"Also upload the rendered PNG to S3 and return presigned_url, which you can pass as source_url to update_clueprint (files=[{path, source_url}])."New value: +"Also upload the rendered image to S3 and return presigned_url, which you can pass as source_url to update_clueprint (files=[{path, source_url}]). Uploads ONE image, so it works with a single frame or a grid but not with layout:'separate' — there is one URL and several frames."
    • removedInput schema / properties / render / properties / timestamp
      Removed value: -{
      -  "description": "The moment within the clip to render, in seconds (default 0).",
      -  "minimum": 0,
      -  "type": "number"
      -}
    • addedInput schema / properties / render / properties / timestamps
      Added value: +{
      +  "anyOf": [
      +    {
      +      "minimum": 0,
      +      "type": "number"
      +    },
      +    {
      +      "items": {
      +        "minimum": 0,
      +        "type": "number"
      +      },
      +      "maxItems": 15,
      +      "minItems": 1,
      +      "type": "array"
      +    }
      +  ],
      +  "description": "Which moment(s) of the clip to render, in seconds (default 0). Pass several — [0.5, 2, 4] — and they come back as one labelled grid costing about a quarter of what those frames cost individually. Prefer several whenever you are checking rather than reading: entry animations mean no single moment shows everything, so one frame is usually the wrong question. A grid cell is a third of the width, so for small print pass a single number. Max 15; duplicates dropped, sorted into time order."
      +}
    • changedInput schema / properties / select / description
      Previous value: -"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'words' (per-word VO timings). Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, geo, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['words'] to sync visuals to narration; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."New value: +"Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none.\n  Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes). Each section returns the key it is named after. Rendering is `render`, not a value here.\n  Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id.\n  Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO.\n  Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this."
    • changedInput schema / properties / select / items / enum
      Previous value: -[
      -  "elements",
      -  "words",
      -  "elements.name",
      -  "elements.geo",
      -  "elements.x",
      -  "elements.y",
      -  "elements.width",
      -  "elements.height",
      -  "elements.start_time",
      -  "elements.end_time",
      -  "elements.rotation",
      -  "elements.keyframes",
      -  "elements.textdata",
      -  "elements.shapedata",
      -  "elements.imagedata",
      -  "elements.videodata",
      -  "elements.zoomdata",
      -  "elements.codedata",
      -  "elements.parent_id"
      -]New value: +[
      +  "elements",
      +  "voiceover_words",
      +  "groups",
      +  "elements.name",
      +  "elements.element_type",
      +  "elements.x",
      +  "elements.y",
      +  "elements.width",
      +  "elements.height",
      +  "elements.start_time",
      +  "elements.end_time",
      +  "elements.rotation",
      +  "elements.keyframes",
      +  "elements.textdata",
      +  "elements.shapedata",
      +  "elements.imagedata",
      +  "elements.videodata",
      +  "elements.zoomdata",
      +  "elements.codedata",
      +  "elements.parent_id"
      +]
  6. First observed

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and destructiveHint annotations, the description discloses critical behavioral details: full reads are large and can exhaust context, render and select are independent, a t=0 render can be misleading due to entry animations, default image widths and token costs, and exactly what shape the returned data takes. This is exceptionally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is genuinely complex and the length is earned. It is front-loaded with the purpose and a context-cost warning, and uses well-organized code examples. There is some redundancy with input-schema descriptions, particularly around render and animations, but overall it is structured and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers what the tool returns, including element shape, voiceover_words null semantics, groups, busy, element_ids, render defaults, and cost implications. Since there is no output schema, this narrative carries the full burden, and it does so comprehensively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaningful usage semantics: element_ids is an independent axis, select projections are demonstrated with concrete results, render flags are explained, and keyframes round-trip without reshaping. Some of this duplicates the already-rich schema descriptions, which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read one clip' and enumerates exactly what is returned (elements, voiceover, background, transition). It makes the tool's scope unambiguous, though it does not explicitly name or contrast a sibling tool such as get_project or get_clueprint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use different select/render modes, with task-oriented examples like 'to fix a layout', 'a timing pass', and 'to sync visuals to narration'. It also warns that an omitted select is fine for a first look but expensive to repeat. It does not explicitly state when not to use this tool versus sibling get_* tools, but the intended usage is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.