Get clip details
get_clipRead one clip: its elements (positions/sizes in canvas pixels), voiceover (text, voice, duration, voiceover_volume), background and transition. Pass render to also get a PNG of the frame.
ASK FOR WHAT YOU NEED. A full read is large — on a dense clip the per-word voiceover array and the element type_data blobs dominate it, and repeated full reads are the main way a long session runs out of context. select returns exactly the parts you name:
select: ['elements.x','elements.y','elements.width','elements.height'] → geometry only, to fix a layout select: ['elements.name','elements.start_time','elements.end_time'] → a timing pass select: ['voiceover_words'] → word timings only, to sync visuals to narration select: ['elements.textdata','voiceover_words'] → rewrite copy against the VO select: ['elements'] → whole element rows, no words select: ['groups'] → group rows only, to get a group_id for update_groups select: [] → no JSON at all (pair with render for the PNG alone — smallest read) (omit select) → everything; fine for a first look, expensive to repeat
render is the other output, and it is separate from select: select shapes the JSON, render produces a PNG.
render: {} → the frame at t=0 render: { timestamps: 2.5 } → the frame 2.5s into the clip render: { timestamps: [0.5, 2, 4] } → those three moments as ONE labelled grid render: { timestamps: [...], layout:'separate'} → the same moments as full-size images (~4x the tokens) render: { save: true } → also uploads the frame and returns presigned_url render: { max_width: 1280 } → a sharper frame when you must read small print select: [], render: {} → the PNG alone, no JSON select: ['elements'], render: {} → element rows AND the frame
A single frame renders 960px wide by default — legible for this design system and about half the tokens of a 1280 frame. A GRID defaults to 1280, because that is the width of the whole grid and a third of 960 would leave each cell unreadable. Either way max_width caps the image you get back; raise it only to read genuinely small print.
Omitting render renders nothing. timestamps, layout and save live inside it because they only mean anything for a render — there is no way to ask for them without asking for the image.
element_ids is the other axis: it picks WHICH element rows come back, independently of select. Combine them for the leanest read — e.g. element_ids: ['el_9'], select: ['elements.x','elements.y'].
Element shape: universal wrapper fields (id, element_type, name, x, y, width, height, start_time, end_time, rotation) plus type-specific data (textdata/shapedata/imagedata/videodata/zoomdata) plus an optional keyframes array when animated. Keyframes come back in the same flat wire shape add_elements takes — { timestamp, positionX?, positionY?, width?, height?, interpolation? } in canvas pixels — so you can round-trip read → edit → update_elements without reshaping.
Clip-level fields include transition (the current transition object — sibling of the update_clips transition arg; null if none) and voiceover_words (per-word timestamps, {word, start, end, punctuated_word} with start/end in SECONDS; null on clips with no transcription).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| render | No | Render the clip as an image. Presence of this object IS the request to render — omit it and nothing is rendered. `{}` renders at t=0. Independent of `select`, which only shapes the JSON: pair `select: []` with `render: {}` for the PNG alone (smallest read). CHECKING YOUR WORK: pass `timestamps` with SEVERAL mid-clip moments, not the t=0 default — text and image elements have entry animations (a ~0.4s slide/fade by default), so at t=0 they have not arrived yet and a correct edit renders as an empty frame, while shapes have no entry animation and do show at t=0. That mix is what makes a single t=0 render actively misleading: some elements appear and others don't. A list comes back as one grid for about a quarter of the tokens of the same frames separately, so checking several moments is the cheap option, not the expensive one. `animations: false` draws everything settled if you would rather not pick moments at all. | |
| select | No | Ask for exactly the JSON you want, GraphQL-style. Omit for everything; pass [] for none. Sections: 'elements' (whole element rows), 'voiceover_words' (per-word VO timings; returns the key of the same name, holding `{word, start, end, punctuated_word}` with start/end in SECONDS — null on a clip with no transcription), 'groups' (group rows: id, name, parent_id, bounds_px, anchor_px, keyframes), 'busy' (generations still writing to this clip, as `[{entity_path, job_type}]` — EMPTY means nothing is pending, which is how you know a voiceover or AI image has landed; it is the same lock that would refuse your write, so a non-empty list also tells you what not to touch yet). Each section returns the key it is named after. Rendering is `render`, not a value here. Per-key: 'elements.<key>' projects element rows to just those keys (id is always kept). Keys: name, element_type, x, y, width, height, start_time, end_time, rotation, keyframes, textdata, shapedata, imagedata, videodata, zoomdata, codedata, parent_id, dropShadows. y_top is text-only and present ONLY when alignment is 'center' — the unambiguous TOP edge, which is exactly the case where `y` is NOT the top but the vertical CENTRE. To put that position back, send it as `y` with y_anchor:'top'. Examples: ['elements.x','elements.y','elements.width','elements.height'] to read geometry; ['elements.name','elements.start_time','elements.end_time'] for a timing pass; ['voiceover_words'] to sync visuals to narration; ['groups'] to resolve a group id for update_groups; [] with render:{} returns the PNG with no JSON (smallest read); ['elements.textdata','voiceover_words'] to rewrite copy against the VO. Mixing 'elements' with 'elements.<key>' returns whole rows. Use element_ids to choose WHICH rows — that is independent of this. | |
| clip_index | Yes | Zero-based clip index | |
| project_id | Yes | The project ID | |
| element_ids | No | WHICH element rows to return — all others are dropped. Independent of `select`, which chooses the sections/keys. Use it to re-inspect just what you added or updated; most add_elements/update_elements already echo the element's resolved layout, so often you don't need this at all. |