Skip to main content
Glama

inset_add

Idempotent

Overlay a clip into a rectangular preview area of a recording, tracking its camera motion, with configurable fades and audio control.

Instructions

Draw a clip into a rectangle of the recording, following its camera — a render inside the app's preview.

rect is where the recording shows what the inset replaces, in the recording's own pixels, and must be the clip's shape. The inset moves and zooms with the recording's reframe windows, so a push into the preview fills the frame with the clip itself, at full sharpness. The span starts at a word, phrase or event of clip_id and ends at one, at a length, or at the clip's end; it plays at 1x and must lie inside one continuous, 1x stretch of the recording (no cut, no retimed span under it).

It fades in and out by default, can dim the recording around it, and plays its own audio with the music bed out underneath (dipped, with a duck) unless mute; level="speech" measures it and sets gain_db. The reply echoes where it plays and dest, where its rect lands in the canvas; plan=true writes nothing. Any inset routes export through the MLT writer, and export's reply gives each inset's rect at its first and last frame.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
dimNoDarken the recording around the inset, 0 (none, the default) to 1 (black); 0.55 reads well.
muteNoPlay none of the asset's audio (and leave the bed alone).
pathNoThe project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing.
planNoResolve the whole call and report what it would do, writing nothing. Prefer it over doing the thing and undoing it.
rectYes[x0, y0, x1, y1] in the recording's OWN pixels: where the recording shows the thing the inset replaces (a preview pane). Must be the asset's shape, within 1%.
afterNoA forward cursor over a phrase's matches: any match at or before this word index is skipped. -1, the default, means from the start.
assetYesThe clip to draw, by clip_id: the render an agent made, in a launch clip. Needs picture.
enterNoHow it appears: `fade` or `none` (a cut). Default fade.
eventNoStart on this event of clip_id: `name`, or `name#k` when the name repeats. Usually the moment the recording's own preview starts playing, which locks the two.
leaveNoHow it goes: `fade` or `none`. Default fade.
levelNo'speech' measures the span the inset plays, once, and records the gain_db that brings its speech to -18 dBFS RMS, the launch clip's film level. Not with gain_db.
phraseNoStart on this phrase's FIRST word, resolved against clip_id's transcript.
src_inNoSeconds into the asset the inset starts from. Default 0.
clip_idYesThe recording the inset is drawn into — its camera (reframe windows) is what the inset follows, and its words or events address the span.
gain_dbNoThe asset's own audio level in dB. Default 0. The music bed goes out under it, or dips under it when the bed has a duck.
secondsNoEnd this long after the start.
positionNoWhere in the stack it goes: 0 is the bottom, omitted is the top.
enter_easeNoThe fade's curve: linear, ease, ease-in or ease-out. Default ease.
leave_easeNoThe fade's curve: linear, ease, ease-in or ease-out. Default ease.
occurrenceNoDisambiguate a phrase by count when it matches more than once, **1-based** in transcript order among the matches after `after`: 1 is the first, 2 the second. Unset, an ambiguous phrase is refused — listing every candidate's range and text — rather than guessed at.
word_indexNoThe word the inset starts on. One of word_index, phrase or event.
until_eventNoEnd on this event of clip_id.
until_phraseNoEnd as this phrase's LAST word ends.
enter_secondsNoHow long the fade in takes. Default 0.4.
leave_secondsNoHow long the fade out takes. Default 0.4.
until_word_indexNoEnd as this word ends. At most one of until_word_index, until_phrase, until_event or seconds; none runs to the asset's end.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.37.0
    • changedInput schema / properties / gain_db / default
      Previous value: -0New value: +null
    • changedInput schema / properties / gain_db / description
      Previous value: -"The asset's own audio level in dB. Default 0. The music bed goes out under it."New value: +"The asset's own audio level in dB. Default 0. The music bed goes out under it, or dips under it when the bed has a duck."
    • changedInput schema / properties / gain_db / type
      Previous value: -"number"New value: +[
      +  "number",
      +  "null"
      +]
    • addedInput schema / properties / level
      Added value: +{
      +  "default": null,
      +  "description": "'speech' measures the span the inset plays, once, and records the gain_db that brings its speech to -18 dBFS RMS, the launch clip's film level. Not with gain_db.",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
  2. Addedv0.36.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, it discloses side effects and rendering behavior: it writes unless `plan=true`, fades by default, dims the recording, lowers the music bed or ducks, routes through the MLT writer, and returns rect locations. It also states a hard constraint (must lie inside one continuous 1x stretch). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but organized into purpose, geometry/span constraints, and effects/plan/export, with the core definition front-loaded. Every sentence carries operational meaning, and the length is proportionate to the tool's 26-parameter complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema and output schema, the description covers the essential behavioral model: span anchoring, reframe behavior, audio handling, planning mode, and export behavior. An agent has enough context to call this tool correctly without needing to infer hidden side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all 26 parameters, so the baseline is 3. The description adds value by explaining relationships between parameters (`rect` must match the clip's shape and follows reframe windows; `level="speech"` measures and sets `gain_db`; `plan=true` writes nothing), which helps an agent choose combinations correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific action ('Draw a clip into a rectangle of the recording'), identifies the resource (a clip/asset inside a recording), and distinguishes the behavior from list/remove siblings (inset_ls, inset_rm). It also adds the distinguishing 'following its camera' detail that sets insets apart from static overlays.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for adding a camera-following picture-in-picture inset and states important placement constraints, but it does not explicitly name alternatives or say when not to use it. The only guidance about choosing an approach is internal: prefer `plan=true` over doing the thing and undoing it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools