Skip to main content
Glama

cue_add

Idempotent

Places a visual asset at a chosen word or event in a transcript, starting the shot exactly there. Handles phrase lookup, source in-point, and cut effects like dissolve or punch.

Instructions

Add a picture cue: from word_index of clip_id onward, show asset.

Or from an event of clip_id, for a recording with no words.

Source-addressed like a word range — asset is an opaque key or path, not checked against disk here; build_shots resolves it, the same way assemble_scream.py's CUES table did by hand. Refused if a cue already sits at that exact word; cue_rm it first to replace it. Echoes the resolved word plus three either side, the same convention every word-indexed tool follows.

Addressed by word_index or phrase (exactly one) — a phrase binds to its first word ("from this word onward"). after/occurrence disambiguate a phrase matching more than once; a resolved phrase is stored alongside the word index, additive metadata cue_reresolve can re-derive after a re-record.

src_start pins where inside asset the shot reads from: seconds in that asset's own source time, which is exactly the number describe_ls reports for a window. This is how a moment you found with describe gets placed — without it the shot reads from wherever the per-asset cursor had got to, which is right for re-using a clip and wrong for showing the thing you searched for.

It is an in-point and never a range: the out-point stays derived from the next cue through the edit, so a later cut still renumbers the shot correctly. The cost is a refusal instead of a rewind — if the shot's length runs past the end of the asset from that in-point, build_shots and the picture lane report it rather than quietly showing the asset's opening seconds instead. Shorten the shot with another cue, or pin earlier. A card takes no src_start; a held frame has no playhead.

dissolve and punch are the cut's effects, as cue_set sets them.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathNoThe project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing.
afterNoA forward cursor over a phrase's matches: any match at or before this word index is skipped. -1, the default, means from the start.
assetNoWhat to show from that word onward: a registered clip id, or `card:<name>` for a card. An opaque key here, resolved by `build_shots` rather than checked against disk now.
eventNoStart the picture on this event of clip_id instead of a word: `name`, or `name#k` when the name repeats — a screen recording's logged moments, for a clip with no transcript. Not with word_index or phrase.
punchNoA scale punch from the cut, about the canvas centre: 1.1 zooms in 10%. Not with a dissolve on the same cue.
phraseNoAddress the cue by what is said instead of by index. It binds to the phrase's **first** word — "from this word onward".
clip_idYesThe transcript the cue is addressed against — the VO on a voiceover project, not the footage being shown. `asset` is what gets seen.
dissolveNoCrossfade into this cue over this many seconds: the incoming shot's frames before its in-point fade in, reaching the shot on the cue's word. Where the clip has none (an unpinned first use), the outgoing shot fades out from the word instead.
src_startNoWhere inside `asset` the shot reads from, in that asset's own source seconds — the number `describe_ls` reports for a window. An in-point and never a range: unpinned, the shot reads from wherever the per-asset cursor had got to, which is right for re-using a clip and wrong for showing the thing you searched for. A card takes none.
occurrenceNoDisambiguate a phrase by count when it matches more than once, **1-based** in transcript order among the matches after `after`: 1 is the first, 2 the second. Unset, an ambiguous phrase is refused — listing every candidate's range and text — rather than guessed at.
punch_easeNoThe punch's curve. Default ease-out.
punch_modeNo`in` (default) zooms to `punch` and holds for the shot; `settle` starts there and eases back.
word_indexNoThe word the picture starts on, in `clip_id`'s transcript. Give this or `phrase`, not both.
dissolve_easeNoThe crossfade's curve: linear (the default), ease, ease-in or ease-out.
punch_secondsNoHow long the punch moves. Default 0.2.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv0.43.0
    • addedInput schema / properties / dissolve
      Added value: +{
      +  "default": null,
      +  "description": "Crossfade into this cue over this many seconds: the incoming shot's frames before its in-point fade in, reaching the shot on the cue's word. Where the clip has none (an unpinned first use), the outgoing shot fades out from the word instead. ",
      +  "type": [
      +    "number",
      +    "null"
      +  ]
      +}
    • addedInput schema / properties / dissolve_ease
      Added value: +{
      +  "default": null,
      +  "description": "The crossfade's curve: linear (the default), ease, ease-in or ease-out.",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
    • addedInput schema / properties / punch
      Added value: +{
      +  "default": null,
      +  "description": "A scale punch from the cut, about the canvas centre: 1.1 zooms in 10%. Not with a dissolve on the same cue. ",
      +  "type": [
      +    "number",
      +    "null"
      +  ]
      +}
    • addedInput schema / properties / punch_ease
      Added value: +{
      +  "default": null,
      +  "description": "The punch's curve. Default ease-out.",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
    • addedInput schema / properties / punch_mode
      Added value: +{
      +  "default": null,
      +  "description": "`in` (default) zooms to `punch` and holds for the shot; `settle` starts there and eases back.",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
    • addedInput schema / properties / punch_seconds
      Added value: +{
      +  "default": null,
      +  "description": "How long the punch moves. Default 0.2.",
      +  "type": [
      +    "number",
      +    "null"
      +  ]
      +}
  2. Changed1 schema field changedv0.37.0
    • addedInput schema / properties / event
      Added value: +{
      +  "default": null,
      +  "description": "Start the picture on this event of clip_id instead of a word: `name`, or `name#k` when the name repeats — a screen recording's logged moments, for a clip with no transcript. Not with word_index or phrase.",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
  3. Changed29 schema fields changedv0.25.0
    • addedInput schema / properties / after / description
      Added value: +"A forward cursor over a phrase's matches: any match at or before this word index is skipped. -1, the default, means from the start."
    • removedInput schema / properties / after / title
      Removed value: -"After"
    • removedInput schema / properties / asset / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / asset / description
      Added value: +"What to show from that word onward: a registered clip id, or `card:<name>` for a card. An opaque key here, resolved by `build_shots` rather than checked against disk now."
    • removedInput schema / properties / asset / title
      Removed value: -"Asset"
    • addedInput schema / properties / asset / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • addedInput schema / properties / clip_id / description
      Added value: +"The transcript the cue is addressed against — the VO on a voiceover project, not the footage being shown. `asset` is what gets seen."
    • removedInput schema / properties / clip_id / title
      Removed value: -"Clip Id"
    • removedInput schema / properties / occurrence / anyOf
      Removed value: -[
      -  {
      -    "type": "integer"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / occurrence / description
      Added value: +"Disambiguate a phrase by count when it matches more than once, **1-based** in transcript order among the matches after `after`: 1 is the first, 2 the second. Unset, an ambiguous phrase is refused — listing every candidate's range and text — rather than guessed at."
    • removedInput schema / properties / occurrence / title
      Removed value: -"Occurrence"
    • addedInput schema / properties / occurrence / type
      Added value: +[
      +  "integer",
      +  "null"
      +]
    • removedInput schema / properties / path / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / path / description
      Added value: +"The project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing."
    • removedInput schema / properties / path / title
      Removed value: -"Path"
    • addedInput schema / properties / path / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedInput schema / properties / phrase / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / phrase / description
      Added value: +"Address the cue by what is said instead of by index. It binds to the phrase's **first** word — \"from this word onward\"."
    • removedInput schema / properties / phrase / title
      Removed value: -"Phrase"
    • addedInput schema / properties / phrase / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedInput schema / properties / src_start / anyOf
      Removed value: -[
      -  {
      -    "type": "number"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / src_start / description
      Added value: +"Where inside `asset` the shot reads from, in that asset's own source seconds — the number `describe_ls` reports for a window. An in-point and never a range: unpinned, the shot reads from wherever the per-asset cursor had got to, which is right for re-using a clip and wrong for showing the thing you searched for. A card takes none."
    • removedInput schema / properties / src_start / title
      Removed value: -"Src Start"
    • addedInput schema / properties / src_start / type
      Added value: +[
      +  "number",
      +  "null"
      +]
    • removedInput schema / properties / word_index / anyOf
      Removed value: -[
      -  {
      -    "type": "integer"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / word_index / description
      Added value: +"The word the picture starts on, in `clip_id`'s transcript. Give this or `phrase`, not both."
    • removedInput schema / properties / word_index / title
      Removed value: -"Word Index"
    • addedInput schema / properties / word_index / type
      Added value: +[
      +  "integer",
      +  "null"
      +]
    • removedInput schema / title
      Removed value: -"cue_addArguments"
  4. First observedv0.24.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already indicating read-write (readOnlyHint=false) and idempotent (idempotentHint=true), the description clearly adds critical behavioral context: the refusal behavior when a cue already exists at that exact word, the in-point semantics (never a range, out-point derived from next cue), and the refusal instead of rewind when the shot runs past the asset's end. It also explains the resolution of `asset` by `build_shots` rather than checking disk, and the additive metadata `cue_reresolve` can re-derive. These details go far beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every paragraph earns its place, each addressing a distinct aspect: basic usage, alternative addressing, source addressing, in-point semantics, and effects. It is front-loaded with the core purpose and immediately gives the alternative. While dense, it uses structured paragraphs and precise terminology, making it efficient for an agent to parse. There is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 15 parameters, 100% schema coverage, a clear output schema, and rich annotations, the description fills all major gaps. It covers what the tool does, how to address cues (word, phrase, event), what the parameters mean in context, and critical edge-case behaviors. The description integrates with the ecosystem by referencing sibling tools like `cue_rm`, `build_shots`, and `describe_ls`, and explains the output convention ('Echoes the resolved word plus three either side'). Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage with detailed per-parameter descriptions. The tool description adds significant semantic nuance beyond those: it explains the concept of 'from this word onward' for phrase binding, the in-point semantics for src_start, and the interplay between dissolve and punch. However, some parameters like punch_ease, punch_mode, and punch_seconds are only slightly extended in the description; the schema already covers them. The description's extra value is in clarifying the overall addressing model, which the schema only hints at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a crisp specification: 'Add a picture cue: from `word_index` of `clip_id` onward, show `asset`', and immediately provides a functional alternative for event-based cues. It clearly distinguishes this from sibling tools like cue_set, cue_rm, and cue_ls by naming the specific action and resource. The resource ('picture cue'), the addressing scheme, and the effect are all explicit, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive usage guidance: it explains when to use word_index vs phrase, when to use event (for recordings with no words), how to replace existing cues ('cue_rm it first'), and how to place a searched moment with src_start. It even contrasts with sibling tools like cue_set (which sets effects) and build_shots (which resolves the asset). The when-not conditions (e.g., 'a card takes no src_start') and alternatives (e.g., 'cue_rm it first') are explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools