Skip to main content
Glama
leancoderkavy

Premiere Pro MCP Server

Plan Active Speaker Reframe

plan_active_speaker_reframe
Read-onlyIdempotent

Plan vertical reframes that follow the active speaker or build stacked/split layouts from word timelines and speaker regions. Returns keyframes and apply routes without altering Premiere.

Instructions

Plan an active-speaker vertical reframe (Motion Scale/Position keyframes that follow whoever is talking) or a static stacked/split layout from a word timeline and static speaker regions. Returns framings, switches, keyframes, and apply routes. Local-only; never changes Premiere.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
layoutNoauto (default) picks active_speaker for 3+ speakers or a region wider than 0.6, stacked for exactly 2 speakers.
headroomNoFraction of the crop height reserved above the region top; defaults to 0.12.
frame_rateNoSequence frame rate used to snap times to frames; defaults to 30.
ease_framesNoFrames to ease each switch with bezier keyframes; 0 (default) emits hold keyframes (hard cuts).
source_frameYesSource clip frame size in pixels, e.g. 1920x1080 or 3840x2160.
target_frameNoTarget sequence frame size in pixels; defaults to 1080x1920.
word_timelineYesCaller-supplied word-timed transcript for one Premiere source item. Words must be ordered by start time and bound to the transcript revision returned by get_clip_transcript_uxp. Different labeled speakers may overlap; same-speaker words may not, even when another speaker is between them.
speaker_regionsYesNormalized (0..1) face/body rectangle of every speaker that appears, measured in the source frame.
min_hold_secondsNoNever switch faster than this; shorter turns are absorbed. Defaults to 1.5.
switch_lead_secondsNoSwitch this long before the speaker starts; defaults to 0.15.
base_video_track_indexNoVideo track holding the source clip in the target sequence; defaults to 0.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYesWhether the tool completed successfully.
dataNoTool-specific result data when ok is true; on failure, diagnostic detail when the tool provides it.
toolYesThe registered MCP tool name.
errorNoFailure detail when ok is false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.18.6
    • changedOutput schema / properties / data / description
      Previous value: -"Tool-specific result data when ok is true."New value: +"Tool-specific result data when ok is true; on failure, diagnostic detail when the tool provides it."
  2. Changed1 schema field changedv1.16.0
    • changedInput schema / properties / word_timeline / description
      Previous value: -"Caller-supplied word-timed transcript for one Premiere source item. Words must be ordered by start time and bound to the transcript revision returned by get_clip_transcript_uxp."New value: +"Caller-supplied word-timed transcript for one Premiere source item. Words must be ordered by start time and bound to the transcript revision returned by get_clip_transcript_uxp. Different labeled speakers may overlap; same-speaker words may not, even when another speaker is between them."
  3. Addedv1.14.9

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/destructiveHint=false, so 'Local-only; never changes Premiere' reinforces rather than establishes safety. The genuinely additive disclosure is that it 'Returns framings, switches, keyframes, and apply routes' — telling the agent this planner hands off to a separate apply step, which is non-obvious workflow context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler: the first front-loads the purpose and both output modes, the second covers the return and the local-only constraint. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter, nested-object planning tool with an output schema, the description covers the essential framing: inputs (word timeline, speaker regions), outputs (framings, switches, keyframes, apply routes), and safety. Nothing critical is missing, though it could have noted that a word timeline bound to a transcript revision is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 11 parameters, including the layout enum semantics, defaults, and bounds, so the schema already carries the parameter burden. The description only gestures at the layout alternatives ('active-speaker ... or a static stacked/split') without adding syntax, constraints, or default behavior beyond what the schema documents. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Plan an active-speaker vertical reframe') and clarifies the mechanism (Motion Scale/Position keyframes that follow whoever is talking), distinguishing it from the plain 'auto_reframe_sequence' sibling. It also names the alternate output mode (static stacked/split layout) up front, so an agent can tell immediately what family of tool this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the required source material ('from a word timeline and static speaker regions') and the non-destructive posture, which implies a plan-then-apply workflow. However it never explicitly says when to pick this over siblings like auto_reframe_sequence or plan_speaker_checkerboard, and the active_speaker vs stacked/split choice is left to the layout enum rather than routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools