Sequencer
Server Details
Create AI images, video, and audio, quote generation costs, and edit Sequencer projects.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 151 tools
Many tools have clearly distinct purposes (e.g., generate_image vs generate_video), and descriptions are detailed. However, with 151 tools, several overlap significantly—such as generate_audio vs generate_audio_track, upscale_video vs enhance_video_media, and multiple overlay-addition tools—creating potential for misselection despite guidance.
All tool names follow a consistent snake_case verb_noun pattern (e.g., create_edit, get_shot, delete_audio_track). Minor variations like get_edit_full or add_remotion_overlay still adhere to the same convention, making the set highly predictable.
With 151 tools, the server is extremely over-scoped for any single MCP integration. The guideline flags 50+ tools as an extreme mismatch, and this count is three times that threshold, making tool selection and maintenance impractical.
The surface covers a vast range of operations across media, edits, scenes, shots, assets, workflows, and productions, but notable gaps exist: no delete_edit, delete_media, delete_workspace, delete_workflow, or delete_folder. These missing lifecycle operations could cause agent dead ends for cleanup tasks.
Available Tools
151 toolsadd_audio_trackBInspect
Add an existing workspace audio media item to the edit timeline with precise placement, trim, fades, volume automation, channel, and optional shot pinning.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| volume | No | Base volume from 0 to 1 | |
| channel | No | Audio lane/channel index | |
| mediaId | Yes | Audio MediaDocV2 ID to place on the timeline | |
| trimEnd | No | Source trim end in seconds | |
| duration | No | Timeline clip duration in seconds | |
| startTime | No | Timeline start time in seconds | |
| trimStart | No | Source trim start in seconds | |
| fadeInTime | No | Fade-in length in seconds | |
| autoChannel | No | Pick the first non-overlapping channel automatically | |
| fadeOutTime | No | Fade-out length in seconds | |
| workspaceId | Yes | The workspace ID | |
| pinnedToShotId | No | Shot ID to pin this audio to, or null to unpin | |
| volumeKeyframes | No | Volume automation curve | |
| relativeStartTime | No | Offset in seconds from the pinned shot start |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It says 'add', but does not disclose what happens on channel collisions, how autoChannel interacts with an explicit channel, what relativeStartTime requires (a pinned shot), permission needs, or whether the call is idempotent. For a 15-parameter mutation, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the action front-loaded and no filler. It is efficient, though cramming six feature groups into one clause makes it read as a feature list rather than guiding the caller.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation with no annotations and no output schema, the description is adequate at the surface level but omits the cross-parameter interactions (autoChannel vs channel, relativeStartTime requiring pinnedToShotId) and any failure/error behavior an agent would want before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 15 parameters in detail. The description groups them into themes (placement, trim, fades, volume automation, channel, shot pinning), which is a useful mental model but adds no syntax or interaction detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (add) and resource (audio track on edit timeline) and scopes the source as an 'existing workspace audio media item', which implicitly separates it from generate_audio_track. It does not name any sibling explicitly, but the resource and the source constraint are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the word 'existing' — the agent can infer this places already-uploaded audio rather than synthesizing it, but the description gives no explicit when-to-use, prerequisites, or when to prefer generate_audio_track instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_media_overlayBInspect
Add an image or video media overlay to the edit timeline with precise transform, border radius, border, shadow, clipping, speed ramp, volume, animations, and keyframes.
| Name | Required | Description | Default |
|---|---|---|---|
| exit | No | Exit animation, or null to remove | |
| size | No | Normalized overlay size | |
| type | Yes | Overlay media type | |
| enter | No | Entrance animation, or null to remove | |
| width | No | Legacy shorthand for size.width | |
| border | No | Border styling, or null to remove | |
| editId | Yes | The edit ID | |
| height | No | Legacy shorthand for size.height | |
| shadow | No | Media overlay drop shadow, or null to remove | |
| volume | No | Video overlay audio volume | |
| clipEnd | No | Video overlay source trim end in seconds | |
| mediaId | Yes | Workspace MediaDocV2 ID to use as the overlay source | |
| opacity | No | Opacity from 0 to 1 | |
| trackId | No | Overlay track ID (auto-creates a track if not provided) | |
| duration | No | Timeline duration in seconds | |
| position | No | Normalized overlay position | |
| rotation | No | Rotation in degrees | |
| clipStart | No | Video overlay source trim start in seconds | |
| keyframes | No | Transform keyframes relative to item start | |
| positionX | No | Legacy shorthand for position.x | |
| positionY | No | Legacy shorthand for position.y | |
| speedRamp | No | Variable speed curve for video overlays | |
| startTime | No | Timeline start time in seconds | |
| workspaceId | Yes | The workspace ID | |
| borderRadius | No | Corner radius in 720p logical pixels | |
| playbackSpeed | No | Flat playback speed for video overlays |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but discloses no side effects, permissions, or behavioral traits. It does not mention that a track is auto-created (per schema), nor what happens to existing overlays, how mutations are reversible, or any auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core purpose and then enumerates capabilities. It is efficient and contains no filler, though the long feature list is a minor detraction from crispness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 26-parameter mutation tool with nested objects, no annotations, and no output schema, the description is far too thin. It should cover required parameters, track auto-creation, side effects, and return behavior, but only lists high-level feature categories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description enumerates feature categories (transform, border radius, shadow, etc.) but adds no syntax, defaults, or meaning beyond what the schema already provides for each parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Add') and resource ('image or video media overlay') scoped to the 'edit timeline'. It clearly distinguishes itself from siblings like add_text_overlay, add_audio_track, and add_remotion_overlay by specifying the media type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no alternatives, and no prerequisites. It lists capabilities but never says when to choose this tool over add_text_overlay or add_remotion_overlay.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_nodeBInspect
Add a node to a workflow. Valid node types: constant_text, constant_media_image, constant_media_video, constant_media_audio, constant_number, media_image, media_video, media_audio, llm_text, upscale_video, video_speed, video_stitch, composition, clone_voice, media_viewer, return_data, string_join
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Optional custom label for the node | |
| config | No | Node configuration/data values | |
| nodeType | Yes | Type of node to add (e.g., constant_text, media_image, llm_text) | |
| workflowId | Yes | Workflow ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It implies an additive mutation but does not disclose side effects, permission requirements, whether the node is persisted immediately, what happens to existing workflow structure, or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then supplies the valid node type list. It contains no filler, and every element is useful for correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema fully documents the five parameters, and the description adds the complete node type enumeration. However, with no annotations and no output schema, it omits behavioral details such as side effects, prerequisites, and return value, leaving meaningful gaps for a mutation tool with a nested config object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds substantial value by enumerating all valid nodeType values, going beyond the schema's brief examples and helping the agent avoid invalid node types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Add a node to a workflow." This clearly distinguishes it from delete_node, update_node, and connect_nodes by action, but it does not explicitly name or contrast those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use add_node versus alternatives such as update_node or connect_nodes, nor does it state prerequisites like needing an existing workflow or workspace. Usage is only implied by the verb "add."
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_overlay_trackCInspect
Create a new overlay track for text, image, or video overlays. Higher order tracks render above lower order tracks.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Track display name | Overlay |
| order | No | Stack order; higher renders on top | |
| editId | Yes | The edit ID | |
| locked | No | Whether the track is locked | |
| visible | No | Whether the track is visible | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the render-order semantics but says nothing about required permissions, whether a new track is empty, what identifier is returned for later item placement, or whether ordering collisions are resolved. For a mutation tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the creation purpose front-loaded and the ordering rule immediately after. Nothing is wasted, though the ordering sentence is partly redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description would ideally state what creating a track returns (e.g., a track ID needed by add_overlay_item). It is adequate for a simple creation call but leaves the return contract and lifecycle context unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so name, order, locked, visible and the two IDs are already documented. The description's stacking statement merely restates the schema's 'Stack order; higher renders on top' and adds no new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (overlay track) plus the kinds of overlay content it supports (text, image, video). It distinguishes the track-level tool from item-level siblings such as add_text_overlay and add_media_overlay, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to create a new track versus reusing an existing one, and no mention of the alternative add_audio_track or of the update_overlay_track/delete_overlay_track lifecycle. The only conditional information is the stacking behavior, which is not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_remotion_overlayAInspect
Create a custom Remotion React animation layer on the edit timeline. Compiles and test-renders it before saving. Call get_remotion_layer_guide first.
| Name | Required | Description | Default |
|---|---|---|---|
| exit | No | Exit animation, or null to remove | |
| size | No | Normalized overlay size | |
| enter | No | Entrance animation, or null to remove | |
| width | No | Legacy shorthand for size.width | |
| border | No | Border styling, or null to remove | |
| editId | Yes | ||
| height | No | Legacy shorthand for size.height | |
| shadow | No | Media overlay drop shadow, or null to remove | |
| opacity | No | Opacity from 0 to 1 | |
| trackId | No | ||
| duration | No | Timeline duration in seconds | |
| position | No | Normalized overlay position | |
| rotation | No | Rotation in degrees | |
| keyframes | No | Transform keyframes relative to item start | |
| positionX | No | Legacy shorthand for position.x | |
| positionY | No | Legacy shorthand for position.y | |
| startTime | No | Timeline start time in seconds | |
| workspaceId | Yes | ||
| borderRadius | No | Corner radius in 720p logical pixels | |
| remotionConfig | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one meaningful behavioral trait: the layer 'Compiles and test-renders it before saving,' telling the agent there is validation with a possible failure path before persistence. That is genuinely useful. It still omits auth/permission requirements, what happens on compile or render failure, whether the item can be edited afterward, and any latency expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, followed by behavioral disclosure and the prerequisite. Zero filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex, 20-parameter tool with nested objects and no output schema, so the description needs to carry more weight than a simple CRUD tool. It covers purpose, key behavior, and a prerequisite, but leaves gaps around failure semantics of the compile/test-render step, whether the created overlay is addressable by a returned id (relevant given the sibling update_overlay_item / delete_overlay_item tools), and timeline placement requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80% across 20 parameters, so the schema already documents nearly everything (size, position, enter/exit, keyframes, legacy shorthands, remotionConfig.source). The description adds no parameter-level meaning, which is acceptable given the high coverage — baseline 3 is correct here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Create a custom Remotion React animation layer on the edit timeline.' This is enough to distinguish it from the sibling add_text_overlay and add_media_overlay tools, since it is explicitly a Remotion React layer rather than a preset overlay. It stops short of differentiating further (e.g. vs. update_overlay_item or add_overlay_track), but the resource identity is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call get_remotion_layer_guide first' gives an explicit prerequisite, which is real routing guidance. However, it says nothing about when to choose this tool over the other overlay-creation siblings, nor what conditions make a Remotion layer the right choice. Usage is implied (custom React animation) but not contrasted against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_shot_media_refBInspect
Add a fine-grained media reference to a shot: image, video, startFrame, endFrame, or referenceFrame. Can also update legacy imageMediaId/videoMediaId for compatibility.
| Name | Required | Description | Default |
|---|---|---|---|
| role | Yes | How this media is used in the shot | |
| order | No | Tab/order within the shot; auto-appends if omitted | |
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID | |
| mediaId | Yes | Referenced MediaDocV2 ID | |
| workspaceId | Yes | The workspace ID | |
| versionOverride | No | Optional media version override | |
| replaceExistingRole | No | Remove existing refs with the same role before adding this one | |
| alsoSetPrimaryMediaId | No | For image/video roles, also set imageMediaId/videoMediaId |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden for a mutation tool, yet it does not state side effects on existing refs, permission/auth needs, or reversibility. The legacy-compatibility note overlaps with the schema's alsoSetPrimaryMediaId description, so little new behavioral context is added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the core purpose front-loaded and the role enumeration immediately useful. No filler or redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema is rich and fully documented, which offsets the missing output schema, but for a nine-parameter mutation tool with zero annotations the description is thin on consequences and edge cases. It is minimally adequate but leaves gaps an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented, giving a baseline of 3. The description's mention of roles and legacy media ids restates what the schema already provides rather than adding syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Add a fine-grained media reference to a shot") and enumerates the exact role values, so an agent understands the operation immediately. It does not explicitly distinguish itself from close siblings like update_shot_media_ref or delete_shot_media_ref, leaving that to the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the verb and the role list; there is no explicit when-to-use versus update_shot_media_ref or add_media_overlay, and no stated prerequisites. The one contextual hint (legacy compatibility) is a behavior note rather than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_text_overlayAInspect
Add a text overlay to the video timeline. Use a presetId to auto-apply a preset style (cinematic-title, lower-third, subtitle, chapter-card, quote, social, breaking-news, credits-roll), or provide custom textConfig. The text field overrides the preset default text. Auto-creates an overlay track if none exist.
| Name | Required | Description | Default |
|---|---|---|---|
| exit | No | Exit animation | |
| size | No | Normalized overlay size; overrides width/height if provided | |
| text | Yes | The text to display | |
| color | No | Text color as hex (e.g. #FFFFFF) or rgba (overrides preset) | |
| enter | No | Entrance animation | |
| width | No | Width (normalized 0–1 of canvas). Uses preset default if not specified. | |
| border | No | Overlay item border styling | |
| editId | Yes | The edit ID | |
| height | No | Height (normalized 0–1 of canvas). Uses preset default if not specified. | |
| shadow | No | Overlay item drop shadow | |
| opacity | No | Overall opacity 0–1 (default 1) | |
| trackId | No | Overlay track ID (auto-creates a track if not provided) | |
| duration | No | How long to show the text (seconds, uses preset default if not specified) | |
| fontSize | No | Font size in points (overrides preset) | |
| position | No | Normalized center position; overrides positionX/positionY if provided | |
| presetId | No | Text preset ID (cinematic-title, lower-third, subtitle, chapter-card, quote, social, breaking-news, credits-roll). Provides default styling, animation, and position. | |
| rotation | No | Rotation in degrees | |
| fontStyle | No | Font style (overrides preset) | |
| keyframes | No | Transform keyframes relative to item start | |
| positionX | No | Horizontal position (0 = left, 0.5 = center, 1 = right). Uses preset default if not specified. | |
| positionY | No | Vertical position (0 = top, 0.5 = center, 1 = bottom). Uses preset default if not specified. | |
| startTime | No | When to show the text overlay (seconds from edit start) | |
| textAlign | No | Text alignment | |
| fontFamily | No | Font family (overrides preset). Available: Inter, Outfit, Playfair Display, Bebas Neue, Roboto Mono | |
| fontWeight | No | Font weight as 100-900, normal, or bold (overrides preset) | |
| lineHeight | No | Line-height multiplier | |
| lineStyles | No | Per-line style overrides keyed by zero-based line index | |
| textConfig | No | Full TextEffectConfig override/extension | |
| strokeColor | No | Text stroke/outline color | |
| strokeWidth | No | Text stroke/outline width | |
| workspaceId | Yes | The workspace ID | |
| borderRadius | No | Overlay item corner radius in 720p logical pixels | |
| letterSpacing | No | Letter spacing in points | |
| textAnimation | No | Per-character/word/line text animation | |
| textShadowBlur | No | Text shadow blur radius | |
| backgroundColor | No | Text background color | |
| textShadowColor | No | Text shadow color | |
| textShadowOffsetX | No | Text shadow horizontal offset | |
| textShadowOffsetY | No | Text shadow vertical offset | |
| backgroundPaddingX | No | Text background horizontal padding | |
| backgroundPaddingY | No | Text background vertical padding | |
| backgroundBorderRadius | No | Text background corner radius |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the auto-creation of an overlay track and the preset/text override precedence, but says nothing about permissions, failure behavior, or what is returned for a mutation of this size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and then the preset/custom choice and the track side effect. No filler, though the parenthetical preset list is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 42-parameter mutation with no annotations and no output schema, the description covers the primary workflow adequately since the schema documents all parameters. It is still thin on behavior (no return value, no permissions/error context) given the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema lacks: the presetId-to-style mapping, the fact that text overrides the preset default text, and the presetId-or-textConfig alternative. That interaction between parameters is not expressed anywhere in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a text overlay to the video timeline'), which clearly separates it from media/audio/remotion overlay siblings. It does not name a sibling explicitly (e.g. add_media_overlay), so it stops short of the 5-level differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the two invocation routes (presetId vs custom textConfig) and enumerates the preset IDs, which is helpful. However it never states when to use this tool versus siblings like add_media_overlay, update_overlay_item, or add_overlay_track, and gives no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_transitionBInspect
Add a single timeline transition between adjacent shots without replacing the whole transition list.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Optional stable transition ID | |
| type | Yes | FFmpeg xfade transition name, e.g. fade, dissolve, wipeleft, slideleft | |
| editId | Yes | The edit ID | |
| duration | Yes | Transition overlap duration in seconds | |
| toShotId | No | Incoming shot ID | |
| startTime | Yes | Nominal boundary time in seconds | |
| boundaryId | No | Boundary ID in the form fromShotId->toShotId | |
| fromShotId | No | Outgoing shot ID | |
| workspaceId | Yes | The workspace ID | |
| replaceExistingBoundary | No | Replace existing transition for the same boundaryId/fromShotId/toShotId |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden for a mutation tool. It conveys additivity but says nothing about required permissions, what happens on a boundary conflict, or side effects; it also doesn't reconcile the additive framing with the schema's replaceExistingBoundary defaulting to true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero waste; the key differentiator ('without replacing the whole transition list') is front-loaded and immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no annotations and no output schema, the description is thinner than ideal, though the 100% schema coverage compensates for parameter detail. An agent gets the operation's intent but little about side effects or failure conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters including type, duration, startTime, and the boundary identifiers. The description adds no parameter-level meaning beyond that, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Add a single timeline transition between adjacent shots') and clarifies the additive scope versus a wholesale replacement. It is distinguishable from delete_transition/update_transition by the 'add a single' framing, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without replacing the whole transition list' implies when to prefer this over a bulk-transition operation, but no alternative is named and no condition or prerequisite (e.g. shots must be adjacent) is stated. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_export_ipBInspect
Run IP and copyright analysis on the completed exported video. Analyzes the actual export artifact when possible, caches the result on the export doc, and returns findings plus instructions for generating the compliance report.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| exportId | No | The export ID to analyze. If omitted, the latest completed export for the edit is analyzed. | |
| workspaceId | Yes | The workspace ID | |
| forceRefresh | No | Re-run analysis even if this export already has a cached IP result |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose real behavior beyond the schema: analysis targets the actual export artifact 'when possible' (implying a fallback), the result is cached on the export doc, and the return includes findings plus report-generation instructions. It omits what happens if no completed export exists, cost/latency, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, and each sentence adds distinct information (what is analyzed/cached, what is returned). No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no annotations and no output schema, the description does cover the return ('findings plus instructions for generating the compliance report') and caching semantics. It is still incomplete on preconditions (no completed export, in-flight exports), failure behavior, and its relationship to generate_ip_report.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so editId, exportId, workspaceId, and forceRefresh are already fully documented, including the default-to-latest-export behavior. The description adds no parameter-level detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair: 'Run IP and copyright analysis on the completed exported video.' An agent can distinguish this from generic analysis or generation tools. It does not name the obvious sibling generate_ip_report, so differentiation is only implied via 'instructions for generating the compliance report.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'on the completed exported video' implies a precondition (the export must exist and be finished), and the schema notes a default of the latest completed export. However, there is no explicit when-to-use versus generate_ip_report or when-not guidance, so routing between the two compliance-related tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
change_audio_track_voiceBInspect
Create a new voice-changed version of an audio media item and optionally keep the existing timeline track pointing at it.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| modelId | Yes | Voice-change/speech-to-speech model ID | |
| trackId | Yes | Audio track whose media should be voice-changed | |
| voiceId | No | Provider voice ID | |
| workspaceId | Yes | The workspace ID | |
| voiceReferenceUrl | No | Reference audio URL for clone/reference mode |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the result is a NEW asset (non-destructive) and that the track can optionally be repointed, which are real behavioral traits. However, it omits whether the operation is long-running/async, whether it costs credits, and what happens to the source media.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action, the result, and the optional side effect with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation/generation tool with no annotations and no output schema, the description is adequate but incomplete: it explains the create-plus-optional-track behavior yet says nothing about the return value, async nature, or failure modes an agent would need to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (including modelId, voiceId, voiceReferenceUrl) are already documented in the schema. The description adds no format, constraint, or dependency detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (a new voice-changed version of an audio media item), and clarifies the operation is additive rather than an in-place edit. It implicitly distinguishes itself from the change_video_voice sibling by scoping to audio, but never names that alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the near-identical change_video_voice sibling or when to prefer generate_audio_track. The only usage hint is the optional track-repointing behavior, which is about mechanics, not selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
change_video_voiceAInspect
Run the video voice-change pipeline: extract source audio, apply speech-to-speech voice change, remux the result into the source video, and save a completed media version.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | Video MediaDocV2 ID to process | |
| modelId | Yes | Voice-change/speech-to-speech model ID | |
| voiceId | No | Provider voice ID | |
| videoUrl | No | Override source video URL; otherwise resolves from active media version | |
| workspaceId | Yes | The workspace ID | |
| voiceReferenceUrl | No | Reference audio URL for clone/reference mode |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that this is a multi-stage mutation that produces a new completed media version, but omits whether the operation is long-running/async, whether the original media version is altered, what permissions or credits are required, and how failures mid-pipeline are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the operation first and then the pipeline stages. No filler, no repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation pipeline with no annotations and no output schema, the description covers what happens but not how the caller should behave: no async/polling expectations, no indication of what the completed media version returns, and no differentiation from sibling voice/lip-sync tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (including voiceId, videoUrl override, and voiceReferenceUrl) are already documented in the schema. The description adds no syntax, format, or dependency guidance (e.g., clone vs. reference mode) beyond that baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run the video voice-change pipeline') and then enumerates the exact pipeline stages, which lets an agent distinguish it from change_audio_track_voice (audio-only) and lip_sync_video_media without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no routing against obvious alternatives such as change_audio_track_voice or lip_sync_video_media. The agent must infer from the name that this operates on a whole video rather than an audio track.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_asset_statusAInspect
Check the approval status of all assets associated with an edit. Used to determine if the user has approved/rejected each asset before proceeding to scene and shot creation.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID to check assets for | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the safety profile; 'check' implies a read, but it never states read-only/non-mutating behavior, permissions, or error behavior. It does disclose the return semantic (per-asset approved/rejected status), which is the main behavioral value added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core purpose front-loaded and the workflow rationale second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with no output schema, the description covers purpose, scope, and the shape of the result (approval status per asset). It could say slightly more about the returned status values or ordering, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so editId and workspaceId are already documented in the schema. The description adds no extra meaning about their format, required-ness, or relationship beyond 'associated with an edit' — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check), resource (asset approval status), and scope (all assets associated with an edit). This scope wording separates it from single-asset siblings like get_asset and from generic list_assets, which does not surface approval state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the decision context: use it to determine whether the user approved/rejected each asset before proceeding to scene and shot creation. That is a clear when-to-use signal, though it names no alternative (e.g. get_asset for a single asset) and states no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connect_nodesAInspect
Connect two nodes in a workflow. The sourceSocketId MUST be an output of the source node, and the targetSocketId MUST be an input of the target node. Use the exact socket IDs returned by add_node.
| Name | Required | Description | Default |
|---|---|---|---|
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID | |
| sourceNodeId | Yes | The node to connect FROM | |
| targetNodeId | Yes | The node to connect TO | |
| sourceSocketId | Yes | The OUTPUT socket ID on the source node (e.g., "out_img", "val", "result") | |
| targetSocketId | Yes | The INPUT socket ID on the target node (e.g., "prompt", "in_img", "media") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses a validation rule (sockets must match node direction), which is behavioral context beyond a mere label, but it says nothing about failure behavior on invalid sockets, whether duplicate connections are rejected or replaced, permission/auth requirements, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, the core action front-loaded, and every sentence adds a distinct constraint. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-required-parameter mutation tool with no annotations and no output schema, the description covers the socket semantics well but omits error behavior, duplicate-connection handling, and what the tool returns on success. Adequate but with visible gaps an agent would likely need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema already defines each field. The description adds a cross-parameter constraint the schema does not state – sourceSocketId must be an output of the source node and targetSocketId an input of the target node – which helps the agent avoid an invalid pairing. It still says nothing about the workflowId/workspaceId relationship.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Connect two nodes in a workflow') and is immediately distinguishable from siblings like add_node, update_node, and delete_connection. The follow-on sentences narrow the scope to a socket-level linkage, so the agent knows exactly what operation is being performed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: use the exact socket IDs returned by add_node, and the sockets must be an output of the source / input of the target. It references the sibling that produces the required IDs (add_node) but does not mention delete_connection for undoing a link or any ordering/prerequisite beyond add_node.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_video_to_hdrconvert video to hdrADestructiveInspect
Convert an existing video to HDR using a catalog model that supports video_hdr. Select output format in modelSettings when supported. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-readOnly, openWorld, non-idempotent, destructive operation, and the description adds that it returns a tracked task with a playable result when complete. It does not warn that this is a paid generation, nor mention cost capping or idempotency-key behavior despite the tool being destructive and expensive. With annotations carrying the safety profile, this is a reasonable 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the conversion action and the model prerequisite, with no filler. It is slightly terse given the tool's complexity, but nothing wastes space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the schema covers most parameters, so the description is adequate for invocation. However, for a 21-parameter, destructive, paid generation tool, it omits cost/idempotency guidance and makes no attempt to orient the agent among the many video-mutation siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 81%, so the schema already documents most parameters and the baseline is 3. The description adds only marginal meaning by pointing at modelSettings for output-format selection; it says nothing about prompt, mediaInputs, idempotencyKey, maxCostUsd, or referenceImages semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Convert an existing video to HDR') and adds a real constraint (the catalog model must support video_hdr). It does not explicitly contrast with near-neighbors such as enhance_video_media or edit_video, so an agent must still infer the boundary, but the HDR specificity is strong enough to identify the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an actionable prerequisite: 'Use get_model_catalog to choose a supported model and its settings,' which routes the agent to the sibling needed before calling. There is no explicit when-not guidance or comparison to the other video-mutating siblings, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_assetAInspect
Create a new character, location, prop, or legacy text-only style asset in a workspace. Character, location, and prop prompts generate a reference image. Style assets do not generate images; routine projects should store their global visual style in edit.style instead. existingMediaUrl may attach an explicit user-supplied reference image.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Asset name | |
| type | Yes | Type of asset to create | |
| editId | No | Edit ID — REQUIRED when creating assets within an edit context. Assets without editId will NOT appear in the editor Assets tab. | |
| prompt | No | Prompt for generating a character, location, or prop reference image. For a legacy style asset, this is normalized into its text description and no image is generated. Ignored if existingMediaUrl is provided. | |
| modelId | No | Model ID for character, location, or prop image generation. Pass the user's preferred model from MODEL_PREFERENCES if set. Defaults to Nano Banana Pro (google-gemini-3-image). Ignored for style assets and when existingMediaUrl is provided. | |
| folderId | No | Folder to place the asset in | |
| description | No | Asset description. For style assets, this must be one short sentence describing only the visual aesthetic shared by every shot; do not list characters, locations, props, actions, or shot-specific objects. | |
| workspaceId | Yes | The workspace ID | |
| existingMediaUrl | No | Download URL of an existing image to use as the asset reference image (e.g. from a user-attached file). When provided, skips image generation and uses this image directly. | |
| voiceDescription | No | For character assets: voice/accent/personality description used for speech generation, e.g. "slow casual Gen Z woman, vocal fry, Valley girl / LA accent". | |
| existingMediaFileName | No | Original file name for existingMediaUrl. Used when copying the uploaded image into workspace media storage. | |
| existingMediaMimeType | No | MIME type for existingMediaUrl. Used when copying the uploaded image into workspace media storage. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose key side effects: character/location/prop creation triggers reference-image generation, style assets do not, and supplying existingMediaUrl skips generation entirely. It does not mention permission requirements, credit/cost implications of image generation, or whether duplicate names are allowed, which are relevant for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the purpose front-loaded, followed by the behavioral distinction and the optional-parameter note. Every sentence adds distinct information with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter creation tool with no output schema and no annotations, the description covers the essential behavioral distinctions (image generation vs. text-only, edit.style alternative, existing-media path). The remaining gaps, such as the editId-in-editor requirement and voiceDescription, are already documented in the schema, so the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 12 parameters, so the schema already explains prompt, modelId, editId, and the existingMedia* fields in detail. The description reinforces the prompt and existingMediaUrl semantics but adds little beyond what the schema states; baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (asset) and enumerates the four asset subtypes it accepts, which matches the type enum exactly. An agent can distinguish this from siblings like generate_image, create_shot, or create_folder without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear routing rule for the ambiguous case: style assets do not generate images, and routine projects should store global visual style in edit.style instead. It also clarifies when existingMediaUrl applies (user-supplied reference image). It does not, however, address when to prefer update_asset over create_asset or when an asset belongs in an edit context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_editAInspect
Create a new edit (project) in a workspace. Aspect ratio is required because it controls every generated shot; use 9:16 for vertical/social films unless the user asks for another format. Defaults are Nano Banana Pro (google-gemini-3-image) for still images, Gemini Omni 1.1 Flash (google-gemini-omni-1-1) for videos, and Cartesia Sonic 3.5 (cartesia-sonic-3-5) for speech. Returns the new edit ID.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Title for the new edit | |
| aspectRatio | Yes | Required project aspect ratio. Use 9:16 for vertical/social films unless the user explicitly asks otherwise. | |
| workspaceId | Yes | The workspace ID | |
| defaultImageModel | No | Default still-image model ID. Defaults to Nano Banana Pro (google-gemini-3-image). | |
| defaultVideoModel | No | Default video model ID. Defaults to Gemini Omni 1.1 Flash (google-gemini-omni-1-1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that aspect ratio is required and controls every generated shot, lists default model choices (useful for resource allocation), and states the return value ('Returns the new edit ID'). It doesn't mention side effects like whether it consumes credits or what permissions are needed, but it covers more than typical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single coherent paragraph of three sentences, front-loaded with the core purpose. It packs multiple useful details (aspect ratio rationale, model defaults, return value) without egregious redundancy, though the aspect ratio guidance is repeated in both description and schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, no output schema, and no annotations, the description is fairly complete: it explains the required aspect ratio, provides defaults for models, and states the return ID. It lacks discussion of error conditions or workspace permissions, but covers the essentials for a creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds marginal value by reiterating the aspect ratio guidance and model defaults, but doesn't introduce new syntax or constraints beyond what's in the schema. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a new edit (project) in a workspace') and distinguishes from siblings like create_production and create_shot by naming the 'edit' entity. The parenthetical '(project)' clarifies the domain term.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides guidance on aspect ratio selection ('use 9:16 for vertical/social films unless the user asks for another format'), which is a usage heuristic. However, it doesn't differentiate when to use create_edit versus alternatives like create_production or create_workspace, nor does it state prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_folderCInspect
Create a new folder in a workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Folder name | |
| workspaceId | Yes | The workspace ID | |
| parentFolderId | No | Parent folder ID (null for root) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it says almost nothing beyond the bare operation. It omits permission requirements, name-uniqueness or collision behavior, whether parentFolderId defaults to root, and whether the operation is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence that is front-loaded with the verb and resource and contains no filler. It is efficient, though its brevity comes at the cost of the missing details noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter create operation with full schema coverage and no output schema, the description is minimally viable. Gaps remain around permissions, name constraints, and parent-folder handling that an agent would benefit from knowing before calling a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (name, workspaceId, parentFolderId) are already documented in the schema. The description adds no additional meaning such as naming rules or parent-folder semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (folder) with a clear scope (in a workspace), so an agent understands what the tool does. It does not, however, differentiate itself from nearby siblings such as list_folders or move_item_to_folder, which handle related folder operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of related tools like list_folders or move_item_to_folder. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_productionADestructiveInspect
Generate a product ad, 3 to 5 shot story, or localized image creative as an editable project. Provide authored shot prompts. For product ads pass the actual product image as a reference; use extract_product_image with a product URL or upload first. For localization supply the original creative, target language as a parameter, and instructions to preserve the product and layout. Returns tracked tasks; poll get_production_status.
| Name | Required | Description | Default |
|---|---|---|---|
| definition | Yes | ||
| parameters | No | ||
| workspaceId | Yes | ||
| referenceImages | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write/destructive/open-world/non-idempotent profile, so the description's job is the rest. It adds genuinely useful behavior: the call returns tracked tasks and the caller must poll get_production_status, and the output is an editable project rather than a finished render. It does not explain cost implications or why destructiveHint is set for a create operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core capability, then branches per mode, and closes with the async polling note. Every clause is informative, though the middle section is dense enough that the three modes run together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the modes, prerequisites, and the async return path, which is important given there is no output schema. For a nested-schema, four-parameter creation tool with zero parameter descriptions in the schema, however, key invocation details (valid model strings, type/resolution/aspectRatio behavior, workspace scoping) remain undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters (including a nested `definition` object), so the description carries the burden. It clarifies intent for `referenceImages` (product ads) and `parameters` (localization target language) and the authored shot prompts, but says nothing about `workspaceId`, `model` values, `type`, `aspectRatio`, or shot duration limits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource and enumerates the three supported modes (product ad, 3-5 shot story, localized creative), which map directly to the `kind` enum. It does not, however, contrast itself against closely-related siblings like run_production_template or create_edit, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use context: product ads require passing the real product image as a reference and using extract_product_image or an upload first; localization requires the original creative, target language as a parameter, and preservation instructions. No explicit when-not-to-use or named alternative (e.g., run_production_template) is supplied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_sceneCInspect
Create a new scene in an edit.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Scene name | |
| order | No | Position in scene order | |
| editId | Yes | The edit ID | |
| locationRef | No | Asset ID of a location asset to associate | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers almost nothing: it does not say where a new scene is inserted (order default), whether locationRef is optional association, what permissions are required, or what is returned. For a mutation tool with zero annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single six-word sentence with no filler, front-loaded with the action. It is efficient but borders on under-specification rather than true conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description leaves out ordering behavior, defaults, and side effects. The schema covers parameter naming, but the agent still lacks the behavioral context needed to invoke this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (name, order, editId, locationRef, workspaceId) are already documented in the schema. The description adds no syntax, defaults, or semantics beyond it, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Create a new scene") and scopes it to an edit, so the operation is unambiguous. It does not distinguish itself from siblings like create_shot, create_edit, or duplicate_shot, but the core action is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus create_shot or create_edit, and no prerequisites stated. "in an edit" faintly implies an edit context is needed, but the required workspaceId/editId relationship is left entirely to the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_shotBInspect
Create a new shot in an edit. Optionally provide a prompt to auto-create a linked MediaDocV2 for image generation. When videoPrompt is omitted, it falls back to the visual image prompt only. Dialogue is never appended automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| order | No | Position in shot order | |
| editId | Yes | The edit ID | |
| prompt | No | Image prompt for AI generation (visual description ONLY — do NOT include dialogue text). Will create a linked MediaDocV2. | |
| volume | No | Shot video audio volume | |
| clipEnd | No | Source clip trim end in seconds; defaults to duration | |
| sceneId | Yes | The scene ID this shot belongs to | |
| dialogue | No | Visible on-screen character speech for this shot. Stored separately for intentional lip-sync/TTS use. Do not use for off-screen narration or captions. | |
| duration | No | Editorial shot duration in seconds. Five seconds is only a compatibility fallback. For planned edits, choose and pass the duration that fits this shot's narration or story beat. | |
| clipStart | No | Source clip trim start in seconds | |
| mediaRefs | No | Fine-grained shot media refs for image/video/start/end/reference frame tabs | |
| speedRamp | No | Variable playback speed curve | |
| startTime | No | Timeline start time in seconds; usually derived by the editor | |
| rawDuration | No | Full raw source duration in seconds | |
| videoPrompt | No | Video/motion prompt for camera movement and action. Include speech only when visible on-screen character lip-sync is intentionally required. Never include off-screen narration. | |
| workspaceId | Yes | The workspace ID | |
| imageMediaId | No | Existing image media reference | |
| taggedAssets | No | Asset IDs to tag in this shot | |
| videoMediaId | No | Existing video media reference | |
| playbackSpeed | No | Shot playback speed multiplier | |
| hideStartFrame | No | Hide/remove the Start Frame tab | |
| linkEndToNextStart | No | Auto-use next shot start frame as this shot end frame | |
| linkStartToPrevEnd | No | Auto-use previous shot end frame as this shot start frame |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose genuinely non-obvious behavior: supplying prompt auto-creates a linked MediaDocV2, videoPrompt omission falls back to the visual image prompt, and dialogue is never auto-appended. It stops short of covering permissions, failure behavior, or ordering side effects on the edit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and followed by two conditional behaviors. Every sentence adds information and none is filler, though the fallback sentence is terse enough to need the schema for full meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter mutation tool with nested objects, no annotations, and no output schema, the description covers only the prompt-related surface. Required IDs, default/derived fields (startTime, duration), and sibling selection are left entirely to the schema, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 22 parameters, establishing a baseline of 3. The description adds modest cross-parameter meaning (prompt vs videoPrompt fallback, dialogue separation) but does not clarify the nested mediaRefs/speedRamp structures or the linkStartToPrevEnd-style options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a new shot in an edit.' An agent immediately knows what the tool produces. However, it never distinguishes itself from close siblings such as create_shot_with_media or duplicate_shot, so the agent must infer the boundary from names alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through the optional prompt/videoPrompt behavior, giving some sense of when to supply generation prompts. But it offers no explicit when-to-use versus create_shot_with_media, no prerequisites, and no mention that the three IDs are required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_shot_with_mediaAInspect
Create a shot with an image prompt in one call. Resolves tagged assets (characters, locations) to get descriptions and reference images, applies the prompt formula from edit settings, then either triggers generation immediately (if all reference images are ready) or queues for dependency-driven generation.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Image model ID (uses the edit default or Nano Banana Pro if not specified) | |
| order | No | Shot order within the scene (auto-appends to end if not provided) | |
| editId | Yes | The edit ID | |
| prompt | Yes | Image generation prompt (visual description ONLY). Use @tags for character/location references (e.g. "@John walks through @CityPark"). Never include dialogue, narration, or spoken text here because it may render visibly in the image. | |
| sceneId | Yes | The scene this shot belongs to | |
| dialogue | No | Visible on-screen character speech for this shot. Stored separately for intentional lip-sync/TTS use. Do not use for off-screen narration or captions. | |
| duration | No | Editorial shot duration in seconds. Five seconds is only a compatibility fallback. For planned edits, choose and pass the duration that fits this shot's narration or story beat. | |
| aspectRatio | No | Fallback aspect ratio only used if the edit has no aspectRatio. Shots normally inherit the edit aspect ratio. | |
| videoPrompt | No | Video/motion prompt describing camera movement and action for video generation. Include speech only when visible on-screen character lip-sync is intentionally required. Never include off-screen narration. | |
| workspaceId | Yes | The workspace ID | |
| taggedAssets | No | Assets to tag in this shot (characters, locations, props, styles) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and discloses important traits: asset resolution, prompt-formula application, immediate generation when references are ready, and queuing when they are not. It omits permissions, cost implications, or failure behavior, but covers the core conditional execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and followed by the execution logic. Every clause adds useful information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter creation/generation tool with no output schema and no annotations, the description explains the key workflow and conditional generation path. It could go further by describing return values, permissions, or cost behavior, but it is sufficient for correct invocation based on the rich input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameter documentation is already thorough in the schema. The description adds only high-level context about tagged assets and the prompt formula, not syntax or edge cases beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, resource, and scope: creating a shot with an image prompt in a single call. It also explains the internal workflow that distinguishes it from a plain create_shot sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The first sentence and workflow imply when to use the tool, but there is no explicit guidance on when to prefer it over siblings like create_shot or generate_shot_video, nor any stated exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_workflowAInspect
Create a new workflow in a workspace. After creating, use add_node to add nodes and connect_nodes to wire them up.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Workflow title | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully signals this is a mutation step that precedes node creation, but omits whether the new workflow's identifier is returned (critical for chaining into add_node), any permission requirements, and idempotency behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the follow-up guidance second. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description should ideally say what is returned (notably the new workflow ID needed by add_node). The chaining hint partially compensates, but the return-value gap leaves it merely adequate for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (title, workspaceId), so the schema already documents them. The description adds no format or constraint detail beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Create a new workflow') and scopes it to a workspace, which cleanly separates it from get_workflow, list_workflows, and execute_workflow. It stops short of naming any sibling explicitly, so it earns a solid 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear procedural context: after creation, use add_node and connect_nodes. That tells the agent how this tool fits in the workflow-building sequence, though it offers no when-not-to-use guidance or alternative creation paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_workspaceCInspect
Create a new workspace (project container).
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Title for the new workspace |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: not whether the caller needs permissions, not whether the new workspace becomes active/selected, not whether the operation is idempotent, and not what happens to existing state. 'Create' implies mutation, but none of the mutation's consequences are described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the verb and resource lead, and the parenthetical clarification is the only extra word. Nothing is padded or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should say more than it does for a creation tool: what the created workspace looks like on return, whether it is scoped to a user/account, and how it relates to sibling containers. As written, an agent can call it but cannot predict the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (title) exists and schema description coverage is 100%, so the schema fully documents it. The description adds no format, length, or uniqueness constraints beyond the schema, which is the expected baseline when the schema already does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new workspace') and adds a clarifying gloss ('project container') that tells an agent what a workspace actually is. It does not, however, differentiate itself from the many sibling creation tools such as create_production, create_folder, or create_edit, so an agent must infer which container to create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to create a workspace versus a production, folder, or scene, and no prerequisites or sequencing hints. The agent gets a definition of the object but nothing about when this tool is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_assetAInspect
Delete an asset from a workspace. Referenced media documents are NOT deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | The asset ID to delete | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses a key side-effect boundary: referenced media documents are not deleted, which is important for a destructive operation. However, it omits whether deletion is reversible, permission requirements, and error behavior for referenced assets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The scoping caveat is appropriately placed second. It could be slightly more efficient but wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description is minimal but covers the core action and one important side-effect boundary. It leaves gaps around permissions, reversibility, and failure modes, which an agent might need for safe operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (workspaceId, assetId) are already fully documented in the schema. The description adds no parameter-specific meaning beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (asset from a workspace), and crucially clarifies a distinction from other delete siblings: referenced media documents are NOT deleted. This scoping detail differentiates it from more invasive delete tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for asset deletion but offers no explicit when-to-use guidance, alternatives (e.g., update_asset vs delete_asset), or prerequisites such as whether the asset must be unreferenced or reviewed first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_audio_trackAInspect
Delete an audio track from the edit timeline. The underlying media document is preserved.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| trackId | Yes | Audio track ID to delete | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It helpfully discloses that the underlying media document is preserved, but it omits whether the deletion is reversible, what permissions are required, and what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that are front-loaded and free of waste. The preservation note earns its place by clarifying destructive scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a three-parameter mutation tool with full schema coverage and no output schema. However, the lack of any usage guidance or reversibility/permission context leaves gaps for an agent deciding whether and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents workspaceId, editId, and trackId. The description adds no syntax or format detail beyond what the schema provides, meriting the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete an audio track') and scopes it to 'the edit timeline', which identifies the domain and distinguishes it from media-level deletions like delete_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives named. Siblings like update_audio_track and list_audio_tracks exist but the description doesn't route the agent to them, nor does it state prerequisites such as the track first existing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_connectionCInspect
Delete a connection (edge) between nodes.
| Name | Required | Description | Default |
|---|---|---|---|
| edgeId | No | Edge ID to delete | |
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID | |
| sourceNodeId | No | Alternative: delete all edges from this source node | |
| targetNodeId | No | Alternative: delete all edges to this target node |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It says 'delete' but discloses nothing about reversibility, required permissions, whether removing by source/target node deletes multiple edges at once, or behavior when the edge does not exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. However, for a destructive tool with three alternative deletion modes, this brevity shades into under-specification rather than tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with no annotations and no output schema, where the description omits the crucial point that edgeId and the node-based parameters are competing deletion strategies. An agent could pass conflicting parameters without warning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already explains edgeId, workspaceId, sourceNodeId and targetNodeId. The description adds only the conceptual framing that a connection links two nodes; it adds no syntax or selection logic beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Delete a connection (edge)') and clarifies the ambiguous term 'connection' with 'edge'. It is distinguishable from delete_node and connect_nodes, though it doesn't explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus delete_node, delete_transition, or connect_nodes, and no indication that the three delete modes (by edgeId, by sourceNodeId, by targetNodeId) are mutually exclusive alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_nodeCInspect
Delete a node from a workflow.
| Name | Required | Description | Default |
|---|---|---|---|
| nodeId | Yes | ||
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full burden of behavioral disclosure, yet provides none. It doesn't state whether deletion is permanent, what happens to connected edges, whether it requires specific permissions, or what side effects occur. 'Delete' implies destruction but nothing beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, economical sentence that is front-loaded with the action. There is no wasted text, but the brevity reflects under-specification rather than optimal conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, three-parameter operation with no annotations and no output schema, the description is completely inadequate. It omits essential information about scope, side effects, permissions, and error conditions, leaving the agent without the context needed to invoke it safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has three required parameters with only 33% description coverage (only workspaceId is described, as 'The workspace ID'). The description adds no meaning for nodeId or workflowId, and doesn't explain the relationship between them or the hierarchical scoping (workspace > workflow > node). It fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a node from a workflow'), which is marginally clearer than a tautology. However, it provides no differentiation from the many sibling delete_* tools (delete_connection, delete_transition, delete_audio_track, etc.), and doesn't clarify the context (workflow graph node vs. other node types). It's essentially a minimal restatement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidance whatsoever. The description doesn't say when to use this tool, what conditions must be met (e.g., node must exist, workflow must be editable), or what alternatives exist (e.g., delete_connection for edges).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_overlay_itemCInspect
Delete an overlay item from a track.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| itemId | Yes | The overlay item ID to delete | |
| trackId | Yes | The overlay track ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Delete' implies destruction, but the description does not state whether the deletion is permanent, whether it cascades, what permissions are needed, or what happens to dependent track data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is a single, front-loaded sentence with no wasted words. It is efficient, though its brevity contributes to gaps in behavioral and usage detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with four required parameters, no annotations, and no output schema, the description is too sparse. It omits side effects, permission requirements, and any indication of how the call relates to sibling item-management tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all four required IDs documented in the input schema. The description adds no parameter meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete an overlay item from a track.' This clearly distinguishes the item-level deletion from the sibling delete_overlay_track, though it does not explicitly name that alternative or contrast the two scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives, and no prerequisites are given. The agent must infer that this is a destructive edit operation and that update_overlay_item or move_overlay_item_to_track should be used instead for non-deletion changes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_overlay_trackAInspect
Delete an overlay track. Set deleteItems=true to also delete all items on the track; otherwise deletion fails if items exist.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| trackId | Yes | The overlay track ID | |
| deleteItems | No | Whether to delete all items on the track too | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the key behavioral trait — the call fails if items exist unless deleteItems=true — and that a cascade delete of child items occurs. It omits other relevant behavior: irreversibility, required permissions, and side effects on the parent edit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the core action front-loaded and the failure condition placed immediately after the parameter that controls it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive call with no output schema and three schema-documented required IDs, the description covers the critical trap an agent must avoid (silent failure vs. cascade deletion). It stops short of warning that this delete is destructive and likely irreversible, which would be valuable given no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description goes beyond the schema by explaining the consequence of deleteItems (failure when omitted, cascade deletion when set), which the schema's one-line description does not convey. Required IDs remain schema-documented only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ("Delete an overlay track") with no ambiguity, and the resource name distinguishes it from nearby siblings such as delete_overlay_item, delete_audio_track, and delete_transition. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives genuine conditional guidance for invoking it: set deleteItems=true to remove items too, otherwise the call fails. However, it offers no guidance on when to choose this tool over alternatives (e.g., delete_overlay_item for individual items, or update_overlay_track), so usage is implied rather than fully routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_sceneBInspect
Delete a scene from an edit. Note: shots referencing this scene will become orphaned.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| sceneId | Yes | The scene ID to delete | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the key side effect that referencing shots become orphaned, which is genuinely valuable context beyond the schema, but it omits whether deletion is permanent/undoable, whether it cascades, or whether special authorization is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the core action is front-loaded and the side-effect caveat follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter delete with no output schema and no annotations, the description covers the action and the main side effect, but leaves out irreversibility, permissions, and downstream reference handling. It is adequate but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents workspaceId, editId, and sceneId. The description adds no format, constraint, or relationship detail beyond implying that a scene belongs to an edit, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Delete a scene from an edit'), which clearly separates it from siblings like delete_shot, delete_asset, or update_scene. It does not explicitly name an alternative tool, so it stops short of full sibling routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives (e.g., update_scene to detach shots first) and no prerequisites such as required permissions. The orphan note describes a consequence, not a usage condition, so an agent gets no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_shotBInspect
Delete a shot. The linked media documents are NOT deleted (they persist independently in the workspace).
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID to delete | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one genuinely useful non-obvious trait: deletion does not cascade to linked media, which persists in the workspace. However, it omits irreversibility, permission requirements, and what happens to shot ordering or scene references after deletion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the important caveat immediately after. Nothing is wasted and nothing critical is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, a destructive delete tool should ideally state irreversibility, error behavior, and any auth requirements. The description covers the key non-cascade semantics but leaves the rest of the destructive-operation profile unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — all three required parameters (workspaceId, editId, shotId) are documented in the schema itself. The description adds no additional parameter meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("Delete a shot") and immediately scopes the blast radius by clarifying that linked media documents are not removed. It does not explicitly name a sibling to contrast with, but the non-cascade note implicitly separates it from tools like delete_shot_media_ref or delete_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives (e.g., delete_scene, delete_shot_media_ref, or moving a shot instead of deleting it), nor any prerequisites such as permissions or required prior state. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_shot_media_refBInspect
Delete shot media refs matching a mediaId and optional role. Optionally clears imageMediaId/videoMediaId when they point at the removed media.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | Optional role to narrow removal | |
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID | |
| mediaId | Yes | Media ID to remove from mediaRefs | |
| workspaceId | Yes | The workspace ID | |
| clearPrimaryMediaId | No | Clear imageMediaId/videoMediaId if it matches the removed media |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the cascading side effect of clearing imageMediaId/videoMediaId, but that sentence largely restates the schema's clearPrimaryMediaId description, and it says nothing about irreversibility, error behavior when no ref matches, or authorization requirements for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core delete action front-loaded and the side effect second. No filler, though the second sentence duplicates the schema's clearPrimaryMediaId wording rather than earning its place with new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter destructive tool with no annotations and no output schema, the description covers the primary action and the main cascade, but omits irreversibility, failure modes, and permission requirements. Adequate but with clear gaps given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented, and the enum values for role are enumerated in the schema. The description references mediaId, role, and the clear behavior but adds no format, default, or interaction detail beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete shot media refs') plus the matching criteria (mediaId and optional role), which cleanly separates it from add_shot_media_ref and update_shot_media_ref. It does not explicitly name those siblings, so an agent must infer the distinction from the verb, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the matching semantics (delete refs for a given mediaId, optionally narrowed by role), but there is no explicit when-to-use guidance, no mention of when to prefer update_shot_media_ref, and no stated prerequisites such as required permissions or edit/shot ownership.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_transitionBInspect
Delete one timeline transition by transitionId or boundaryId.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| boundaryId | No | Boundary ID to delete | |
| workspaceId | Yes | The workspace ID | |
| transitionId | No | Transition ID to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden for a destructive mutation. 'Delete' implies irreversibility, but the description does not state required permissions, whether the deletion can be undone, or what happens to the timeline after removal. It adds only the two identification modes, which is thin for a delete operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler that conveys verb, resource, cardinality, and targeting options. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Parameters are fully covered by the schema, so the main gap is behavioral: with no annotations and no output schema, the description should say more about the consequences of deletion and error conditions. It is adequate but leaves the agent guessing about a destructive operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented, giving a baseline of 3. The description does add one useful relational detail — that a transition can be targeted by either transitionId or boundaryId — but it does not explain when to prefer one over the other.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete one timeline transition') with the scope clearly bounded to a single item. It does not explicitly differentiate itself from the sibling tools update_transition and add_transition, but the verb+resource pair is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no mention of prerequisites, and no reference to alternatives such as update_transition or add_transition. Usage is only inferable from the tool name itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_shotAInspect
Duplicate a shot with all media refs, trimming, playback, frame-linking, prompts, dialogue, and tags. Media documents are referenced, not copied.
| Name | Required | Description | Default |
|---|---|---|---|
| order | No | Optional destination order | |
| editId | Yes | The edit ID | |
| shotId | Yes | Shot ID to duplicate | |
| sceneId | No | Optional destination scene ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The key behavioral disclosure — 'Media documents are referenced, not copied' — is valuable since it tells the agent the duplication is shallow, avoiding wasted copies. However it says nothing about permissions, side effects on the source shot, or whether the new shot gets a fresh ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and lists the carried-over attributes compactly, with the shallow-copy caveat as a tight second sentence. Zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a shallow-copy tool with a fully-covered schema and no output schema. It does not cover where the duplicate lands by default, required permissions, or return value, which for a creation tool with no annotations leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including the optional destination sceneId and order. The description adds nothing about parameter behavior, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Duplicate a shot') and enumerates exactly which attributes carry over (media refs, trimming, playback, frame-linking, prompts, dialogue, tags). It is immediately distinguishable from update_shot, create_shot, and split_shot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb, but there is no explicit when-to-use vs create_shot (which also produces a new shot) or split_shot. The default destination (same scene as source) and required workspace/edit context are not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageAInspect
Edit an existing image using AI. Provide the media ID of the image to edit and a text prompt describing what to change. Creates a new version on the media doc — the original is preserved. Use sparingly — only when the user explicitly asks to modify an existing image. Uses Nano Banana Pro (google-gemini-3-image).
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Image edit model ID. Uses the default if omitted. | |
| mediaId | Yes | Media ID of the existing image to edit | |
| editPrompt | Yes | Instructions for what to change in the image (e.g., "make the sky more orange", "add a hat to the character") | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses non-destructiveness ('Creates a new version on the media doc — the original is preserved') and names the underlying model. It does not mention latency, cost, or rate limits, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, zero redundancy. Purpose comes first, then required inputs, then a behavioral guarantee, then a usage constraint — well front-loaded and highly readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation-style tool with no annotations and no output schema, the description covers purpose, inputs, non-destructiveness, and usage limits. It does not explain the return value (a new media version ID would be useful) but is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds narrative context around mediaId ('the image to edit') and editPrompt ('describing what to change'), and notes the model parameter has a default, which is slightly more than the schema alone provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Edit an existing image using AI', with the required inputs (media ID + text prompt) named. It clearly distinguishes itself from siblings like generate_image (new image) and edit_video (video editing) by scoping to existing images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance: 'Use sparingly — only when the user explicitly asks to modify an existing image.' This is a strong, actionable exclusion criterion that routes the agent correctly and prevents over-invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_videoedit videoADestructiveInspect
Edit existing footage using a source video and prompt. Creates a separate result and preserves the original clip. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, destructiveHint=true, idempotentHint=false, so the safety/mutation profile is covered structurally. The description adds real value by stating the source clip is preserved and that a tracked task with a playable result is returned, but it says nothing about cost exposure, failure/credit behavior, or what the destructive annotation actually implies for the caller.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core operation and followed by the two facts an agent most needs (non-destructive output, model catalog pointer, async result). No filler, though the closing return-value sentence is mildly redundant given the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter async generation tool with a documented output schema and full annotations, the description covers the essentials: what it does, the non-destructive guarantee, and where to get valid model/settings values. The dense optional parameters (resolution, aspectRatio, referenceImages, idempotencyKey) are left entirely to the schema, which is acceptable at this coverage level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 81%, so the schema already documents model, sourceMediaId, prompt, mediaInputs, scaleFactor, and the rest; the baseline of 3 applies. The description only echoes the source-video-plus-prompt inputs and points at get_model_catalog, adding no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (edit existing footage from a source video plus prompt) and adds the important scope fact that it produces a separate result while the original is preserved. It does not explicitly differentiate itself from close siblings such as enhance_video_media, expand_video, or upscale_video, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete routing rule — call get_model_catalog to pick a supported model and its settings — which is useful. But it never says when to choose this tool over generate_video, enhance_video_media, or upscale_video, so the agent must infer the boundary from names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enhance_video_mediaCInspect
Run a video editor enhancement pass on a workspace video media item. Supports upscale, HDR, output codec, bit depth, scale factor, and arbitrary model parameters.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | Video MediaDocV2 ID to process | |
| modelId | Yes | Primary video enhancement model ID | |
| bitDepth | No | Output bit depth | |
| videoUrl | No | Override source video URL; otherwise resolves from the active media version | |
| enableHdr | No | Whether to chain/apply HDR conversion | |
| hdrModelId | No | HDR model ID | |
| extraParams | No | Additional raw V3 model parameters | |
| outputCodec | No | Output codec, e.g. h264, h265, prores_422, exr_piz | |
| scaleFactor | No | Upscale factor | |
| workspaceId | Yes | The workspace ID | |
| outputFormat | No | Output format override for sequence/HDR models, e.g. exr | |
| videoEncoder | No | Encoder override for models that expose it | |
| enableUpscale | No | Whether this pass should upscale | |
| transferFunction | No | HDR transfer function, e.g. hlg |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing beyond the operation being a processing pass. It doesn't say whether the job is synchronous or asynchronous, whether it consumes credits, whether it mutates the existing media or produces a new asset, or what model-parameter behavior to expect from extraParams.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the operation front-loaded and no filler. The second sentence is somewhat of a feature dump that duplicates the schema, keeping it just short of maximum efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex tool: 14 parameters, a free-form nested extraParams object, no output schema, and no annotations. The description supplies none of the missing context an agent needs to invoke it confidently, such as job lifecycle, result handling, or how the enhancement pass relates to the resulting media version.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the 14 parameters are already documented in the schema, and the baseline is 3. The description's feature list (upscale, HDR, output codec, bit depth, scale factor, arbitrary model parameters) largely restates schema fields rather than adding format, default, or interaction semantics beyond them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Run a video editor enhancement pass on a workspace video media item') and enumerates the enhancement capabilities it covers. It is clear on its own, but it never distinguishes itself from close siblings like upscale_video, convert_video_to_hdr, and expand_video, which the listed capabilities directly overlap with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-routing guidance. The second sentence reads as a capability list ('Supports upscale, HDR, output codec...') rather than a condition that selects this tool over upscale_video or convert_video_to_hdr, so the agent must infer when this composite pass is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_workflowAInspect
Check graph execution availability. This legacy tool cannot run graphs. Use run_production_template for reusable productions or the authenticated workflow API for published graph endpoints.
| Name | Required | Description | Default |
|---|---|---|---|
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. To its credit it discloses the most consequential trait — the tool is legacy and non-functional for execution. However it omits anything about the return payload, authentication, error behavior, or whether calling it has side effects, which matters for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste, and the deprecation warning is front-loaded before the alternatives. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must do more than redirect. It adequately warns the agent off and points to replacements, but leaves the return value, failure modes, and the meaning of the second required parameter unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Two required parameters at 50% schema coverage: workspaceId is documented in the schema, workflowId is not. The description adds no meaning to either — no format, no relationship between workspace and workflow, no indication that IDs must reference an existing graph. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ("Check graph execution availability") and immediately clarifies the misleading name by declaring this a legacy tool that cannot run graphs. It distinguishes itself from siblings by naming run_production_template and the workflow API. Point deducted because the tool name still implies execution, and the description never says what the availability check actually returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly redirects the agent: use run_production_template for reusable productions, or the authenticated workflow API for published graph endpoints. That is concrete routing guidance. It stops short of stating the narrow condition under which this tool itself should still be called.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expand_videoexpand videoADestructiveInspect
Run the Reframe preset on a complete workspace video up to 120 seconds and 500 MB. Outpaint a new aspect ratio, join provider chunks, and restore the original audio. Supports runway-aleph-2, google-gemini-omni-1-1, luma-ray-3-2-reframe, wan-2-2-vace-reframe, and beeble-switchframe. SwitchFrame requires a workspace image in mediaInputs.reference_image_uri. Call once per requested source/format/settings variant. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| duration | No | Expected complete source duration in seconds, used to verify the approved Reframe plan. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false, openWorldHint=true). The description adds value beyond them: the 120s/500MB envelope, that output arrives as a tracked task with a playable result when complete, and which models are supported. It does not explain what 'destructive' means here or the cost/retry semantics, which the schema's idempotencyKey description partly carries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and hard limits, then capabilities, then prerequisites and model list. Five sentences, no filler; the model enumeration is long but directly actionable for the required 'model' parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter, nested-schema, open-world generation tool, the description supplies the constraints (time/size), model compatibility, prerequisite for SwitchFrame, async return shape, and the catalog lookup for settings. With a 22-property schema and an output schema present, nothing essential is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already high (82%), so the baseline is 3; the description still adds meaning by naming the five valid models and the beeble-switchframe requirement on mediaInputs.reference_image_uri, plus the duration/size limits tied to the duration parameter. It does not clarify the many optional URL-vs-mediaId alternatives, which the schema covers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run the Reframe preset on a complete workspace video'), then enumerates the actual operations (outpaint aspect ratio, join provider chunks, restore audio). This clearly separates it from siblings like edit_video, generate_video, and enhance_video_media without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable context: 'Call once per requested source/format/settings variant', 'Use get_model_catalog to choose a supported model', and the SwitchFrame prerequisite for mediaInputs.reference_image_uri. It does not explicitly contrast with the closest siblings (edit_video, upscale_video, enhance_video_media), so routing between them still requires inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_editAInspect
Start a production export for an edit. Creates an export record, calls export-project-v2, and returns the final video preview, download, and project URLs when the synchronous export completes. Use these links for final delivery.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| exportId | No | Optional existing export document ID to use | |
| editTitle | No | Optional denormalized edit title for the export record | |
| watermark | No | Whether to render a watermark | |
| exportType | No | Export format | mp4 |
| workspaceId | Yes | The workspace ID | |
| exportSettings | No | Fine-grained codec, resolution, framerate, HDR, upscale, and audio settings |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does add real behavioral detail: it creates an export record, invokes export-project-v2, runs synchronously, and returns preview/download/project URLs. It omits permissions, cost/credit implications, and failure modes, so the disclosure is useful but incomplete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and followed by mechanism and usage. No filler, though the mechanism sentence ('calls export-project-v2') is arguably internal detail that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly steps in to name the return values (video preview, download, and project URLs) and describes the synchronous flow. For a 7-parameter tool with a nested settings object, this is nearly complete, missing only cost/auth context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters including the nested exportSettings object. The description adds no parameter-level meaning beyond the schema, which is the baseline 3 when structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a production export for an edit') and distinguishes itself from read-oriented siblings like get_export and list_exports by describing it as a starting/production action. An agent can tell what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing sentence 'Use these links for final delivery' implies the tool's purpose (final delivery output) but never states when to choose it over get_export or list_exports, nor any prerequisites. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_product_imageBRead-onlyIdempotentInspect
Find a product image from a product URL. Prefers Product structured data; page_preview is a candidate to inspect before using. If extraction fails, upload the product image or use a direct image URL.
| Name | Required | Description | Default |
|---|---|---|---|
| productUrl | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so safety behavior is covered. The description adds useful extraction behavior about preferred structured data and failure handling, but it does not explain permissions, rate limits, or what happens during partial extraction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core action. The phrase about page_preview is somewhat cryptic, but the overall three-sentence structure is efficient and avoids unnecessary repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the basic purpose, extraction preference, and failure fallback, which is useful for a read-only annotated tool. However, with no output schema and no parameter descriptions, it leaves workspaceId semantics and return behavior unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain either required parameter meaningfully. It implies productUrl from 'product URL' but says nothing about workspaceId, leaving half the required parameters undocumented beyond the schema's raw types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: find a product image from a product URL. It is clear enough to distinguish extraction from generation or editing tools, but it does not explicitly name sibling alternatives or fully clarify what 'find' returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives some when-to-use context, including preferring Product structured data and falling back to upload or direct image URL if extraction fails. However, it does not name the fallback tools or specify when this tool should be avoided in favor of siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_audioGenerate Sequencer audioADestructiveInspect
Generate AI audio through Sequencer. Use this whenever the user asks to make, create, generate, or render speech, voiceover, narration, music, a song, or a sound effect with Sequencer or names an audio/music model/provider Sequencer supports: Cartesia, ElevenLabs, Fish Audio, Lyria, Suno, or another audio model. If no speech model is specified, Sequencer uses Cartesia Sonic 3.5 (cartesia-sonic-3-5). An active skill's explicit model still takes priority. If no workspaceId is known, omit it and the server will use the user default workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model ID, for example cartesia-sonic-3-5 for Cartesia Sonic 3.5, elevenlabs-tts-v3 for speech, lyria-002 for music, or another ID returned by get_model_catalog. Uses cartesia-sonic-3-5 if not specified. | |
| prompt | Yes | Text to speak (for TTS) or music prompt | |
| voiceId | No | Voice ID for TTS models | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| stylePrompt | No | Delivery and accent direction for supported speech models, for example: warm Argentine Spanish with natural Rioplatense pronunciation. | |
| workspaceId | No | Optional workspace ID. If omitted, Sequencer uses the user default/personal workspace. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the paid/generative nature is covered structurally. The description adds the model-resolution precedence ('An active skill's explicit model still takes priority') and the workspace-defaulting behavior, but those largely restate the model and workspaceId schema descriptions, and it never states outright that this starts a billable generation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and trigger conditions are front-loaded in the first two sentences, which is the right ordering for a router-style tool. The trailing sentences about default models and workspaceId partly duplicate the schema, adding mild redundancy, but the text stays free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description covers provider coverage, default model fallback, skill precedence, and workspace defaulting. It leaves cost behavior and the interaction between voiceId/stylePrompt and specific model families to the schema, which is acceptable but not exhaustive for a paid, destructive generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters are already documented in the schema, making 3 the baseline. The description reinforces model defaults and workspaceId omission, and adds the skill-priority rule, but contributes no syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Generate AI audio through Sequencer') and enumerates the modalities it covers: speech, voiceover, narration, music, song, sound effects. It does not, however, differentiate itself from close siblings such as generate_audio_track or generate_video_audio, so an agent still has to infer which generator to pick for timeline- or video-scoped audio.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong positive routing: 'Use this whenever the user asks to make, create, generate, or render speech, voiceover, narration, music, a song, or a sound effect with Sequencer or names an audio/music model/provider Sequencer supports.' It also names the supported providers and the default model resolution rule. What is missing is any when-not guidance relative to the adjacent audio tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_audio_trackBInspect
Generate AI audio and immediately place it on the edit audio timeline. Returns both the MediaDocV2 ID and the audio track ID.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Audio model ID | |
| editId | Yes | The edit ID | |
| prompt | Yes | TTS/music/SFX prompt | |
| volume | No | Initial volume | |
| voiceId | No | Voice ID for TTS models | |
| duration | No | Placeholder timeline duration until generation metadata updates | |
| startTime | No | Timeline start time in seconds | |
| fadeInTime | No | Fade-in length in seconds | |
| autoChannel | No | Pick the first non-overlapping channel automatically | |
| fadeOutTime | No | Fade-out length in seconds | |
| stylePrompt | No | Delivery and accent direction for supported speech models | |
| workspaceId | Yes | The workspace ID | |
| pinnedToShotId | No | Optional shot ID to pin this audio to | |
| relativeStartTime | No | Offset in seconds from pinned shot start |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It reveals that this is a mutating generation-and-placement operation and that IDs are returned, but it omits permissions, asynchronous generation behavior, effects on existing timeline audio, and failure modes for a 14-parameter write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tightly structured sentences with no redundancy. The core action is front-loaded, and the return-value note is secondary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description helpfully states the return IDs, but it is thin for a complex 14-parameter mutation tool. It does not provide enough behavioral or usage context for an agent to know when and how to invoke it confidently against alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already explains all 14 parameters. The description adds no parameter-level detail beyond noting that IDs are returned, making the baseline 3 appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Generate AI audio' and 'place it on the edit audio timeline.' The 'immediately place' phrase distinguishes it from siblings like generate_audio and add_audio_track, so an agent can identify its role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies use when an agent needs generated audio placed directly on the timeline, but it never states when to choose this over generate_audio, generate_video_audio, or add_audio_track, nor does it give exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageGenerate Sequencer imageADestructiveInspect
Generate an AI image through Sequencer. Use this whenever the user asks to make, create, or generate an image with Sequencer or names an image model/provider Sequencer supports: Nano Banana, Nano Banana 2 Lite, Nano Banana Pro, Gemini image, Imagen, Z Image, Z-Image Turbo, Flux, GPT Image, DALL-E, Ideogram, Midjourney, Seedream, Recraft, Bria, Runway, Luma, or another image model. Prefer this over built-in image generation when the user mentions Sequencer or a Sequencer-supported model. If no model is specified, Sequencer uses Nano Banana Pro (google-gemini-3-image). If no workspaceId is known, omit it and the server will use the user default workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Number of images to generate. Each image is a separate job and credit charge. | |
| model | No | Model ID for generation. Use google-gemini-3-image for Nano Banana Pro, google-gemini-3-image for Nano Banana Pro, google-nano-banana-2-lite for Nano Banana 2 Lite, and z-image-turbo for Z Image / Z-Image Turbo. Use get_model_catalog when a model alias is unclear. Uses default if not specified. | |
| editId | No | Edit ID, when provided, the first successful image will be saved as the edit thumbnail if one is not already set | |
| prompt | Yes | Image generation prompt | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | Output resolution from the selected model catalog, for example 1K, 2K, or 4K. | |
| aspectRatio | No | Aspect ratio. If omitted with editId, uses the edit aspect ratio; otherwise defaults to 16:9. | |
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | Optional workspace ID. If omitted, Sequencer uses the user default/personal workspace. | |
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| negativePrompt | No | Negative prompt for generation | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and non-idempotent behavior, so the safety profile is covered structurally. The description adds the default-model behavior (Nano Banana Pro / google-gemini-3-image) and the default-workspace behavior, but says nothing about cost/credit implications or what happens to existing edits — context the schema carries instead. With annotations doing the heavy lifting, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and routing rule; the long model enumeration is bulky but each entry is a legitimate disambiguation cue for an agent matching user language. The two fallback defaults are placed at the end where they belong.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 23-parameter, nested-object tool with an output schema and near-total schema coverage, the description supplies the missing decision layer — when to pick this tool and what defaults apply — without duplicating field documentation. Only the cost/credit aspect of a paid generation is left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 96%, so the baseline is 3, but the description adds genuinely new semantics: the concrete default model ID when none is specified, the default workspace fallback, and a pointer to get_model_catalog for alias resolution. That is real value beyond the schema's own field text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ("Generate an AI image through Sequencer") and enumerates the exact model families it can target, which cleanly separates it from edit_image, render_image, generate_video, and upscale_image in the sibling list. An agent can identify the tool's job without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ("whenever the user asks to make, create, or generate an image... or names an image model/provider"), names the alternative to prefer against (built-in image generation), and provides a fallback path (use get_model_catalog when an alias is unclear). This is genuine when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ip_reportAInspect
Generate a formal HTML IP compliance report from analyze_export_ip findings. Uploads it to Cloud Storage, saves a report reference in Firestore, and returns a report download card marker.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| summary | Yes | Summary paragraph from analyze_export_ip | |
| exportId | No | The export ID that was analyzed | |
| findings | Yes | JSON string of the findings array from analyze_export_ip | |
| editTitle | No | Title of the edit/export | Untitled Export |
| limitations | No | Known analysis limitations from analyze_export_ip | |
| overallRisk | Yes | Overall risk assessment | |
| workspaceId | Yes | The workspace ID | |
| analysisScope | No | What the analysis actually covered | export_video_full |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden and does disclose meaningful side effects: it uploads to Cloud Storage, writes a report reference to Firestore, and returns a download card marker. That is substantive for a write/mutation-style tool. It still omits permission requirements, idempotency, and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with what the tool produces and immediately followed by the side effects and return artifact. No filler or restated name/title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no annotations, and no output schema, the description usefully discloses the persistence targets and the returned artifact ('report download card marker'), which substitutes for a return-value schema. It is close to complete, though it could state prerequisites and error/rate-limit behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 9 well-documented parameters (including two enums), so the schema already does the heavy lifting. The description adds no parameter-level detail beyond implicitly pointing at the analyze_export_ip outputs, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Generate a formal HTML IP compliance report') and names the upstream data source ('from analyze_export_ip findings'), which cleanly separates it from the sibling analyze_export_ip that only produces findings. An agent can identify this tool without opening its schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'from analyze_export_ip findings' implies the tool is used as a follow-up step after analysis, which is useful contextual routing. However, there is no explicit when-to-use/when-not statement, no prerequisite spelled out (e.g. 'call analyze_export_ip first'), and no mention of alternatives or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_shot_videoAInspect
Generate a video for one existing edit shot and attach the queued media to that shot immediately. Use this for project and timeline builds. It preserves the shot still as the source frame, replaces the prior primary video reference, and keeps the result inside the edit instead of presenting it as standalone media.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Video model ID. Uses the edit default when omitted. | |
| editId | Yes | The edit ID | |
| prompt | No | Motion prompt. Defaults to the shot videoPrompt when omitted. | |
| shotId | Yes | The shot that should receive the generated video | |
| volume | No | Attached shot-video volume. Keep 0 when narration or a separate mix owns the audio. | |
| duration | No | Generated source duration in seconds. It must cover the shot editorial duration. | |
| resolution | No | Output resolution supported by the selected model, for example 720p or 1080p. | |
| aspectRatio | No | Aspect ratio. Uses the edit default when omitted. | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose key behavior: it 'preserves the shot still as the source frame, replaces the prior primary video reference, and keeps the result inside the edit.' This tells the agent about source frame handling, that a prior primary video reference is replaced (mutation side effect), and the scoping. It does not mention rate limits, error conditions, or whether other shot media refs are affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and scope, followed by behavioral effects. No filler words and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generation tool with 9 params, full schema coverage, and no output schema, the description covers purpose, usage context, and key side effects. It is slightly short on explicit alternatives to route against siblings like generate_video and generate_video_audio, but overall completeness is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters with semantics (model, prompt, duration, volume, etc.). The description adds no parameter-specific detail beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Generate a video for one existing edit shot') and immediately distinguishes scope ('attach the queued media to that shot immediately'), separating it from standalone video tools like generate_video. The phrase 'inside the edit instead of presenting it as standalone media' explicitly differentiates from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: 'Use this for project and timeline builds,' which guides when to use it versus standalone generation. However, it never names specific alternatives (generate_video, run_video_workflow) or states exclusions, so it lacks the explicit when-not guidance needed for a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoGenerate Sequencer videoADestructiveInspect
Generate an AI video through Sequencer. Use this whenever the user asks to make, create, generate, animate, or render a video/clip/animation with Sequencer or names a video model/provider Sequencer supports: MiniMax H3, PrunaAI P-Video, Veo, Seedance, Happy Horse, Kling, Runway, Luma, Sora, Hailuo, PixVerse, Google Omni, or another video model. For project builds, use this only after the user has approved the shot still images/storyboard. If no model is specified, Sequencer uses Gemini Omni 1.1 Flash (google-gemini-omni-1-1). If no workspaceId is known, omit it and the server will use the user default workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model ID for generation, for example google-gemini-omni-1-1, token360-seedance-2-fast, google-veo-3-1-fast, kling-v2-1, runway-gen4, luma-ray, sora, hailuo, seedance, or another ID returned by get_model_catalog. Uses default if not specified. | |
| editId | No | Edit ID for resolving project defaults such as aspectRatio and defaultVideoModel. | |
| prompt | Yes | Video generation prompt | |
| duration | No | Generated source duration in seconds. Five seconds is only a compatibility fallback. For project shots, choose and pass the shortest model-supported duration that covers the intended editorial beat, and state the same length in the video prompt. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | Output resolution supported by the selected model, for example 480p, 720p, or 1080p. | |
| aspectRatio | No | Aspect ratio. If omitted with editId, uses the edit aspect ratio; otherwise defaults to 16:9. | |
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | Optional workspace ID. If omitted, Sequencer uses the user default/personal workspace. | |
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Media ID of source image for image-to-video | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety and side-effect profile is covered. The description adds useful defaulting behavior (Gemini Omni 1.1 Flash fallback, server-side default workspace), but says nothing about paid generation cost, latency, or failure/recovery behavior beyond what the schema's maxCostUsd and idempotencyKey fields already state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose, trigger phrasings and prerequisite are front-loaded before the fallback details, and every sentence carries information. The long model-provider enumeration is verbose but does real selection work by matching user-named providers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter, one-required-parameter tool with an output schema and 95% schema coverage, the description supplies the missing high-level context: what triggers it, the storyboard-approval gate, and the default-resolution rules. Cost, rate limits and async/return behavior are left to the schema and output schema, which is acceptable here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 95% across 22 parameters, so the schema already carries the parameter semantics and the baseline is 3. The description's only added parameter value is the default-model and omitted-workspaceId behavior, which is largely restated from the schema descriptions of `model` and `workspaceId`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource (generate an AI video through Sequencer) and the second enumerates the user phrasings and model providers it covers, which is genuinely useful for matching. It stops short of naming its nearest sibling (generate_shot_video) so the agent must infer the boundary between standalone generation and shot generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use trigger (any request to make/create/animate/render a video, or naming a supported model) plus a real prerequisite for project builds: only after shot stills/storyboard are approved. It also resolves two common unknowns (default model, omitted workspaceId). No explicit when-not guidance or named alternative tool is offered, so it is strong context rather than full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_video_audioBInspect
Generate or replace audio for a video media item. Supports add/replace audio modes and prompt-driven video-to-audio models.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | Prompt describing the desired audio, music, foley, or ambience | |
| mediaId | Yes | Video MediaDocV2 ID to process | |
| modelId | Yes | Video-to-audio model ID | |
| videoUrl | No | Override source video URL; otherwise resolves from active media version | |
| audioMode | No | Whether generated audio replaces or adds to original audio | replace |
| extraParams | No | Additional raw V3 parameters | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does not disclose whether this is an asynchronous job, how long it takes, whether it costs credits, what happens to the original audio track in "add" vs "replace" mode, or whether the operation is reversible — all material for a generative mutation on existing media.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action front-loaded before the supporting capability statement. Slightly more could be said about mode semantics, but nothing present is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The input schema is rich and fully documented, which covers parameter-level needs. However, with no output schema and no annotations, the description omits the post-invocation picture (job handle, polling, resulting media version) that an agent needs to chain this generation call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters are already documented, including the audioMode enum and the videoUrl override. The description only echoes the mode concept and the model-driven nature of the call, adding no format, default, or interaction detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb pair (generate or replace) and a specific resource (audio for a video media item), which is far more precise than the generic siblings. It does not, however, name or distinguish itself from close siblings such as generate_audio, generate_audio_track, or add_audio_track, so an agent must still infer which tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Supports add/replace audio modes and prompt-driven video-to-audio models" implies the situations the tool covers, and the video-to-audio framing hints that a video media item is the precondition. There is no explicit when-to-use/when-not statement and no routing guidance against generate_audio, add_audio_track, or lip_sync_video_media, which are all plausible alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_assetBInspect
Get detailed information about a specific asset including all media references.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | The asset ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but provides little behavioral context beyond the implied read operation. It does not disclose auth requirements, error behavior, rate limits, or the exact nature of the returned 'detailed information' beyond mentioning media references.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the core purpose without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema and no annotations, the description is minimally adequate. It states the purpose and mentions media references, but leaves undefined what other fields 'detailed information' includes, which could matter for an agent deciding whether to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both required parameters (workspaceId, assetId) are fully documented in the schema. The description adds no additional meaning about parameter usage or constraints, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Get) and resource (asset), and adds scope ('specific asset') and content ('including all media references'). It does not explicitly distinguish this tool from sibling alternatives like list_assets or get_media, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as list_assets, get_media, or check_asset_status. The description only implies usage for fetching details on a known asset, without stating conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_available_toolsBInspect
Get a list of all tools available on the Sequencer MCP server, organized by category. Shows what the AI assistant can do. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Tool category to list | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses an auth trait ('No authentication required') and 'Get a list' implies a read-only operation, but it says nothing about return format, pagination, or result size for what is likely a very large listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose and the scoping constraint. 'Shows what the AI assistant can do' is mildly redundant filler, but overall there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-required-parameter discovery tool with no output schema, the description covers purpose, organization, and auth. The main gap is that with no output schema, the shape of the returned list (names, descriptions, categories) is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single category parameter is fully documented with an enum and default in the schema. The description's 'organized by category' hints at the parameter's role but adds no syntax or semantics beyond what the schema already supplies, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a list of all tools available on the Sequencer MCP server') plus the organizing principle ('by category'). However, it does not distinguish itself from similarly-scoped siblings such as get_sequencer_capabilities, get_feature_list, or get_server_info, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative routing. 'Shows what the AI assistant can do' gestures at a discovery purpose but is a restatement of purpose rather than guidance, and none of the overlapping capability-listing siblings are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_balanceAInspect
Get the authenticated user's current credit balance, subscription plan, and billing summary. Use this when the user asks about their credits, balance, or plan.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It conveys that the read is scoped to the 'authenticated user' (implying an auth requirement) and implies a read-only operation via 'Get', which is genuinely useful. However, it says nothing about failure modes, whether unauthenticated calls error, or freshness of the balance, so the behavioral picture is only partially filled in.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The payload (what is returned) comes first and the routing cue second, which is the right front-loading for a zero-argument getter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no parameters, so the description must at least name the returned fields, which it does (balance, plan, billing summary). It is nearly complete for a simple read; only error/auth behavior and the relationship to sibling billing tools are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. The description correctly does not invent inputs, confirming the parameterless nature.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') and enumerates the exact resources returned: credit balance, subscription plan, and billing summary. It is unambiguous on its own, but it does not differentiate itself from adjacent siblings such as get_pricing_info or get_subscription_help, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Use this when the user asks about their credits, balance, or plan.' That is clear positive guidance, but there is no negative guidance or named alternative for overlapping billing/plan questions, so routing between this and get_pricing_info remains a guess.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_creative_tipsCInspect
Get creative tips for a specific type of content.
| Name | Required | Description | Default |
|---|---|---|---|
| contentType | Yes | Type of content you are creating |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It says nothing about whether tips are static or generated, how large the response is, or whether any permissions or limits apply. 'Get creative tips' conveys only a read intent and nothing else.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with no filler, and the resource is front-loaded. It is appropriately short, though there is so little content that conciseness is partly a byproduct of under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should ideally hint at what the returned tips look like and how they differ from other guide tools. For a simple one-parameter lookup this is adequate, but the return content is left entirely unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single contentType parameter is fully documented with an enum in the schema, so the baseline is 3. The phrase 'for a specific type of content' merely restates the schema and adds no new format or constraint detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Get creative tips') and the enum-backed scope ('a specific type of content'). However, it does nothing to distinguish itself from the many sibling guidance tools such as get_image_prompt_guide, get_video_prompt_guide, and get_storytelling_guide, leaving the agent to guess which guide applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of any alternative. The description never says when an agent should reach for creative tips versus the prompt guides or storytelling guide that sit alongside it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_editCInspect
Get detailed information about a specific edit, optionally including scenes and shots.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID | |
| includeShots | No | Include shot data | |
| includeScenes | No | Include scene data |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read ('Get') but says nothing about permissions, response size, whether requesting shots/scenes is expensive, or how this differs in payload from get_edit_full.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the resource front-loaded and the optional behavior appended. No filler, though it is arguably too terse for a tool with a near-identical sibling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with full schema coverage and no output schema, the description is minimally adequate. The significant omission is disambiguation from get_edit_full, which an agent in this 130-tool namespace genuinely needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema. The description's 'optionally including scenes and shots' merely restates includeScenes/includeShots without adding format or cost semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get detailed information about a specific edit') and even flags the optional expansion of scenes/shots. However, it does not differentiate itself from the obviously related sibling get_edit_full, leaving the agent unsure which 'get' to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus get_edit_full, list_edits, or inspect_edit_frame. The word 'optionally' hints at the include flags but the description never states the condition under which scenes/shots should be requested.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_edit_fullAInspect
Get complete project state in one call: edit settings, all scenes, all shots with resolved media prompts/status. This is the most efficient way to understand the full edit.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID | |
| includeMediaVersions | No | Set true only when all historical media versions are required. Default false keeps large projects compact while still returning active media URLs, prompts, and status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that this is a single-call aggregate reader and that media prompts/status come resolved, and the schema's includeMediaVersions note hints at payload-compaction behavior. However it says nothing about read-only safety, permissions, or how large a 'complete state' response can get for a big project.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The scope and payload are front-loaded before the efficiency claim, so an agent can stop reading after the first sentence and still know what it gets back.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only aggregate fetch with no output schema, the description adequately covers what is returned (settings, scenes, shots, resolved media prompts/status). It is nearly complete, missing only sibling differentiation and any caution about response size, which touches on usage rather than capability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so editId, workspaceId, and includeMediaVersions are already documented in the schema, including the 'set true only when all historical media versions are required' guidance. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get complete project state') and enumerates the returned parts: edit settings, all scenes, all shots with resolved media prompts/status. It is clear what the tool does, but it never distinguishes itself from siblings like get_edit, list_scenes, or list_shots, so an agent must infer that this is the aggregate version.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing sentence ('most efficient way to understand the full edit') implies when to reach for it, but there are no explicit exclusions or named alternatives. An agent is left to guess when a narrower call such as get_edit or list_shots would be cheaper or sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_experience_gameARead-onlyIdempotentInspect
Read a game and its proposal, approval status and revision, organized project files, compiled HTML, current version, and asset:// references. Read before editing. Preserve existing files and assets.
| Name | Required | Description | Default |
|---|---|---|---|
| gameId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and non-open-world behavior. The description adds valuable return-content detail beyond the annotations by listing proposal, approval status, revision, project files, compiled HTML, version, and asset:// references.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the read action and its returned components, and it avoids unnecessary filler. The final preservation instruction is brief, though somewhat less essential for a read-only tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema, the description usefully enumerates the returned content and notes 'Read before editing.' However, it leaves the required workspaceId/gameId semantics unexplained, which is a meaningful gap given 0% schema description coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for workspaceId and gameId, yet the description does not explain either parameter's format, pattern, or role. It only implies that a game is identified, leaving the required workspace scoping undocumented in both description and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read a game' and enumerates the game's proposal, approval status and revision, project files, compiled HTML, current version, and asset references. The singular, ID-based read is distinguishable from list_experience_games and save_experience_game, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context with 'Read before editing,' which tells the agent when this tool is appropriate. It does not name alternatives or state when not to use it, so it falls short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_experience_worldBRead-onlyIdempotentInspect
Read a world and its placed characters, story, visual theme, and game lineup to build a game that fits it.
| Name | Required | Description | Default |
|---|---|---|---|
| worldId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds what the read returns (characters, story, theme, lineup), but says nothing about permissions, workspace scoping, or size/pagination of the payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. Slightly run-on in its enumeration, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates what is read back, which covers the return-value burden. However, for a 2-parameter tool with 0% schema coverage it leaves parameter meaning entirely unexplained, so it is only partially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required parameters, and the description never mentions workspaceId or worldId. Notably worldId is a bare $ref to workspaceId, so an agent gets no clarification of the distinction or format — the description does nothing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('a world') and enumerates the payload: placed characters, story, visual theme, game lineup. It is distinguishable from sibling reads like get_experience_game or list_experience_games, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing phrase 'to build a game that fits it' implies the intended context (feeding a game-creation flow), but there is no explicit when-to-use, when-not-to-use, or named alternative among the many get_* / list_experience_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exportAInspect
Get a single export record for an edit, including status, URLs, settings, upscale/HDR state, and IP analysis if present.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| exportId | Yes | The export document ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It helpfully enumerates what the record contains (status, URLs, settings, upscale/HDR state, conditional IP analysis), which is the main behavioral signal. It says nothing about auth requirements, error/not-found behavior, or whether the IP analysis is costly to fetch, leaving meaningful gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste; the resource and its distinguishing scope ('single export record for an edit') lead, and the returned-field list follows. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description usefully previews the return contents, which is the right compensation. It is still short of complete for a lookups tool lacking any permissions, error, or conditional-fetch guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the three required identifiers are each documented in the schema. The description adds no parameter-level meaning beyond what the schema provides, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (a single export record) scoped to an edit, and enumerates the fields returned (status, URLs, settings, upscale/HDR state, IP analysis). This implicitly distinguishes it from list_exports, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the three required IDs (workspaceId/editId/exportId) and the 'single record' framing, so an agent can infer it is the by-ID retrieval path. However, there is no explicit guidance on when to prefer this over list_exports, analyze_export_ip, or export_edit, which are closely related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_feature_listBInspect
Get a structured list of all Sequencer platform features organized by category. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Feature category to retrieve | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that no authentication is required, which is real behavioral context for an agent, but it says nothing about rate limits, response size, or caching. The auth disclosure earns partial credit but leaves notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The core action is stated first, followed by a useful constraint, and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-required-parameter read tool, the description adequately conveys what it returns ('structured list ... organized by category') and notes the auth situation. With no output schema, a brief note on the shape of the returned categories would round it out, but nothing essential for calling it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the category parameter and its enum values are fully documented in the schema. The description's phrase 'organized by category' only loosely echoes that filtering exists and adds no syntax or default detail beyond the schema's default='all'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a structured list of all Sequencer platform features organized by category'), which is clear and concrete. However, it does not differentiate itself from closely related info siblings such as get_sequencer_capabilities, get_platform_overview, or get_available_tools, leaving the agent to guess which feature-listing tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no mention of alternatives, despite several overlapping siblings that enumerate platform capabilities. The only contextual hint is 'No authentication required,' which is a precondition, not usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_image_prompt_guideBInspect
Get comprehensive guidance on writing effective AI image generation prompts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no parameters, the description carries the full burden, yet it only says the tool provides 'comprehensive guidance.' It does not describe what the guidance covers, its format, or its length, so it adds little behavioral context beyond the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the action and resource. Nothing is wasted, though it is arguably too terse to be maximally useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter documentation retrieval tool with no output schema, the description is minimally adequate. It does not say what the guide contains or how it differs from the several other prompt/guide tools in the sibling list, leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to document. The baseline of 4 applies; the description correctly implies a no-argument retrieval call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Get) and resource (guidance on AI image generation prompts), which is clear. However, it does not distinguish itself from close siblings like get_video_prompt_guide, get_creative_tips, or get_usage_guide, so an agent must infer the difference from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_video_prompt_guide or get_remotion_layer_guide. The agent gets no signal about which guide to fetch for a given task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mediaGet Sequencer mediaARead-onlyIdempotentInspect
Get detailed information about a specific media object. For completed audio, directMediaUrl and media.url are the playable file. audioAccess.waveformVisualizationUrl is JSON amplitude data for UI rendering only. Call listen_to_audio when the agent needs to hear the file.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | The media ID | |
| workspaceId | Yes | The workspace ID | |
| presentation | No | Use timeline for internal edit media that should stay inside the editor instead of appearing as a standalone chat deliverable. | standalone |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). Beyond that, the description adds genuinely non-obvious response semantics: directMediaUrl/media.url is the playable file while audioAccess.waveformVisualizationUrl is UI-only JSON amplitude data. That prevents an agent from misusing a field as a playable source. No auth or rate-limit context, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler, front-loaded with the core action, then the field semantics that matter most, then the sibling routing. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description usefully clarifies the ambiguous media URLs. For a 3-param read tool with full schema coverage and annotations, this is close to complete; only the lack of explicit differentiation from other getters leaves a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the enum values for presentation and their meaning, so the schema already carries the parameter burden. The description adds nothing about mediaId/workspaceId/presentation syntax, which is the expected baseline when the schema is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get detailed information about a specific media object.' An agent knows this is a single-item detail fetch, distinct from list_media. It does not explicitly contrast with other single-item getters like get_asset, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one explicit routing rule: 'Call listen_to_audio when the agent needs to hear the file,' which is the key adjacent-tool decision for this resource type. However, it offers no guidance on when to prefer this over get_asset or list_media, so the when-not side is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_catalogAInspect
Get available Sequencer AI models for a given modality (image, video, audio, music). Use this when the user asks what models are available or names a model/alias and you need the exact model ID. Common aliases include Nano Banana Pro, Nano Banana 2, Nano Banana 2 Lite, Gemini image, Imagen, Z Image, Z-Image Turbo, MiniMax H3, PrunaAI P-Video, Happy Horse, Google Omni, Flux, GPT Image, Seedream, Recraft, Bria, Runway, Luma, Sora, Veo, Kling, Hailuo, Seedance, PixVerse, ElevenLabs, Fish Audio, Suno, and Lyria.
| Name | Required | Description | Default |
|---|---|---|---|
| modality | Yes | The modality to get models for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden and does disclose the alias-to-ID mapping purpose and that it returns available models. It does not, however, describe the return shape (no output schema) or note that some aliases may not exist for every modality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are tight and front-loaded, but the long list of ~30 aliases is bulky and largely redundant for selecting or invoking the tool. It earns partial credit for utility, but the enumeration bloats the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup tool with full schema coverage, the description is complete enough to select and call it. Missing only a note about what the response includes (model IDs vs. metadata), which is minor given no output schema is required to be described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and modality is an enum already fully self-documenting, so the description adds little beyond listing the four modality values inline. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (available Sequencer AI models), scoped to a modality. Clearly distinguishes itself from siblings like get_sequencer_capabilities or get_available_tools by focusing on model names/IDs for generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: when the user asks what models are available or names an alias and you need the exact model ID. That is a precise, actionable trigger tied to a concrete task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_node_definitionsBInspect
Get detailed schema for all available node types including inputs, outputs, and default values.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Filter by category: input, generation, processing, output, utility | |
| nodeType | No | Get definition for a specific node type |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose return content (inputs, outputs, defaults), which is useful. But it says nothing about read-only safety, output volume, pagination, or whether results are cached/static for this schema-fetching operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler that states the resource and its return payload. It is efficient, though it packs the enumeration into a trailing clause rather than structuring it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameter-light read tool with no output schema, the description covers what is returned, and the schema covers the inputs. It falls short on sibling disambiguation and on how the two filter parameters interact, which an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two parameters, so the schema fully documents category and nodeType. The description adds no parameter detail (e.g., whether category and nodeType combine or are mutually exclusive), so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (node definitions) and describes the payload (inputs, outputs, default values), so an agent knows exactly what it retrieves. However, it offers no differentiation from the very close sibling list_available_nodes, leaving the agent to guess which to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_available_nodes or the sequencer guides. Usage is only implied by the name, which is not enough to route between near-duplicate siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_page_linkAInspect
Get information about a page and provide a clickable link the user can optionally click.
⚠️ DO NOT USE THIS TOOL when users explicitly ask to "go to", "take me to", "open", "show me", or "navigate to" a page. For those requests, use navigate_to_page instead - it will automatically navigate them.
USE THIS TOOL ONLY when:
Answering questions ABOUT a page (e.g., "what can I do on the billing page?")
Suggesting a page the user MIGHT want to visit (but didn't explicitly ask to go there)
Providing reference links in your response
The tool returns a NAV_OFFER tag that renders as a clickable button (user must click to navigate).
Available pages by category:
MARKETING: home, studio, pricing, about, careers, feedback, extension, branding DASHBOARD: projects, images, videos, audio, workflows, billing EDITOR: editor, workflowEditor, flowRunner, publicFlow MODEL ARENA: compare, compareImageToVideo, compareTextToVideo, compareTextToImage, compareElementToVideo, challenges COMMUNITY: explore, jobBoard, tutorials, tutorialShortcut TUTORIALS: tutorial3DRepositioning, tutorialVoiceNormalization, tutorialBlending, tutorialCharacters, tutorialUpscaling, tutorialHooks, tutorialWriting, tutorialRealEstate EVENTS: events, filmFestival LEGAL: privacy, terms CODES: affiliate, codes, claimCode ADMIN: admin, adminArena, adminFestival, adminNewsletters
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | The page to navigate to. Common requests: pricing (costs), projects (dashboard), tutorials (help), compare (model arena), editor (project editing), workflows (automation) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that the tool does not navigate, that it emits a NAV_OFFER tag rendered as a clickable button, and that the user must click to actually navigate. That is the key behavioral trait an agent needs to avoid falsely telling a user they were moved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the critical warning, then scoped by rules, then the page reference. The structure is excellent, but the bulleted page list largely repeats the enum already present in the schema, adding bulk without new routing information beyond the category grouping.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, yet the description explains the return mechanism (NAV_OFFER tag rendered as a button) and the full callable surface. Nothing an agent needs to select and invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum already documents every allowed page value, so the baseline is 3. The description adds value beyond the schema by grouping the values into semantic categories (MARKETING, DASHBOARD, EDITOR, ADMIN, etc.), which helps intent matching, though it does not explain value formatting since the values are opaque identifiers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get information about a page and provide a clickable link') and explicitly distinguishes itself from the sibling navigate_to_page, including the exact phrasing that selects the other tool. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit DO NOT USE rule with quoted trigger phrases, names the alternative tool and what it does, then enumerates three concrete USE-only scenarios. This is the textbook when/when-not/alternative structure.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_platform_overviewAInspect
Get a comprehensive overview of what Sequencer is, including its value proposition, target audience, and key differentiators. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one useful behavioral trait: no authentication is required. It also names the categories of content returned. However, it says nothing about response format, size, or whether the overview is static versus user-specific, which leaves real gaps for a tool with zero structured behavioral hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and followed by the content scope and the no-auth note. Every clause adds information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates by naming the content categories the overview contains and stating the auth requirement. That is sufficient for an agent to decide to call it, though the absence of any sibling differentiation is the one remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the schema is empty and there is nothing for the description to disambiguate. Baseline 4 applies for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') and resource ('overview of what Sequencer is') and enumerates the content covered: value proposition, target audience, and key differentiators. It is clear what the tool returns, but it never distinguishes itself from the many similar informational siblings such as get_sequencer_capabilities, get_feature_list, or get_usage_guide, so an agent still has to guess which orientation tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: this reads like an orientation/intro call, and 'No authentication required' hints it can be invoked before login. There is no explicit statement of when to prefer it over get_sequencer_capabilities or get_feature_list, and no exclusions, leaving the agent to infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pricing_infoBInspect
Get pricing information including subscription tiers, credit costs, and billing FAQs. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | Specific pricing tier to retrieve | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose one genuinely non-obvious trait ('No authentication required'), which an agent could not infer from the schema, and the 'Get' verb implies read-only. However it says nothing about return format (structured tiers vs prose FAQ), freshness, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource and contents, with the auth precondition tacked on efficiently. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description is the only source of behavioral context. It handles the auth precondition and enumerates content categories, but leaves the agent unable to distinguish this from get_balance or get_subscription_help, which is a real gap in a tool set with many overlapping 'get info' siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema description coverage is 100% with an explicit enum plus default. The description adds nothing beyond the schema about the tier filter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('pricing information'), then enumerates contents: subscription tiers, credit costs, and billing FAQs. It does not distinguish itself from near-neighbors like get_subscription_help, get_balance, or get_support_faq, so the agent cannot tell which one to reach for from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use context and names no alternatives, despite several siblings (get_balance, get_subscription_help, get_support_faq) that plausibly overlap. 'No authentication required' is a useful precondition but is not routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_production_statusBRead-onlyIdempotentInspect
Read production generation progress and output media IDs, including any failed shots.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds one behavioral detail beyond the schema – that failed shots are included in the result – but says nothing about progress format, polling frequency, or what happens for an unknown runId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and names the return content. No waste, though it is terse enough that it omits context an agent would want.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description does useful work by stating what is returned (progress, media IDs, failed shots), and annotations cover the safety profile. However, with 0% parameter documentation and no usage or sibling routing guidance, the definition is only minimally adequate for a two-required-param status tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters (workspaceId, runId), and the description does not mention either parameter or add any meaning such as ID format or scope. With two undocumented required params, the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Read') plus the resource ('production generation progress and output media IDs') and even names part of the payload ('failed shots'). It is clear what the tool returns, though it does not explicitly differentiate itself from nearby siblings like get_production_template or check_asset_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided: it does not say to poll this after create_production/generate_shot_video, nor when to prefer it over check_asset_status. Usage is only implied by the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_production_templateBRead-onlyIdempotentInspect
Read a production template or one of its immutable versions.
| Name | Required | Description | Default |
|---|---|---|---|
| version | No | ||
| templateId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered without the description. The description adds one genuine behavioral detail beyond the annotations: that versions are immutable, which is useful context. It does not cover error behavior for a missing or invalid version, so it remains thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the verb and resource lead immediately. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, 0% schema description coverage, and three undocumented parameters, the description carries the burden of explaining inputs and return expectations and does not. A read tool for a versioned resource should at least say what a versioned read returns or how version omission behaves.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters, so the description must compensate. It only loosely implies the meaning of 'version' via 'immutable versions' and says nothing about workspaceId or templateId, leaving two required identifiers entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Read') and resource ('production template or one of its immutable versions'), making the scope clear. It does not explicitly name or contrast with siblings like list_production_templates or run_production_template, so sibling differentiation is only implied by the singular 'a production template'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as list_production_templates (for discovery) or validate_production_template. The phrase 'or one of its immutable versions' hints that a versioned read is possible but never explains when to supply a version or what the absence of one means.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_proposalAInspect
Get the current proposal content for an edit. Returns the full markdown plus reviewUrl. External MCP clients should show the markdown inline in chat first and treat reviewUrl as optional.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| proposalId | No | Specific proposal ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden. It usefully discloses the return shape (full markdown plus reviewUrl) and how to surface it, but says nothing about permissions, whether the read is scoped by workspace, or what happens when proposalId is omitted. Adequate for a read tool but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the operation before the presentation advice. No filler, though the client-display sentence sits slightly awkwardly next to the core definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by naming the return fields (markdown and reviewUrl). Combined with 100% schema coverage on the inputs, an agent has what it needs to call the tool, though the optional proposalId remains ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all three parameters and the baseline is 3. The description adds no meaning beyond 'for an edit' and does not clarify the relationship between the required editId and the optional proposalId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the current proposal content for an edit'), which an agent can distinguish from the mutating siblings save_proposal and update_proposal. It does not name those siblings explicitly, but the read verb makes the boundary clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the context ('for an edit') and gives presentation guidance to external MCP clients (show markdown inline, reviewUrl optional), but never states when to prefer this over update_proposal/save_proposal or what conditions select it. Usage is inferred rather than specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_remotion_layer_guideAInspect
Get the contract, supported APIs, and example for creating custom React animation layers.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the returned content (contract, APIs, example) but says nothing about side effects, permissions, or output format; the read-only nature is only implied by 'Get'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that names the verb, resource, and contents without waste. It is appropriately sized for a simple guide-retrieval tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter guide tool with no output schema, the description identifies what documentation is returned (contract, supported APIs, example). It does not specify format or length, but that is a minor omission for this tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to clarify. Per the rubric, zero params yields a baseline of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('remotion layer guide'), and enumerates the payload: contract, supported APIs, and example. The focus on 'custom React animation layers' distinguishes it from sibling action tools like add_remotion_overlay.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for creating custom React animation layers' provides clear context for when to invoke this guide, but it does not name alternatives or exclusions (e.g., when to use add_remotion_overlay instead). This is clear context without explicit when-not guidance, matching a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sceneCInspect
Get detailed information about a specific scene.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| sceneId | Yes | The scene ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Get' implies a read-only operation, but nothing confirms it, and there is no disclosure of what 'detailed information' actually contains, whether it is expensive, or whether the workspace/edit hierarchy must exist first.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though its brevity reflects under-specification as much as economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and three required identifiers, the description is only minimally sufficient. It never explains the return shape or the containment relationship among workspaceId, editId, and sceneId, but the fully-covered schema keeps it viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three required parameters are documented in the schema (though only tautologically, e.g. 'The scene ID'). The description adds no scoping detail such as the workspace→edit→scene containment, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('a specific scene') plus the scope of the payload ('detailed information'), so an agent knows this is a single-entity read. It does not, however, distinguish itself from siblings like list_scenes or get_edit_full, leaving the boundary implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus list_scenes, get_edit, or get_shot, and no prerequisites are stated. The only usable signal is the singular 'specific scene', which the agent must infer means 'after you already have a sceneId'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sequencer_capabilitiesAInspect
Call this first when deciding whether Sequencer can help with a user request. It maps common user intents and model aliases to the right Sequencer MCP tools, including image, video, audio, music, storyboard, edit, workspace, and media workflows.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the tool's nature as a first-call intent-mapping oracle rather than a mutating operation, which is genuinely useful context. It does not state that it is read-only, whether it has no side effects, or anything about rate limits or response shape beyond the conceptual mapping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, tightly written, with the critical action-first directive ("Call this first") front-loaded. No filler or restated title content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, annotation-free, output-schema-free discovery tool, the description adequately conveys what the agent gets back (intent/alias to tool mapping) and when to call it. It could be slightly more complete by clarifying how it relates to the other catalog-style siblings, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero parameters, so per the rubric the baseline is 4. There is nothing for the description to explain, and it correctly avoids inventing parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific function: a discovery/router tool that maps user intents and model aliases to the correct Sequencer MCP tool, and enumerates the covered domains (image, video, audio, music, storyboard, edit, workspace, media). It is clearly not a generator or editor. However, it doesn't distinguish itself from other discovery siblings like get_available_tools, get_feature_list, or get_usage_guide, so sibling differentiation is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Call this first when deciding whether Sequencer can help with a user request" gives an explicit trigger condition and even an ordering rule. It stops short of naming alternatives or when-not-to-use conditions (e.g., when a concrete tool is already known), leaving some routing ambiguity against the other list/get-capability siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sequencer_skillLoad Sequencer skillARead-onlyIdempotentInspect
Load the complete public instructions for a Sequencer skill before using it. Supports motion design, films, webpage videos, games, restaurant menu videos, news videos, architectural tours and article podcasts. Read-only, requires no account access and generates no media.
| Name | Required | Description | Default |
|---|---|---|---|
| skillId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and non-destructive behavior, so the description is not the only source. It still adds value by asserting 'requires no account access and generates no media' — the auth-free nature is not captured by the annotations, though the no-media claim largely restates readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action before the domain list and the read-only guarantees. The supported-domain enumeration is longer than strictly necessary but is informative rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only loader with no output schema, the definition covers purpose, timing, and safety properties adequately. The remaining gap is that the returned 'instructions' format and any length or usage constraints are unspecified, but this is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single skillId parameter has 0% schema description coverage, but the enum itself enumerates all valid values. The description partially compensates by naming the supported domains (motion design, films, webpage videos, games, restaurant menu videos, news videos, architectural tours, podcasts), which loosely map to enum entries, but it never names the parameter or explains that one of these identifiers must be supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Load the complete public instructions for a Sequencer skill.' An agent can distinguish this from most siblings, though it does not explicitly contrast with the close sibling get_sequencer_capabilities, leaving some ambiguity about which Sequencer-related tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'before using it' gives a timing cue, implying this is a prerequisite step prior to invoking a skill. However, there is no explicit when-not guidance and no reference to get_sequencer_capabilities or other discovery siblings, so the agent must infer the selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_server_infoCInspect
Get information about this MCP server and its capabilities.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses almost nothing: not whether this is a safe read, what fields come back (server name, version, tool list?), whether it is cached or static, or whether any auth is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. It is efficient, though the brevity here comes partly from under-specification rather than from disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and a crowded family of overlapping discovery tools, the description should at minimum say what information is returned and how it differs from get_available_tools/get_feature_list. As written it leaves the agent unable to know what it will get back or when it is the correct choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics for the description to add; the baseline of 4 applies. Nothing in the description misleads about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a verb and resource ('get information about this MCP server'), which is more than a tautology, but 'and its capabilities' is vague and the description does nothing to separate it from near-identical siblings such as get_available_tools, get_feature_list, get_platform_overview, or get_model_catalog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no stated trigger, and no mention of any alternative tool. An agent facing a dozen sibling discovery tools gets no help deciding whether this one or get_feature_list/get_available_tools is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_shotBInspect
Get detailed information about a specific shot including its media references.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Get' implies a read, but the description says nothing about error behavior for missing IDs, permission requirements, or whether the media references are IDs, URLs, or embedded objects. Only a minimal content hint is offered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The key noun phrase ('a specific shot') and the scope hint ('including its media references') are both stated up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-ID read tool with no output schema and no annotations, the description is minimally viable: it says what entity is fetched and one aspect of the return. It does not describe the response shape or the relationship between the three required IDs, leaving gaps an agent must discover at call time.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all three parameters (workspaceId, editId, shotId), so the schema already documents them. The description adds no syntax, format, or hierarchical relationship detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get detailed information about a specific shot.' The added phrase 'including its media references' hints at content scope. It does not differentiate itself from siblings like get_scene, get_edit_full, or list_shots, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus list_shots (for enumerating shots) or get_edit_full (for a whole edit). No prerequisites or context given. The agent must infer usage purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_storytelling_guideCInspect
Get a comprehensive storytelling framework guide for creating compelling narratives.
| Name | Required | Description | Default |
|---|---|---|---|
| framework | No | Specific framework to get | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies read-only static content via the word 'guide', but does not disclose the return format, whether the guide is static or generated, or any auth constraints. Beyond the noun 'guide' it adds little behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the resource is named immediately. It is appropriately sized, though it is arguably too terse to carry the missing usage and behavioral context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, non-required, read-only guide tool with no output schema, the description is minimally viable. It does not explain what the guide actually contains or how the six framework options differ, which an agent deciding among frameworks would want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the 'framework' enum and its default are already fully documented in the schema. The description does not mention the parameter or add any meaning (e.g., how 'all' differs from a single framework), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Get a ... storytelling framework guide'), which is clearly distinct from other get_*_guide siblings like get_video_prompt_guide or get_remotion_layer_guide. However, it never names or contrasts those siblings, so differentiation relies on the topic noun alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for creating compelling narratives' states a purpose, not a usage guideline. There is no indication of when to call this versus other guide tools (e.g., get_creative_tips, get_image_prompt_guide) or any prerequisite/context for invoking it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_subscription_helpBInspect
Get help with subscription management including upgrading, downgrading, canceling, and payment issues. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | Specific subscription action to get help with | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses one important trait ('No authentication required'), which helps the agent know it can call this without credentials. However, it does not state whether the tool is read-only, what it returns (text, links, etc.), or any other side-effect behavior, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The purpose is front-loaded, and the key behavioral fact ('No authentication required') is placed immediately after, making the description easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, rich schema coverage, and absence of annotations or an output schema, the description is minimally adequate. It states the tool's scope and the no-auth condition, but omits what the tool returns and fails to resolve the 'canceling' vs enum mismatch, leaving the agent with some uncertainty.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single enum parameter is documented, so baseline would be 3. However, the description introduces a mismatch: it lists 'canceling' as a supported action, but the enum contains no 'cancel' value (only upgrade, downgrade, changePayment, contactSupport, all). This could mislead an agent into passing an invalid action.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get help') and resource ('subscription management') and enumerates relevant sub-actions (upgrading, downgrading, canceling, payment issues). It clearly tells an agent what the tool is for, but does not differentiate it from sibling tools like get_support_faq or get_pricing_info, which could also be relevant for subscription-related questions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (subscription management help) and adds a useful condition ('No authentication required'), but gives no explicit guidance on when to use this tool instead of alternatives like get_support_faq or request_login. The agent must infer the appropriate scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_support_faqAInspect
Get frequently asked support questions and answers. Use this to help users with common issues. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Support topic category | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses 'No authentication required,' which is genuine behavioral context, and 'Get' implies a read-only operation. But it says nothing about result size, pagination, or caching/staleness of FAQ content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool returns. No wasted prose, though the middle sentence ('help users with common issues') is generic filler that adds little routing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, single-optional-parameter read tool with no output schema, the description is nearly complete: it states purpose, usage context, and the notable no-auth property. The main gap is that it neither hints at the return shape nor differentiates from related help/info siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single enum parameter and a default, so the schema already fully documents 'topic.' The description adds no syntax, default-behavior, or filtering semantics beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb+resource are explicit: 'Get frequently asked support questions and answers.' It clearly states what is returned. However, it does nothing to distinguish itself from closely related siblings such as get_subscription_help, get_pricing_info, get_feature_list, and get_usage_guide, all of which could plausibly cover 'common issues.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to help users with common issues' implies a usage context but names no alternatives and no exclusions. An agent cannot tell from this text whether to route a billing question here or to get_subscription_help/get_pricing_info.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_text_presetsAInspect
Get the available text overlay presets. Each preset provides a complete styling template (font, color, animations, position) for common text overlay styles like cinematic titles, lower-thirds, subtitles, breaking news, etc. Use the preset ID with add_text_overlay to quickly create styled text.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the shape of the returned data (a complete styling template covering font, color, animations, position), which is useful behavioral context, but says nothing about read-only safety, pagination, or any auth/permission requirements. Adequate but with clear gaps for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose, then the payload description, then the actionable next step. No filler, no repetition of the tool name as a definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by explaining what a preset contains and giving example categories, which is what an agent needs to decide whether to call it. It stops short of noting ordering, count, or stability of IDs, but for a static preset list this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline this scores 4. The description adds no parameter detail because none exists, and it correctly spends its words on return-value semantics instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the available text overlay presets') and immediately characterizes what a preset is, distinguishing it from sibling read tools. The description of preset contents (font, color, animations, position) and examples (cinematic titles, lower-thirds, subtitles, breaking news) make the resource concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent how to act on the result: 'Use the preset ID with add_text_overlay to quickly create styled text,' which names the downstream sibling and the condition for using it. No exclusions or when-not guidance are given, but for a simple enumerating read tool the routing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usage_guideAInspect
Get the current usage guide for the Sequencer MCP server, including authentication, tool categories, generation workflows, and response conventions. Use this when integrating a new MCP client or when an agent is unsure how to operate Sequencer.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It helpfully lists the kinds of content returned (auth, categories, workflows, response conventions), which implies a safe read-only fetch and partially compensates for the absent output schema, but it says nothing about how often the guide changes, whether it is large, or how it relates to the other guide-style tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-formed sentences with no filler; the scope of the guide is front-loaded and the usage condition follows immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter documentation-retrieval tool with no output schema, the description is nearly self-sufficient: it tells the agent what the payload covers, which is the key missing structured information. Only the relationship to the many sibling guide/overview tools remains unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing further for the description to disambiguate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (get the current usage guide) and enumerates its contents: authentication, tool categories, generation workflows, response conventions. It is clear on its own, but does not differentiate itself from closely related siblings such as get_server_info, get_available_tools, or get_platform_overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the triggering situations: integrating a new MCP client, or an agent being unsure how to operate Sequencer. It gives clear context but names no alternative (e.g. when to prefer get_available_tools or get_sequencer_capabilities instead), so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_prompt_guideAInspect
Get comprehensive guidance on writing effective AI video generation prompts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It implies a read-only, side-effect-free reference fetch ('get ... guidance'), which is low risk, and 'comprehensive' hints at volume, but it says nothing about the shape, size, or format of what comes back and no output schema exists to fill that gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though it is arguably under-sized rather than optimally sized given the guidance opportunities available.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only guide fetch with no output schema, the description is minimally sufficient. It does not say what topics the guidance covers or how it relates to the other guide tools, leaving the agent with only a general sense of scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and it correctly avoids inventing parameter language.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('get ... guidance on writing ... prompts') with a clear domain qualifier ('AI video generation'). It is clear what the tool returns, but it never distinguishes itself from the near-identical sibling get_image_prompt_guide or from get_creative_tips/get_storytelling_guide, so an agent must infer the split.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can guess it should call this before composing a video generation prompt. There is no explicit when-to-use trigger, no when-not-to-use, and no mention of the adjacent guide tools it overlaps with (get_image_prompt_guide, get_creative_tips).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workflowBInspect
Get detailed information about a workflow including all nodes and connections.
| Name | Required | Description | Default |
|---|---|---|---|
| workflowId | Yes | The workflow ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the result includes all nodes and connections, which is useful return-content context, but it does not state that the operation is read-only, what permissions are required, or any rate limits or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the tool's purpose and scope with zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation with two well-documented parameters and no output schema, the description adequately explains what the tool returns. It could be improved with usage guidance or notes on permissions, but it is largely complete for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both required parameters (workspaceId, workflowId) are fully documented in the schema. The description adds no parameter-level detail, which is acceptable given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('workflow'), and specifies scope ('detailed information including all nodes and connections'). This distinguishes it from a mere summary, but it does not explicitly name sibling tools like get_workflow_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as get_workflow_summary or list_workflows. The description only states what it returns, leaving the agent to infer usage contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workflow_building_guideAInspect
Get comprehensive guidance on building node-based workflows, including available node types, socket connections, and common workflow patterns. CALL THIS FIRST when building complex workflows.
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | No | Specific workflow pattern to get | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Get comprehensive guidance' implies a safe read-only informational tool, and the content scope is stated, but there is no explicit safety, auth, or rate-limit context. For a simple guide-retrieval tool this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core purpose is front-loaded, and the directive to call it first is clearly placed after the content summary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only one optional parameter, the description sufficiently explains what the tool returns and when to use it. It could be more complete by distinguishing the guide from overlapping sibling tools, but it covers the essentials for this simple retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the enum 'pattern' parameter is already documented in the schema. The description does not mention the parameter or explain how choosing a workflow pattern affects the returned guidance. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: getting guidance on building node-based workflows, and spells out the content categories (node types, socket connections, common patterns). It does not explicitly differentiate itself from siblings like get_node_definitions or list_available_nodes, which also cover node information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context with 'CALL THIS FIRST when building complex workflows,' which tells the agent when to invoke it. However, it names no alternatives and gives no exclusions, such as when a simpler node lookup is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workflow_summaryBInspect
Get a human-readable summary of a workflow's structure and data flow.
| Name | Required | Description | Default |
|---|---|---|---|
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It discloses the output is a human-readable summary of structure and data flow, which implies a read-only operation, but it omits permissions, side effects, error behavior, and rate limits, leaving clear gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero wasted words. It is appropriately sized for a simple get tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the return value (a summary), which is useful since there is no output schema. However, it lacks usage context and parameter semantics, so it is only minimally complete for an agent choosing among many sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (workspaceId documented, workflowId not). The description adds no parameter-level meaning beyond the schema; it does not clarify either ID's role or format, so it only partially compensates for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get' plus 'human-readable summary of a workflow's structure and data flow.' It is clear what the tool does, but it does not explicitly differentiate itself from sibling tools like get_workflow or list_workflows, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no prerequisites, and no alternatives. It does not tell the agent when to choose this over get_workflow, so the score is 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workspaceBInspect
Get detailed information about a specific workspace including member count.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses only that the result includes a member count; it says nothing about read-only safety, auth/permission requirements, or error behavior for an unknown ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the resource and payload hint front-loaded and no wasted words. It is appropriately sized for a one-parameter lookup tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-ID read tool with full schema coverage and no output schema, the description covers the essentials. The main gap is the lack of routing guidance against the sibling list/search workspace tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already documents 'workspaceId'. The description adds no format or sourcing detail beyond what the schema provides, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get detailed information about a specific workspace') and even names a piece of the payload ('member count'). An agent can distinguish it from list_workspaces/search_workspace by the singular, ID-scoped framing, though the description never explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: fetch by a known workspaceId. The description does not say when to prefer this over search_workspace or list_workspaces, nor any prerequisites, leaving the agent to infer selection from the required ID parameter alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_media_from_urlAInspect
Import a publicly reachable image, video, or audio URL into Sequencer workspace media and return a mediaId. Use this when a file is already hosted at an HTTP(S) URL. Local disk paths such as C:/... are not reachable by the MCP server; use upload_media_base64 or the REST /v1/upload endpoint for local files.
| Name | Required | Description | Default |
|---|---|---|---|
| fileName | No | Optional file name. If omitted, derived from the URL. | |
| folderId | No | Optional destination folder ID | |
| mimeType | No | Optional MIME type. If omitted, inferred from the file name or HTTP content-type. | |
| sourceUrl | Yes | Publicly reachable HTTP(S) URL to import | |
| workspaceId | Yes | The workspace ID | |
| presentation | No | Use timeline when the import is internal production media for an edit and should not create a standalone chat preview. | standalone |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool imports a publicly reachable URL, returns a mediaId, and that local paths are not reachable, but it omits permissions, rate limits, idempotency, and error behavior. This is partial behavioral context, so a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and return value, then the usage condition and the alternative for local files. Every sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage, alternatives, and return value for a tool with no output schema and full schema coverage. It lacks behavioral details like authentication or async behavior, but those are secondary for correct invocation, so a 4 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters. The description adds no parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Import' and resource 'publicly reachable image, video, or audio URL into Sequencer workspace media', plus the return value 'mediaId'. It distinguishes itself from sibling upload_media_base64 by explicitly routing local files elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use ('when a file is already hosted at an HTTP(S) URL'), when not to use ('Local disk paths such as C:/... are not reachable'), and names alternatives ('use upload_media_base64 or the REST /v1/upload endpoint'). This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_edit_frameAInspect
Render and inspect the final composed frame of an edit at an exact timeline time. This is the fast visual spot-check tool for typography, overlays, framing, style consistency, transition boundaries, and whether the picture matches the intended beat. It renders the real timeline composition without exporting the entire video or mixing audio. Use watch_video or review_edit when motion, pacing, or sound must be judged across time.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit/project ID whose composed timeline frame should be inspected | |
| workspaceId | Yes | The workspace ID | |
| timestampSeconds | Yes | Exact timeline position in seconds, for example 12.5 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses that it renders the real timeline composition, avoids exporting the entire video, and does not mix audio — all meaningful cost/scope traits. It is missing only auth/permission requirements and how the inspected frame is delivered back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: what it does, what it is for, and what it is not for. The core purpose is front-loaded and the routing guidance follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read/render tool with no annotations and no output schema, the description covers purpose, scope limits, and alternatives well. The main residual gap is not saying what the caller receives (frame image, URL, or analysis), which matters given there is no output schema to explain the return.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all three params are documented in the schema, so the baseline is 3. The description reinforces 'exact timeline time' but adds no format or constraint detail beyond what the schema already provides (e.g., the 12.5 seconds example).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Render and inspect the final composed frame of an edit at an exact timeline time.' It immediately distinguishes itself from siblings by naming watch_video and review_edit as the tools for motion/pacing/sound.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames this as the 'fast visual spot-check' for static attributes (typography, overlays, framing, transitions) and states the when-not condition: use watch_video or review_edit when motion, pacing, or sound must be judged across time. Alternatives and selection criteria are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_storyboardAInspect
Build one ordered contact sheet from every completed shot starting frame in an edit and load it into visual context. Use this after storyboard image generation and before the first shot video. Judge the sequence as a whole for continuity, composition, defects, style, full-bleed framing, shot variety, and duration rhythm. Re-run it after changing a shot image, shot order, or timing.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit whose storyboard should be inspected | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully reveals that the operation loads imagery into visual context and lists evaluation criteria, but says nothing about permissions, cost, rate limits, or whether the underlying data is modified. Adequate but incomplete for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place, with the core action front-loaded followed by timing, criteria, and re-run conditions. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, no-output-schema workflow tool, the description covers purpose, workflow timing, judging criteria, and re-run conditions well. It stops short of describing what the loaded contact sheet looks like or any constraints, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both workspaceId and editId are already documented in the schema. The description adds no format, syntax, or scoping detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Build one ordered contact sheet from every completed shot starting frame in an edit and load it into visual context.' This clearly separates it from the single-frame sibling inspect_edit_frame by scoping to the whole sequence, though it never names that alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong lifecycle placement ('after storyboard image generation and before the first shot video') and explicit re-run triggers ('after changing a shot image, shot order, or timing'). It lacks a named alternative or a 'when not to use' clause, which keeps it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lip_sync_video_mediaBInspect
Run lip sync on a video media item using a supplied audio URL. Queues the V3 job against the existing MediaDocV2.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | Video MediaDocV2 ID to process | |
| modelId | Yes | Lip-sync model ID | |
| audioUrl | Yes | Audio URL to drive the lip sync | |
| videoUrl | No | Override source video URL; otherwise resolves from active media version | |
| extraParams | No | Additional raw V3 parameters | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose that the invocation queues an asynchronous V3 job rather than producing a result inline, which is genuinely useful. It omits permissions/auth needs, cost, expected duration, and how to poll the queued job.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the execution model. Every clause carries information; the only minor deduction is that the second sentence is compressed to the point of jargon ('V3 job', 'MediaDocV2') without unpacking it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation-style tool with a nested free-form extraParams object and no output schema, the description should say what the call returns (e.g., a job or media ID) and how the queued job is tracked. It covers the core action but leaves result handling and the extraParams escape hatch unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains mediaId, modelId, audioUrl, videoUrl, extraParams, and workspaceId. The description only reinforces the media/audio relationship and adds nothing about the videoUrl override or extraParams passthrough semantics, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Run lip sync on a video media item using a supplied audio URL.' It also identifies the mechanics ('Queues the V3 job against the existing MediaDocV2'), which separates it from generic voice/edit tools. It does not explicitly name a sibling it is not, so it falls short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives such as change_video_voice or generate_video_audio, which are the closest siblings an agent would weigh against this. 'Against the existing MediaDocV2' implies a precondition but never states it as a rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_assetsCInspect
List assets in a workspace. Can filter by type (character, location, prop).
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter by asset type | |
| limit | No | Maximum number of assets to return | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read-only listing but never says so, and it omits pagination behavior despite a limit parameter defaulting to 50, nor does it say whether results are ordered or scoped to a workspace the caller must have access to.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core operation front-loaded and zero filler. The second sentence is slightly redundant with the schema's type enum, which keeps it from being maximally efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with a fully documented three-parameter schema, the description covers the essentials. With no output schema, though, it should at least hint at what is returned or how the 50-item default limit behaves, which it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so workspaceId, limit, and the type enum are already fully documented in the schema. The description only restates the type filter with the same enum values, adding no syntax or behavioral detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List assets in a workspace') plus the filterable dimension, so the operation is unambiguous. However, it never distinguishes itself from closely related siblings such as list_media, list_folders, or get_asset, so an agent still has to guess which listing tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Can filter by type' describes a capability, not a condition for selecting this tool. There is no guidance on when to prefer list_assets over get_asset (single asset) or list_media, nor any mention of prerequisites like needing a valid workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_audio_tracksBInspect
List all audio timeline tracks in an edit, with resolved audio media URLs/status when available.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that audio media URLs/status are 'resolved ... when available', implying results may be incomplete or pending, but it omits pagination, ordering, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the scope ('in an edit') plus the return caveat are packed efficiently. Slightly terse given the behavioral burden placed on it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema, the description gives a reasonable sense of what comes back, including the conditional resolution of media URLs. It still leaves return ordering, pagination, and failure behavior unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both editId and workspaceId are already documented in the schema. The description adds no extra meaning about these parameters, which is the expected baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('audio timeline tracks in an edit'), which an agent can separate from sibling 'list_overlay_tracks'. It does not explicitly name the sibling it differs from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in an edit' implies the tool is scoped to one edit's audio tracks, but there is no statement of when to prefer this over list_overlay_tracks, list_media, or get_edit_full, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_nodesAInspect
List all node types available for building workflows. Use get_node_definitions for detailed schema.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. 'List' strongly implies a non-mutating read, which is the key safety signal, but the description says nothing about the category filter's effect, result shape, or whether the list is paginated/complete. Adequate for a simple catalog read, thin given zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the primary purpose front-loaded and the routing hint second. Ideal size for a simple listing tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should say more about what a 'node type' looks like in the response (name, id, category?) and that the category argument narrows results. It covers the core purpose but leaves the return shape and the sole parameter unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'category' parameter or that results can be filtered. The enum values (input, generation, processing, output, utility, all) are largely self-explanatory, but an agent reading only the prose would not know filtering exists at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all node types available for building workflows'), which is exactly distinguishable from the neighboring get_node_definitions. An agent can tell immediately this is the discovery/catalog call rather than the detailed-schema call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to the sibling tool: 'Use get_node_definitions for detailed schema.' That is a clear alternative and the condition that selects it. It stops short of a full when/when-not statement (e.g. no mention of when to prefer this over get_workflow_building_guide).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_editsCInspect
List all edits (projects) in a workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of edits to return | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers almost nothing: no indication that this is a read-only operation, no pagination behavior, no statement of what happens when the workspace has many edits. Worse, 'all edits' mildly conflicts with the schema's limit parameter defaulting to 20, so the description overstates the completeness of the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler and the resource front-loaded. It is efficient, though the word 'all' is the one term that should have been qualified rather than kept.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter list tool with no output schema and no annotations, this is minimal but barely viable. The agent still lacks the return shape and pagination behavior, and the 'all' versus limit=20 tension is unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented (workspaceId and limit), and the baseline is 3. The phrase 'in a workspace' reinforces the required workspaceId but adds no syntax or semantics beyond the schema, and it never mentions the limit control.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List all edits') and usefully disambiguates the term by equating edits with projects. It does not, however, distinguish itself from siblings like get_edit, get_edit_full, or list_workspaces, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this listing versus get_edit, get_edit_full, search_workspace, or review_edit. The only implicit context is the workspace scoping, and there is no statement of prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
listen_to_audioAInspect
Load a completed Sequencer audio item into the agent context so it can hear narration, pacing, pauses, pronunciation, and delivery before timing visuals.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | The completed audio media ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose real behavior — the audio is loaded into the agent's context (not just referenced) and the prerequisite that the item be 'completed' — but says nothing about permissions, size/duration limits, or failure behavior for what is otherwise an unannotated operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler, front-loading the action and resource before the intent. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter load tool with no output schema and no annotations, the description covers the action, the target, the prerequisite, and the payoff (hearing the audio to judge delivery). It does not clarify side effects or the exact form the loaded audio takes, but it is largely sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so mediaId and workspaceId are already documented in the schema. The description adds no syntax, format, or constraint detail beyond that, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (load) and resource (completed Sequencer audio item) and clarifies the intent: perceiving narration, pacing, pauses, pronunciation, and delivery. This distinguishes it from watch_video and get_media conceptually, though it never names an alternative sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before timing visuals' implies the workflow stage at which this should be called, which is useful implied guidance. However, there is no explicit when-not condition and no named alternative (e.g., watch_video vs listen_to_audio), leaving routing partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_experience_gamesBRead-onlyIdempotentInspect
List saved playable games in a workspace, with playtest links and current version IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world scope, so safety is covered. The description adds useful context about what is returned (playtest links, current version IDs), but says nothing about pagination, ordering, or whether it lists across the workspace or only saved ones beyond the word 'saved'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with the output detail trailing. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with rich annotations and no output schema, the description covers purpose, scope, and the shape of returned data. Only the absence of usage routing keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single workspaceId parameter, but the description implies the workspace scoping ('in a workspace'). It doesn't add format or validation meaning beyond that, so it only partially compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List saved playable games') and adds scope ('in a workspace') plus return content ('playtest links and current version IDs'). An agent can distinguish it from siblings like get_experience_game and save_experience_game, though the singular/plural relationship is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use framing or alternatives are given. The agent must infer from the name that this lists games while get_experience_game fetches one and save_experience_game creates one; the description never routes it that way.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_exportsBInspect
List prior exports for an edit, including status, format, version index, URLs, and export settings.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum exports to return | |
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It mentions what fields are returned (status, format, version index, URLs, export settings), which is helpful, but does not state that it is read-only, whether it paginates, how results are sorted, or any authorization requirements. Significant behavioral gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence that front-loads the scope ('List prior exports for an edit') and then lists the returned fields. Every part earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no output schema and no annotations, the description successfully conveys the return contents. It is largely complete, though it could mention pagination behavior or how the limit parameter affects results to be fully thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (workspaceId, editId, limit). The description adds no additional meaning beyond the schema; it merely implies the edit context, which is already captured by the required editId parameter. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (list) and resource (prior exports) scoped to an edit, and enumerates the fields returned. It is clear enough to distinguish from general edit operations, but does not explicitly differentiate itself from sibling tools like get_export or export_edit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided. The description implies it is for viewing prior exports, but does not explain when to choose this over get_export (single export details) or export_edit (creating an export), and offers no exclusions or context on typical invocation scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_foldersCInspect
List folders in a workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | The workspace ID | |
| parentFolderId | No | Parent folder ID to list children of (null for root) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and supplies almost none. It does not state that the operation is read-only, whether results are paginated or ordered, or what permission level is required to enumerate a workspace's folders.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the verb and resource. It is efficient, though the brevity comes partly from under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter list tool this is minimally viable, but with no output schema the description could usefully say what is returned (folder IDs, names, nesting) and whether the hierarchy is flat or recursive. The parentFolderId root case is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself already explains workspaceId and the parentFolderId root-vs-child semantics. The description adds no syntax, format, or behavioral detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (folders) scoped to a workspace, so the agent knows exactly what operation is performed. It does not differentiate itself from nearby siblings like create_folder or move_item_to_folder, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives, nor any stated prerequisites or exclusions. The agent must infer that this is the read counterpart to create_folder purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mediaCInspect
List media objects in a workspace. Can filter by type (image, video, audio) and status.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter by media type | |
| limit | No | Maximum number of media to return | |
| status | No | Filter by generation status | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state that this is a read-only operation, does not mention pagination despite a limit parameter defaulting to 20, and does not describe ordering or what a result contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the resource front-loaded and the filtering capability immediately after. Nothing is wasted, though the brevity comes partly from under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no output schema and no annotations, the description covers the basics but omits the return shape, ordering, and pagination behavior. An agent can call it, but cannot predict what comes back or how to page through results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each of the four parameters described inline including enum values, so the baseline is 3. The description restates the type and status filters but adds no syntax or semantic detail beyond the schema, and ignores limit and workspaceId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List media objects') scoped to a workspace, which is unambiguous. However, it never distinguishes itself from closely named siblings like list_assets, get_media, or search_workspace, so an agent must infer which lister applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of filterable fields hints at how to narrow results, but there is no when-to-use guidance, no exclusions, and no pointer to alternatives such as list_assets or get_media. The agent is left to guess which listing tool fits a given request.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_overlay_itemsBInspect
List all overlay items in a specific overlay track. Returns full item details including text config, position, timing, and animations.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| trackId | Yes | The overlay track ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose that full item details are returned (text config, position, timing, animations), which is useful behavioral context, but says nothing about pagination, ordering, permissions, or the read-only nature of the call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the primary action first, the return payload second. Nothing wasted, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description's summary of returned fields is doing necessary work. For a three-ID read tool with no annotations, this is close to sufficient, though a note on ordering/pagination would complete it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three required IDs are documented in the schema, so the baseline of 3 applies. The description adds no per-parameter meaning beyond restating that the track must be an overlay track.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (overlay items) scoped to a specific overlay track, which cleanly separates it from list_overlay_tracks (tracks) and get_* variants. It does not explicitly name or route against siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as get_edit_full or list_overlay_tracks. The agent must infer that this is the way to enumerate items within a single track.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_overlay_tracksBInspect
List all overlay tracks in an edit. Overlay tracks contain text, image, and video items layered on top of the main video timeline.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. 'List' implies a safe read, and the second sentence usefully discloses what the returned tracks contain (text, image, and video items), but there is no mention of pagination, ordering, permissions, or return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, purpose front-loaded ahead of the clarifying definition. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with full schema coverage and no output schema, the description supplies the domain concept an agent needs and is small enough to be honest. Minor gaps (pagination, ordering) are not critical here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two required identifiers (workspaceId, editId), so the schema already documents both parameters. The description adds no syntax or constraint detail beyond that, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (List) plus resource (overlay tracks) and a scope qualifier ('in an edit'). The second sentence defines what an overlay track is, which helps distinguish it from list_audio_tracks and list_overlay_items, though it does not name those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use statement, no prerequisites, and no routing to alternatives such as list_overlay_items (which lists the items inside these tracks) or list_audio_tracks. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_production_templatesBRead-onlyIdempotentInspect
List reusable production templates in a workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description only adds that results are scoped to a workspace, with no mention of pagination, ordering, or result volume.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the scope qualifier front-loaded and no wasted words. It is not under-specified to the point of being useless, though it is spartan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no output schema, the essential purpose is conveyed, but nothing is said about return shape, ordering, or limits. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter with 0% schema description coverage. The phrase 'in a workspace' hints that workspaceId scopes the listing, which is minimal but real added meaning over the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (reusable production templates) with a workspace scope. It is distinguishable from singular/mutating siblings like get_production_template, run_production_template, and save_production_template, but it never explicitly names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this list tool versus get_production_template for a single template, or how it relates to run/save/validate variants. The agent must infer the choice from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scenesAInspect
List all scenes in an edit, ordered by sequence.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the 'ordered by sequence' return ordering, but says nothing about pagination, result size, whether full scene objects or summaries are returned, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; scope and ordering are stated immediately with no redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-required-param read-only list tool with no output schema and no annotations, the description covers the essentials of what is returned and in what order. It falls short only on pagination/return-shape expectations, which the absence of an output schema makes slightly more relevant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both required params (workspaceId, editId), so the schema already documents them fully. The description adds no format, constraint, or scoping detail beyond the schema's own field descriptions, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (scenes) scoped to 'in an edit' and adds the ordering behavior. It is clearly distinguishable from the singular get_scene and from list_shots/list_edits, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the 'in an edit' phrase signals you must already have an editId, but there is no explicit when-to-use guidance and no mention of when to prefer get_scene (single) or list_shots (children) instead. Nothing is misleading, but nothing routes the agent either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_shotsBInspect
List all shots in an edit, optionally filtered by sceneId. Shots reference media objects instead of embedding generation data.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| sceneId | No | Filter to shots in a specific scene | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does disclose one genuinely useful data-model trait: shots reference media objects rather than embedding generation data. However, it omits pagination, ordering, and permission requirements for a list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action and scope front-loaded and the data-model caveat placed after it. Every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with a fully documented schema and no output schema, the description covers purpose, scope, and one structural note. Minor gaps remain around result ordering/pagination, but nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including sceneId as a filter. The description's mention of the sceneId filter adds no syntax or semantics beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (shots) with a clear scope qualifier ('in an edit'), which distinguishes it from the singular get_shot and from list_scenes. It does not explicitly name a sibling alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no conditions for preferring this over get_shot or list_scenes, and no prerequisites. 'Optionally filtered by sceneId' is parameter behavior, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workflowsCInspect
List all workflows in a workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers almost none: it does not state that the operation is read-only, whether results are paginated, what the default ordering is, or what happens with the limit parameter. 'List all' even conflicts mildly with the presence of a default limit of 20.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the scope stated up front and no filler. It is arguably too terse for the parameters it leaves unexplained, but it wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter list tool with no output schema and no annotations, the description is minimally adequate. The pagination/limit semantics are the one real gap an agent would need in order to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: workspaceId is documented in the schema and reinforced by 'in a workspace', but limit has no description anywhere, leaving page-size behavior undefined. The description adds little beyond the schema, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List all workflows') and scopes it to a workspace, which is unambiguous on its own. It does not however distinguish itself from nearby siblings like get_workflow or get_workflow_summary, so the agent must infer the list-vs-single distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no statement of prerequisites or of how this differs from get_workflow, execute_workflow, or create_workflow. The listing intent is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspace_membersBInspect
List all members of a workspace, including their roles and emails.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the returned fields (roles, emails) but says nothing about permissions required, pagination/result limits, or whether non-members can call it, leaving real gaps for a multi-tenant tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and scope come first and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list with one required parameter and no output schema, the description is adequate about what is returned, but it omits pagination behavior, ordering, and scope/permission notes that an agent would need to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (workspaceId) with 100% schema description coverage, so the schema already documents it. The description adds no format or sourcing guidance beyond that, which matches the baseline 3 for well-covered schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (workspace members) plus the returned fields (roles, emails). It is easy to distinguish from search_workspace, but the description never names or contrasts siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer it is for enumerating workspace membership, but there is no statement of when to prefer it over search_workspace or get_workspace, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesAInspect
List all workspaces accessible to the authenticated user. Use this when the user asks to choose, inspect, or connect a Sequencer workspace, or when a non-generation tool requires workspaceId. For standalone generate_image, generate_video, and generate_audio calls, do not call this just to find a workspace; omit workspaceId and Sequencer will use the default workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of workspaces to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose genuinely useful behavior — that omitting workspaceId falls back to a default workspace for generation tools — but says nothing about pagination, ordering, what happens with zero workspaces, or permission requirements beyond 'authenticated user'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each doing distinct work: purpose, positive trigger, negative trigger plus fallback. The routing information is front-loaded and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description never indicates what a returned workspace contains or which field an agent should pass as workspaceId — a small but real gap for a lookup tool whose output feeds other calls. Everything else an agent needs to decide whether to call it is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'limit' parameter, so the schema already documents it fully. The description adds no format, default, or range information beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'List all workspaces accessible to the authenticated user.' An agent can distinguish this from get_workspace, create_workspace, or search_workspace without opening a schema, and the scoping phrase clarifies whose workspaces are returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger conditions ('user asks to choose, inspect, or connect a Sequencer workspace' and 'a non-generation tool requires workspaceId') and the exclusion ('for standalone generate_image, generate_video, and generate_audio calls, do not call this'), plus the alternative behavior (omit workspaceId and the default is used). This is textbook when/when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_item_to_folderBInspect
Move a media, edit, or asset item to a folder (or to root by passing null folderId).
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ID of the item to move | |
| folderId | No | Target folder ID (null to move to root) | |
| itemType | Yes | Type of item to move | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden for a mutation tool. It discloses only that passing null moves the item to root; it says nothing about permissions, whether the move is reversible, side effects on the source folder, or failure/error behavior. The single disclosed trait duplicates the schema, leaving most behavioral context missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action and the key edge case (null folderId) with zero padding. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is thin: it covers the action and one parameter edge case but omits result expectations, permission requirements, and error conditions. The fully documented input schema partly compensates, making this minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all four parameters documented including the enum and the null-folderId convention, so the baseline is 3. The description restates the item types and the null-to-root behavior but adds no format, syntax, or constraint details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Move a media, edit, or asset item to a folder') and enumerates the item types it operates on. It is clear what the tool does, but it offers no differentiation from the similar move siblings (move_shot_to_scene, move_overlay_item_to_track), which an agent must distinguish by inference from the itemType enum.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites (e.g. needing an existing folder or workspace), and no named alternative among the many move_* siblings. The only usage hint is the parenthetical about null folderId, which is a parameter convention rather than selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_overlay_item_to_trackAInspect
Move an overlay item from one overlay track to another, preserving all timing, transform, animation, and styling data.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| itemId | Yes | Overlay item ID to move | |
| workspaceId | Yes | The workspace ID | |
| sourceTrackId | Yes | Current overlay track ID | |
| targetTrackId | Yes | Destination overlay track ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral load. It does disclose valuable semantics — the move is non-destructive and preserves timing, transform, animation, and styling data. However, it is silent on what happens at the source track, placement/index within the target track, required permissions, and failure modes, which are meaningful gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the operation, the source, the destination, and the key guarantee. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter mutation with no output schema and no annotations, the description is adequate but not complete: it omits positioning within the target track, ordering side effects, auth/permission requirements, and error behavior. The preservation guarantee is the one substantive addition over structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all five required parameters are documented in the schema, so the baseline is 3. The description adds no syntax, format, or constraint detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (move), resource (overlay item), and both endpoints (from one overlay track to another), which cleanly separates it from siblings like update_overlay_item or move_shot_to_scene. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description — this is the tool for relocating an existing overlay item between tracks — but there is no explicit when-to-use guidance, no mention of prerequisites (both tracks must exist), and no routing against alternatives such as update_overlay_item or delete/ re-add workflows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_shot_to_sceneBInspect
Move a shot from one scene to another, optionally setting its order in the target scene.
| Name | Required | Description | Default |
|---|---|---|---|
| order | No | Position in the target scene (appended to end if not specified) | |
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID to move | |
| workspaceId | Yes | The workspace ID | |
| targetSceneId | Yes | The target scene ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It does not state whether this requires specific permissions, whether the move is reversible, what happens to existing shot order in the source scene, or any side effects. The only behavioral hint is that order is optional and defaults to appending, which is minimal for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero waste, covering the core action and optional parameter without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is too sparse. It fails to cover permissions, side effects, source scene behavior, or return information, leaving significant gaps for an agent to call it safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters including the optional order with its default behavior. The description adds no parameter meaning beyond stating that order is optional, which is redundant with the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (move), resource (shot), and scope (from one scene to another), which is unambiguous and distinguishable from siblings like reorder_shots or update_shot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when relocating a shot between scenes, but provides no explicit when-to-use guidance or alternatives like reorder_shots or update_shot. The optional 'order' parameter is mentioned but not contrasted with reorder_shots, leaving the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_paywallAInspect
Open the paywall or subscription upgrade modal. Use this when an action requires a Pro subscription (e.g. generating certain media) and the user is not subscribed. In Sequencer/Spark, the frontend will show the upgrade modal. In external MCP clients, this tool cannot open Sequencer UI; give the user the Sequencer app URL and explain they need to upgrade there.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Brief explanation of why a subscription is needed, e.g., "generate videos", "use the Pro image model", "access premium features" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the crucial client-dependent behavior (frontend shows a modal; external MCP clients cannot open Sequencer UI and must be given the app URL instead). It doesn't cover error/return behavior, but the environment-dependent semantics are the key trait and are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the trigger condition, then the client-specific caveat. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a single-param tool with no output schema: purpose, when-to-use, and the important environment-dependent outcome are all covered. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single 'reason' parameter with examples. The description adds no additional meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Open the paywall or subscription upgrade modal') that is unambiguous and clearly distinct from siblings like get_pricing_info or get_subscription_help. An agent can tell precisely what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the triggering condition ('when an action requires a Pro subscription... and the user is not subscribed') and then branches on client context, telling the agent exactly what to do in external MCP clients versus Sequencer/Spark.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_sequencer_workspaceSequencer StudioBRead-onlyInspect
Open the Sequencer workspace browser to view projects, recent creations, and shot previews. Read-only. Creation starts only after a user chooses a task and approves its plan and cost.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | No | ||
| startingTask | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true/destructiveHint=false, and the description reinforces 'Read-only' while adding a genuinely new behavioral fact: creation is gated on the user choosing a task and approving plan and cost. It stops short of saying what 'opening' actually does to the client state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and scope. The approval-gating sentence earns its place as workflow context, though it could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and an ambiguous 'open' verb, the description does not clarify whether this returns a page link, a workspace object, or just triggers UI navigation. For a 2-param zero-required tool with no annotations covering that gap, a bit more is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 2 parameters, and the description never mentions workspaceId or startingTask. The phrase 'chooses a task' loosely gestures at the startingTask enum, but none of the 11 enum values or the workspaceId format are explained in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open the Sequencer workspace browser') plus what the user sees (projects, recent creations, shot previews). It is clearly distinct from get_workspace/list_workspaces, though it never names a sibling to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence implies this is the entry point before any creation flow, but it does not say when to prefer this over navigate_to_page, get_sequencer_capabilities, or get_workspace, nor any prerequisites. Usage is only inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quote_generationQuote Generation CostARead-onlyIdempotentInspect
Get a read-only USD quote for an image, video, or audio generation before requesting approval. Select an exact model ID from get_model_catalog and specify all billable settings. Reports usage-based estimates and missing inputs explicitly. Creates no media and spends no credits.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| model | Yes | Exact model ID from get_model_catalog. | |
| goFast | No | ||
| duration | No | Output duration in seconds. | |
| frameCount | No | ||
| resolution | No | ||
| aspectRatio | No | ||
| outputWidth | No | ||
| modelVariant | No | Public variant setting from get_model_catalog, when required. | |
| outputHeight | No | ||
| generateAudio | No | ||
| characterCount | No | ||
| outputImageCount | No | ||
| referenceImageCount | No | ||
| referenceVideoDurations | No | Duration of each billable reference or source video, in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so 'creates no media and spends no credits' is largely redundant with the structured safety profile. The one genuinely new behavioral claim is 'Reports usage-based estimates and missing inputs explicitly', which tells the agent the response surfaces gaps — useful, but the rest adds little beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: purpose, prerequisite/routing, and behavioral guarantees, with the read-only cost-estimate framing front-loaded. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, multi-modality estimation tool with no output schema, the description is minimally adequate: it explains intent and safety but leaves the agent guessing which parameters matter for a given media type. It also does not indicate whether unsupported parameters are ignored or rejected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 15 parameters at only 27% schema description coverage, the description should compensate but does not: it only says to 'specify all billable settings' without indicating which of duration/frameCount/resolution/aspectRatio/characterCount/etc. apply to which modality. The model-ID pointer repeats what the schema already documents for the 'model' property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource — 'Get a read-only USD quote for an image, video, or audio generation' — so the agent immediately knows this estimates cost rather than producing media. However, it never differentiates itself from cost-adjacent siblings like get_pricing_info or get_balance, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the context of use ('before requesting approval') and a concrete prerequisite ('Select an exact model ID from get_model_catalog'), which is actionable routing guidance. It does not name any alternative tool or exclusion case, so it isn't a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_image_backgroundremove image backgroundADestructiveInspect
Remove an image background using a catalog model that supports remove_background. The selected model determines transparency support. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered structurally. The description adds modest context: model-dependent transparency support and asynchronous completion ('Returns a tracked task... when complete'). It does not discuss cost, credit consumption, or reversibility, which is a real gap given a paid destructive generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action before the model-selection and return-behavior details. No filler, though the final sentence restates what the output schema already conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter async image-generation-adjacent tool with an output schema and full annotations, the description covers the essentials: what it does, model dependency, and task-based return. It is slightly thin on cost/authorization context given destructiveHint=true and the paid-generation nature, but nothing critical for invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (81%) across 21 params, so the schema already documents most arguments. The description only reiterates the model/get_model_catalog relationship already in the schema and adds the transparency-support nuance, which is marginal added meaning. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Remove an image background') and immediately scopes it to catalog models that support remove_background. This clearly distinguishes it from the sibling remove_video_background and remove_video_background_media tools, so an agent can select correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one actionable routing instruction ('Use get_model_catalog to choose a supported model and its settings'), which is genuinely useful for picking the required model param. However, it offers no when-to-use/when-not guidance relative to the many sibling image/video editing tools, and no prerequisites beyond the model lookup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_video_backgroundremove video backgroundADestructiveInspect
Remove a video background when an active catalog model supports remove_background and video input. Select an alpha-capable output format in modelSettings when supported. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context beyond annotations: model capability gating, an instruction to select an alpha-capable output format, and the fact that the call returns a tracked task and playable result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the operation and condition, then prerequisites, then return behavior. Each sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 21-parameter tool with nested objects and an output schema, the description orients the agent on model dependency, output format selection, and async return behavior. It omits sibling differentiation and does not mention idempotency or cost controls, but those are largely covered by schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 81%, so the schema already documents most parameters. The description adds a small amount of meaning for outputFormat and modelSettings by noting alpha-capable output, but it does not provide syntax or constraints beyond the schema for the 21 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Remove a video background') and adds a clear capability condition ('when an active catalog model supports remove_background and video input'). However, it does not distinguish this tool from the sibling remove_video_background_media, so an agent must infer the difference from names alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use: the selected model must support remove_background and video input, and get_model_catalog should be used to choose a supported model and settings. It stops short of explicitly stating when to prefer this over remove_video_background_media or remove_image_background.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_video_background_mediaBInspect
Remove or matte the background of a workspace video media item. Defaults to robust-video-matting and supports green-screen, alpha-mask, and foreground-mask outputs.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | Video MediaDocV2 ID to process | |
| modelId | No | Background removal model ID. Defaults to robust-video-matting. | robust-video-matting |
| videoUrl | No | Override source video URL; otherwise resolves from the active media version | |
| outputType | No | Replicate robust_video_matting output_type | green-screen |
| extraParams | No | Additional raw V3 model parameters | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose some behavior: the default matting model and the three output modes. It omits important traits for a heavy video job, such as whether the operation is asynchronous, how long it takes, cost, required permissions, or whether it mutates the source media.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and configuration detail second. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, nested-object, no-annotation, no-output-schema mutation tool, the description covers the purpose and output modes but says nothing about return values, async behavior, permissions, or the extraParams escape hatch. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, including mediaId, modelId, videoUrl, outputType, and extraParams. The description only restates the model default and output types at a high level, adding little beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('remove or matte the background of a workspace video media item') plus the supported output flavors, so an agent can tell it is video-background removal rather than image removal. It does not, however, distinguish itself from the sibling remove_video_background, leaving ambiguity about which of the two to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the sibling remove_video_background or remove_image_background, no prerequisites, and no exclusions. The mention of defaults and output types is configuration, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_imageShow Sequencer imageARead-onlyIdempotentInspect
Use this after get_media reports that a Sequencer image is completed. It displays completed image media inline in ChatGPT.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | Completed image media ID returned by generate_image or get_media | |
| workspaceId | Yes | Workspace containing the completed image |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds a genuinely useful behavioral fact: the image is shown inline in ChatGPT rather than returned as a URL. It does not discuss permissions or failure modes, but those are covered adequately by the annotations and the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, the usage precondition front-loaded ahead of the effect, with no filler. It is tight and readable, though marginally terse given the omission of any sibling disambiguation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations declaring safety, the description only needs to convey the precondition and the rendering effect, both of which it does. The remaining gap is the lack of explicit contrast with render_media, a minor omission for a completion-gated display tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both required parameters (mediaId, workspaceId) are already documented in the schema, and the description adds no syntax or format detail beyond the 'completed image' qualification. Baseline 3 is appropriate when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('displays') and resource ('completed image media') with the scope narrowed to Sequencer images, so the agent knows it is a rendering operation. It does not explicitly differentiate itself from the sibling render_media, which is the main residual ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition: use it only after get_media reports the Sequencer image is completed. It stops short of naming an alternative tool (e.g. render_media) or stating the 'when not to use' case, so it is clear context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_mediaView MediaARead-onlyIdempotentInspect
Display one or several Sequencer image, video, and audio items in one inline carousel. The player monitors queued and generating media until every item is ready.
| Name | Required | Description | Default |
|---|---|---|---|
| mediaId | Yes | First media ID, or a comma-separated list of media IDs for clients using the earlier single-field schema | |
| workspaceId | Yes | Workspace containing the media | |
| additionalMediaIds | No | Up to seven additional media IDs to show in the same carousel |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, closed-world and non-destructive behavior, so the safety profile is covered. The description adds genuinely new behavior: the player continues monitoring queued and generating media until every item is ready, which tells the agent this call waits on readiness rather than failing fast.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action and its multi-item carousel scope front-loaded before the readiness behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and annotations cover safety. Purpose plus readiness semantics are complete enough to invoke correctly; only edge behavior (e.g., what happens if an item never becomes ready) is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema fields even document the legacy comma-separated mediaId form and the max of seven additional IDs. The description adds no parameter-level detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (display), the resource (Sequencer image/video/audio items), and the presentation mode (one inline carousel). This clearly separates it from the sibling render_image, which serves a single image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'One or several ... in one inline carousel' implies when to pick this over single-item rendering, and the readiness-monitoring sentence hints at use with pending media. However, it never states an explicit when-not condition or names an alternative tool, leaving usage inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reorder_shotsBInspect
Bulk reorder shots within an edit. Efficiently updates all shot order values in a single batch write.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| shotOrders | Yes | Array of shot ID and new order pairs | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses useful mutation semantics ('single batch write', updating all order values at once), but says nothing about permissions, atomicity, partial-failure behavior, or what happens to shots omitted from the list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences that convey the operation and its batch nature without padding. 'Efficiently' is mild promotional filler, keeping it just below a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch mutation tool with no annotations and no output schema, the description is minimally adequate: it explains the operation but omits return shape, error/conflict handling, and whether a complete shot list is required. Nothing critical is stated incorrectly, but meaningful gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters and the nested shotId/order pairs are already documented in the schema. The description adds only that the array represents 'all' order values, which is marginal beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('reorder shots') plus scope ('within an edit', 'bulk'), which is more precise than the generic siblings like update_shot or move_shot_to_scene. It does not explicitly name those alternatives, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Bulk' and 'all shot order values in a single batch' imply the tool is for multi-shot reordering rather than single updates, which hints at when to use it. However, it names no alternatives (e.g. update_shot, move_shot_to_scene) and states no conditions or exclusions, leaving the routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_asset_reviewAInspect
Request the user to review and approve/reject specific assets. For external MCP clients, returns inlineReview assets with image URLs/markdown that should be shown directly in chat; reviewUrl is only a secondary Sequencer link. In Spark, the in-app review modal may still open. Recommended after the initial character/location pass in a proposal build before expensive shot generation; optional for later incremental edits or explicit user-directed asset creation.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| assetIds | Yes | Array of asset IDs to review | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses the return shape (inlineReview assets with image URLs/markdown to render in chat, reviewUrl as a secondary Sequencer link) and the client-dependent behavior (Spark may open an in-app modal). It omits whether the call blocks or is async and any state/permission implications, but the behavioral disclosure is well above average for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then behavior, then usage timing in four dense sentences with essentially no filler. It is slightly long, but each clause (return format, secondary link, modal behavior, workflow timing) carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with 100% schema coverage and no output schema, the description supplies the missing behavioral and workflow context: what is returned, how to surface it, and when to invoke it. Only the sync/async nature of the request and any follow-up expectations are left unstated, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so workspaceId, editId, and assetIds are already documented in the schema. The description adds no format, constraint, or ordering detail beyond that, which matches the baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Request the user to review and approve/reject specific assets'), making the action unambiguous. It implicitly separates itself from sibling tools like set_asset_approval_status (agent-sets status) and review_edit, but never names an alternative explicitly, so it falls just short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance ('Recommended after the initial character/location pass in a proposal build before expensive shot generation') and a when-optional condition ('optional for later incremental edits or explicit user-directed asset creation'). Clear context is provided, but no alternative tool is named for the case where the agent should do something else.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_loginAInspect
Request the user to log in or register to continue. Use this when an action requires authentication and the user is not logged in. In Sequencer/Spark, the frontend will show a login modal and resume after login. In external MCP clients such as Codex, Claude Desktop, Claude Code, or Antigravity, this tool cannot open a modal; tell the user to reconnect or re-authorize the Sequencer MCP server instead, and do not claim a modal opened.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Brief explanation of why login is needed, e.g., "create a project", "save your workflow", "access your media" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses that behavior differs by client, that a modal appears and execution resumes in Sequencer/Spark, that no modal can open in external MCP clients, and explicitly warns not to claim a modal opened. This is exactly the kind of environment-dependent side-effect disclosure annotations would not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and trigger before the environment-specific caveat, and every sentence earns its place. The third sentence is dense but necessary to prevent an incorrect claim of a modal opening.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, non-mutating gating tool with no output schema, the description covers purpose, trigger, environment-dependent behavior, and the correct fallback action. Nothing an agent needs to call and act on this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the 'reason' parameter is already documented with examples in the schema, so the description need not add parameter detail. It provides no additional meaning beyond what the schema already states, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('request the user to log in or register') and clearly identifies this as an authentication-gating tool, distinct from the many content-mutation siblings. An agent can tell what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition ('when an action requires authentication and the user is not logged in') and then branches on environment, naming the client-specific alternatives (Sequencer/Spark vs. Codex/Claude Desktop/Claude Code/Antigravity). This is unambiguous when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_editAInspect
Export and analyze the current edit, then return a structured video review report. Prefer this over watch_video when the user asks to watch, review, analyze, or improve a film. The backend handles video analysis and returns provider-agnostic JSON with issues and recommended edit actions.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | quick returns the highest-impact issues; deep returns a more thorough report | quick |
| focus | No | Optional review focus | all |
| editId | Yes | The edit/project ID to export and review | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure. It usefully states that the backend performs the analysis and returns provider-agnostic JSON containing issues and recommended edit actions, which tells the agent the output shape and that analysis is server-side. However, it omits operationally important traits for a video-analysis call: expected latency/async behavior, cost implications of exporting and analyzing, and whether it mutates or locks the edit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it does, when to prefer it, and what comes back. Front-loaded with the action and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by describing the return as a structured report with issues and recommended actions. It is nearly complete for a two-required-param tool; only the async/cost dimension of a heavy video-analysis operation is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with mode and focus already documented via enums and descriptions, so the baseline is 3. The description adds no extra meaning about how mode or focus change the review, leaving the schema to do all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific compound action — export and analyze the current edit — and names the deliverable (structured video review report). The sibling distinction from watch_video is made explicit, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('user asks to watch, review, analyze, or improve a film') and names the competing tool (watch_video) with a 'prefer this over' directive. This is the strongest form of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_production_templateADestructiveInspect
Generate a saved production template into a new editable Sequencer project. Each run spends credits and has tracked output tasks. Poll get_production_status. To rerun, call again; previous projects remain preserved.
| Name | Required | Description | Default |
|---|---|---|---|
| version | No | ||
| parameters | No | ||
| templateId | Yes | ||
| workspaceId | Yes | ||
| referenceImages | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true and idempotent=false, and the description reinforces this with concrete consequences: each run spends credits, produces tracked output tasks, and does not overwrite prior projects. That credit-spend and preservation context is genuinely additive beyond the annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action, then credit/task behavior, then follow-up polling, then rerun semantics. Every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and 0% schema description coverage on 5 parameters including a nested object, the description covers lifecycle and cost behavior well but leaves parameter usage entirely undocumented, which is a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, including a nested referenceImages object with url/label/mediaId, yet the description mentions none of them. It does not compensate for the coverage gap by explaining version, parameters, or reference image semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate a saved production template into a new editable Sequencer project') and clearly distinguishes itself from siblings like get_production_template, save_production_template, and validate_production_template by describing the instantiation/execution behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable operational guidance: poll get_production_status after running, and call again to rerun while prior projects remain preserved. It does not explicitly state prerequisites such as needing a templateId from list_production_templates, but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_video_workflowCInspect
Run a Sequencer workflow API endpoint from the video editor context. Inputs can include video/audio/image URLs and arbitrary scalar fields.
| Name | Required | Description | Default |
|---|---|---|---|
| inputs | Yes | Workflow API inputs keyed by input label and/or node ID | |
| workflowId | Yes | Workflow ID with API access enabled | |
| workspaceId | Yes | Workspace used for permission checking |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It never states that this is a mutation that kicks off a potentially long-running/expensive render, whether execution is synchronous or async, what auth or permission model applies, or what happens to the workspace on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and no filler. Slightly under-specified rather than over-long, but structurally clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a workflow-invocation tool with no annotations, no output schema, and a free-form nested inputs object, the description omits what the call returns (job ID, outputs, artifacts), async/latency behavior, and the relationship to the sibling execute_workflow. This is a material gap for an agent needing to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a little context about inputs containing URLs and arbitrary scalar fields, but this is largely redundant with the schema's "keyed by input label and/or node ID" note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ("Run") and a specific resource (a Sequencer workflow API endpoint) and scopes it to the video editor context. However, it does not distinguish itself from the very similar sibling execute_workflow, leaving the agent to guess which execution tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus execute_workflow, run_production_template, or other run/execute siblings. There are no prerequisites, no exclusions, and no indication of the required API-access workflow precondition beyond the schema field description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_experience_gameAInspect
Build ONLY a game whose current proposal was approved by the user. Read get_experience_game, follow its approved proposal, and pass proposalRevision and expectedVersionId. Supply organized files (index.html, styles.css, scripts/*.js, optional JSON/docs) with local stylesheet/classic script references, or one complete HTML source. Use ordered classic scripts and shared namespaces, no imports/requires. The server bundles code to a single playable HTML under 256 KB and stores the editable package separately (up to 64 files / 1 MB). Media belongs in existing asset:// IDs, canvas or inline art. Include responsive start, pause, help, score/progress, win/loss and restart UI plus keyboard/touch controls. External URLs, network APIs, workers, embeds, eval, and parent access are blocked. Match the approved worldId. Returns the Games tab playtest link. Never paste implementation code in the chat.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| title | Yes | ||
| gameId | Yes | ||
| prompt | No | ||
| source | No | ||
| worldId | No | ||
| entryPoint | No | ||
| description | Yes | ||
| workspaceId | Yes | ||
| proposalRevision | Yes | ||
| expectedVersionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations beyond a generic write profile, the description carries the full burden and delivers rich behavior: server-side bundling to a single playable HTML under 256 KB, editable package storage limits (64 files / 1 MB), allowed media sources (asset:// IDs, canvas, inline art), a hard sandbox (no external URLs, network APIs, workers, embeds, eval, parent access), code constraints (classic scripts, shared namespaces, no imports/requires), required UI surface (start, pause, help, score, win/loss, restart, keyboard/touch), worldId matching, and the return value (Games tab playtest link). This is far beyond what any annotation provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded imperative constraint ('Build ONLY a game whose current proposal was approved'), then constraint-dense sentences with no filler. The single block packs many independent rules (bundling, sandbox, UI, file limits), which slightly hurts scanability but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with 11 params, 0% schema coverage, no output schema, and no explanatory annotations, the description covers prerequisites, input alternatives, constraints, sandbox boundaries, UI requirements, and the return value. An agent has what it needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains proposalRevision and expectedVersionId usage, the files vs source alternative, entryPoint-implied structure, and worldId matching, going meaningfully beyond the raw schema. It does not explain workspaceId/gameId/prompt/description semantics, so it is not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (build/save a game) with a clear precondition ('current proposal was approved') and explicitly distinguishes itself from the sibling get_experience_game by telling the agent to read it first. An agent can immediately tell this tool persists/builds a game while save_experience_game_proposal is the proposal step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('ONLY a game whose current proposal was approved by the user'), a required prerequisite flow (read get_experience_game, follow its approved proposal, pass proposalRevision/expectedVersionId), and an explicit prohibition ('Never paste implementation code in the chat'). Routing to the prerequisite sibling is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_experience_game_proposalAInspect
FIRST step for creating or revising a game. Save a proposal document and a visible Games tab entry, then STOP for human approval. Markdown must have headings starting Concept, Gameplay, Controls, Screens, Files, and Playtest. Describe audience, objective/core loop/levels/win/loss, desktop/touch controls, visual/audio direction, start/pause/help/score/game-over/restart UI, code organization and assets, and concrete checks. Do not include game code. For an existing game read it first and pass expectedVersionId and expectedProposalRevision. A revised proposal always needs fresh approval.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| gameId | No | ||
| prompt | No | ||
| worldId | No | ||
| description | Yes | ||
| workspaceId | Yes | ||
| proposalMarkdown | Yes | ||
| expectedVersionId | No | ||
| expectedProposalRevision | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a mutation (readOnlyHint=false) that is neither idempotent nor destructive. The description adds real value beyond that: it creates a visible Games tab entry, requires stopping for human approval, forbids embedding game code, and specifies the optimistic-concurrency parameters for revisions. It does not cover permissions or failure behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the essential 'first step' framing and the stop-for-approval instruction. It is dense and the markdown content requirements form a long clause, but each sentence carries actionable information (required headings, concurrency fields, no code, fresh approval).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no output schema, the description supplies the workflow, the artifact produced, the approval gate, and the concurrency contract an agent needs. It is let down only by the unexplained ancillary parameters (prompt, worldId, workspaceId), which no other structured field documents.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must carry the load. It explains proposalMarkdown in depth (required headings, content, no code) and mentions expectedVersionId/expectedProposalRevision, and implies gameId for revisions, but leaves workspaceId, title, description, prompt, and worldId entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (save an experience game proposal) and pins down its position in the workflow as the FIRST step for creating or revising a game. It clearly distinguishes itself from siblings like save_experience_game, save_proposal, and update_proposal by describing the proposal+Games-tab artifact and the approval gate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to STOP for human approval and that revisions require reading the existing game first and passing expectedVersionId/expectedProposalRevision, plus that a revised proposal always needs fresh approval. This gives strong when-to-use context, though it never names the sibling tools (get_experience_game, save_experience_game, update_proposal) an agent should consider instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_production_templateBDestructiveInspect
Save a versioned production template for product ads, 3 to 5 shot stories, or image localization. Templates use {{parameter}} placeholders and produce editable projects when run.
| Name | Required | Description | Default |
|---|---|---|---|
| definition | Yes | ||
| templateId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false, and openWorldHint=false, so the safety profile is covered. The description usefully adds that templates are versioned and that running them produces editable projects, but it omits the destructive implication of supplying templateId (overwriting/versioning an existing template), which is exactly where the destructiveHint matters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded with the essential fact (what is saved, what it contains, what it yields). No filler, no restatement of the tool name beyond what is needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A destructive mutation tool with three parameters, a nested definition object, no output schema, and zero schema description coverage needs more than two sentences. The description leaves the create-vs-update semantics, the response shape, and most definition fields unexplained, so the agent cannot call this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, yet it only partially compensates: it explains the {{parameter}} placeholder convention used in shot prompts and roughly what belongs in the definition. It says nothing about workspaceId, about templateId being the update path, or about resolution/aspectRatio/type, and its '3 to 5 shot stories' framing conflicts with the schema's 1–8 shot bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save a versioned production template') and enumerates the supported kinds (product ads, short stories, image localization), which maps cleanly onto the definition.kind enum. It does not, however, distinguish itself from the close siblings validate_production_template, run_production_template, or get_production_template, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no mention of alternatives or sequencing (e.g., validating a template first, or getting a template before re-saving it). It also never explains the create-vs-update branch implied by the optional templateId, which is the main decision an agent faces here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_proposalAInspect
Save a proposal document for user review before implementing a complex project. The proposal MUST include a Models section listing the friendly names of the planned image/still and video models, a "💰 Estimated Cost" section, and a concrete numbered scene/shot or beat list. Default project models are Nano Banana Pro (google-gemini-3-image) for character/location/shot still images and Gemini Omni 1.1 Flash (google-gemini-omni-1-1) for videos after still-image approval. Keep those raw IDs in tool arguments and use friendly model names in user-facing proposal text. Do not use source-file references such as "from script.md" or vague language such as "approximately X shots" instead of enumerating the planned work. Display the returned proposalMarkdown inline in chat and wait for the user to use the inline Accept or Decline controls. The reviewUrl is secondary and must never be opened automatically. Do not respond with only a link or a relative URL.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID to save the proposal to | |
| threadId | No | Optional chat thread ID | |
| sourceMedia | No | Uploaded media from the original Spark request that must remain available after proposal approval | |
| workspaceId | Yes | The workspace ID | |
| proposalMarkdown | Yes | The full proposal content in Markdown format, including concrete numbered scenes/shots or beats. If based on a script or beat sheet, extract and enumerate the planned shots; do not merely reference the file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior beyond the schema: the proposal is saved for user review, the returned proposalMarkdown must be shown inline, the agent must wait for Accept/Decline, and reviewUrl must never be auto-opened. It omits permissions, failure modes, and what happens on decline, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then requirements, then display behavior, in a logical order. It is long and somewhat dense with formatting rules, but nearly every sentence carries actionable guidance, with only mild redundancy around the model-ID versus friendly-name rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by explaining what is returned (proposalMarkdown, reviewUrl) and how to handle it, plus the full content contract. What remains thin is edge-case behavior (errors, decline path), which is reasonable to omit here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description genuinely extends proposalMarkdown semantics — mandatory sections, enumerated shots rather than file references, and the rule to keep raw model IDs in arguments while using friendly names in user text. That is real added meaning beyond the schema string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Save) and resource (a proposal document) and adds scope ("for user review before implementing a complex project"). It does not explicitly differentiate itself from close siblings like update_proposal, get_proposal, set_proposal_status, or save_experience_game_proposal, so an agent must infer the boundary from context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context — call this before implementing a complex project — and spells out required proposal content (Models section, cost section, numbered scene/beat list). It stops short of naming alternatives or stating when not to use it (e.g., versus update_proposal), so it lacks explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_urlAInspect
Read a public HTML web page and return its visible text, metadata, and direct image URLs. Use this only when the user supplies a webpage URL whose content or images must inform a Sequencer project. Do not use it for media files, JSON, waveform URLs, or internal generation output. Use get_media and listen_to_audio for generated audio. This does not render client-side JavaScript and does not import or generate media.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public HTTP(S) page URL to inspect | |
| imageLimit | No | Maximum number of page image URLs to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose key traits: it does not render client-side JavaScript and does not import or generate media. It does not mention rate limits, truncation of long pages, or error behavior, so a small gap remains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what it does, then usage conditions, then exclusions and alternatives. Every sentence carries distinct routing or behavioral information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only fetch with no output schema, describing the returned fields (visible text, metadata, image URLs) plus the JS-rendering limitation is nearly complete; only limits on returned text size or pagination are unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both url and imageLimit are fully documented in the schema. The description only implicitly references image URLs and adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read/scrape) and resource (public HTML web page) plus the payload it returns: visible text, metadata, and direct image URLs. This clearly separates it from media-fetching and generation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (user supplies a webpage URL whose content or images must inform a Sequencer project) and when not to (media files, JSON, waveform URLs, internal generation output), naming two alternatives, get_media and listen_to_audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_workspaceAInspect
Search existing workspace media by prompt, file name, description, and tags. This searches metadata, not image pixels. Use a concise subject keyword such as "cat". Follow nextCursor until hasMore is false when the user asks for all matches. Never create an edit while searching.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Optional media type | |
| limit | No | ||
| query | Yes | Subject or metadata terms to match | |
| cursor | No | nextCursor from the previous search page | |
| workspaceId | Yes | Workspace to search |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it discloses meaningful traits: metadata-only matching, pagination semantics (follow nextCursor until hasMore is false), and a safety-style prohibition against creating an edit while searching. Auth requirements and rate limits remain unstated, but the key behavioral profile is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then adds scope, query style, pagination, and a caution in compact sentences. Each sentence earns its place, though the closing 'Never create an edit while searching' reads as a somewhat cryptic aside.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains the relevant return behavior (hasMore/nextCursor) so pagination is clear. For a 5-param search tool with no annotations it covers purpose, scope, query style, and pagination adequately; only auth/permission context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the baseline is 3; the description adds practical meaning by advising a concise subject keyword for 'query' and explaining the nextCursor/hasMore pagination contract. This goes slightly beyond the schema's terser 'Subject or metadata terms to match'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search existing workspace media') plus the exact metadata fields it matches against (prompt, file name, description, tags). The line 'This searches metadata, not image pixels' usefully scopes it apart from any visual/semantic search, though no sibling tool (e.g. list_media) is named to anchor the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides query-formulation advice ('Use a concise subject keyword such as "cat"') and pagination guidance, plus a caution against creating an edit during search. However, it never states when to choose this over list_media/get_media or other retrieval siblings, so tool-selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_asset_approval_statusAInspect
Approve, reject, or reset one or more assets from chat. External MCP clients should call this after showing inline asset images/descriptions and receiving user approval or revision feedback, instead of requiring the user to open Sequencer.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional approval comment or rejection/edit note from the user | |
| editId | No | Optional edit ID. When provided, the response includes aggregate approval status for that edit. | |
| status | Yes | New approval status | |
| assetIds | Yes | Asset IDs to update | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the intended call context and that it operates on a batch of assets, but says nothing about required permissions, whether status changes are reversible, what 'reset' (pending) implies for downstream review state, or how note/editId affect the outcome.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action verb and resource, with no padding. The second sentence is longer than needed but still earns its place by explaining the intended agent workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations and no output schema, so the description must cover behavior; it covers the workflow trigger well but omits permission requirements, side effects on review state, and any mention of the aggregate status returned when editId is supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including the status enum. The description adds only loose mapping ('approve, reject, or reset' → approved/rejected/pending) and never mentions editId or its aggregate-status behavior, so it does not exceed baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action set (approve, reject, reset) on a specific resource (assets) with batch scope (one or more). It is distinguishable from neighbors like request_asset_review and check_asset_status, though it does not explicitly name them to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear calling condition: invoke after presenting inline asset images/descriptions and receiving user approval or revision feedback, and as a substitute for sending the user into Sequencer. No explicit when-not conditions or named alternative tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_proposal_statusAInspect
Approve, reject, or reset a saved proposal. Use this in external MCP clients when the user approves or rejects a proposal in chat, so the Sequencer edit state stays in sync without relying on a Spark-only UI modal.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional short note about why the proposal was approved, rejected, or reset | |
| editId | Yes | The edit ID | |
| status | Yes | New proposal status | |
| proposalId | No | Specific proposal ID. If omitted, the active proposal for the edit is used. | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the sync intent and the external-client/chat trigger, but omits what happens to the proposal after the call, whether the change is reversible, permission requirements, and how 'reset' maps to the 'pending' enum value. Adequate but with clear behavioral gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action before the usage context. Efficient with no filler, though the second sentence packs several ideas (external clients, chat trigger, sync, modal avoidance) into one long clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers purpose and trigger but leaves out prerequisites, side effects, and return behavior. The 100%-covered schema handles parameters, so the remaining gap is behavioral rather than structural.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including the status enum and the optional proposalId fallback. The description's approve/reject/reset wording loosely maps to the enum but adds no syntax or format detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (approve, reject, reset) and the resource (a saved proposal), so the operation is unambiguous. It does not explicitly distinguish itself from close siblings like update_proposal or save_proposal, but the status-setting scope is clear from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage context: external MCP clients, when the user approves or rejects a proposal in chat, to keep Sequencer edit state in sync without the Spark-only modal. It stops short of naming the alternative tool to use when that condition doesn't hold, so it's clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
split_shotAInspect
Split a shot into two adjacent shot records at a time offset from the shot start, preserving media refs and playback metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| shotId | Yes | Shot ID to split | |
| splitTime | Yes | Seconds from shot start where the split occurs | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that media refs and playback metadata are preserved across the split, but it is silent on mutation-safety essentials: whether the original shot is mutated in place or replaced by two new records, whether the operation is reversible, what IDs result, and whether any permission or lock is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that identifies the action, the result, and the key constraint (offset from shot start) with no filler. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the definition is adequate but incomplete: it does not describe the shape of the result (e.g., the two resulting shot IDs or which one retains the original ID), which the agent needs since no output schema exists to fill that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including splitTime as seconds from shot start. The description's 'time offset from the shot start' merely restates that definition without adding format, units, or boundary semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Split a shot into two adjacent shot records') plus the mechanism ('at a time offset from the shot start'), which clearly distinguishes it from siblings like duplicate_shot, create_shot, or update_shot. It never explicitly names a sibling or scoping condition, so it lands just short of the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool for dividing an existing shot at a timestamp, but there is no when-to-use vs. when-not guidance and no mention of alternatives such as duplicate_shot or reorder_shots. Nothing is misleading, but nothing routes the agent either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trigger_auto_layoutAInspect
Trigger automatic layout of nodes in the workflow. This arranges nodes left-to-right based on data flow. Call this after adding all nodes and connections.
| Name | Required | Description | Default |
|---|---|---|---|
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully describes the layout algorithm ('left-to-right based on data flow'), but does not disclose side effects such as whether existing manually-positioned nodes are overwritten, whether the operation is reversible, or any permission requirements. Adequate behavioral context, but significant gaps for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the action, the mechanism, and the timing. The core purpose is front-loaded with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutating tool with no annotations and no output schema, the description covers what it does and when to call it, which is the essential minimum. It omits destructiveness/reversibility of layout changes and says nothing about the undocumented workflowId parameter, leaving notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (workspaceId is documented, workflowId is not), and the description adds no parameter-level meaning beyond the implicit mention of 'the workflow'. It fails to compensate for the undocumented workflowId parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Trigger) and resource (automatic layout of nodes in the workflow), and further clarifies the effect: nodes arranged left-to-right based on data flow. This is clearly distinguishable from siblings like add_node, connect_nodes, or update_node without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance: 'Call this after adding all nodes and connections.' That is a clear usage condition. It does not name alternatives, but for an auto-layout operation there is no competing sibling, so the lack of exclusions is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_assetCInspect
Update an existing asset.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name | |
| assetId | Yes | The asset ID | |
| description | No | New description. For style assets, this must stay one short sentence about the global aesthetic shared by every shot. | |
| workspaceId | Yes | The workspace ID | |
| primaryMediaId | No | New primary media ID | |
| voiceDescription | No | For character assets: update the voice/accent/personality description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether this is a partial or full replace, what happens to omitted fields, whether permissions or asset state (e.g. approved assets) constrain edits, or whether the change is reversible. The schema hints at special rules for style and character assets, but the description surfaces none of that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and contains zero filler, but its brevity reflects under-specification rather than disciplined concision for a six-parameter mutation tool. It reads more like a title than a usable definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and asset-type-specific rules embedded in the schema, the description is too thin to be complete. It does not say which asset types are supported, what the update semantics are, or what the caller gets back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the property descriptions already explain each field, including the notable constraints for description on style assets and voiceDescription on character assets. The description adds nothing beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb and resource ("Update an existing asset"), so the basic operation is unambiguous. However, it gives no differentiation from the many other update_* siblings (update_shot, update_media_tags, update_scene, update_workspace), and it does not indicate what kinds of assets exist or what aspects can be updated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no pointer to alternatives such as get_asset or set_asset_approval_status. The agent must infer entirely from the name that this is the correct tool for modifying an asset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_audio_trackCInspect
Update an audio timeline track: placement, trim, fades, volume, keyframes, channel, and shot pinning.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| volume | No | Base volume from 0 to 1 | |
| channel | No | Audio lane/channel index | |
| mediaId | No | Replace with a different audio MediaDocV2 ID | |
| trackId | Yes | Audio track ID | |
| trimEnd | No | Source trim end in seconds | |
| duration | No | Timeline clip duration in seconds | |
| startTime | No | Timeline start time in seconds | |
| trimStart | No | Source trim start in seconds | |
| fadeInTime | No | Fade-in length in seconds | |
| autoChannel | No | When timing changes, move to first non-overlapping channel | |
| fadeOutTime | No | Fade-out length in seconds | |
| workspaceId | Yes | The workspace ID | |
| pinnedToShotId | No | Shot ID to pin this audio to, or null to unpin | |
| volumeKeyframes | No | Volume automation curve | |
| relativeStartTime | No | Offset in seconds from the pinned shot start |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Update' signals mutation, but the description does not explain partial-update semantics, required permissions, whether omitted fields are preserved, how volumeKeyframes interact with volume, or any side effects. It mainly restates editable categories already visible in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is a single front-loaded sentence with no filler or repetition. Every phrase maps to a meaningful editable aspect of the audio track.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter mutation tool with no annotations and no output schema, the one-sentence description is not complete enough. It omits required context behavior, partial update semantics, defaults, and the effect of operations like keyframe replacement or shot pinning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter already has its own description. The tool description adds only a grouped list of editable categories and no additional syntax, defaults, or constraints beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Update') and resource ('audio timeline track'), and enumerates editable aspects such as placement, trim, fades, volume, and keyframes. It is clearly distinct from add_audio_track and delete_audio_track, though it does not explicitly differentiate itself from update_overlay_track or other update siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not guidance, no prerequisites, and no mention of alternatives such as add_audio_track or update_overlay_track. Usage is only implied by the verb and resource.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_editAInspect
Update an existing edit (project). Supports partial updates — only provided fields are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | One short sentence describing only the visual aesthetic shared by every shot (e.g., "cinematic noir, dramatic lighting, subtle film grain"). Do not include characters, locations, props, actions, or shot-specific details. | |
| title | No | New title for the edit | |
| editId | Yes | The edit ID to update | |
| videoModel | No | Default model ID for shot video generations | |
| aspectRatio | No | Default aspect ratio for new media | |
| speechModel | No | Default model ID for speech generation. Defaults to Cartesia Sonic 3.5 (cartesia-sonic-3-5) for new edits. | |
| transitions | No | Complete replacement list of timeline transition boundary objects. Use fromShotId/toShotId for robust adjacency-based transitions. | |
| workspaceId | Yes | The workspace ID | |
| styleAssetIds | No | Legacy array of explicitly supplied style-reference asset IDs. Routine projects use the edit.style text and should leave this unset. | |
| shotImageModel | No | Default model ID for shot still/image generations | |
| sceneImageModel | No | Default model ID for scene/location image generations | |
| defaultImageModel | No | Default model ID for new image generations. Also updates shotImageModel, sceneImageModel, and characterImageModel unless those are provided separately. Defaults to google-gemini-3-image for new edits. | |
| defaultVideoModel | No | Default model ID for new video generations. Also updates videoModel unless videoModel is provided separately. Defaults to google-gemini-omni-1-1 for new edits. | |
| characterImageModel | No | Default model ID for character image generations |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the genuinely useful partial-update contract (unmentioned fields are untouched), but says nothing about permissions, side effects of cascading model fields, or the fact that `transitions` is a destructive full replacement rather than a merge.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the partial-update behavior second. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter mutation tool with no annotations and no output schema, the description is thin: it omits any note on the cascading defaults (defaultImageModel/defaultVideoModel), the complete-replacement semantics of `transitions`, or auth expectations. The rich schema compensates for much of this, but the description itself leaves behavioral gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 14 parameters, so the schema already documents each field's meaning and defaults. The description adds only the generic partial-update rule and no per-parameter semantics beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Update an existing edit') and disambiguates the jargon with '(project)', which separates it from update_shot, update_scene, update_transition, and update_workspace. It does not explicitly name any sibling it is not, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The partial-update statement implies you can call it to patch a single field, which is useful usage context. However, there is no guidance on when to use this versus update_shot/update_scene, nor any prerequisite or caution about the whole-list replacement behavior of `transitions`.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_media_tagsAInspect
Add, remove, or replace searchable tags on an existing workspace media item. Use only when the user asks to tag or retag media. This updates metadata; it does not tag an asset into an edit shot.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | add | |
| tags | Yes | Tags to add, remove, or replace | |
| mediaId | Yes | Exact media item ID | |
| workspaceId | Yes | Workspace containing the media |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It usefully states that this is a metadata mutation and clarifies the scope boundary, but it never warns that mode="replace" discards existing tags (destructive semantics) or mentions permission requirements and tag-count limits. For an unannotated mutation tool, that is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scope, followed by the usage trigger and the disambiguation clause. No filler and nothing repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description still covers purpose, usage trigger, and the critical scope boundary, which is most of what an agent needs. It is only short of the replace-mode semantics and failure/permission behavior expected of an unannotated mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the enum mode parameter is undocumented in the schema, but the description's "Add, remove, or replace" maps directly onto the three enum values, supplying the meaning the schema omits. The tags/mediaId/workspaceId parameters are already described in the schema, so no further compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb set (add, remove, replace), a precise resource (searchable tags), and a precise target (an existing workspace media item). It also distinguishes itself from the asset/shot-tagging siblings by explicitly stating it does not tag an asset into an edit shot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear positive trigger ("Use only when the user asks to tag or retag media") and a when-not condition (not for tagging into an edit shot), which is strong guidance. It stops short of naming the actual alternative tool (e.g. update_shot_media_ref / add_shot_media_ref) the agent should route to instead, leaving that inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_nodeCInspect
Update a node configuration.
| Name | Required | Description | Default |
|---|---|---|---|
| config | Yes | Updated node configuration | |
| nodeId | Yes | ||
| workflowId | Yes | ||
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether the update merges or replaces the existing config, not what permissions are needed, not whether changes are reversible or destructive, and not what the call returns. For a mutation on a nested config object this is a critical omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and contains no filler, which is good. But its brevity here reflects under-specification rather than disciplined conciseness, so it is only adequate rather than exemplary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with four required parameters, a nested free-form config object, no annotations, and no output schema needs substantial explanation of behavior, parameter expectations, and results. The description supplies none of this and is wholly inadequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: workspaceId and config have descriptions while nodeId and workflowId have none. More importantly, the nested `config` object is declared with `additionalProperties: {}` and no structure, and the description adds nothing about what configuration keys are valid or expected. The description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Update a node configuration" gives a clear verb (update) and resource (node configuration), which distinguishes it from add_node and delete_node by implication. However, it never says what a node is in this workflow/graph context or what aspect of configuration is being changed, so the agent gets only a bare minimum of differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives among the many sibling mutation tools (add_node, delete_node, connect_nodes, get_node_definitions). The agent must infer everything about invocation context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_overlay_itemCInspect
Update an existing overlay item. Can change text, position, timing, animation, opacity, and more.
| Name | Required | Description | Default |
|---|---|---|---|
| exit | No | Exit animation, or null to remove | |
| size | No | Normalized overlay size | |
| text | No | New text content | |
| type | No | Change item type when also supplying compatible fields | |
| color | No | New text color (hex or rgba) | |
| enter | No | Entrance animation, or null to remove | |
| width | No | Legacy shorthand for size.width | |
| border | No | Border styling, or null to remove | |
| editId | Yes | The edit ID | |
| height | No | Legacy shorthand for size.height | |
| itemId | Yes | The overlay item ID | |
| shadow | No | Media overlay drop shadow, or null to remove | |
| volume | No | Video overlay audio volume | |
| clipEnd | No | Video overlay trim end in seconds, or null to clear | |
| mediaId | No | MediaDocV2 ID for image/video overlays, or null to remove | |
| opacity | No | Opacity from 0 to 1 | |
| trackId | Yes | The overlay track ID | |
| duration | No | Timeline duration in seconds | |
| fontSize | No | New font size | |
| position | No | Normalized overlay position | |
| rotation | No | Rotation in degrees | |
| clipStart | No | Video overlay trim start in seconds | |
| fontStyle | No | New font style | |
| keyframes | No | Transform keyframes relative to item start | |
| positionX | No | Legacy shorthand for position.x | |
| positionY | No | Legacy shorthand for position.y | |
| speedRamp | No | Variable speed curve for video overlays, or null to clear | |
| startTime | No | Timeline start time in seconds | |
| textAlign | No | New text alignment | |
| fontFamily | No | New font family | |
| fontWeight | No | New font weight | |
| lineHeight | No | New line-height multiplier | |
| lineStyles | No | Per-line text style overrides | |
| textConfig | No | Full TextEffectConfig partial update | |
| strokeColor | No | New text stroke color | |
| strokeWidth | No | New text stroke width | |
| workspaceId | Yes | The workspace ID | |
| borderRadius | No | Corner radius in 720p logical pixels | |
| letterSpacing | No | New letter spacing | |
| playbackSpeed | No | Flat playback speed for video overlays | |
| textAnimation | No | New text animation | |
| remotionConfig | No | Update animation code or editable settings, merged with the existing composition | |
| textShadowBlur | No | New text shadow blur | |
| backgroundColor | No | New text background color | |
| textShadowColor | No | New text shadow color | |
| textShadowOffsetX | No | New text shadow horizontal offset | |
| textShadowOffsetY | No | New text shadow vertical offset | |
| backgroundPaddingX | No | New text background horizontal padding | |
| backgroundPaddingY | No | New text background vertical padding | |
| backgroundBorderRadius | No | New text background corner radius |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose whether this is a partial update, how omitted fields behave, permission requirements, reversibility, or return behavior. Listing mutable field categories is not meaningful behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the verb and resource front-loaded. However, 'and more' is filler and the definition is under-specified for a 50-parameter tool. Structurally efficient but not fully earning its brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 50-parameter mutation tool with no annotations and no output schema, the description is far too thin. It omits partial-update semantics, required identifier roles, null-to-remove behavior, and side effects, leaving the agent to infer behavior from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 50 parameters, so the schema already documents each parameter. The description lists categories such as text, position, timing, animation, and opacity, but adds no syntax, constraints, or per-parameter meaning beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (overlay item), and lists example mutable fields. It distinguishes itself from add/delete/list siblings by verb, but only implicitly differentiates from update_overlay_track by saying 'item' rather than 'track'. Clear, though sibling differentiation is not explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, no prerequisites such as required workspace/edit/track/item IDs, and no mention of partial-update behavior. Usage is only implied by 'update an existing overlay item'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_overlay_trackBInspect
Update overlay track metadata: name, stack order, visibility, and lock state.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Track display name | |
| order | No | Stack order; higher renders on top | |
| editId | Yes | The edit ID | |
| locked | No | Whether the track is locked | |
| trackId | Yes | The overlay track ID | |
| visible | No | Whether the track is visible | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says 'update' but does not disclose partial-vs-full update behavior, permission requirements, reversibility, or what happens to omitted fields, leaving a significant transparency gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. It efficiently identifies the operation and the kinds of metadata being changed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema is rich and fully documented, and no output schema exists, so parameter coverage is adequate. However, with no annotations and no usage or behavioral context, the description is only minimally complete for a mutation tool with seven parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema documents each parameter including required IDs. The description adds little beyond restating a few updatable field names ('stack order' for order is a slight clarification), so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: update overlay track metadata. The listed fields clarify the operation, but it does not explicitly distinguish itself from sibling tools such as update_overlay_item or update_audio_track.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, prerequisites, or alternatives. It implies updating an existing track, but an agent must infer that from the name and schema alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_proposalCInspect
Update an existing proposal to track execution progress.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit ID | |
| proposalId | No | Specific proposal ID to update | |
| workspaceId | Yes | The workspace ID | |
| proposalMarkdown | Yes | Updated proposal content with execution steps marked complete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and largely fails. It implies a mutation but does not state whether proposalMarkdown replaces the entire document, whether edits are reversible, whether editId is an optimistic-concurrency token, or what permissions are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is efficient, though the brevity comes at the cost of the missing behavioral detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description should cover permissions, overwrite semantics, and the meaning of the editId token. None of that is present, so an agent cannot call it correctly with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (including the required editId/workspaceId/proposalMarkdown) are already documented in the schema. The description adds no syntax, format, or constraint detail beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (proposal) plus an intent (track execution progress), which is more than a tautology. However, it does not differentiate from siblings like save_proposal or set_proposal_status, leaving the agent to infer which mutation path applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus save_proposal (create/overwrite) or set_proposal_status (status transitions). The phrase 'track execution progress' hints at a scenario but names no alternative or precondition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_sceneCInspect
Update an existing scene.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name | |
| order | No | New order position | |
| editId | Yes | The edit ID | |
| sceneId | Yes | The scene ID | |
| locationRef | No | Asset ID for location | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not state permission requirements, whether the update is partial or full, irrevocability, or concurrency implications. 'Update an existing scene' implies mutation but discloses nothing beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no redundancy. It is front-loaded, but its brevity comes at the cost of useful context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no annotations and no output schema, the description is too thin. It omits behavioral details and does not explain that it supports partial updates via optional fields. The schema provides parameter names, but the description fails to compensate for the missing annotation layer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is documented in the input schema. The description itself adds no parameter meaning. Baseline 3 applies because the schema already does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource ('Update an existing scene'), which is more specific than a pure tautology. However, it does not distinguish update_scene from sibling mutation tools like update_shot, update_edit, or update_workspace, or from scene-specific siblings such as get_scene, delete_scene, or list_scenes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool, when not to, or which alternatives exist. With many 'update_*' siblings, an agent receives no routing signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_shotBInspect
Update an existing shot. Can change scene, order, timing, media refs, dialogue, motion prompt, and tagged assets. Only provided fields are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| order | No | New order position | |
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID | |
| volume | No | Volume level (0-1) | |
| clipEnd | No | Clip end time, or null to clear | |
| sceneId | No | Move shot to a different scene | |
| dialogue | No | Visible on-screen character speech for intentional lip-sync/TTS. Do not use for off-screen narration or captions. | |
| duration | No | New duration in seconds | |
| clipStart | No | Clip start time | |
| mediaRefs | No | Complete replacement list of fine-grained media refs | |
| speedRamp | No | Variable speed curve, or null to clear | |
| startTime | No | Timeline start time in seconds; usually derived by editor | |
| rawDuration | No | Full source duration in seconds | |
| videoPrompt | No | Video/motion prompt for camera movement and action | |
| workspaceId | Yes | The workspace ID | |
| imageMediaId | No | New image media reference, or null to clear | |
| taggedAssets | No | Updated tagged asset IDs | |
| videoMediaId | No | New video media reference, or null to clear | |
| playbackSpeed | No | Playback speed multiplier | |
| hideStartFrame | No | Hide/remove the Start Frame tab | |
| linkEndToNextStart | No | Auto-use next shot start frame as this shot end frame | |
| linkStartToPrevEnd | No | Auto-use previous shot end frame as this shot start frame |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It discloses the partial-update semantics, but says nothing about permission/auth requirements, whether clearing values (clipEnd/speedRamp/imageMediaId accept null) is reversible, or the destructive nature of full-list replacement fields like mediaRefs and taggedAssets. For a 22-parameter mutation tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and ending with the most operationally important constraint. The field enumeration is somewhat redundant against a 100%-covered schema, which keeps it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter mutation tool with no annotations and no output schema, the description is thin: it omits sibling routing, permission requirements, and the replace-not-merge behavior of list fields. The partial-update sentence is the one strong element, but overall an agent would need to open the schema and guess at alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 22 parameters, which sets the baseline at 3. The description adds a useful conceptual grouping (scene, order, timing, media refs, dialogue, motion prompt, tagged assets) but omits several parameters entirely (speedRamp, playbackSpeed, linkEndToNextStart, hideStartFrame, volume), so it does not exceed the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource ('Update an existing shot') followed by an enumeration of the mutable field groups. An agent immediately knows this mutates a shot. However, it never distinguishes itself from siblings that overlap heavily (move_shot_to_scene, reorder_shots, update_shot_media_ref, split_shot), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Only provided fields are changed' is a genuine invocation guideline that tells the agent this is a partial/PATCH-style update, so no full read-modify-write is required. But there is no guidance on when to prefer this tool over the scene/reorder/media-ref siblings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_shot_media_refBInspect
Update a specific shot media ref by matching mediaId and optional role. Can change role, order, mediaId, and version override.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | Current role to match | |
| order | No | Replacement order | |
| editId | Yes | The edit ID | |
| shotId | Yes | The shot ID | |
| mediaId | Yes | Current media ID to match | |
| newRole | No | Replacement role | |
| newMediaId | No | Replacement media ID | |
| workspaceId | Yes | The workspace ID | |
| versionOverride | No | Replacement version override, or null to clear | |
| alsoSetPrimaryMediaId | No | For image/video roles, also update imageMediaId/videoMediaId |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the matching mechanism and the mutable fields, but omits side effects such as the default-true alsoSetPrimaryMediaId behavior, permission requirements, and whether the change is reversible or idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the matching semantics are front-loaded and the changeable fields follow. Slightly compressed, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a ten-parameter mutation tool with no annotations and no output schema, the description is serviceable but thin: it never warns about the alsoSetPrimaryMediaId side effect or confirm whether unspecified fields are preserved. The rich schema covers parameters, but behavioral context is under-supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all ten parameters, including the match-vs-replacement distinction (mediaId/newMediaId, role/newRole). The description's field list ('role, order, mediaId, and version override') largely restates schema content, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Update a specific shot media ref') and clarifies the matching semantics (mediaId plus optional role). It is clearly distinguishable from add_shot_media_ref and delete_shot_media_ref by the 'update' verb, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'by matching mediaId and optional role' phrasing implies how to target an existing ref, which is useful implied usage. But there is no explicit when-to-use vs when-not, and no pointer to the add/delete siblings for the create/remove cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_transitionBInspect
Update one timeline transition by transitionId or boundaryId.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | New transition type | |
| editId | Yes | The edit ID | |
| duration | No | New duration in seconds | |
| toShotId | No | New incoming shot ID | |
| startTime | No | New nominal boundary time | |
| boundaryId | No | Boundary ID to update | |
| fromShotId | No | New outgoing shot ID | |
| workspaceId | Yes | The workspace ID | |
| transitionId | No | Transition ID to update | |
| newBoundaryId | No | New boundary ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies mutation but does not state whether the update is partial or full-replace, what happens to unmentioned fields, permission requirements, or reversibility. For a 10-parameter write tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler. It is efficient, though for a 10-parameter mutation the sizing is arguably too lean rather than appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, no annotations, and no output schema, a one-line description is inadequate. It omits partial-update semantics, the effect on adjacent shots/boundaries, and any prerequisite context an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a genuinely useful relationship: that a transition is addressed by transitionId OR boundaryId. The schema lists both without clarifying they are alternative identifiers, so this is a real contribution.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (timeline transition) with scope ('one'), and adds the two identification keys (transitionId or boundaryId). This distinguishes it from add_transition and delete_transition, though it does not name those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no explanation of when to key by transitionId versus boundaryId. The agent must infer the correct context entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_workspaceCInspect
Update a workspace. Only the workspace owner can update.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | Workspace icon (emoji) | |
| color | No | Workspace color (hex or named) | |
| title | No | New workspace title | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, but it only discloses the owner-only permission requirement. It says nothing about whether the update is partial or full, whether omitted fields are preserved, or what the result contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The purpose comes first and the permission constraint second, which is efficient, though it is arguably too terse given the gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is far too thin. It omits partial-vs-full update semantics, reversibility, return value, and error behavior, leaving significant gaps an agent needs filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the four parameters (icon, color, title, workspaceId) is already documented in the schema. The description adds nothing about parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource (update a workspace), which is unambiguous on its own. However it does not differentiate from the many sibling update_* tools or explain what fields are updatable, leaving the reader to infer from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is the authorization constraint ('Only the workspace owner can update'). There is no when-to-use vs alternatives, no mention of get_workspace/list_workspaces for reading, and no context on when updating is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_media_base64AInspect
Upload a local image, video, or audio file into Sequencer by passing base64 file contents. This is the MCP-native local-file path when the client can read the file bytes. For large videos, prefer the REST /v1/upload multipart endpoint to avoid MCP payload limits.
| Name | Required | Description | Default |
|---|---|---|---|
| fileName | Yes | Original file name, e.g. idle_loop_silent.mp4 | |
| folderId | No | Optional destination folder ID | |
| mimeType | Yes | MIME type, e.g. video/mp4 | |
| base64Data | Yes | Base64-encoded file contents. Data URL prefixes are accepted. | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses the MCP payload-size constraint and the base64 transport mechanism, but says nothing about auth/permission needs, what happens on duplicate or failed uploads, or what the call returns after the bytes land.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information: what it does, when it applies, and when to route elsewhere. The most decision-relevant content (local-file path) is front-loaded with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should say more about what an agent gets back (an asset/media ID?) and whether upload kicks off processing. It covers transport, scope, and limits adequately but leaves the post-call state undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (workspaceId, fileName, mimeType, base64Data, folderId) are already documented in the schema, including the data-URL prefix note. The description adds no per-parameter meaning beyond that, which is the expected baseline here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a precise verb and resource ('Upload a local image, video, or audio file into Sequencer') plus the mechanism (base64 contents). It implicitly separates itself from the URL-based import path by calling itself the 'local-file path', but never names the sibling tool (e.g. import_media_from_url), so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the enabling condition ('when the client can read the file bytes') and an explicit exclusion with a named alternative ('For large videos, prefer the REST /v1/upload multipart endpoint to avoid MCP payload limits'). This is textbook when-to-use/when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_imageupscale imageADestructiveInspect
Upscale an existing image using a catalog model that supports upscale_image. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=false, and openWorld=true, so safety is covered. The description usefully adds that the operation returns a tracked task plus a playable result when complete, and that the model must support the operation — but does not restate idempotencyKey retry semantics that live only in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the operation, then the model prerequisite, then the outcome. Minor redundancy between 'a catalog model that supports upscale_image' and 'Use get_model_catalog', but nothing wasteful enough to harm selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and rich annotations, the description needn't explain returns, and it covers the model dependency and result semantics. For a 21-parameter, nested-object tool the coverage is adequate, though it could say more about the modelSettings/mediaInputs flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 81%, so the schema itself documents most parameters. The description only reinforces the model-selection dependency and says nothing about scaleFactor, resolution, or cost controls beyond what the schema already states, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Upscale an existing image') that cleanly separates it from siblings like upscale_video, edit_image, and enhance_video_media. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to get_model_catalog to select a supported model and its settings, which is the key prerequisite. It does not, however, contrast itself against near-alternatives like edit_image or enhance_video_media, so when-not-to-use is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_videoupscale videoADestructiveInspect
Upscale an existing video using a catalog model that supports upscale_video. Use get_model_catalog to choose a supported model and its settings. Returns a tracked task and playable result when complete.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Active model ID supporting this operation. Use get_model_catalog. | |
| prompt | No | Edit or enhancement instructions. | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| resolution | No | ||
| aspectRatio | No | ||
| endFrameUrl | No | ||
| mediaInputs | No | Named media slots from get_model_catalog, for example reference_image_uri or mask_url. | |
| scaleFactor | No | Upscaling multiplier when supported by the selected model. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| workspaceId | No | ||
| outputFormat | No | Output format when supported by the selected model. | |
| generateAudio | No | Explicitly enable or disable generated video audio, when supported. | |
| modelSettings | No | Model-specific controls using keys and options from get_model_catalog inputConstraints.slots. | |
| sourceMediaId | Yes | Completed workspace source media. The original is preserved. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. | |
| sourceImageUrl | No | Public HTTP(S) source image/start-frame URL. Use mediaId for workspace images. | |
| sourceVideoUrl | No | Public HTTP(S) source-video URL. | |
| endFrameMediaId | No | Workspace image for the final frame, when supported by the selected model. | |
| referenceImages | No | Images to guide generation. Use labels such as person or product and refer to them as @person or @product in the prompt. Each item accepts a workspace mediaId or a public HTTP(S) URL. | |
| sourceImageMediaId | No | Workspace image to use as the source image/start frame. | |
| sourceVideoMediaId | No | Workspace video to edit or transform, when supported by the selected model. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (openWorldHint, destructiveHint, non-idempotent), and the description adds that the call is asynchronous, returning a tracked task with a playable result. It says nothing about charges for a paid generation beyond what the schema's maxCostUsd and idempotencyKey fields already imply. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and then the prerequisite and the return shape. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 21 parameters, nested objects, and an output schema present, the description adequately covers the operation, the required catalog lookup, and the async return. It does not disambiguate the many overlapping source inputs (sourceMediaId vs sourceVideoMediaId vs sourceVideoUrl), though the schema handles that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 81%, above the high-coverage threshold, so the schema already documents model, prompt, cost, scaleFactor, media slots, and idempotency. The description adds only the model-catalog pointer, which the schema's 'model' description already repeats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upscale an existing video') and scopes it to catalog models that support the operation. It does not, however, distinguish itself from near-siblings like upscale_image, enhance_video_media, expand_video, or convert_video_to_hdr, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete prerequisite: consult get_model_catalog to pick a supported model and its settings. That is useful routing, but there is no when-to-use-this-vs-alternatives guidance, no statement of when upscaling is inappropriate, and no mention of cost/preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_production_templateARead-onlyIdempotentInspect
Render a template with parameters and estimate generation cost without spending credits. Missing price data is reported explicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| version | No | ||
| parameters | No | ||
| templateId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety is clear. The description adds valuable context: rendering occurs without spending credits and missing price data is reported explicitly, which goes beyond annotations. It doesn't detail rate limits or auth needs, but those are less critical given annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that front-load the core purpose and then add a key behavioral trait. No extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main action and a crucial cost-related aspect, but with 4 parameters at 0% schema coverage, it leaves gaps in parameter explanation. No output schema exists, so describing return values isn't required, but the description could be more complete about what 'render' entails or what the output contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning all parameters (workspaceId, templateId, version, parameters) lack documentation in the schema. The description mentions 'parameters' and 'render a template' but doesn't explain required parameters, their formats, or the nested 'parameters' object, failing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (render/validate) and resource (template) and clarifies it estimates generation cost. It contrasts with sibling run_production_template by emphasizing cost estimation without spending credits, though the name 'validate_production_template' might suggest more than just rendering, creating slight ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly tells when to use it: before running a template to estimate costs. However, it doesn't explicitly name alternatives like run_production_template or quote_generation, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
watch_videoAInspect
Export the current project as a low-quality review video and return its URL for AI analysis. Use this when the user explicitly asks to watch, review, or analyze the project, or when the user approved video generation for an active finished-video skill that requires a final review-and-edit pass. The video is exported at 480p/15fps for fast, cost-efficient viewing. Returns a video URL that the agent should fetch and inspect before applying concrete edit fixes.
| Name | Required | Description | Default |
|---|---|---|---|
| editId | Yes | The edit/project ID to export and review | |
| workspaceId | Yes | The workspace ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the 480p/15fps low-quality profile, the cost rationale, that a URL is returned, and that the agent should fetch and inspect it before editing. It omits whether the export is synchronous or takes time, whether it persists as an artifact, and idempotency/repeat-call behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and result in the first sentence, followed by usage conditions and then behavior. Three sentences, each doing work, though the phrasing is slightly verbose in the middle sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description correctly explains the return value (a video URL to fetch and inspect). For a two-parameter review-export tool the coverage is nearly complete, with only export timing/persistence left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the schema already documents editId and workspaceId. The description implies the project/edit context ('current project') but adds no syntax or format detail beyond it, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb, resource and output: export the current project as a low-quality review video and return its URL. It is clearly separable from siblings like export_edit (full-quality export) and review_edit, and it resolves the misleading tool name 'watch_video' by explaining the actual action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions: the user explicitly asks to watch/review/analyze, or video generation was approved for an active finished-video skill needing a final review pass. It does not name a sibling alternative or a when-not-to-use case, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
151 tool updates
- First observed
add_audio_track - First observed
add_media_overlay - First observed
add_node - First observed
add_overlay_track - First observed
add_remotion_overlay - First observed
add_shot_media_ref - First observed
add_text_overlay - First observed
add_transition - First observed
analyze_export_ip - First observed
change_audio_track_voice - First observed
change_video_voice - First observed
check_asset_status - First observed
connect_nodes - First observed
convert_video_to_hdr - First observed
create_asset - First observed
create_edit - First observed
create_folder - First observed
create_production - First observed
create_scene - First observed
create_shot - First observed
create_shot_with_media - First observed
create_workflow - First observed
create_workspace - First observed
delete_asset - First observed
delete_audio_track - First observed
delete_connection - First observed
delete_node - First observed
delete_overlay_item - First observed
delete_overlay_track - First observed
delete_scene - First observed
delete_shot - First observed
delete_shot_media_ref - First observed
delete_transition - First observed
duplicate_shot - First observed
edit_image - First observed
edit_video - First observed
enhance_video_media - First observed
execute_workflow - First observed
expand_video - First observed
export_edit - First observed
extract_product_image - First observed
generate_audio - First observed
generate_audio_track - First observed
generate_image - First observed
generate_ip_report - First observed
generate_shot_video - First observed
generate_video - First observed
generate_video_audio - First observed
get_asset - First observed
get_available_tools - First observed
get_balance - First observed
get_creative_tips - First observed
get_edit - First observed
get_edit_full - First observed
get_experience_game - First observed
get_experience_world - First observed
get_export - First observed
get_feature_list - First observed
get_image_prompt_guide - First observed
get_media - First observed
get_model_catalog - First observed
get_node_definitions - First observed
get_page_link - First observed
get_platform_overview - First observed
get_pricing_info - First observed
get_production_status - First observed
get_production_template - First observed
get_proposal - First observed
get_remotion_layer_guide - First observed
get_scene - First observed
get_sequencer_capabilities - First observed
get_sequencer_skill - First observed
get_server_info - First observed
get_shot - First observed
get_storytelling_guide - First observed
get_subscription_help - First observed
get_support_faq - First observed
get_text_presets - First observed
get_usage_guide - First observed
get_video_prompt_guide - First observed
get_workflow - First observed
get_workflow_building_guide - First observed
get_workflow_summary - First observed
get_workspace - First observed
import_media_from_url - First observed
inspect_edit_frame - First observed
inspect_storyboard - First observed
lip_sync_video_media - First observed
list_assets - First observed
list_audio_tracks - First observed
list_available_nodes - First observed
list_edits - First observed
list_experience_games - First observed
list_exports - First observed
list_folders - First observed
list_media - First observed
list_overlay_items - First observed
list_overlay_tracks - First observed
list_production_templates - First observed
list_scenes - First observed
list_shots - First observed
list_workflows - First observed
list_workspace_members - First observed
list_workspaces - First observed
listen_to_audio - First observed
move_item_to_folder - First observed
move_overlay_item_to_track - First observed
move_shot_to_scene - First observed
navigate_to_page - First observed
open_paywall - First observed
open_sequencer_workspace - First observed
quote_generation - First observed
remove_image_background - First observed
remove_video_background - First observed
remove_video_background_media - First observed
render_image - First observed
render_media - First observed
reorder_shots - First observed
request_asset_review - First observed
request_login - First observed
review_edit - First observed
run_production_template - First observed
run_video_workflow - First observed
save_experience_game - First observed
save_experience_game_proposal - First observed
save_production_template - First observed
save_proposal - First observed
scrape_url - First observed
search_workspace - First observed
set_asset_approval_status - First observed
set_proposal_status - First observed
split_shot - First observed
trigger_auto_layout - First observed
update_asset - First observed
update_audio_track - First observed
update_edit - First observed
update_media_tags - First observed
update_node - First observed
update_overlay_item - First observed
update_overlay_track - First observed
update_proposal - First observed
update_scene - First observed
update_shot - First observed
update_shot_media_ref - First observed
update_transition - First observed
update_workspace - First observed
upload_media_base64 - First observed
upscale_image - First observed
upscale_video - First observed
validate_production_template - First observed
watch_video
Publisher details
- Operator
- Sequencer Media · Publisher source
- Operator website
- https://sequencer.media/plugin · Publisher source
- Vendor relationship
- First-party · Publisher source
- Documentation
- https://sequencer.media/plugin#connect · Publisher source
- Trust center
- Unknown
- Restrictions
- A Sequencer account and OAuth connection are required. AI generation uses Sequencer credits. Review the cost quote before running paid generation. · Publisher source
Related MCP Connectors
Generate and edit images, create videos, quote credit costs, and retrieve private results.
- MorphedOAuthapp.morphed
Create AI images and videos, manage projects and credits, and use workspace campaign context.
- AITOPIAOAuthai.aitopia
Generate and edit images, video and audio: 300+ AI models, video tools, voice cloning, transcripts.
1 AI image, video, audio and text tools for creators. Pay per use with credits.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceGenerate and refine AI images/audio/video through natural conversation.409Apache 2.0

Popcraft MCPofficial
AlicenseNot gradedqualityBmaintenanceEnables AI agents and MCP clients to generate and edit video, images, 3D, music, and sound effects, upload or reuse media, preview credit costs, and follow production workflows through a single hosted connector.MIT- AlicenseBqualityBmaintenanceExecution control layer for AI agents - Reserve, execute, burn/refund pattern for media generation162MIT
- AlicenseAqualityCmaintenanceAI image and video generation, editing, and region repair via Gemini, OpenAI, and Grok1170 npm5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.