Skip to main content
Glama

Script to video

script_to_video

Convert a narration script into an editable video with AI-generated visuals and captions, ideal for explainer videos about a minute or longer. Pass actorEntityId for avatar narration.

Instructions

Preferred for narrated / informational / explainer videos from text, especially ~1 minute or longer. Turn a verbatim narration script into an editable video with AI-generated visuals and captions. Prefer this over storyboard_to_video unless the user wants a short shot-directed storyboard. For avatar narration, pass actorEntityId and optionally set avatarQuality.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
styleNoVisual style for generated images. Write a full, strict paragraph in the same form as the app defaults below: name the medium, texture, and palette, then lock composition. Do not pass a short label such as "watercolor", "flat art", or "cinematic". Image models fill the frame with extra objects, readable text, charts, diagrams, tables, and labels unless the style forbids that. The still then looks crowded and hard to use. Every style must keep the picture simple. Include this composition lock verbatim: A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. No on-image text, letters, labels, captions, charts, diagrams, tables, legends, or infographic layout unless the user explicitly asked for one specific word or number on screen. Copy an app default in full (those already include the uncluttered-subject lock), then add the no-text / no-diagram sentence if it is missing. A custom style is allowed only when it is equally long and strict, and includes the same composition lock. Omit this field for the Realistic default. Watercolor: Loose watercolor illustration, visible brushstrokes, soft color bleeds, paper texture, muted palette. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Realistic: Photorealistic photograph, natural lighting Whiteboard: Minimalist whiteboard explainer style, simple line drawings, marker sketch aesthetic, clean white background, subtle accent colors. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. 3D cartoon: 3D cartoon render, rounded forms, soft lighting, vibrant colors, clean matte materials, smooth stylized characters Flat art: Corporate memphis flat art illustration, simple geometric shapes, bold colors, clean composition, white background with sparse subtle geometric accents such as small dots, lines, or shapes scattered in the margins. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no busy background patterns. Paper: Paper collage illustration, torn edges, layered cut paper shapes, mixed media texture, handmade craft aesthetic, matte paper finish. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Ink paint: Traditional Japanese sumi-e ink wash painting, expressive black ink brushstrokes, varied tonal gradations from deep black to soft gray, rice-paper texture, minimalist composition. No text, letters, calligraphy, kanji, signatures, or red seal stamps unless explicitly required by the subject. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Anime: Anime illustration, cel-shaded style, vibrant colors, clean linework, soft gradient backgrounds Editorial: Editorial illustration, conceptual art, bold geometric shapes, sophisticated color palette, negative space, minimalist composition, strong silhouette Isometric: Isometric 3D illustration, clean vector style, soft gradient background, matte pastel colors, simple geometric objects, no characters. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Claymation: Claymation style, soft clay figurines, plasticine texture, handmade stop-motion aesthetic, warm studio lighting, shallow depth of field Education: Educational infographic style, simple icons, pastel colors, white background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Chalkboard: Simple, minimalist, white chalk line drawings on dark green chalkboard, hand-drawn sketch style, chalk dust texture. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Boho: Simple graphic with only a few elements. Boho linocut poster. Textured organic strokes with crisp lines and a white textured background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Poster: Risograph print aesthetic, halftone texture, limited duotone palette, paper grain, high contrast, bold flat colors
scriptYesNarration script, spoken verbatim.
qualityNoGeneration quality. Omit to use workspace settings.
voiceIdNoCatalog display name (e.g. Matilda) or voice id from list_tts_voices.
languageNoOutput language as a BCP-47 code, such as en, es, or fr.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
actorEntityIdNoId of an ACTOR entity (vg_enti_...) with an image reference. When set, narration is delivered by that actor avatar.
avatarQualityNoAvatar generation quality tier. Applies when actorEntityId is provided. Omit to use workspace settings.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.2.1

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description usefully adds that the output is editable and includes AI visuals and captions, plus the avatar narration path. However, it never discloses that generation is long-running/async (implied only in the autoExport schema text), nor any cost or duration limits, which is the most important missing behavioral fact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the selection guidance before the capability statement and the avatar tip. No filler, and the highest-value routing information (preferred use case and the sibling exclusion) comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% schema description coverage, the description need not explain return values or parameters. It covers purpose, the alternative tool, and the avatar variant, which is everything an agent needs to choose and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the style field's semantics are exhaustively documented in the schema, so the schema does the heavy lifting. The description only adds a light steer ('pass actorEntityId and optionally set avatarQuality'), which largely restates what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Turn a verbatim narration script into an editable video') with explicit scope ('narrated / informational / explainer videos', '~1 minute or longer'). An agent can distinguish this from storyboard_to_video without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with a clear preference statement, then names the sibling alternative and the exact condition that selects it ('Prefer this over storyboard_to_video unless the user wants a short shot-directed storyboard'). It also routes the avatar case to specific parameters, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.