Skip to main content
Glama

@videogen/mcp

A Model Context Protocol (MCP) server that exposes the full VideoGen API to any MCP client (Cursor, Claude Desktop, Windsurf, etc.). Your agent can generate videos from scripts, produce images / voiceovers / music / avatars, upload files, export and remix projects, and manage runs — all authenticated with your own API key.

Learn more about the VideoGen MCP server.

Find the hosted server on Smithery.

It ships in two transports. Tool surface is the same except for ChatGPT Apps commerce policy (see below):

  • Local (stdio) — the client launches @videogen/mcp as a subprocess and reads the key from VIDEOGEN_API_KEY.

  • Remote (Streamable HTTP) — a hosted, multi-tenant HTTP server. Clients connect over the network and authenticate per request with Authorization: Bearer <api-key>; no key lives on the server.

Host surfaces (remote): /mcp (and stdio) is the STANDARD surface: when the user is out of credits or needs a plan feature, tools and guidance walk them into get_app_deep_link (OPEN_UPGRADE / OPEN_PURCHASE_CREDITS / OPEN_ENABLE_TOP_UPS). /mcp/chatgpt is the ChatGPT Apps surface: those commerce deep links are omitted, billing errors are rewritten to “manage your VideoGen account,” and copy never says purchase / buy / upgrade / top-ups (OpenAI Plugins digital-goods policy). See .cursor/rules/chatgpt-mcp-no-commerce.mdc and mcp/src/hostSurface.ts.

This is distinct from the hosted documentation MCP at https://docs.videogen.io/_mcp/server, which only lets clients read the API docs. This server actually executes the API.

How it works

  • Wraps the official @videogen/sdk and calls the live API through the SDK's authenticated passthrough.

  • The local transport reads your key from VIDEOGEN_API_KEY; the remote transport reads it from each request's Authorization: Bearer header and builds a fresh, per-request server bound to that team (stateless — no session state is shared between requests).

  • The key never leaves your machine except in requests to the VideoGen API (local), or is forwarded only to the VideoGen API for the duration of the request and never persisted (remote).

  • Long-running operations (workflows, media tools, project exports) use composite tools. The hosted HTTP transport returns the run/execution id immediately; use the corresponding get_* tool to continue polling. The local stdio transport waits for completion by default.

Related MCP server: Sora 2 MCP Server

Configuration

Get an API key from app.videogen.io/api. Both transports expose the same tools; the remote server is recommended.

Nothing to install or update. Point any MCP client that supports the Streamable HTTP transport at the hosted endpoint and send your API key as a bearer token:

{
  "mcpServers": {
    "videogen": {
      "url": "https://mcp.videogen.io/mcp",
      "headers": {
        "Authorization": "Bearer sk_videogen_live_..."
      }
    }
  }
}

The remote server:

  • Accepts MCP JSON-RPC messages via POST /mcp (stateless — a fresh server per request).

  • Reads the API key or OAuth access token from the Authorization: Bearer header. The standard /mcp endpoint requires credentials for every request and returns 401 with a WWW-Authenticate challenge to start OAuth. The ChatGPT-specific /mcp/chatgpt endpoint serves tool discovery unauthenticated and returns its OAuth challenge on a protected tool result because ChatGPT does not start OAuth from the standard transport challenge. The credential is forwarded only to the VideoGen API and never stored.

  • Exposes GET /health for load-balancer / Cloud Run startup probes.

  • Handles CORS preflight (OPTIONS) so browser-based clients can connect.

  • Supports three upload paths. For small assets (images, logos, short audio), upload_file takes base64-encoded contents inline (fileData). For large files, create_file_upload returns { fileId, uploadUrl }; the client PUTs the raw bytes to that short-lived pre-signed URL (no Authorization header) and then calls get_file with { fileId, wait: true } to wait for processing. In ChatGPT (an MCP Apps host), open_uploader renders an in-chat upload widget so the user can pick a file directly: the widget itself calls create_file_upload, PUTs the bytes client-side, and reports back only the resulting vg_file_... id, so the pre-signed URL is never surfaced to the model. Either way you get a vg_file_... id to pass to other tools. The local (stdio) server instead uploads by local filePath. (The server never fetches a caller-supplied URL, so there is no SSRF surface.)

Local (stdio) — Cursor / Claude Desktop

Runs as a subprocess launched by your client with npx. Reads the API key from the VIDEOGEN_API_KEY environment variable, which never leaves your machine except in requests to the VideoGen API:

{
  "mcpServers": {
    "videogen": {
      "command": "npx",
      "args": ["-y", "@videogen/mcp"],
      "env": {
        "VIDEOGEN_API_KEY": "sk_videogen_live_..."
      }
    }
  }
}

Environment variables

Variable

Required

Default

Description

VIDEOGEN_API_KEY

local only

—

Your VideoGen API key (local server). On the remote server the key travels in the Authorization header instead.

VIDEOGEN_BASE_URL

no

https://api.videogen.io

Override the upstream API base URL (e.g. for local development). Applies to both transports. When unset, the remote server resolves the upstream API per deployment environment (dev/prerelease/prod); local runs default to the public prod API.

VIDEOGEN_OAUTH_ISSUER

no

—

Remote server only. Full OAuth 2.1 issuer URL. When set (or derived from the var below), the server advertises OAuth protected-resource metadata (RFC 9728) and a resource_metadata 401 challenge so MCP clients can discover the authorization server and run account linking. Must match the issuer that the upstream API (VIDEOGEN_BASE_URL) validates tokens against.

VIDEOGEN_OAUTH_SUPABASE_PROJECT_URL

no

—

Remote server only. Supabase project base URL; the issuer is derived as ${url}/auth/v1. Ignored when VIDEOGEN_OAUTH_ISSUER is set.

VIDEOGEN_OPENAI_APPS_CHALLENGE_TOKEN

no

—

Remote server only. Public OpenAI Plugins / ChatGPT Apps domain-verification token. When set, GET /.well-known/openai-apps-challenge returns that exact value as text/plain. When unset, that path is a non-JSON-RPC 404.

Tools

Workflows (end-to-end video)

script_to_video, voiceover_to_video, slideshow_to_video, storyboard_to_video, list_workflow_runs, get_workflow_run, cancel_workflow_run

Media tools

generate_image, generate_video_clip, text_to_speech, generate_sound_effect, generate_music, generate_avatar, vectorize_image, remove_image_background, remove_video_background, upscale_image, upscale_video, image_3d_effect, list_tool_executions, get_tool_execution, cancel_tool_execution

Projects

list_projects, get_project, export_project, get_project_export, remix_project, list_project_remix_actions

Files

upload_file, create_file_upload, get_file, list_files, open_uploader

open_uploader is a ChatGPT App widget (remote server only): it renders an in-chat file picker (a React component served as an MCP UI resource) so a ChatGPT user can attach a file without pasting a link. It is a no-op on clients that don't render MCP Apps UI — those use upload_file / create_file_upload instead.

Entities

list_entities, create_entity, get_entity, update_entity, archive_entity, add_entity_reference, remove_entity_reference

Create ACTOR / PRODUCT / VISUAL_STYLE entities, attach uploaded image references, then pass vg_enti_... ids into workflows and generate_avatar (actorEntityId, storyboard entity attachments, etc.).

Resources & account

list_tts_voices, list_languages, get_me, get_app_deep_link

Guidance (docs for agents)

Operational manuals exposed as MCP resources and mirror tools (no API credential required):

Resource URI

Mirror tool

guidance://getting-started

get_getting_started_guidance

guidance://async-tasks

get_async_tasks_guidance

guidance://workflows

get_workflows_guidance

guidance://tools-vs-workflows

get_tools_vs_workflows_guidance

Call the matching get_*_guidance tool before non-trivial setup, polling, workflow/remix/export, or tools-vs-workflows decisions. Many hosts never auto-attach resources; the tools are the reliable path. Content is grounded in the public docs at docs.videogen.io but written for MCP tool usage (including hosted wait caps).

Creative MCP tools expose user intent rather than the lower-level developer API request shape. Use style for a full, strict visual-style paragraph (medium, texture, palette, then a simple composition lock) and aspectRatio ({ width, height } units, e.g. { width: 16, height: 9 }) for output dimensions. Omitted workflow styles use the app Realistic look (Photorealistic photograph, natural lighting). Do not pass a short label such as watercolor. Image models pack the frame with text, charts, and diagrams unless the style keeps the picture simple. remix_project accepts curated edits: CAPTIONS, TRANSITIONS, CONVERT_IMAGES_TO_VIDEOS, and ZOOM. CONVERT_IMAGES_TO_VIDEOS generates AI video clips from stills (expensive). ZOOM is cheap Ken Burns camera motion. watermarkMode and endScreenMode are not exposed; MCP always sends AUTO (Free-plan output includes the VideoGen watermark; VideoGen Pro removes it).

Workflow, media-generation, and export tools start the operation and poll internally. On the hosted (Streamable HTTP) server, the wait window is capped under Cloudflare's proxy timeout, so long generations return a still-running snapshot instead of a 524. Use get_tool_execution, get_workflow_run, or get_project_export to continue polling.

For generate_avatar, provide audioFileId and an ACTOR entity via actorEntityId. You may set avatarQuality to LOW, STANDARD, HIGH, or MAX. script_to_video, slideshow_to_video, and CHANGE_NARRATOR accept the same optional actor fields.

See the full tool reference for every tool's parameters and its REST-endpoint mapping.

Development

pnpm build           # bundle to dist/ (dist/index.js, dist/http.js, and the widget)
pnpm typecheck       # type-check
pnpm lint:fix        # lint

# Smoke test the local (stdio) server:
VIDEOGEN_API_KEY=sk_videogen_live_... node dist/index.js

# Smoke test the remote (HTTP) server:
PORT=8080 node dist/http.js
#   Health:  curl http://localhost:8080/health
#   MCP:     curl -X POST http://localhost:8080/mcp \
#              -H 'Authorization: Bearer sk_videogen_live_...' \
#              -H 'Content-Type: application/json' \
#              -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

VIDEOGEN_BASE_URL overrides the upstream API for both transports (e.g. point at a local API during development).

Deployment (remote server)

The remote server deploys to Cloud Run through the standard monorepo CI/CD, alongside api / backend / external / frontend:

  • Image: ci-cd/docker/mcp/Dockerfile — a minimal multi-stage Node build (no ffmpeg / gcloud / VPC tooling; the server only makes outbound HTTPS calls to its per-environment VideoGen developer API — dev.api.videogen.io / prerelease.api.videogen.io / api.videogen.io, all of which exist).

  • Service registry: mcp is registered in ci-cd/utils.ts (CLOUD_RUN_SERVICES).

  • Build config: ci-cd/config.build.ts (CLOUD_RUN_SERVICE_TO_BUILD_CONFIG_MAP.mcp).

  • Deploy config: ci-cd/config.deploy.ts (mcp) — 1 CPU / 2 GiB, concurrency 80, 60m request timeout (tool calls long-poll workflows/exports), allowUnauthenticated: true (auth is per-request via the caller's API key, not GCP IAM).

  • Generated pipelines: ci-cd/cloudbuild/mcp.build.yaml and the mcp entries in the GitHub build/deploy workflow matrices are produced by the CI/CD generators — regenerate them (do not hand-edit) after changing the build/deploy config.

Manual GCP steps (one-time)

CI/CD builds and deploys the Cloud Run service, but a couple of steps must be done by hand in GCP the first time:

  1. Custom domain / DNS. The service is reachable at its generated *.run.app URL immediately. To serve it at mcp.videogen.io, create a Cloud Run domain mapping (or add it behind the existing load balancer) and add the corresponding DNS record. Update the url in the client config above once the domain resolves.

  2. Verify the first deploy. After the first successful deploy, confirm GET https://<service-url>/health returns 200 and that a POST /mcp with a valid bearer token lists tools.

ChatGPT connector (OAuth) checklist

After deploying the remote MCP server for an environment (e.g. DEV → https://dev.mcp.videogen.io/mcp):

  1. Confirm discovery + auth modes:

    • GET https://<mcp-host>/.well-known/oauth-protected-resource/mcp returns 200 with resource ending in /mcp and authorization_servers equal to the MCP origin (pathless — Cursor workaround).

    • GET https://<mcp-host>/.well-known/oauth-protected-resource/mcp/chatgpt returns 200 with resource ending in /mcp/chatgpt and authorization_servers equal to the Supabase Auth issuer (…/auth/v1).

    • GET https://<mcp-host>/.well-known/oauth-authorization-server returns 200 with issuer equal to the MCP origin and authorize/token/register endpoints on Supabase Auth.

    • npx -y @modelcontextprotocol/inspector@1.0.0 --cli https://<mcp-host>/mcp/chatgpt --transport http --method tools/list lists ~35+ tools (including open_uploader as noauth and API tools as oauth2).

    • Every anonymous POST against /mcp, including initialize and tools/list, returns HTTP 401 + WWW-Authenticate so Cursor / Claude start OAuth immediately.

    • Anonymous tools/call get_me against /mcp/chatgpt returns HTTP 200 with a tool-result _meta["mcp/www_authenticate"] challenge (ChatGPT).

  2. In the ChatGPT app (e.g. VideoGen (DEV)), set the MCP URL to https://<mcp-host>/mcp/chatgpt (not /mcp). Copy the exact OAuth callback URL (https://chatgpt.com/connector/oauth/{callback_id}) into that environment's Supabase Auth OAuth client redirect allowlist.

  3. Prefer DCR in the ChatGPT connector builder (Supabase advertises registration_endpoint; it does not advertise CIMD / client_id_metadata_document_supported). After ChatGPT DCR's a client for our app, add that exact client_id to TRUSTED_OAUTH_CLIENT_IDS in core/src/logic/oauth/trustedOAuthClientIds.ts so the consent screen treats it as verified.

  4. ChatGPT Settings → the app → Refresh → Scan Tools → complete the OAuth prompt when shown.

  5. Start a new conversation, select the app from the tools menu, and call a read tool (e.g. get_me). Do not reuse an old chat after metadata changes.

STANDARD_TESTS runs anonymous ChatGPT discovery against /mcp/chatgpt automatically via mcp-http (pnpm --filter @videogen/api mcp:test -- --transport http --env <env> --server-mode remote), which invokes the Inspector CLI before the authenticated full smoke and asserts open_uploader advertises noauth and API tools advertise oauth2. That scheme check fails against a host that has not yet been redeployed with this metadata.

Available Tools

49 tools
add_entity_referenceAdd entity referenceAInspect

Attach an uploaded file (vg_file_...) as a reference on an entity. Images work for every entity type. Slideshow themes may also attach a PDF or PowerPoint. Built-in entities cannot have references added. For new PRODUCT/ACTOR entities, call this right after create_entity with isDefault: true so the entity has a usable thumbnail and generation reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileIdYesImage file id (vg_file_...) to attach as a reference. Upload first via upload_file / create_file_upload / open_uploader.
entityIdYesEntity id (vg_enti_...).
isDefaultNoWhen true, make this the primary reference (thumbnail). Defaults to false.
descriptionNoOptional description of this reference image.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesDisplay name.
entityIdYesEntity id (vg_enti_...).
createdAtYesUnix created-at timestamp.
isBuiltInNoTrue for VideoGen catalog entities that cannot be modified.
updatedAtYesUnix updated-at timestamp.
entityTypeYesACTOR, PRODUCT, VISUAL_STYLE, or SLIDESHOW_THEME.
referencesYesAttached reference images.
actorConfigNoVoice/avatar summary for ACTOR entities; null for other types.
descriptionYesDescription (empty when unset).

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the generic mutation profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds real behavioral context beyond that: entity-type restrictions, accepted file formats, and the ordering dependency on create_entity. It stops short of disclosing idempotency, overwrite behavior, or failure modes, so it is strong but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action before constraints and the workflow tip. No filler, and the most important scoping rule (built-in entities excluded) is stated plainly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a non-destructive mutation tool with full schema coverage and annotations, the description covers constraints, file-type rules, and call ordering, which is everything required to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by tying isDefault to a concrete workflow step (setting the usable thumbnail/generation reference right after entity creation), which adds selection meaning beyond the schema's plain 'primary reference (thumbnail)' phrasing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (attach) and resource (uploaded file as a reference on an entity), with the file id prefix (vg_file_...) making the input type unambiguous. It is clearly separable from siblings like remove_entity_reference, upload_file, and create_entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use rules: images work for all entity types, PDF/PowerPoint only for slideshow themes, built-in entities cannot take references, and new PRODUCT/ACTOR entities should call this immediately after create_entity with isDefault: true. This is exactly the routing and prerequisite guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_entityArchive entityA
Destructive
Inspect

Archive an entity so it no longer appears in lists or pickers. Built-in entities cannot be archived.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityIdYesEntity id (vg_enti_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
archivedYesAlways true on success.
entityIdYesArchived entity id.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds useful behavioral context beyond annotations: archived entities disappear from lists and pickers, and built-in entities cannot be archived. It could disclose more about reversibility or permission requirements, but it meaningfully exceeds the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and immediately followed by the most important constraint. Every sentence earns its place and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, destructive mutation tool with full schema coverage, annotations, and an output schema, the description covers the key behavior and a critical caveat. It is nearly complete, though it leaves some operational details unstated, such as whether archiving is reversible or what permissions are required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single required parameter entityId is fully documented in the schema as an entity ID with a vg_enti_ prefix example. The description does not add syntax, format, or meaning beyond what the schema already provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource, 'Archive an entity,' and states the observable effect: it no longer appears in lists or pickers. It also distinguishes archive behavior from sibling retrieval and listing tools by focusing on removal from visibility, making the intent clear without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating the effect and a key exclusion: built-in entities cannot be archived. However, it does not explicitly say when to choose archive_entity over update_entity, remove_entity_reference, or other mutation tools, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_tool_executionCancel tool executionB
Destructive
Inspect

Request cancellation of an in-progress tool execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolExecutionIdYesTool execution id (vg_tool_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation profile is covered. The description adds a modest but real nuance by saying 'Request' cancellation, implying best-effort/asynchronous behavior rather than guaranteed termination, and 'in-progress' bounds which executions are eligible — but it never states what happens to partial output or whether cancellation is final.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is efficient. It is borderline under-specified rather than over-concise — the brevity costs it the usage and behavioral detail it could have carried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the destructive profile. Still, for a mutation/cleanup tool the description lacks timing semantics (immediate vs deferred cancellation), idempotency, and any note on required permissions, leaving an agent with the minimum needed to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter and schema description coverage is 100% (the schema documents the vg_tool_... id format), so the description's single sentence adds no parameter detail. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (cancel) plus specific resource (tool execution), and it scopes the target to an in-progress execution, which distinguishes it from get_tool_execution/list_tool_executions. It does not, however, mention the near-identical sibling cancel_workflow_run, so an agent must infer the boundary from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given. With cancel_workflow_run sitting in the same toolset, the description offers nothing to route between cancelling a tool execution versus a workflow run, and no prerequisites (e.g. needing a live execution id) are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_workflow_runCancel workflow runB
Destructive
Inspect

Request cancellation of an in-progress workflow run.

ParametersJSON Schema
NameRequiredDescriptionDefault
workflowRunIdYesWorkflow run id (vg_work_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this mutates state. The description adds one genuinely useful nuance beyond that: 'Request' implies the cancellation is asynchronous and may not take effect immediately. It still omits idempotency, whether in-flight work is discarded, and permission requirements for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste; the destructive/async action leads. It is arguably under-specified rather than over-long, but as structure goes it is efficient and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool with an output schema and explicit destructive annotations, the description covers the essentials but leaves meaningful gaps: no outcome/return behavior, no note on whether a cancelled run can be resumed, and no routing to sibling cancel/get tools. Adequate, not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the single workflowRunId parameter is documented with its 'vg_work_...' prefix format. The description adds no syntax or constraint detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (cancel) and resource (workflow run) plus a scope qualifier ('in-progress'), which clearly distinguishes it from list/get workflow-run tools. It does not, however, differentiate itself from the nearby cancel_tool_execution sibling, which an agent could plausibly confuse for the same intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus get_workflow_run, list_workflow_runs, or cancel_tool_execution, and no prerequisites or exclusions (e.g., only pending/running runs can be cancelled). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_entityCreate entityAInspect

Create an ACTOR (character), PRODUCT (product/object), VISUAL_STYLE, or SLIDESHOW_THEME entity. After create, attach at least one reference with add_entity_reference (upload the file first). Slideshow themes may attach an image or a PDF / PowerPoint. Use the returned entityId as actorEntityId on generate_avatar / script_to_video, a product/style reference in storyboard scenes, or slideshowThemeEntityId on slideshow_to_video.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesDisplay name for the entity.
entityTypeYesACTOR = consistent character; PRODUCT = product/object; VISUAL_STYLE = look/style for generated images; SLIDESHOW_THEME = shared slide design system for a slideshow deck.
descriptionNoOptional longer description.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesDisplay name.
entityIdYesEntity id (vg_enti_...).
createdAtYesUnix created-at timestamp.
isBuiltInNoTrue for VideoGen catalog entities that cannot be modified.
updatedAtYesUnix updated-at timestamp.
entityTypeYesACTOR, PRODUCT, VISUAL_STYLE, or SLIDESHOW_THEME.
referencesYesAttached reference images.
actorConfigNoVoice/avatar summary for ACTOR entities; null for other types.
descriptionYesDescription (empty when unset).

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the mutation/safety profile is covered structurally. The description adds genuinely useful behavior: creation alone is insufficient (a reference must follow), and the type-to-downstream mapping. It stops short of stating idempotency or failure/limit behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler, front-loaded with the create action followed by the required follow-up step and the downstream consumption. Every clause carries distinct operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers the remaining unknowns: valid types, the mandatory reference attachment, file upload ordering, and where entityId is reused. Nothing an agent needs to call and chain this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so entityType, name, and description are already documented with rich enum comments. The description still adds value by linking each entity type to the downstream parameter it feeds (actorEntityId, product/style reference, slideshowThemeEntityId), which the schema does not say. Slightly above the baseline for that extra mapping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Create an ACTOR... PRODUCT... VISUAL_STYLE... SLIDESHOW_THEME entity') and enumerates exactly what can be created, distinguishing it from sibling mutation tools like update_entity or archive_entity. An agent immediately knows this is the entity-creation entry point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the post-create requirement ('attach at least one reference with add_entity_reference (upload the file first)'), names the prerequisite tool and ordering, and spells out where the returned entityId is consumed (generate_avatar, script_to_video, storyboard scenes, slideshow_to_video). This is a full workflow route, not just a hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_file_uploadCreate file uploadAInspect

Start an upload for a large file, or when file bytes cannot be inlined. Returns { fileId, uploadUrl }. PUT the raw file bytes to uploadUrl with NO Authorization header (it is a short-lived pre-signed URL). Then call get_file with { fileId, wait: true } to wait until processing finishes, and pass the returned fileId to workflows, tools, logos, or B-roll. For small files, prefer upload_file.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoFile type. Inferred after processing when omitted.
displayNameYesDisplay name for the file.
isTemporaryNoWhen true, the file is temporary (guaranteed available for 24 hours, not analyzed for search). Defaults to false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileIdYesFile id to use after uploading bytes (vg_file_...).
uploadUrlYesPre-signed URL: PUT raw file bytes here with no Authorization header.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only establish that this is a non-destructive, non-open-world write. The description adds critical operational behavior the annotations cannot: the return shape { fileId, uploadUrl }, the requirement to PUT raw bytes with NO Authorization header to a short-lived pre-signed URL, and the follow-up get_file call to await processing. That is exactly the extra context that makes invocation succeed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the triggering condition and return contract, then the transport instruction and the follow-up call. Every sentence carries unique operational value; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description fully covers the end-to-end flow (create → PUT bytes → get_file wait → consume fileId), which is what an agent needs to chain this multi-step upload correctly. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (type, displayName, isTemporary) are already documented including enum values and defaults. The description adds no parameter-level detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start an upload for a large file, or when file bytes cannot be inlined') and explicitly distinguishes itself from the sibling upload_file ('For small files, prefer upload_file'). An agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions (large file, or bytes cannot be inlined), an explicit alternative for the opposite case (small files → upload_file), and the required follow-up sequence (PUT bytes, then get_file with wait:true). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_projectExport projectAInspect

Export a project to an MP4 and return its status or download URL. Renders often take a few minutes. Tell the user that wait up front.

ParametersJSON Schema
NameRequiredDescriptionDefault
qualityNoExport quality. Omit for the default.
projectIdYesProject id to export.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileNoHydrated export file metadata.
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdYesExport id (vg_expo_...).
projectIdNoExported project id (present after polling).
downloadUrlNoSigned MP4 download URL when succeeded; otherwise null.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoExported MP4 file id when succeeded.
thumbnailUrlNoSigned thumbnail URL when ready.
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.
thumbnailUrlExpiresAtNoUnix expiry for thumbnailUrl.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the generic safe-mutation profile (not read-only, not destructive, not open-world). The description adds genuinely new behavioral context: the operation is asynchronous, latency is minutes-scale, and the response may be a status rather than a finished file. It does not mention whether the export can be cancelled or re-run, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action, and every clause carries information (return shape, latency, user-facing messaging). The trailing 'Tell the user that wait up front' is slightly awkwardly phrased but earns its place as an instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is correctly omitted, and the annotations cover the safety profile. The async/latency disclosure plus the return-shape hint make this sufficient for correct invocation, though a note about follow-up polling via get_project_export or task listing would have closed the loop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with only two parameters (projectId and an enum quality with its own 'Omit for the default' note), so the schema already carries the semantics. The description adds nothing about parameter meaning; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Export a project to an MP4') plus the return shape ('status or download URL'), which is far more than the bare title. It implicitly distinguishes itself from the sibling get_project_export (which retrieves an existing export) but never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The timing note ('Renders often take a few minutes. Tell the user that wait up front') gives real operational context for a long-running job. However, it offers no guidance on when to choose this tool over get_project_export or the async-task/guidance siblings, so usage selection is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_avatarGenerate avatarAInspect

Generate a talking-head avatar video from an ACTOR entity and an uploaded audio file. Pass actorEntityId and optionally set avatarQuality. Typically takes a few minutes (longer for longer audio). Tell the user that wait up front.

ParametersJSON Schema
NameRequiredDescriptionDefault
audioFileIdYesAudio file id for the avatar to lip-sync.
actorEntityIdYesId of a built-in stock actor or an ACTOR entity (vg_enti_...) with an image reference.
avatarQualityNoAvatar generation quality tier. Applies when actorEntityId is provided. Omit to use workspace settings.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly=false, destructive=false, openWorld=false). The description adds genuinely new behavioral context: generation takes 'a few minutes (longer for longer audio)' and the agent should tell the user about the wait up front. That latency/UX disclosure is exactly what annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then optional parameter, then the latency caveat and user-facing instruction. Nothing is wasted, though the final instruction sentence is more UX coaching than tool semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and annotations carry the safety profile. The latency expectation is supplied, leaving sibling differentiation as the only meaningful gap. Adequate for a 3-parameter async generation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents actorEntityId, audioFileId and the avatarQuality enum tiers. The description only restates that actorEntityId is passed and avatarQuality is optional, adding no format or constraint detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate a talking-head avatar video from an ACTOR entity and an uploaded audio file.' That is clearly distinguishable from generic siblings like generate_video_clip or prompt_to_video_clip, though it never names an alternative to route against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It notes the two required inputs implicitly but gives no when-to-use or when-not-to-use guidance, and does not distinguish itself from the other video-generation siblings such as generate_video_clip or script_to_video. The only routing signal is the input type (actor + audio).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageGenerate imageAInspect

Generate an image from a text prompt, optionally conditioned on source images (image-to-image) and actor, product, or visual-style entity ids. Typically takes 15–60 seconds. Tell the user that wait up front.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of the image to generate.
qualityNoGeneration quality. Omit to use workspace settings.
entityIdsNoOptional actor, product, or visual-style entity ids (vg_enti_...) used as identity/reference.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
imageFileIdsNoOptional reference image file ids.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (not read-only, not open-world, not destructive), so the description's latency disclosure ('typically takes 15–60 seconds') and the explicit instruction to tell the user about the wait add real behavioral value an agent cannot infer from structured fields. It stops short of covering cost, failure modes, or whether generation is cancellable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core purpose and the conditioning modes, followed by the single operational caveat that matters to an agent. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter generative tool with a nested aspect-ratio object and an output schema, the description covers purpose, modes, and latency, and the output schema relieves it of explaining return values. It is nearly complete, though it omits any mention of cost, concurrency, or how generation interacts with workflow/tool-execution siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so prompt, quality, entityIds, aspectRatio, and imageFileIds are already documented in the schema itself. The description restates the conditioning inputs at a high level without adding syntax, constraints (e.g., the 4-image cap), or defaults beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Generate') plus resource ('an image') with the two operating modes spelled out: text-to-image and image-to-image conditioned on source images/entity ids. This distinguishes it from siblings like generate_video_clip, generate_avatar, and upscale_image without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by describing optional conditioning modes, but never states when to prefer this over siblings (e.g., vs. generate_avatar for actor likeness, or upscale_image for refinement). No exclusions or prerequisites are given; the only imperative is the latency warning to relay to the user.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_motion_graphicGenerate motion graphicAInspect

Generate an animated motion graphic video from a text prompt. Best for precise text animations (typing effects, kinetic typography, lower thirds) that stock or generated footage can't express. Outputs a transparent WebM overlay by default; set transparentBackground to false for an opaque MP4. Optionally pass reference media file ids to display or animate. Typically takes 2–5 minutes because VideoGen writes animation code and then renders it; complex prompts can take longer. Tell the user that wait before starting and keep polling calmly — a healthy in-progress job is expected.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of the animated motion graphic.
fileIdsNoOptional reference media file ids.
entityIdsNoOptional actor, product, or visual-style entity ids (vg_enti_...).
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
subToolModesNoOptional per-capability controls for generated images, video clips, voiceover, and stock media. Omit to use AUTO for every capability.
durationSecondsNo
transparentBackgroundNoRender a transparent WebM for use as an overlay on other video or images. Defaults to true. Set to false for an opaque MP4.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations covering safety (readOnlyHint=false, destructiveHint=false), the description adds rich behavioral disclosure: default transparent WebM output vs opaque MP4, 2–5 minute render time because VideoGen writes animation code then renders, longer for complex prompts, and polling expectations. This is exactly the extra context the annotation set cannot provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then gives use-case, output format, references, timing, and polling guidance in tight sentences. Every sentence earns its place; no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 7-param tool with nested objects, the description covers use case, output format, timing expectations, and polling behavior, and an output schema exists so return values needn't be explained. The few uncovered params are fully documented in the schema, leaving no critical gap for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents most parameters (including the transparentBackground default and aspectRatio semantics). The description reinforces transparentBackground and fileIds ('reference media file ids to display or animate') but omits entityIds, subToolModes, and durationSeconds. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Generate an animated motion graphic video from a text prompt') and immediately distinguishes it from siblings by scoping it to precise text animations (typing effects, kinetic typography, lower thirds) that stock/generated footage can't express. An agent can tell this apart from generate_video_clip or prompt_to_video_clip without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use guidance ('Best for precise text animations...') and contrasts implicitly against stock/generated footage alternatives. It also adds operational guidance (wait 2–5 minutes, keep polling calmly). It stops short of naming specific sibling tools or explicit when-not-to-use conditions, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_musicGenerate musicAInspect

Generate a music track from a text prompt. Typically takes 1–5 minutes depending on track length. Tell the user that wait up front.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesGenre, mood, instrumentation, and tempo.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false); the description adds the critical long-running latency trait and an explicit UX instruction about surfacing the wait. It does not clarify whether the call blocks or returns an async handle.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the latency constraint, then the actionable instruction to the agent. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need no explanation, and the latency trait is disclosed. The main remaining gap is whether this is an async/task-based call that later surfaces via list_tool_executions, which matters given the sibling task-management tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter with 100% schema coverage, and the schema description ('Genre, mood, instrumentation, and tempo') is actually richer than the description's generic 'text prompt'. Baseline 3 applies since the schema already does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (music track) from a text prompt, which cleanly separates it from siblings like generate_sound_effect, text_to_speech, and generate_video_clip. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives operational guidance (1-5 minute wait, tell the user up front) but no tool-selection guidance explaining when to prefer this over generate_sound_effect or other audio siblings. Usage is implied by the resource name rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_sound_effectGenerate sound effectCInspect

Generate a sound effect from a text prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of the sound effect.
durationSecondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnlyHint=false, openWorldHint=false, destructiveHint=false), and the description adds nothing beyond that: no latency, cost, async/job behavior, or output format context. For a generative tool with an output schema it could still note operational traits, but it repeats only the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no waste, and the key information is front-loaded. It is appropriately sized, though it is so short that it borders on under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and the tool only has two parameters, so the description needn't explain much. However, given three audio-generating siblings and an undocumented parameter, an agent lacks the routing and parameter guidance needed to call this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: the prompt parameter is documented in the schema and echoed by the description ('text prompt'), but durationSeconds has no description anywhere and the description does not compensate with range or default information. The description adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate a sound effect') plus the input modality ('from a text prompt'), so the agent knows exactly what the tool produces. It does not differentiate itself from the closely related siblings generate_music and text_to_speech, which is the only gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over generate_music or text_to_speech, despite both being present in the sibling list and producing audio. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_video_clipGenerate video clipAInspect

Generate a video clip from a text prompt, source images, source videos, spokenDialogue, or reference audio. quality is optional (LOW, STANDARD, HIGH, or MAX). Typically takes 1–3 minutes (HIGH/MAX can be longer). Tell the user that wait up front and keep polling calmly.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptNoDescription of the video to generate.
qualityNoVideo generation quality. Omit to use workspace settings.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
audioFileIdsNoOptional reference audio file ids used to lip-sync from a recording. Use spokenDialogue to have the model speak a line it generates itself.
imageFileIdsNoOptional reference image file ids.
videoFileIdsNoOptional reference video file ids.
generateAudioNoWhether the result must include generated audio.
spokenDialogueNoExact line the subject should speak as native lip-synced speech. The model synthesizes the voice from this text.
durationSecondsNoOptional clip length in whole seconds (1 to 30). Omit or pass null for Auto.
startFrameFileIdNoOptional opening-frame still file id. Used as the first frame of the clip.
voiceDescriptionNoNatural-language description of the voice that speaks spokenDialogue. Used when spokenDialogue is set.
suppressBackgroundMusicNoWhen true, the clip will not include a musical soundtrack. Spoken dialogue and environmental sound are still allowed. Use this when background music will be added separately.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (not read-only, not destructive, closed-world), so the bar is lower. The description adds genuinely non-structured behavior: expected latency of 1-3 minutes with HIGH/MAX running longer, and the polling expectation. It could go further on failure/retry behavior, but this is solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then optionality, then operational timing. The final 'keep polling calmly' sentence is informal but earns its place as behavioral guidance. No significant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, all-optional generation tool with a full schema and an output schema, the description supplies the missing non-schema essentials: input modalities and expected latency. Sibling disambiguation is the one notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (quality enum, aspectRatio units, audioFileIds vs spokenDialogue, durationSeconds, suppressBackgroundMusic) is already documented in the schema. The description repeats the quality enum values and names some input types without adding format or interaction detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (video clip) and enumerates the accepted input modalities (text prompt, images, videos, spokenDialogue, reference audio). It does not, however, distinguish itself from the sibling prompt_to_video_clip or storyboard_to_video, which appear to overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives practical operating guidance (expect 1-3 minutes, keep polling calmly, warn the user up front), which is useful when-context. But it never states when to pick this tool over prompt_to_video_clip, storyboard_to_video, or script_to_video, so routing between siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_async_tasks_guidanceAsync tasks guidanceA
Read-only
Inspect

How to handle async workflows, tool executions, and exports: statuses, hosted vs local wait caps, polling with get_* tools, and when webhooks apply outside MCP. Call before polling or when a start tool returns a still-running snapshot. Equivalent to reading the guidance://async-tasks MCP resource — provided as a tool for clients that don't read resources directly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
markdownYesFull guidance document in markdown.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds real context beyond that: it enumerates the content covered (statuses, hosted vs local wait caps, polling, webhook scope) and discloses that it is an equivalent substitute for the `guidance://async-tasks` resource for clients that don't read resources. It says nothing about the shape of what is returned, though an output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: scope first, trigger second, resource equivalence last. No restatement of the title and every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter documentation tool with an output schema, the description supplies purpose, trigger, and content coverage. Return-value detail is correctly left to the output schema, so nothing necessary for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case. The description correctly implies a parameterless, read-only lookup and adds no misleading parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact subject matter (async workflows, tool executions, exports: statuses, wait caps, polling, webhooks) and frames the tool as documentation rather than an action, which cleanly separates it from the action siblings like list_tool_executions and get_tool_execution. It never distinguishes itself from the other guidance siblings (get_workflows_guidance, get_tools_vs_workflows_guidance, get_getting_started_guidance), so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is stated as an explicit trigger: "Call before polling or when a start tool returns a still-running snapshot." That is a concrete when-to-use condition tied to observable agent state, not an implication the agent has to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_entityGet entityA
Read-only
Inspect

Fetch one entity by id, including its reference images. Built-in catalog entities are included and have isBuiltIn true.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityIdYesEntity id (vg_enti_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesDisplay name.
entityIdYesEntity id (vg_enti_...).
createdAtYesUnix created-at timestamp.
isBuiltInNoTrue for VideoGen catalog entities that cannot be modified.
updatedAtYesUnix updated-at timestamp.
entityTypeYesACTOR, PRODUCT, VISUAL_STYLE, or SLIDESHOW_THEME.
referencesYesAttached reference images.
actorConfigNoVoice/avatar summary for ACTOR entities; null for other types.
descriptionYesDescription (empty when unset).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds useful return context (reference images are bundled, built-in entities are included and flagged), which is genuine value beyond annotations, but it says nothing about error behavior for unknown ids or auth needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by scope notes. Nothing is wasted, though the second sentence is more about return shape than invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it still volunteers the isBuiltIn/reference-image behavior. For a single-param read tool this is essentially complete; only usage routing to list_entities is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter at 100% schema description coverage, so the schema already documents entityId and its vg_enti_ prefix format. The description's 'by id' adds no syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (one entity by id), and adds scope detail: reference images included, built-in catalog entities returned with isBuiltIn true. It implicitly separates itself from list_entities by saying 'one entity', but never names the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'fetch one entity by id' — the agent can infer this is the single-item lookup versus list_entities. There is no explicit when-to-use guidance, no prerequisites, and no exclusions (e.g., archived entities, permissions).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_fileGet fileA
Read-only
Inspect

Fetch a file by id with freshly hydrated (non-expired) signed URLs for its thumbnail, preview, and download renditions. Set wait: true to poll until the file finishes processing — use this right after PUTting bytes to a create_file_upload URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoWhen true, poll until the file finishes processing and a rendition is ready. Defaults to false (a single fetch).
fileIdYesFile id (vg_file_...).
timeoutMsNoMaximum time to wait for processing before giving up, in milliseconds.
pollIntervalMsNoHow often to poll while waiting, in milliseconds.

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeNoFile type once processing has determined it.
scopeNoFile scope.
fileIdYesFile id (vg_file_...).
hlsSourceNo
descriptionNoFile description when analyzed.
displayNameNoDisplay name for the file.
downloadUrlNoSigned download URL when ready.
publicHlsUrlNo
thumbnailUrlNoSigned thumbnail URL when available.
previewSourceNo
downloadSourceNo
sourceToolTypeNo
transcriptTextNoPlain transcript text when available.
durationSecondsNoDuration in seconds for video/audio; null for images.
thumbnailSourceNo
publicPlaybackIdNo
downloadUrlExpiresAtNoUnix expiry for downloadUrl.
sourceToolExecutionIdNo
thumbnailUrlExpiresAtNoUnix expiry for thumbnailUrl.
isPublicPreviewEnabledNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/non-destructive/non-openWorld, so the safety profile is covered. The description adds genuinely useful behavior beyond that: URLs are 'freshly hydrated (non-expired)', and wait mode polls with a give-up condition implied by timeoutMs. It does not describe error behavior when a file id is unknown, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what is fetched and what it returns, then the wait/escalation guidance. No filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and it covers the key usage sequence (post-upload hydration + polling). Minor gap: no note on auth/permission requirements or not-found handling, but for a read tool with full schema and annotations this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in-schema, which sets the baseline at 3. The description reinforces wait's semantics ('poll until the file finishes processing') but adds nothing about timeoutMs or pollIntervalMs beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (file by id), and goes further by naming exactly what comes back: freshly hydrated signed URLs for thumbnail, preview, and download renditions. This clearly separates it from list_files and the upload/create sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Set wait: true to poll until the file finishes processing — use this right after PUTting bytes to a create_file_upload URL.' It names the sibling and the sequencing condition that determines when to use this tool and when to use wait mode.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_getting_started_guidanceGetting started guidanceA
Read-only
Inspect

How to authenticate, verify with get_me, introduce VideoGen with workflow example prompts, choose workflows vs media tools, and follow VideoGen id conventions. Call when connecting, onboarding, the user asks what VideoGen can do, or how to set up the API or MCP. Equivalent to reading the guidance://getting-started MCP resource — provided as a tool for clients that don't read resources directly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
markdownYesFull guidance document in markdown.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/non-destructive/non-open-world, so the bar is lower, and the description still adds meaningful context: it is equivalent to the guidance://getting-started resource and exists as a fallback for clients that cannot read resources directly. That explains why the tool exists and why invoking it has no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, front-loaded sentences with no filler; the content list comes first and the resource-equivalence rationale second. The first sentence is a long enumeration but every item carries information, so it stays within acceptable length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only guidance tool with an output schema already defining the return shape, the description covers purpose, triggers, and the resource-fallback rationale. Nothing an agent needs in order to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema imposes no semantic burden and the description correctly says nothing about inputs. Baseline of 4 applies since there is nothing for the description to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Enumerates concrete content (authentication, get_me verification, workflow example prompts, workflows vs media tools, ID conventions) with a clear verb+resource framing, so an agent knows exactly what payload it gets. It does not, however, differentiate itself from overlapping siblings like get_workflows_guidance and get_tools_vs_workflows_guidance, which cover adjacent topics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions: connecting, onboarding, 'what can VideoGen do', setting up the API or MCP. That is strong when-to-use coverage. What is missing is when-not-to-use and routing against the sibling guidance tools that cover a subset of this material.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_meGet accountA
Read-only
Inspect

Fetch the account and team behind the API key (apiKeyId, apiKeyNickname, email, displayName, teamId). Use it as a connection test.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
emailYesAccount email.
teamIdYesTeam id the API key belongs to.
apiKeyIdYesId of the API key used for this request.
displayNameYesAccount display name, or null.
apiKeyNicknameYesNickname of the API key.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that the identity is derived from the API key rather than parameters, which is useful, but it says nothing about rate limits, error behavior, or caching. Baseline 3 with annotations carrying the load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and scoped with a usage note. No filler, no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description only needs to convey purpose and when to call it, which it does. Listing the returned fields is mildly redundant against the output schema but harmless; nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The field names listed in the description (apiKeyId, email, teamId, etc.) are return values already covered by the output schema, not parameters, so they add no parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (the account and team behind the API key), and names the returned identity fields. No sibling tool overlaps with this purpose, so an agent can identify it immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives a use case: 'Use it as a connection test.' That is a clear context for invocation. It does not, however, describe when NOT to use it or any alternative, though no sibling competes for this function.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_projectGet projectB
Read-only
Inspect

Fetch metadata and the shareable URL for a single project.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesProject id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
titleYesProject title.
statusYesHigh-level project status.
createdAtYesUnix timestamp when the project was created.
projectIdYesProject id (vg_proj_...).
updatedAtYesUnix timestamp when the project was last updated.
projectUrlYesDeep link to open the project in the VideoGen editor.
aspectRatioYes
assistantIdYesAssistant conversation id, or null for older projects.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds that a shareable URL is returned, which is mild extra context, but nothing about permissions, caching, or error behavior beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that is front-loaded with the verb and resource. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and annotations carry the safety profile, so the remaining need is minimal. It is adequate for a simple one-parameter read, though it does not mention what to do when a project is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One required parameter with 100% schema description coverage, so the schema already documents projectId. The description adds only that the target is a 'single project', which does not meaningfully extend the schema's semantics; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Fetch metadata and the shareable URL for a single project'), making scope clear. It implies differentiation from list_projects by saying 'single project' but never names the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., ID format or where to obtain projectId), and no routing to alternatives such as list_projects or export_project. Usage has to be inferred entirely from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_exportGet project exportA
Read-only
Inspect

Fetch the current status of a project export. Poll until status is succeeded, failed, or cancelled.

ParametersJSON Schema
NameRequiredDescriptionDefault
exportIdYesExport id (vg_expo_...) from export_project.
projectIdYesProject id that owns the export.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileNoHydrated export file metadata.
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdYesExport id (vg_expo_...).
projectIdNoExported project id (present after polling).
downloadUrlNoSigned MP4 download URL when succeeded; otherwise null.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoExported MP4 file id when succeeded.
thumbnailUrlNoSigned thumbnail URL when ready.
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.
thumbnailUrlExpiresAtNoUnix expiry for thumbnailUrl.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, non-destructive, and non-open-world, so safety is settled. The description adds real behavioral value beyond that: this is a polling endpoint with three named terminal states, which tells the agent when to stop calling it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded and the polling instruction immediately after. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering the safety profile, the description needs only to explain the polling loop and terminal conditions, which it does explicitly. Nothing required to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (exportId, projectId) are already documented, including the exportId format and its origin. The description adds no additional parameter meaning; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the current status of a project export') and implicitly distinguishes itself from the sibling export_project by scoping to status retrieval rather than export creation. An agent can tell what it returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operational guidance: poll repeatedly until status reaches 'succeeded', 'failed', or 'cancelled'. It doesn't name export_project as the prerequisite alternative, though the schema's exportId description ('from export_project') covers the pairing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tool_executionGet tool executionB
Read-only
Inspect

Fetch the current status and results of a single tool execution.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolExecutionIdYesTool execution id (vg_tool_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that the result reflects 'current' status, hinting at an in-progress vs completed state, but says nothing about polling cadence, terminal states, or error cases. Modest added value on top of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the verb and resource with zero filler. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values, and the single parameter is fully documented in the schema. For a simple read tool with full annotation coverage, this is nearly complete; only the polling/terminal-state behavior an agent would want is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, and the schema documents it at 100% coverage including the vg_tool_... id format. The description adds no further detail about the identifier's origin or source. Baseline 3 is appropriate when the schema fully carries parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (a single tool execution) plus what is returned (current status and results). The word 'single' implicitly separates it from the sibling list_tool_executions, though the full set of siblings is never named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance. The agent must infer that this is for polling or checking one execution, and nothing tells it when to prefer this over list_tool_executions or cancel_tool_execution. The polling context is only implied by 'current status'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tools_vs_workflows_guidanceTools vs workflows guidanceA
Read-only
Inspect

When to use standalone media tools (automatic model routing for a single asset) versus workflows (full editable professional video). Call when choosing between generate_* tools and workflow tools. Equivalent to reading the guidance://tools-vs-workflows MCP resource — provided as a tool for clients that don't read resources directly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
markdownYesFull guidance document in markdown.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only, non-destructive, and closed-world. The description adds useful context by stating it is equivalent to the guidance://tools-vs-workflows resource and explaining why it exists as a tool, though it does not detail output format or other operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences, front-loading the core purpose first, then usage guidance, then resource equivalence. Every sentence adds value with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and an existing output schema, the description covers purpose, usage, and resource equivalence adequately. The only notable gap is not explicitly distinguishing this guidance tool from the other guidance siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The description appropriately does not waste space explaining parameters that do not exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: guidance on choosing between standalone generate_* tools and workflow tools. It distinguishes the two categories well, but does not differentiate itself from sibling guidance tools like get_workflows_guidance or get_getting_started_guidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Call when choosing between generate_* tools and workflow tools,' providing clear usage context. However, it offers no when-not-to-use guidance and does not route the agent away from other guidance tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workflow_runGet workflow runA
Read-only
Inspect

Fetch the current status and result of a single workflow run.

ParametersJSON Schema
NameRequiredDescriptionDefault
workflowRunIdYesWorkflow run id (vg_work_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds that the response contains the run's 'current status and result', but says nothing about freshness, waiting/polling behavior, or error cases for a run that does not exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with zero filler; the verb, scope, and returned content all land in the first clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not enumerate return fields, and a single-parameter read tool needs little more than this. It is adequately complete, though a note on which sibling to use when the run id is unknown would close the last gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter has 100% schema description coverage including the id format hint ('vg_work_...'), so the schema carries the parameter documentation. The description adds no further meaning about the id or its alternatives, which is the expected baseline when coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Fetch the current status and result of a single workflow run'), making the operation unambiguous. The word 'single' implicitly separates it from the sibling list_workflow_runs, but the split is left for the agent to infer rather than named outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'a single workflow run' (versus a list), which points the agent toward this tool when it already has an id. There is no explicit statement of when to prefer this over list_workflow_runs or any prerequisite/scoping guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workflows_guidanceWorkflows guidanceA
Read-only
Inspect

Canonical run → remix → export flow and when to use each workflow tool. Prefer script_to_video for longer narrated / informational videos; keep storyboard_to_video to ≤3 scenes unless the user asks for more. Call before starting a full video project. Equivalent to reading the guidance://workflows MCP resource — provided as a tool for clients that don't read resources directly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
markdownYesFull guidance document in markdown.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely useful behavior context: it is equivalent to the guidance://workflows MCP resource and exists as a fallback for clients that cannot read resources directly, which explains duplication an agent might otherwise find confusing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences: the routing content first, then the usage constraint, then the resource-equivalence footnote. No filler, and the most actionable material is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain the returned guidance content, and annotations cover the safety profile. It is complete enough to trigger correct use, though a hint about the scope of the guidance (which tools the flow covers) would close the last gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document beyond what the empty schema already shows. Baseline for a parameterless tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete resource (the canonical run → remix → export flow) and states its role as routing guidance across workflow tools, which separates it from generation siblings. It could be sharper about how it differs from get_getting_started_guidance or get_tools_vs_workflows_guidance, which are the closest alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Call before starting a full video project') plus concrete selection rules between two named siblings (script_to_video for longer narrated videos, storyboard_to_video ≤3 scenes). An agent knows both the trigger and the downstream routing decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image_3d_effectImage 3D effectAInspect

Add 3D parallax motion to a still image, producing a video.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileIdYesSource image file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is not read-only, not destructive, and not open-world. The description adds that the output is a video, which is useful but does not cover asynchronous behavior, expected processing time, or any file/asset lifecycle implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action and result, with no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter effect tool with an output schema and explicit annotations, the description is largely sufficient to invoke correctly. It could be more complete by mentioning whether the generated video is synchronous or asynchronous and by routing users away from similar sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the single parameter is already fully documented as 'Source image file id.' The description adds no extra meaning or constraints beyond what the schema provides, which is appropriate but not enriching.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and transformation: 'Add 3D parallax motion to a still image, producing a video.' It clearly distinguishes the tool from generic image or video generation siblings by specifying the input type and the effect applied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case but gives no explicit when-to-use guidance, prerequisites, or alternatives. It does not tell the agent when this effect is preferable to sibling tools like generate_video_clip or generate_motion_graphic.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_entitiesList entitiesA
Read-only
Inspect

List built-in actors, products, visual styles, and slideshow themes plus team entities. Built-in rows have isBuiltIn true and cannot be updated or archived. Filter with entityType when you only need one kind.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
cursorNoPagination cursor from a previous response's `nextCursor`.
entityTypeNoWhen set, only return entities of this type.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
entitiesYesEntities visible to the API key.
nextCursorYesCursor for the next page, or null.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds meaningful context beyond them: built-in rows carry isBuiltIn=true and cannot be updated or archived, telling the agent which results are immutable. It does not discuss pagination behavior, but the safety/immutability disclosure is real added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with what is listed, followed by the immutability caveat and the filter hint. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the entity kinds, the built-in immutability constraint, and filtering. Pagination is left entirely to the schema, which is acceptable, leaving only a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, cursor, and entityType are all documented in the schema. The description only restates the entityType filter condition, adding no format or syntax detail beyond the schema — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific verb (list) and enumerates the resource kinds it returns: built-in actors, products, visual styles, slideshow themes, plus team entities. This clearly separates it from the single-item siblings (get_entity, create_entity), though 'team entities' is left slightly undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers one conditional hint — 'Filter with entityType when you only need one kind' — which is genuine usage guidance. It does not state when to prefer get_entity over listing, nor any prerequisites, so guidance is implied rather than complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_filesList filesB
Read-only
Inspect

List files visible to the current API key.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
cursorNoPagination cursor from a previous response's `nextCursor`.

Output Schema

ParametersJSON Schema
NameRequiredDescription
filesYesFiles visible to the API key.
hasMoreYesWhether another page is available.
nextCursorYesCursor for the next page, or null.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — results are scoped to the current API key's visibility/permissions — but says nothing about pagination behavior or result ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. Every word carries information (action, resource, scope).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema, full annotation coverage, and complete parameter documentation, the only thing left for the description is scope, which it provides. Adequate, though a nod to pagination or sibling routing would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit and cursor are already documented in the schema. The description adds no parameter meaning beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List files") and adds scope ("visible to the current API key"). It is clear on its own, but does not distinguish itself from the file-related siblings get_file, upload_file, or create_file_upload.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of alternatives, even though get_file exists as a natural sibling for fetching a single file. The API-key scoping hints at breadth but is not framed as a selection criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_languagesList languagesA
Read-only
Inspect

List supported languages for narration and captions.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoOptional case-insensitive substring filter across each item's searchable text fields.

Output Schema

ParametersJSON Schema
NameRequiredDescription
languagesYesSupported languages.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds only that the result is the set of supported languages, which is thin context beyond the name; it says nothing about caching, stability, or auth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource and its scope front-loaded. Every word earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and there are no required parameters. Combined with the annotations, everything needed to call the tool correctly is available; only the absence of sibling disambiguation keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single optional 'query' parameter, which is documented as a case-insensitive substring filter. The description adds no additional meaning about the filter, so the schema-describes-everything baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List supported languages') and adds scope ('for narration and captions'), which tells the agent exactly what the enumeration covers. It does not, however, differentiate itself from the nearby sibling list_tts_voices, which an agent could plausibly confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for narration and captions' implies the usage context, so an agent can infer this is a lookup for valid language codes. There is no explicit when-to-use statement, no mention of the alternative list_tts_voices, and no exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_project_remix_actionsList project remix actionsB
Read-only
Inspect

List the status of remix actions applied to a project.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesProject id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
nextCursorYesCursor for the next page, or null.
remixActionsYesRemix actions for the project.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered without the description. The description adds only that the payload is per-project status information, with no note on ordering, scope of 'actions', or whether results can be empty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, stating resource and scope immediately. Nothing is redundant or misplaced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the single parameter is schema-documented. Still missing is any indication of what constitutes a 'remix action' or how results are bounded, which matters for an agent deciding whether this call answers its question.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter (projectId) and schema description coverage is 100%, so the schema already documents it fully. The description merely implies the project scoping and adds no format or identifier guidance beyond that, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb ('List'), resource ('remix actions'), and scope ('applied to a project'), plus the returned facet ('status'). It is distinguishable from the sibling remix_project, which creates rather than reports, though the description never states that contrast explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use condition, no prerequisite, and no pointer to an alternative tool. An agent must infer from the name alone that this is the retrieval counterpart to remix_project.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList projectsA
Read-only
Inspect

List projects. API-created projects only by default; pass includeUiProjects to also include dashboard projects.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
cursorNoPagination cursor from a previous response's `nextCursor`.
selfOnlyNoWhen true, restrict results to items created by the API key owner rather than the whole team.
includeUiProjectsNoInclude projects created in the VideoGen dashboard, not just API-created ones.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
projectsYesProjects, most recently updated first.
nextCursorYesCursor for the next page, or null when hasMore is false.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds a genuinely useful behavioral default (UI projects silently excluded unless opted in) that the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the default-scope constraint front-loaded before the opt-in escape hatch. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values and pagination cursor need not be described. The description covers the one non-obvious behavior (default filtering). Minor omission: no note on result ordering or team-vs-owner scope, though selfOnly is in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (limit, cursor, selfOnly, includeUiProjects) are already documented in the schema. The description only elaborates on includeUiProjects, adding nothing beyond the schema's own wording. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List projects') and immediately clarifies scope: API-created projects only by default. That scope note distinguishes it from a naive reading, though it does not explicitly contrast with siblings like get_project or list_project_remix_actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives actionable conditional guidance: default returns API-created projects, and passing includeUiProjects widens to dashboard projects. It lacks an explicit 'use get_project for a single project' exclusion, but the condition that changes behavior is spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tool_executionsList tool executionsB
Read-only
Inspect

List past tool executions, most recent first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
cursorNoPagination cursor from a previous response's `nextCursor`.
selfOnlyNoWhen true, restrict results to items created by the API key owner rather than the whole team.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
nextCursorYesCursor for the next page, or null.
toolExecutionsYesTool executions, most recent first.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds the ordering guarantee ('most recent first'), which is real behavioral context, but says nothing about pagination behavior, result size, or retention limits despite a cursor parameter existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the verb and resource, with the ordering caveat appended where it matters. No padding, though it is arguably too terse to carry structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a 100%-covered schema, an output schema, and annotations covering the safety profile, the omission of pagination and scope-filter explanation is acceptable. The description supplies only the ordering guarantee, which is the one behavior not expressed in the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, cursor and selfOnly are already fully documented in the schema. The description adds no syntax, defaults, or interpretation for any of them, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List past tool executions') plus the ordering ('most recent first'), which is genuinely useful. It does not, however, distinguish itself from close siblings like get_tool_execution or cancel_tool_execution, so an agent gets no help disambiguating within this family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'past' implies this is for historical/completed executions rather than in-flight ones, which is the only usage signal present. There is no statement of when to prefer this over get_tool_execution or list_workflow_runs, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tts_voicesList text-to-speech voicesB
Read-only
Inspect

List available text-to-speech voices for narration, text_to_speech, and workflows.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
queryNoOptional case-insensitive substring filter across each item's searchable text fields.
cursorNoPagination cursor from a previous response's `nextCursor`.
includeDeprecatedVoicesNoInclude deprecated voices in the results.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
ttsVoicesYesAvailable text-to-speech voices.
nextCursorYesCursor for the next page, or null.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered elsewhere. The description itself adds no behavioral context such as filtering scope, deprecation handling, or pagination behavior, so it contributes almost nothing on this dimension.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though arguably so terse that it leaves room for a little more useful scope detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering the safety profile, the description does not need to explain return values. It is minimally adequate for a simple list tool, but omits filtering/pagination cues that would round out the picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with four well-documented parameters (limit, query, cursor, includeDeprecatedVoices), so the schema already carries parameter meaning. The description adds no additional parameter semantics, which is the expected baseline 3 when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb (List) with a specific resource (text-to-speech voices) and names the consuming features (narration, text_to_speech, workflows). It is distinct from siblings like list_languages or list_files, though it doesn't explicitly contrast itself with any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentioning narration/text_to_speech/workflows implies when an agent would want voices, but there is no explicit when-to-use or when-not-to-use guidance and no named alternative tool. Usage is only weakly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workflow_runsList workflow runsB
Read-only
Inspect

List workflow runs, most recent first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
cursorNoPagination cursor from a previous response's `nextCursor`.
selfOnlyNoWhen true, restrict results to items created by the API key owner rather than the whole team.

Output Schema

ParametersJSON Schema
NameRequiredDescription
hasMoreYesWhether another page is available.
nextCursorYesCursor for the next page, or null.
workflowRunsYesWorkflow runs, most recent first.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false. The description adds the ordering behavior ('most recent first'), which is useful context, but does not discuss pagination behavior, default limits, or scope restrictions beyond what the schema covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It conveys the core action and ordering immediately, which is appropriately concise for a simple list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema and annotation coverage, the description need not explain return values or safety guarantees. However, it omits any scoping or filtering context and does not help an agent distinguish this tool from get_workflow_run, leaving a clear gap for sibling selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all three parameters documented in the schema. The description adds no additional parameter semantics, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('workflow runs') and adds a scope detail ('most recent first'). It does not explicitly distinguish itself from sibling tools like get_workflow_run or list_tool_executions, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternative-tool routing are provided. The description only implies that this is for listing, without explaining when an agent should choose it over get_workflow_run or other list tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prompt_to_video_clipPrompt to videoBInspect

Generate one short AI video clip from a prompt inside an editable project.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of the short video clip.
qualityNoVideo generation quality. Omit to use workspace settings.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
imageFileIdsNoOptional reference image file ids.
durationSecondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is already covered. The description adds that exactly one clip is produced and that it lands in an editable project, but says nothing about generation latency, cost, or the fact that the run can block until an MP4 is ready (a behavior that only appears in the autoExport parameter's schema text).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though its brevity is partly the source of the definition's gaps rather than a virtue of tight writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the mutation profile. But for a six-parameter, nested-object generation tool with an ambiguous sibling, the description omits the usage routing and the long-running/autoExport behavior that an agent needs before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so most parameters are already documented in the schema and the baseline of 3 applies. The description adds nothing about duration, quality tiers, aspect ratio units, or the autoExport workflow, and does not compensate for the one undocumented parameter (durationSeconds).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Generate one short AI video clip from a prompt') and adds a scope qualifier ('inside an editable project'), which is more than a restatement of the name. However, it does not distinguish itself from the near-identical sibling 'generate_video_clip', nor from the other prompt-to-video routes (script_to_video, storyboard_to_video), leaving an agent unable to tell them apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no alternative is named despite several siblings that also produce video from input. The agent is left to infer that this is the one-shot, prompt-driven path versus the storyboard/script paths.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remix_projectRemix projectAInspect

Apply curated edits to an existing project. CONVERT_IMAGES_TO_VIDEOS generates AI video clips from every still and is expensive. Use ZOOM for cheap Ken Burns camera motion. Poll with list_project_remix_actions for status.

ParametersJSON Schema
NameRequiredDescriptionDefault
editsYesCurated edits to apply in order. CONVERT_IMAGES_TO_VIDEOS generates AI video clips from every still and is expensive. Use ZOOM for cheap Ken Burns camera motion.
projectIdYesProject id to edit.
saveAsNewProjectNoApply edits to a copy of the project.

Output Schema

ParametersJSON Schema
NameRequiredDescription
projectIdYesEdited project id (or the duplicate when saveAsNewProject).
projectUrlYesDeep link to the project in the VideoGen editor.
remixActionIdsYesRemix action ids, one per requested action in order.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark the tool as non-read-only and non-destructive, so the mutation risk is already structured. The description adds important operational context beyond annotations: CONVERT_IMAGES_TO_VIDEOS is expensive, and the operation is asynchronous enough to require polling with list_project_remix_actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences are front-loaded with the core action, then a cost warning and the polling instruction. Every sentence earns its place, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and parameter descriptions are complete, so return values need not be explained. The description covers cost and polling status, but it leaves implicit how remix_project relates to other project-creation siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains each edit enum value in detail, including the same expensive/cheap distinction. The description repeats some of that meaning but adds no syntax or format details beyond the structured schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: "Apply curated edits to an existing project." It also names the polling sibling, list_project_remix_actions, but does not explicitly distinguish remix_project from other project-creation or editing siblings like storyboard_to_video or update_entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear guidance on choosing between edit types: CONVERT_IMAGES_TO_VIDEOS is expensive, while ZOOM is cheap Ken Burns motion. It also says to poll with list_project_remix_actions for status, but it does not state when to use remix_project instead of alternative video-generation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_entity_referenceRemove entity referenceAInspect

Detach a reference image from an entity by file id. Built-in entities cannot have references removed.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileIdYesReference image file id (vg_file_...) to remove.
entityIdYesEntity id (vg_enti_...).

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesDisplay name.
entityIdYesEntity id (vg_enti_...).
createdAtYesUnix created-at timestamp.
isBuiltInNoTrue for VideoGen catalog entities that cannot be modified.
updatedAtYesUnix updated-at timestamp.
entityTypeYesACTOR, PRODUCT, VISUAL_STYLE, or SLIDESHOW_THEME.
referencesYesAttached reference images.
actorConfigNoVoice/avatar summary for ACTOR entities; null for other types.
descriptionYesDescription (empty when unset).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the mutation-not-read profile is covered structurally. The description adds a real behavioral constraint (built-in entity restriction) beyond the annotations. It does not state what happens if the file id is not attached, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler. The core action is front-loaded and the failure constraint follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and both params are fully documented. The entity-type restriction is the key extra context; only idempotency/nonexistent-reference behavior is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both parameters documented (fileId vg_file_..., entityId vg_enti_...). The description's 'by file id' adds nothing beyond what the schema already states, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Detach), resource (reference image), and mechanism (by file id) in one sentence. It is clearly distinguished from the sibling add_entity_reference and the other entity tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-not condition ('Built-in entities cannot have references removed'), which tells the agent when the call will fail. However, it names no alternative tool or fallback for the built-in case, so the routing guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_image_backgroundRemove image backgroundCInspect

Remove the background from an image.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileIdYesSource image file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the agent knows this is a non-destructive local mutation. The description adds nothing beyond that: it doesn't say whether a new file is produced, whether the original is preserved, or what the returned artifact is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler or redundancy. It is appropriately sized for a one-parameter tool, though it is thin enough that it borders on under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, for an image transformation tool the description omits the workflow context (uploaded file required, result is a new asset), leaving the agent with only the bare operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter exists and schema coverage is 100%, so the schema fully documents imageFileId as the source image file id. The description adds no format, source, or constraint detail beyond that, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (remove) and resource (image background), so the operation is unambiguous. It does not, however, explicitly distinguish itself from the near-identical sibling remove_video_background; the agent must infer the split from the names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as remove_video_background or vectorize_image, nor any stated precondition (e.g., that the image must first be uploaded and have a file id). Usage is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_video_backgroundRemove video backgroundCInspect

Remove the background from a video.

ParametersJSON Schema
NameRequiredDescriptionDefault
videoFileIdYesSource video file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is not read-only and not destructive, but the description adds nothing behavioral beyond that: no indication of processing time, whether it creates a new file or overwrites, or whether the operation is async. It essentially restates the title.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero filler and the action front-loaded. It is concise, though that conciseness comes at the cost of substance rather than from efficient compression of useful content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a non-read-only media-processing tool the description omits key context such as processing duration, output artifact handling, and any input constraints. Too thin given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With one parameter and 100% schema description coverage, the schema fully documents videoFileId ('Source video file id.'). The description adds no syntax, format, or sourcing detail beyond what the schema already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Remove') and resource ('background from a video'), so the action is unambiguous. It implicitly differentiates from the sibling remove_image_background by media type, though it never names or contrasts that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus alternatives like remove_image_background, nor any prerequisites (e.g., video format, size limits, whether a file must be uploaded first). Usage is only inferable from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

script_to_videoScript to videoAInspect

Preferred for narrated / informational / explainer videos from text, especially ~1 minute or longer. Turn a verbatim narration script into an editable video with AI-generated visuals and captions. Prefer this over storyboard_to_video unless the user wants a short shot-directed storyboard. For avatar narration, pass actorEntityId and optionally set avatarQuality.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoVisual style for generated images. Write a full, strict paragraph in the same form as the app defaults below: name the medium, texture, and palette, then lock composition. Do not pass a short label such as "watercolor", "flat art", or "cinematic". Image models fill the frame with extra objects, readable text, charts, diagrams, tables, and labels unless the style forbids that. The still then looks crowded and hard to use. Every style must keep the picture simple. Include this composition lock verbatim: A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. No on-image text, letters, labels, captions, charts, diagrams, tables, legends, or infographic layout unless the user explicitly asked for one specific word or number on screen. Copy an app default in full (those already include the uncluttered-subject lock), then add the no-text / no-diagram sentence if it is missing. A custom style is allowed only when it is equally long and strict, and includes the same composition lock. Omit this field for the Realistic default. Watercolor: Loose watercolor illustration, visible brushstrokes, soft color bleeds, paper texture, muted palette. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Realistic: Photorealistic photograph, natural lighting Whiteboard: Minimalist whiteboard explainer style, simple line drawings, marker sketch aesthetic, clean white background, subtle accent colors. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. 3D cartoon: 3D cartoon render, rounded forms, soft lighting, vibrant colors, clean matte materials, smooth stylized characters Flat art: Corporate memphis flat art illustration, simple geometric shapes, bold colors, clean composition, white background with sparse subtle geometric accents such as small dots, lines, or shapes scattered in the margins. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no busy background patterns. Paper: Paper collage illustration, torn edges, layered cut paper shapes, mixed media texture, handmade craft aesthetic, matte paper finish. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Ink paint: Traditional Japanese sumi-e ink wash painting, expressive black ink brushstrokes, varied tonal gradations from deep black to soft gray, rice-paper texture, minimalist composition. No text, letters, calligraphy, kanji, signatures, or red seal stamps unless explicitly required by the subject. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Anime: Anime illustration, cel-shaded style, vibrant colors, clean linework, soft gradient backgrounds Editorial: Editorial illustration, conceptual art, bold geometric shapes, sophisticated color palette, negative space, minimalist composition, strong silhouette Isometric: Isometric 3D illustration, clean vector style, soft gradient background, matte pastel colors, simple geometric objects, no characters. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Claymation: Claymation style, soft clay figurines, plasticine texture, handmade stop-motion aesthetic, warm studio lighting, shallow depth of field Education: Educational infographic style, simple icons, pastel colors, white background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Chalkboard: Simple, minimalist, white chalk line drawings on dark green chalkboard, hand-drawn sketch style, chalk dust texture. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Boho: Simple graphic with only a few elements. Boho linocut poster. Textured organic strokes with crisp lines and a white textured background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Poster: Risograph print aesthetic, halftone texture, limited duotone palette, paper grain, high contrast, bold flat colors
scriptYesNarration script, spoken verbatim.
qualityNoGeneration quality. Omit to use workspace settings.
voiceIdNoCatalog display name (e.g. Matilda) or voice id from list_tts_voices.
languageNoOutput language as a BCP-47 code, such as en, es, or fr.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
actorEntityIdNoId of an ACTOR entity (vg_enti_...) with an image reference. When set, narration is delivered by that actor avatar.
avatarQualityNoAvatar generation quality tier. Applies when actorEntityId is provided. Omit to use workspace settings.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description usefully adds that the output is editable and includes AI visuals and captions, plus the avatar narration path. However, it never discloses that generation is long-running/async (implied only in the autoExport schema text), nor any cost or duration limits, which is the most important missing behavioral fact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the selection guidance before the capability statement and the avatar tip. No filler, and the highest-value routing information (preferred use case and the sibling exclusion) comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% schema description coverage, the description need not explain return values or parameters. It covers purpose, the alternative tool, and the avatar variant, which is everything an agent needs to choose and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the style field's semantics are exhaustively documented in the schema, so the schema does the heavy lifting. The description only adds a light steer ('pass actorEntityId and optionally set avatarQuality'), which largely restates what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Turn a verbatim narration script into an editable video') with explicit scope ('narrated / informational / explainer videos', '~1 minute or longer'). An agent can distinguish this from storyboard_to_video without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with a clear preference statement, then names the sibling alternative and the exact condition that selects it ('Prefer this over storyboard_to_video unless the user wants a short shot-directed storyboard'). It also routes the avatar case to specific parameters, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

slideshow_to_videoSlideshow to videoAInspect

Build an editable narrated video from an uploaded PDF or slideshow file. Upload the file first with upload_file, then pass its fileId. For avatar narration, pass actorEntityId and optionally set avatarQuality.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileIdYesUploaded PDF or slideshow file id.
voiceIdNoCatalog display name (e.g. Matilda) or voice id from list_tts_voices.
languageNoOutput language as a BCP-47 code, such as en, es, or fr.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
slideScriptsNoOptional narration for each slide, in order.
actorEntityIdNoId of an ACTOR entity (vg_enti_...) with an image reference. When set, narration is delivered by that actor avatar.
avatarQualityNoAvatar generation quality tier. Applies when actorEntityId is provided. Omit to use workspace settings.
slideshowThemeEntityIdNoOptional id of a SLIDESHOW_THEME entity (vg_enti_...) whose reference board defines the shared slide design system. Omit when converting an uploaded deck's original pages; VideoGen derives a theme from those pages in the background.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the write/mutation profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so safety is covered. The description adds the useful prerequisite ordering and notes the output is 'editable,' but says nothing about async behavior, runtime, or cost beyond what autoExport in the schema already conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: purpose first, then the actionable prerequisite and the optional avatar path. No filler or restated field lists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich 9-parameter schema, an output schema, and annotations, the description only needs to orient the agent — which it does by naming the required upload step. It stops short of routing among the numerous video-generation siblings, but otherwise nothing essential to invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including actorEntityId, avatarQuality, aspectRatio, and autoExport is already documented at the source. The description only restates the fileId prerequisite and the actorEntityId/avatarQuality pairing, adding little beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Build') and resource ('editable narrated video') with the required source ('uploaded PDF or slideshow file'). This clearly separates it from siblings like script_to_video, voiceover_to_video, and storyboard_to_video.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete prerequisite sequence ('Upload the file first with upload_file, then pass its fileId') and a conditional path for avatar narration. It does not, however, contrast this tool against the many other video-generation siblings, so the agent must infer when a slideshow source is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storyboard_to_videoStoryboard to videoAInspect

Build an editable video from an ordered storyboard (frame-by-frame shot list). Every scene needs a visual prompt and may include spoken words. Much more credit-heavy than script_to_video: use at most 3 scenes unless the user explicitly asks for more. Prefer script_to_video for ~1 minute+ narrated / informational videos. If the user has not named a workflow, ask with pros/cons before calling this.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoVisual style for generated images. Write a full, strict paragraph in the same form as the app defaults below: name the medium, texture, and palette, then lock composition. Do not pass a short label such as "watercolor", "flat art", or "cinematic". Image models fill the frame with extra objects, readable text, charts, diagrams, tables, and labels unless the style forbids that. The still then looks crowded and hard to use. Every style must keep the picture simple. Include this composition lock verbatim: A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. No on-image text, letters, labels, captions, charts, diagrams, tables, legends, or infographic layout unless the user explicitly asked for one specific word or number on screen. Copy an app default in full (those already include the uncluttered-subject lock), then add the no-text / no-diagram sentence if it is missing. A custom style is allowed only when it is equally long and strict, and includes the same composition lock. Omit this field for the Realistic default. Watercolor: Loose watercolor illustration, visible brushstrokes, soft color bleeds, paper texture, muted palette. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Realistic: Photorealistic photograph, natural lighting Whiteboard: Minimalist whiteboard explainer style, simple line drawings, marker sketch aesthetic, clean white background, subtle accent colors. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. 3D cartoon: 3D cartoon render, rounded forms, soft lighting, vibrant colors, clean matte materials, smooth stylized characters Flat art: Corporate memphis flat art illustration, simple geometric shapes, bold colors, clean composition, white background with sparse subtle geometric accents such as small dots, lines, or shapes scattered in the margins. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no busy background patterns. Paper: Paper collage illustration, torn edges, layered cut paper shapes, mixed media texture, handmade craft aesthetic, matte paper finish. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Ink paint: Traditional Japanese sumi-e ink wash painting, expressive black ink brushstrokes, varied tonal gradations from deep black to soft gray, rice-paper texture, minimalist composition. No text, letters, calligraphy, kanji, signatures, or red seal stamps unless explicitly required by the subject. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Anime: Anime illustration, cel-shaded style, vibrant colors, clean linework, soft gradient backgrounds Editorial: Editorial illustration, conceptual art, bold geometric shapes, sophisticated color palette, negative space, minimalist composition, strong silhouette Isometric: Isometric 3D illustration, clean vector style, soft gradient background, matte pastel colors, simple geometric objects, no characters. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Claymation: Claymation style, soft clay figurines, plasticine texture, handmade stop-motion aesthetic, warm studio lighting, shallow depth of field Education: Educational infographic style, simple icons, pastel colors, white background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Chalkboard: Simple, minimalist, white chalk line drawings on dark green chalkboard, hand-drawn sketch style, chalk dust texture. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Boho: Simple graphic with only a few elements. Boho linocut poster. Textured organic strokes with crisp lines and a white textured background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Poster: Risograph print aesthetic, halftone texture, limited duotone palette, paper grain, high contrast, bold flat colors
scenesYesOrdered scenes. Every scene requires a visual `prompt`.
qualityNoImage generation quality for scene opening frames. Optional; defaults to HIGH. Does not change video generation quality.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the mutation profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so safety is covered. The description adds genuinely non-structured context: it is 'much more credit-heavy' than script_to_video, there is an implicit scene cap, and the output is editable. It does not disclose runtime/async behavior or failure modes, so it is not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five sentences, each load-bearing: purpose, scene requirement, cost/scene-cap warning, alternative routing, and ambiguity handling. It is front-loaded with the routing-critical cost warning before the fallback advice.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, nested, enum-bearing tool with an output schema available, the description covers the decision-critical information an agent needs (what it builds, cost profile, scene limit, alternative tool). Return values need not be explained because an output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the style field already carries an extremely detailed enumeration and composition-lock guidance, so the schema does most of the work. The description confirms that every scene requires a prompt and may carry spoken words, which mildly reinforces the schema but adds no new syntax or constraints. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Build an editable video from an ordered storyboard') and immediately glosses the resource ('frame-by-frame shot list'). It also names the sibling it is not (script_to_video), so an agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (ordered shot list), when-not (prefer script_to_video for ~1 minute+ narrated/informational videos), a hard operational constraint (at most 3 scenes unless the user asks for more), and a routing rule for ambiguous cases (ask with pros/cons). This is close to a complete decision procedure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speechText to speechBInspect

Convert text into spoken audio using a selectable voice.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to speak.
speedNoSpeech-rate multiplier.
voiceIdYesCatalog display name (e.g. Matilda) or voice id from list_tts_voices.
languageNoISO-639-1 pronunciation language hint.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=false, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond the basic conversion action, such as generation latency, temporary file handling, or voice availability constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently communicates the core action and one key variable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple text-to-speech tool with rich structured fields, the description is mostly complete: output schema handles return values, annotations cover safety, and the schema documents all parameters. The main remaining gap is sibling differentiation, but the core invocation requirements are covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents all four parameters, including text, speed, voiceId, and language. The description's mention of a 'selectable voice' adds no meaning beyond what the schema already provides, making the baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Convert text into spoken audio.' It also mentions a selectable voice, which clarifies scope. However, it does not differentiate this tool from similar siblings such as voiceover_to_video, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance about when to use this tool versus alternatives like voiceover_to_video or when to first call list_tts_voices. Usage is only implied by the tool name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_entityUpdate entityAInspect

Update an entity's display name and/or description. Built-in entities cannot be updated.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoNew display name. Omit to leave unchanged.
entityIdYesEntity id (vg_enti_...).
descriptionNoNew description. Omit to leave unchanged.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesDisplay name.
entityIdYesEntity id (vg_enti_...).
createdAtYesUnix created-at timestamp.
isBuiltInNoTrue for VideoGen catalog entities that cannot be modified.
updatedAtYesUnix updated-at timestamp.
entityTypeYesACTOR, PRODUCT, VISUAL_STYLE, or SLIDESHOW_THEME.
referencesYesAttached reference images.
actorConfigNoVoice/avatar summary for ACTOR entities; null for other types.
descriptionYesDescription (empty when unset).

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-destructive, closed-world write. The description adds a real constraint beyond that: built-in entities reject updates, which is a failure condition the agent would otherwise not know. It still omits permission/auth requirements and what happens on an invalid entityId.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler. The primary action is front-loaded and the constraint follows immediately, so nothing needs to be skimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, the input schema is fully documented, and annotations cover the safety profile, so the description need not explain return values. It supplies the one piece of context the structured fields lack (built-in entities), leaving only minor gaps like auth requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains that name and description are optional ('Omit to leave unchanged') and documents the entityId format. The description only restates the same two fields, adding no syntax or format meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (update) and resource (entity) plus the exact updatable fields (display name, description). It is clearly distinct from its siblings create_entity, get_entity and list_entities, though it does not explicitly name those siblings to reinforce the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Built-in entities cannot be updated' gives a genuine when-not condition, which is useful. However, there is no guidance on when to prefer this tool over archive_entity or create_entity, and no prerequisites or context beyond the built-in restriction are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileUpload fileAInspect

Upload a local file to VideoGen and wait until it is processed. Returns the file with its id (vg_file_...) and signed URLs. Use the returned fileId for voiceover_to_video, slideshow_to_video, logos, or B-roll. To upload a remote asset, download it first and pass its local path.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoFile type. Inferred when omitted.
filePathYesAbsolute path to a local file to upload.
displayNameNoDisplay name for the file. Defaults to the source file name.

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeNoFile type once processing has determined it.
scopeNoFile scope.
fileIdYesFile id (vg_file_...).
hlsSourceNo
descriptionNoFile description when analyzed.
displayNameNoDisplay name for the file.
downloadUrlNoSigned download URL when ready.
publicHlsUrlNo
thumbnailUrlNoSigned thumbnail URL when available.
previewSourceNo
downloadSourceNo
sourceToolTypeNo
transcriptTextNoPlain transcript text when available.
durationSecondsNoDuration in seconds for video/audio; null for images.
thumbnailSourceNo
publicPlaybackIdNo
downloadUrlExpiresAtNoUnix expiry for downloadUrl.
sourceToolExecutionIdNo
thumbnailUrlExpiresAtNoUnix expiry for thumbnailUrl.
isPublicPreviewEnabledNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-destructive write with no open-world access, and the description adds the key behavioral fact those annotations cannot convey: the call blocks until processing finishes. It does not state size limits, failure modes, or expected latency, which would be valuable for an upload that waits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences that lead with the action and proceed to return value and downstream usage; no filler. Slightly dense, but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the local-vs-remote distinction, the blocking behavior, and the returned identifier, and an output schema exists so return formatting needn't be repeated. Missing only guidance on limits or error handling for a tool that waits on processing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter (type, filePath, displayName) is already documented with defaults and inference rules, so the description adds little parameter detail. Baseline 3 is appropriate when the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upload a local file to VideoGen') plus the synchronous processing behavior, which is more than a restatement of the name. It does not, however, differentiate itself from the sibling 'create_file_upload', which an agent could reasonably confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear downstream context ('Use the returned fileId for voiceover_to_video, slideshow_to_video, logos, or B-roll') and an explicit precondition for remote assets ('download it first and pass its local path'). It stops short of saying when to prefer this over the sibling create_file_upload.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upscale_imageUpscale imageBInspect

Increase the resolution of an image.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileIdYesSource image file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is largely covered. The description adds the actual operation behavior (resolution increase), but it does not clarify whether the original image is modified, whether a new asset is created, or whether the operation is synchronous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word contributes to the purpose, making it appropriately sized for a simple one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema, available annotations, and existing output schema, the description is minimally sufficient to identify the operation. However, it omits usage context relative to sibling image tools and does not clarify important mutation-side-effect details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage and documents the single required parameter imageFileId as "Source image file id." The description adds no parameter-specific meaning beyond what the schema already provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: "Increase the resolution of an image." This clearly separates it from siblings such as upscale_video and generate_image, though it does not explicitly contrast it with other image-editing tools like vectorize_image or remove_image_background.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The name implies a straightforward image operation, but the description does not help an agent choose between this and sibling image tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upscale_videoUpscale videoCInspect

Increase the resolution of a video.

ParametersJSON Schema
NameRequiredDescriptionDefault
videoFileIdYesSource video file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds nothing beyond that: it does not disclose that this is likely an async/long-running generation task, whether it costs credits, or what resolution limits or scale factors apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste, which is structurally clean. But its brevity reflects under-specification rather than disciplined conciseness, leaving questions about scale factor and output resolution unanswered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The schema is minimal (one required id) and the annotations cover the safety profile, but for a media-processing tool the description omits any notion of async execution or upscaling options, leaving it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (videoFileId) exists and schema description coverage is 100%, so the schema fully documents it. The description adds no syntax, format, or constraint detail beyond what the schema provides, which is the baseline 3 case when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (increase) and resource (resolution of a video), so the core action is unambiguous. However, it offers no differentiation from the sibling upscale_image or any of the many video-generation siblings, so an agent must infer the distinction on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, no prerequisites, and no mention of alternatives such as upscale_image or the generate_video_clip family. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vectorize_imageVectorize imageBInspect

Convert a raster image into a vector (SVG).

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileIdYesSource image file id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds the output format (SVG), which is not in the schema or annotations, but says nothing about whether processing is async, whether a new file is created, or how to retrieve the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The action and both endpoints of the transformation are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover safety. For a one-parameter tool the description is close to sufficient, though it omits processing model (sync vs async) and input-format constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single parameter (imageFileId), so the schema carries the semantics. The description adds no additional detail such as accepted image formats or size limits, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-and-resource pair (convert a raster image into a vector/SVG) that is unambiguous. No sibling does vectorization, but the description never names or contrasts an alternative, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as upscale_image or image_3d_effect. The agent must infer usage from the purpose statement alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voiceover_to_videoVoiceover to videoBInspect

Build an editable video with AI-generated visuals from an uploaded voiceover audio file.

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoVisual style for generated images. Write a full, strict paragraph in the same form as the app defaults below: name the medium, texture, and palette, then lock composition. Do not pass a short label such as "watercolor", "flat art", or "cinematic". Image models fill the frame with extra objects, readable text, charts, diagrams, tables, and labels unless the style forbids that. The still then looks crowded and hard to use. Every style must keep the picture simple. Include this composition lock verbatim: A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. No on-image text, letters, labels, captions, charts, diagrams, tables, legends, or infographic layout unless the user explicitly asked for one specific word or number on screen. Copy an app default in full (those already include the uncluttered-subject lock), then add the no-text / no-diagram sentence if it is missing. A custom style is allowed only when it is equally long and strict, and includes the same composition lock. Omit this field for the Realistic default. Watercolor: Loose watercolor illustration, visible brushstrokes, soft color bleeds, paper texture, muted palette. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Realistic: Photorealistic photograph, natural lighting Whiteboard: Minimalist whiteboard explainer style, simple line drawings, marker sketch aesthetic, clean white background, subtle accent colors. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. 3D cartoon: 3D cartoon render, rounded forms, soft lighting, vibrant colors, clean matte materials, smooth stylized characters Flat art: Corporate memphis flat art illustration, simple geometric shapes, bold colors, clean composition, white background with sparse subtle geometric accents such as small dots, lines, or shapes scattered in the margins. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no busy background patterns. Paper: Paper collage illustration, torn edges, layered cut paper shapes, mixed media texture, handmade craft aesthetic, matte paper finish. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Ink paint: Traditional Japanese sumi-e ink wash painting, expressive black ink brushstrokes, varied tonal gradations from deep black to soft gray, rice-paper texture, minimalist composition. No text, letters, calligraphy, kanji, signatures, or red seal stamps unless explicitly required by the subject. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Anime: Anime illustration, cel-shaded style, vibrant colors, clean linework, soft gradient backgrounds Editorial: Editorial illustration, conceptual art, bold geometric shapes, sophisticated color palette, negative space, minimalist composition, strong silhouette Isometric: Isometric 3D illustration, clean vector style, soft gradient background, matte pastel colors, simple geometric objects, no characters. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Claymation: Claymation style, soft clay figurines, plasticine texture, handmade stop-motion aesthetic, warm studio lighting, shallow depth of field Education: Educational infographic style, simple icons, pastel colors, white background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Chalkboard: Simple, minimalist, white chalk line drawings on dark green chalkboard, hand-drawn sketch style, chalk dust texture. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Boho: Simple graphic with only a few elements. Boho linocut poster. Textured organic strokes with crisp lines and a white textured background. A clear uncluttered subject centered in the frame, occupying only the middle half of the image, with generous empty margins on all four sides, no background clutter. Poster: Risograph print aesthetic, halftone texture, limited duotone palette, paper grain, high contrast, bold flat colors
fileIdYesUploaded voiceover audio file id.
qualityNoGeneration quality. Omit to use workspace settings.
languageNoOutput language as a BCP-47 code, such as en, es, or fr.
autoExportNoWhen true (the default), the run stays in progress until an MP4 is ready. Use downloadUrl from the result. Set false only if you will call remix_project and then export_project yourself.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
exportIdNoExport id when autoExport succeeded (vg_expo_...).
projectIdYesProject created for this exact workflow attempt (vg_proj_...). A retry creates a different project.
projectUrlYesDeep link to the project created for this exact workflow attempt. Keep it paired with workflowRunId; a retry returns a new URL.
downloadUrlNoSigned MP4 download URL when autoExport succeeded. Give this to the user.
attemptIndexNoCurrent or latest attempt index.
exportFileIdNoFile id of the rendered MP4 when autoExport succeeded.
workflowTypeNoWorkflow type (present after polling).
workflowRunIdYesWorkflow run id (vg_work_...). Keep it paired with this response's projectId and projectUrl; a retry returns a new tuple.
remixActionIdsNoRemix action ids from the start response (empty when none requested).
progressPercentageNoCompletion progress 0-100 (present after polling).
downloadUrlExpiresAtNoUnix expiry for downloadUrl.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so safety is covered. The word 'editable' usefully signals the result is a project rather than a finished file, but the description adds nothing about generation latency, credits, or how the async run behaves (that lives in the schema's autoExport text).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, naming the output and the required input first. It is arguably too terse for a six-parameter, nested-object tool, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the parameter schema is exhaustive, so return values and inputs are adequately covered. However, for a complex multi-stage tool the description omits the editable-project/remix workflow and any note that generation is asynchronous, leaving the agent dependent on schema text to piece the flow together.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including an extremely detailed style parameter and a clear autoExport/aspectRatio explanation, so the schema does all parameter work. The description contributes no additional parameter meaning, which is the expected baseline when coverage is this high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (build an editable video) and identifies the input source (an uploaded voiceover audio file), which distinguishes it from script_to_video, storyboard_to_video, and slideshow_to_video. It stops short of naming those siblings explicitly, so an agent must infer the boundary from the input type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance. The description never says how this differs from the many other video-generating siblings or when a user should prefer it over script_to_video. Some routing context exists only indirectly in the autoExport schema text (which mentions remix_project and export_project), not in the description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 49 tool updatesv2.2.1
    • First observedadd_entity_reference
    • First observedarchive_entity
    • First observedcancel_tool_execution
    • First observedcancel_workflow_run
    • First observedcreate_entity
    • First observedcreate_file_upload
    • First observedexport_project
    • First observedgenerate_avatar
    • First observedgenerate_image
    • First observedgenerate_motion_graphic
    • First observedgenerate_music
    • First observedgenerate_sound_effect
    • First observedgenerate_video_clip
    • First observedget_app_deep_link
    • First observedget_async_tasks_guidance
    • First observedget_entity
    • First observedget_file
    • First observedget_getting_started_guidance
    • First observedget_me
    • First observedget_project
    • First observedget_project_export
    • First observedget_tool_execution
    • First observedget_tools_vs_workflows_guidance
    • First observedget_workflow_run
    • First observedget_workflows_guidance
    • First observedimage_3d_effect
    • First observedlist_entities
    • First observedlist_files
    • First observedlist_languages
    • First observedlist_project_remix_actions
    • First observedlist_projects
    • First observedlist_tool_executions
    • First observedlist_tts_voices
    • First observedlist_workflow_runs
    • First observedprompt_to_video_clip
    • First observedremix_project
    • First observedremove_entity_reference
    • First observedremove_image_background
    • First observedremove_video_background
    • First observedscript_to_video
    • First observedslideshow_to_video
    • First observedstoryboard_to_video
    • First observedtext_to_speech
    • First observedupdate_entity
    • First observedupload_file
    • First observedupscale_image
    • First observedupscale_video
    • First observedvectorize_image
    • First observedvoiceover_to_video

TDQS

B3.4/5.0

Scored across 49 tools

Disambiguation4/5

Most tools have clearly distinct purposes, but there is some overlap: script_to_video, storyboard_to_video, voiceover_to_video, and slideshow_to_video all produce editable videos from different inputs, requiring careful description reading to choose correctly. The guidance tools (get_*_guidance) also overlap somewhat but are clearly differentiated by topic.

Naming Consistency4/5

Naming is mostly consistent with snake_case verb_noun or noun_verb patterns (generate_image, list_projects, get_project, upload_file, create_entity), though some deviations exist like storyboard_to_video, script_to_video (noun_to_noun) and guidance tools prefixed with get_. Still generally predictable and readable.

Tool Count2/5

49 tools is excessive for the apparent scope of a video generation API, even accounting for the broad domain (workflows, media generation, files, entities, projects, account). Many tools could be consolidated or grouped, making the surface heavy and increasing selection difficulty.

Completeness4/5

Coverage is quite comprehensive: creation, listing, getting, updating, archiving, uploading, guidance, and export for projects, workflows, tool executions, files, and entities. Minor gaps exist (e.g., no direct delete for projects/files, no update for file metadata), but core lifecycle operations are present.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers