Shotstack MCP Server
Shotstack MCP server provides 23 tools (12 read-only, 11 confirmed account operations) for rendering, templates, generation, hosted assets, ingest, uploads, and private account/config checks.
Submit renders and check render status.
Create, list, get, update, delete, and render templates with merge fields.
List generation models, generate assets, and get generated assets.
Inspect, transfer, and delete hosted assets, including lookup by render ID.
Ingest sources, list/get/delete sources, and create signed upload URLs stored in a private file.
List configured private account labels without network access.
Run configuration checks via doctor/login CLI utilities.
Use named private accounts and environments with explicit confirmation for mutations; reads remain safe.
Share the same handlers across MCP and CLI, with no second API implementation.
Enables sending rendered videos, images, audio, and other assets to an Amazon S3 bucket as an output destination, with options for region, bucket, prefix, filename, and ACL.
Integrates with Dolby.io audio enhancement to apply presets to ingested source renditions via Shotstack's ingest pipeline.
Supports generating audio assets using ElevenLabs text-to-speech and music models through Shotstack's generation API.
Enables sending rendered videos, images, audio, and other assets to a Google Cloud Storage bucket as an output destination, with options for bucket, prefix, and filename.
Enables sending rendered videos and assets to Google Drive, with optional folder ID and custom filename settings.
Enables sending rendered videos to Vimeo, with options for name, description, folder URI, and privacy settings.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Shotstack MCP Servercheck the status of my latest video render"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Shotstack MCP Server & CLI
Shotstack MCP server and CLI for Codex and AI agents. 23 tools for current rendering, templates, generation models, Serve and Ingest, with private accounts and explicit mutation approval. One shared implementation supplies both binaries and a desktop bundle.
Built and maintained by Navid Moazzez. The complete guide is on navid.me.
The terminal illustrates real command names and approval flow. It is not a recording of a paid provider render. Shotstack already has official CLI/local/hosted MCP products; their Studio and semantic validation are compared below.
Requires Node 22+ and eligible Shotstack API access for account operations. Validation: fixture tests, schema validation and protocol/artifact discovery are separate from provider-account rendering, desktop GUI outcomes and fresh measured task/token evidence. Pending evidence is recorded, without invented success rates or efficiency claims.
Two ways to use it
Command line
npm install -g @thenavidm/shotstack-mcp-cli@latest
shotstack-cli
shotstack-cli list-models --agent
shotstack-cli schema render
shotstack-cli render --payload-file /absolute/private/approved-edit.json --account sandbox --confirm --agent--confirm authorizes the specific requested operation. --agent and --yes do not supply consent or a provider credit cap.
MCP server, for your AI app
codex mcp add shotstack -- npx -y @thenavidm/shotstack-mcp-cli@latestConfigure private credentials/environment first. Ask: “Inspect this existing render and return its status; do not submit another render.” Client and OS details are in INSTALL.md.
Which one
Where you work | Surface |
Codex or an agent with shell access | Local MCP, shared CLI or both |
Desktop chat | Compatible local MCP or .mcpb |
Scripts/CI | CLI or an MCP client |
Remote-URL-only chat | Official provider-hosted MCP |
Related MCP server: openstack-incidents
Features
Capability | CLI command | MCP tool |
Render and status | render / get-render | render / get_render |
Templates | create-template / list-templates / render-template | create_template / list_templates / render_template |
Generation models and jobs | list-models / generate-asset / get-generated-asset | list_models / generate_asset / get_generated_asset |
Hosted assets | get-asset / transfer-asset | get_asset / transfer_asset |
Ingest sources | ingest-source / list-sources / get-source | ingest_source / list_sources / get_source |
Private upload credential | create-upload-url-file | create_upload_url_file |
Account/environment labels | list-accounts | list_accounts |
Configuration checks | doctor / login | CLI utilities |
Contents
Number | Section | What it covers |
1 | Practical tasks | |
2 | MCP/CLI/desktop | |
3 | Keys, environments, credits | |
4 | All client/OS routes | |
5 | Doctor and first read | |
6 | Flags and scripting | |
7 | Actual measurement method | |
8 | Every route/argument | |
9 | Native workflow inputs | |
10 | Status and private credential files | |
11 | Isolated named profiles | |
12 | Approval and direct-call policies | |
13 | Shared architecture and sync | |
14 | Credentials and provider processing | |
15 | Private settings and bounds | |
16 | Upgrade/revoke/uninstall | |
17 | Errors and recovery | |
18 | Official and community tradeoffs | |
19 | Release history and migration | |
20 | Accordion answers |
1. What you can ask it
Inspect my available generation models before choosing a supported asset.
Read this existing render's status; do not submit a duplicate.
Prepare a reviewed timeline and submit only the approved sandbox render.
Save this template, then render it with the merge values I selected.
Ingest this selected public source URL after confirmation.
Create an upload URL, keeping its temporary credential in my private file.
Inspect a hosted asset and confirm its exact ID before deletion.
Actual discovery supplies 23 tools: 12 reads and 11 confirmed account operations. The same handlers serve MCP and CLI. Read-only discovery and direct-call refusal agree. These checks do not establish successful provider rendering or GUI installation.
2. Quick install
npm install -g @thenavidm/shotstack-mcp-cli@latest
shotstack-cli --version
shotstack-cli login
shotstack-cli doctor
shotstack-cli toolsNode 22+ is required for manual installation. The shotstack-2.0.1.mcpb archive bundles production dependencies for a compatible desktop host. Complete setup is in INSTALL.md.
After configuring private local credentials:
codex mcp add shotstack -- npx -y @thenavidm/shotstack-mcp-cli@latest
codex mcp list3. Set up Shotstack access
Private API keys and environments
Sign in to your intended account at app.shotstack.io, open the account menu and choose API Keys.
Select the sandbox or production key for the task. Set SHOTSTACK_ENV=stage or v1 to match. Production is the default; a sandbox key does not turn production URLs into sandbox URLs.
Save the key in an owner-only token file outside repositories. Set SHOTSTACK_TOKEN_FILE to its absolute path. SHOTSTACK_API_KEY in private local settings is the alternative.
Run shotstack-cli doctor, then doctor --network. The network check requests available generation models and does not print their content.
Read the selected schema, review the exact edit/account action and confirm only that operation. Do not submit a production render merely to test installation.
Requests use x-api-key on the fixed api.shotstack.io origin, with /edit, /serve or /ingest and the selected /stage or /v1 prefix. This package does not accept arbitrary API hosts, forward keys through redirects or perform OAuth. login prints setup instructions; it neither saves keys nor creates accounts.
On macOS/Linux, use a private directory (0700) and regular token-only file (0600). Windows users must restrict the file's ACL to their user; POSIX mode checking does not establish Windows ACL protection. Token files cannot be symlinks or exceed 64 KB. They override environment keys and are cached until restart. GUI settings may differ from your shell environment.
Access, plans, credits and limits
The wrapper is free AGPL software. Shotstack access, rendering, generated assets, storage and serving follow the provider's account terms. No account role or OAuth scope bypass is supplied. Request API keys and check your dashboard's balance and current plan before approving charges.
The provider currently documents ten new-account credits valid for 30 days, with no credit card required. Sandbox renders are watermarked, limited to ten minutes and require at least one credit in the balance. AI generation in sandbox still consumes credits. Production rendering is billed by output duration; this wrapper is not a spending or money cap.
Rate limits use a fixed 60-second window per API key, across all plans. Current production/sandbox request limits are Edit 300/150, Serve 600/300 and Ingest 300/120. Wait for the window reset after 429; do not immediately repeat a render.
The local default 150 ms pacing is per account label/process, not a provider quota reservation. Shared keys in several processes still share quota. Mutations never retry automatically. GET 429 retries require an explicit Retry-After of at most ten seconds; missing/longer delays return exit 7. The default maximum is two retries and configurable upper bound five.
Current source limits are 5 GB per file and 10 GB combined source/output disk usage. Local request JSON has a separate 5 MiB cap and responses a 10 MiB cap. The wrapper does not upload local media bytes or automatically collect all pages. Check credit consumption for the selected current model and hosting charges.
Revoke and rotate
Rotate or revoke the intended key through your account's API Keys controls, update private client settings, and restart. Remove the client entry when disconnecting. npm removal does not revoke the provider key, undo renders, remove hosted assets or delete private output files. Keep account records and signed URLs out of GitHub issues and public logs.
4. Connect your client
INSTALL.md provides Codex-first configuration plus optional Claude Code, Claude Desktop archive/manual setup, Cursor, VS Code/Copilot, Windsurf, Zed, Gemini CLI, Cline and Docker. Use Node 22+ on macOS, Windows or Linux; GUI settings and remote development environments need their own accessible private credentials.
Manual MCP launches npx with arguments -y and @thenavidm/shotstack-mcp-cli@latest over stdio. This package has no public HTTP relay. Remote-only clients can use the official https://mcp.shotstack.io/ service with its supported OAuth/API-key setup.
npm installation ships SKILL.md but does not register an agent skill. Add the shipped file through the client's supported skill location. Client approvals and confirm=true are separate; the guard requires the exact requested mutation, not consent inferred from returned content.
5. Check it works
shotstack-cli --version
shotstack-cli doctor
shotstack-cli doctor --network
shotstack-cli list-accounts --agent
shotstack-cli list-models --agentHelp, schemas, discovery and account labels work without a provider key. Network doctor requests GET /edit/{environment}/models without returning its content. Success establishes that account request, not every rendering model or endpoint. Full discovery exposes 23 tools; read-only exposes 12.
For an existing selected render, use get-render --id REAL_RENDER_ID. A job status read should not create a new render. Missing credentials exits 10; invalid input and refused writes exit 2.
6. Output, flags and exit codes
Tool results go to stdout. Errors are JSON on stderr. Reads and generation return structured JSON, so --select can retain nested fields.
shotstack-cli render --help
shotstack-cli schema render
shotstack-cli get-render --id REAL_RENDER_ID --agent --select response.status,response.urlFlag | What it does |
| JSON output |
| Single-line JSON |
| JSON, compact, no input and no color |
| Keep selected fields; dotted paths descend and arrays are traversed |
| Confirm the requested paid media operation |
| Automation switches; none overrides the spending guard |
| Return an accepted job instead of polling |
| Save completed media locally; requires waiting for completion |
Global output flags apply to tool commands. doctor has its own --network option and returns a JSON diagnostic.
Exit code | Meaning | What a script should do |
0 | Success | Read stdout |
2 | Usage, invalid input or a refused write | Fix the input or confirm only the requested action |
3 | Job or local upload file not found | Check the ID/path |
4 | Authentication or entitlement rejected | Check private credential settings and permissions |
5 | API, network or polling failure | Inspect an accepted job before another paid submission |
7 | Rate limited | Wait; do not loop over paid submissions |
10 | Credentials not configured | Complete local setup |
The underscore spelling also works. generate_asset and generate-asset call the same tool. Nested objects use quoted JSON. Arrays of objects use repeated flags, one JSON object at a time.
7. MCP or CLI and token cost
MCP and CLI use the same SDK server, input schemas, handlers and confirmation guard. CLI commands use the SDK's in-memory transport; there is no second API implementation. Choose shell calls for scripts and the local MCP for a stdio AI client.
Measurement | Required evidence |
Eager MCP loading | Actual tool schemas and instructions sent to the model |
Deferred discovery | Actual selected schemas and lookup overhead |
Skill read once | Complete SKILL.md and command discovery |
Recurring skill listing | The installed skill description |
Equivalent successful task | Help/schema, reasoning, requests, output, retries and achieved result |
Fresh Codex measurements are pending. Record model/client/package versions, date, loading settings, input/output usage, latency and equivalent results. Compare an existing-render status task and an approved template/render task with identical account environment and response fields.
Do not estimate tokens from characters, borrow another package's results, infer efficiency from 23 tools or say CLI has zero cost. Local --select trims output after receipt; it does not change provider response size or quota. Provider credits remain separate from model tokens. Claude Code measurements are deferred at the current Codex priority.
8. Every tool and argument
Every route and argument below comes from actual stdio discovery and reviewed current upstream schemas. Read exact nested definitions with shotstack-cli schema COMMAND before constructing a body. Native body names, arrays, unions and enums remain exact.
Tool | Route | Mode |
|
| Confirm requested operation |
|
| Read |
|
| Confirm requested operation |
|
| Read |
|
| Read |
|
| Confirm requested operation |
|
| Confirm requested operation |
|
| Confirm requested operation |
|
| Read |
|
| Confirm requested operation |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm requested operation |
|
| Read |
|
| Confirm requested operation |
|
| Confirm requested operation |
|
| Read |
|
| Read |
|
| Confirm requested operation |
|
| Confirm requested operation |
| Local, no network | Read |
render
shotstack-cli render
Argument | Required | Type | Details |
| No; body and guard rules apply | Timeline | See the full input schema. |
| No; body and guard rules apply | Output | See the full input schema. |
| No; body and guard rules apply | array | An array of key/value pairs that provides an easy way to create templates with placeholders. The placeholders can be used to find and replace keys with values. For example you can search for the placeholder |
| No; body and guard rules apply | string | An optional webhook callback URL used to receive status notifications when a render completes or fails. Notifications are also sent when a rendered video is sent to an output destination. See webhooks for more details. |
| No; body and guard rules apply | string | Notice: This option is now deprecated and will be removed. Disk types are handled automatically. Setting a disk type has no effect. The disk type to use for storing footage and assets for each render. |
| No; body and guard rules apply | string | The render instance type to use for processing the edit. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: timeline, output.
get_render
shotstack-cli get-render
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
create_template
shotstack-cli create-template
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The template name |
| No; body and guard rules apply | Edit | See the full input schema. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: name.
list_templates
shotstack-cli list-templates
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
get_template
shotstack-cli get-template
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
update_template
shotstack-cli update-template
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | The template name |
| No; body and guard rules apply | Edit | See the full input schema. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: name.
delete_template
shotstack-cli delete-template
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
render_template
shotstack-cli render-template
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The id of the template to render in UUID format. |
| No; body and guard rules apply | array | An array of key/value pairs that provides an easy way to create templates with placeholders. The placeholders can be used to find and replace keys with values. For example you can search for the placeholder |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: id.
probe_media
shotstack-cli probe-media
Argument | Required | Type | Details |
| Yes | string | Public HTTPS URL of the selected media file. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
generate_asset
shotstack-cli generate-asset
Argument | Required | Type | Details |
| No; body and guard rules apply | string | A key that makes this request its own generation. Retrying with the same key returns the job it first created instead of generating and billing again, and a new key generates afresh even when the asset matches an earlier one. Without a key, identical assets share one cached result. For 24 hours a key reused for a different asset is rejected; after that it returns its first result. |
| No; body and guard rules apply | GenerationAsset | See the full input schema. |
| No; body and guard rules apply | number | The length, in seconds, of the clip the asset fills. A model that generates to a duration takes it from this value in place of its own duration option. Other models ignore it. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: asset.
get_generated_asset
shotstack-cli get-generated-asset
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
list_models
shotstack-cli list-models
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
get_model
shotstack-cli get-model
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
get_asset
shotstack-cli get-asset
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
delete_asset
shotstack-cli delete-asset
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
get_asset_by_render_id
shotstack-cli get-asset-by-render-id
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
transfer_asset
shotstack-cli transfer-asset
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The file URL to fetch and transfer. |
| No; body and guard rules apply | string | An identifier for the asset which must be provided by the client. The identifier does not need to be unique. |
| No; body and guard rules apply | array | Specify the storage locations and hosting services to send the file to. Items: Destinations. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: url, id, destinations.
ingest_source
shotstack-cli ingest-source
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The URL of the file to be ingested. The URL must be publicly accessible or include credentials. |
| No; body and guard rules apply | Outputs | See the full input schema. |
| No; body and guard rules apply | Destinations | See the full input schema. |
| No; body and guard rules apply | string | An optional webhook callback URL used to receive status notifications when sources are uploaded and renditions processed. |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Use native body flags or one complete payload/payload_file; these cannot be mixed. Required body fields: .
list_sources
shotstack-cli list-sources
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
get_source
shotstack-cli get-source
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
delete_source
shotstack-cli delete-source
Argument | Required | Type | Details |
| Yes | string | Exact resource ID from the selected account. minLength: |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
create_upload_url_file
shotstack-cli create-upload-url-file
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Shotstack account; selects private credentials and stage/v1 environment. |
| No; body and guard rules apply | boolean | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
| Yes | string | New absolute JSON file in a private owner-only directory. Signed upload URL stays out of model output; no overwrite. minLength: |
The output file must be new and absolute, in a private directory. It is reserved exclusively before the API call; the signed URL is never returned in model output. This command creates a temporary upload credential; it does not upload file bytes.
list_accounts
shotstack-cli list-accounts
Argument | Required | Type | Details |
None | No | None | No arguments |
Nested request definitions
The following definitions are shared by the request schemas. oneOf selects one validated asset branch; $ref names point to the definition heading here. Required fields are specific to the selected branch. Any deeper inline shape remains available through schema COMMAND.
Timeline
A timeline represents the contents of a video edit over time, an audio edit over time, in seconds, or an image layout. A timeline consists of layers called tracks. Tracks are composed of titles, images, audio, html or video segments referred to as clips which are placed along the track at specific starting point and lasting for a specific amount of time.
Argument | Required | Type | Details |
| No; body and guard rules apply | Soundtrack | A music or audio soundtrack file in mp3 format. Deprecated - use an AudioAsset clip on its own track instead. |
| No; body and guard rules apply | string | A hexadecimal value for the timeline background colour. Defaults to #000000 (black). |
| No; body and guard rules apply | array | An array of custom fonts to be downloaded for use by the HTML assets. Items: Font. |
| Yes | array | A timeline consists of an array of tracks, each track containing clips. Tracks are layered on top of each other in the same order they are added to the array with the top most track layered over the top of those below it. Ensure that a track containing titles is the top most track so that it is displayed above videos and images. minItems: |
| No; body and guard rules apply | boolean | Disable the caching of ingested source footage and assets. See caching for more details. |
Soundtrack
Notice: The Soundtrack is deprecated, use an AudioAsset clip on its own track instead. This type continues to function; no behaviour change for existing integrations. A music or audio file in mp3 format that plays for the duration of the rendered video or the length of the audio file, which ever is shortest.
Argument | Required | Type | Details |
| Yes | string | The URL of the mp3 audio file. The URL must be publicly accessible or include credentials. minLength: |
| No; body and guard rules apply | string | The effect to apply to the audio file |
| No; body and guard rules apply | number | Set the volume for the soundtrack between 0 and 1 where 0 is muted and 1 is full volume (defaults to 1). |
Font
Download a custom font to use with the HTML asset type, using the font name in the CSS or font tag. See our custom fonts getting started guide for more details.
Argument | Required | Type | Details |
| Yes | string | The URL of the font file. The URL must be publicly accessible or include credentials. |
Track
A track contains an array of clips. Tracks are layered on top of each other in the order in the array. The top most track will render on top of those below it.
Argument | Required | Type | Details |
| Yes | array | An array of Clips comprising of TitleClip, ImageClip or VideoClip. minItems: |
Clip
A clip is a container for a specific type of asset, i.e. a title, image, video, audio or html. You use a Clip to define when an asset will display on the timeline, how long it will play for and transitions, filters and effects to apply to it.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Optional client-generated identifier. Used by client SDKs (e.g. the Shotstack Studio SDK) to reference a clip across edits without relying on its position in the timeline. The render API does not use this field and it does not appear in render output. |
| Yes | Asset | See the full input schema. |
| Yes | Union | The start position of the Clip on the timeline. |
| Yes | Union | The duration the Clip should play for. |
| No; body and guard rules apply | string | Set how the asset should be scaled to fit the viewport using one of the following options: |
| No; body and guard rules apply | Union | Scale the asset to a fraction of the viewport size - i.e. setting the scale to 0.5 will scale asset to half the size of the viewport. This is useful for picture-in-picture video and scaling images such as logos and watermarks. Use a number or an array of Tween objects to create a custom animation. |
| No; body and guard rules apply | number | Set the width of the clip bounding box in pixels. This constrains the width of the clip, overriding the default behavior where clips fill the viewport width. minimum: |
| No; body and guard rules apply | number | Set the height of the clip bounding box in pixels. This constrains the height of the clip, overriding the default behavior where clips fill the viewport height. minimum: |
| No; body and guard rules apply | string | Place the asset in one of nine predefined positions of the viewport. This is most effective for when the asset is scaled and you want to position the element to a specific position. |
| No; body and guard rules apply | Offset | Offset the location of the asset relative to its position on the viewport. The offset distance is relative to the width of the viewport - for example an x offset of 0.5 will move the asset half the viewport width to the right. |
| No; body and guard rules apply | Transition | See the full input schema. |
| No; body and guard rules apply | string | A motion effect to apply to the Clip. |
| No; body and guard rules apply | string | A filter effect to apply to the Clip. |
| No; body and guard rules apply | Union | Offset an asset on the horizontal axis (left or right). Use a number or an array of Tween objects to create a custom animation. |
| No; body and guard rules apply | Transformation | A transformation lets you modify the visual properties of a clip. Available transformations are rotate, skew and flip. Transformations can be combined to create interesting new shapes and effects. |
| No; body and guard rules apply | string | A unique identifier for this clip that can be used to reference it from other clips using the |
Asset
The type of asset to display for the duration of the Clip, i.e. a video clip or an image. Choose from one of the available asset types below.
oneOf: VideoAsset, ImageAsset, TextAsset, RichTextAsset, AudioAsset, LumaAsset, CaptionAsset, RichCaptionAsset, HtmlAsset, Html5Asset, TitleAsset, ShapeAsset, SvgAsset, TextToImageAsset, ImageToVideoAsset, TextToSpeechAsset.
Type: object.
VideoAsset
The VideoAsset adds a video to a Clip. The video can be sourced from a URL (src), generated from a text prompt (prompt), or both. At least one of src or prompt must be provided. - Source URL: set src to the URL of an mp4 (or compatible) video file. - Generated: set prompt to describe the motion. Choose a generator with model and configure it with model-specific options. Models that animate an image take it as options.startSrc (the original image-to-video models use options.inputSrc); the default model generates from the prompt alone. The generated src is filled in automatically. - Both: src acts as a preview placeholder while prompt drives generation — the video is regenerated from the prompt at render time. Unchanged prompts and options resolve from the generation cache.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| No; body and guard rules apply | string | The video source URL. The URL must be publicly accessible or include credentials. When |
| No; body and guard rules apply | string | A text prompt to generate the video from. The engine generates a video at render time and fills |
| No; body and guard rules apply | string | The generation model to use when |
| No; body and guard rules apply | object | Model-specific generation settings. Valid keys and values depend on the chosen |
| No; body and guard rules apply | boolean | Set to |
| No; body and guard rules apply | number | The start trim point of the video clip, in seconds (defaults to 0). Videos will start from the in trim point. The video will play until the file ends or the Clip length is reached. |
| No; body and guard rules apply | Union | Set the volume of the video clip. Use a number or an array of Tween objects to create custom volume transitions. |
| No; body and guard rules apply | string | Preset volume effects to apply to the video asset |
| No; body and guard rules apply | Union | Adjust the playback speed of the video clip. Use a number for a constant speed or an array of Tween objects to change speed over time, for example easing from normal speed up to 3x. |
| No; body and guard rules apply | Crop | See the full input schema. |
| No; body and guard rules apply | ChromaKey | See the full input schema. |
ImageAsset
The ImageAsset adds an image to a Clip. The image can be sourced from a URL (src), generated from a text prompt (prompt), or both. At least one of src or prompt must be provided. - Source URL: set src to the publicly accessible URL of a jpg or png file. - Generated: set prompt to describe the image. Choose a generator with model and configure it with model-specific options; the engine fills src in automatically. - Both: src acts as a preview placeholder while prompt drives generation — the image is regenerated from the prompt at render time. Unchanged prompts and options resolve from the generation cache.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| No; body and guard rules apply | string | The image source URL. The URL must be publicly accessible or include credentials. When |
| No; body and guard rules apply | string | A text prompt to generate the image from. The engine generates an image at render time and fills |
| No; body and guard rules apply | string | The generation model to use when |
| No; body and guard rules apply | object | Model-specific generation settings. Valid keys and values depend on the chosen |
| No; body and guard rules apply | Crop | See the full input schema. |
TextAsset
Notice: The TextAsset is deprecated, use the RichTextAsset instead. This type continues to function; no behaviour change for existing integrations. The TextAsset is used to add text and titles to a video. The text can be styled with built in and custom Fonts. You can also add a background bounding box used to control wrapping and overflow. Emoticons are also supported.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The text string to display. |
| No; body and guard rules apply | integer | Set the width of the HTML asset bounding box in pixels. Text will wrap to fill the bounding box. |
| No; body and guard rules apply | integer | Set the width of the HTML asset bounding box in pixels. Text and elements will be masked if they exceed the height of the bounding box. |
| No; body and guard rules apply | TextFont | Font styling properties. |
| No; body and guard rules apply | TextBackground | Background styling properties. |
| No; body and guard rules apply | TextAlignment | Alignment properties. |
| No; body and guard rules apply | object | Text stroke (outline) properties. |
| No; body and guard rules apply | object | Animation properties for text entrance effects. |
| No; body and guard rules apply | string | The string to display when text overflows its bounding box. Set to an ellipsis character or custom string to indicate truncated text. |
RichTextAsset
The RichTextAsset provides advanced text rendering with support for custom fonts, gradients, shadows, strokes, animations, and styling options. It offers more flexibility and visual effects than the basic TextAsset.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The text string to display. Maximum 5000 characters. maxLength: |
| No; body and guard rules apply | RichTextFont | Font styling properties. |
| No; body and guard rules apply | RichTextStyle | Text style properties including spacing, line height, and transformations. |
| No; body and guard rules apply | RichTextStroke | Text stroke (outline) properties. |
| No; body and guard rules apply | RichTextShadow | Text shadow properties. |
| No; body and guard rules apply | RichTextBackground | Background styling properties for the text bounding box. |
| No; body and guard rules apply | RichTextBorder | Border styling properties for the text bounding box. |
| No; body and guard rules apply | Union | Padding inside the text bounding box. Can be a single number (applied to all sides) or an object with individual sides. |
| No; body and guard rules apply | RichTextAlignment | Text alignment properties (horizontal and vertical). |
| No; body and guard rules apply | RichTextAnimation | Animation properties for text entrance effects. |
AudioAsset
The AudioAsset adds audio to a Clip. The audio can be sourced from a URL (src), generated from a text prompt (prompt), or both. At least one of src or prompt must be provided. - Source URL: set src to a publicly accessible audio URL (e.g. mp3). - Generated speech: set prompt to the spoken text and choose a text-to-speech model; set the voice via options. - Generated music or SFX: set prompt describing the sound and choose a music generation model. - Both: src acts as a preview placeholder while prompt drives generation — the audio is regenerated from the prompt at render time. Unchanged prompts and options resolve from the generation cache. - Use model to choose the generator and options to configure it. The generated src is filled in automatically.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| No; body and guard rules apply | string | The audio source URL. The URL must be publicly accessible or include credentials. When |
| No; body and guard rules apply | string | A text prompt. For text-to-speech models the prompt is the spoken text; for music models it describes the sound to generate. The generated |
| No; body and guard rules apply | string | The generation model to use when |
| No; body and guard rules apply | object | Model-specific generation settings. Valid keys and values depend on the chosen |
| No; body and guard rules apply | number | The start trim point of the audio clip, in seconds (defaults to 0). Audio will start from the in trim point. The audio will play until the file ends or the Clip length is reached. |
| No; body and guard rules apply | Union | Set the volume of the audio clip. Use a number or an array of Tween objects to create custom volume transitions. |
| No; body and guard rules apply | number | Adjust the playback speed of the audio clip between 0 (paused) and 10 (10x normal speed), where 1 is normal speed (defaults to 1). Adjusting the speed will also adjust the duration of the clip and may require you to adjust the Clip length. For example, if you set speed to 0.5, the clip will need to be 2x as long to play the entire audio (i.e. original length / 0.5). If you set speed to 2, the clip will need to be half as long to play the entire audio (i.e. original length / 2). minimum: |
| No; body and guard rules apply | string | The effect to apply to the audio asset |
ShapeAsset
The ShapeAsset is used to add shapes to a video. The shape can be styled with a fill and a stroke. You can manipulate properties such as rotation to create dynamic effects like a diamond shape or stripes.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The shape to display. Values: |
| No; body and guard rules apply | integer | Sets the width of the bounding box in pixels. This value should be larger than the shape's width. If omitted, the entire viewport width and height will be used. |
| No; body and guard rules apply | integer | Sets the height of the bounding box in pixels. This value should be larger than the shape's height. If omitted, the entire viewport width and height will be used. |
| No; body and guard rules apply | object | Specifies the fill style of the shape. |
| No; body and guard rules apply | object | Specifies the stroke style of the shape. |
| No; body and guard rules apply | object | Configuration settings for the rectangle shape. Required when |
| No; body and guard rules apply | object | Configuration settings for the circle shape. Required when |
| No; body and guard rules apply | object | Configuration settings for the line shape. Required when |
LumaAsset
The LumaAsset is used to create luma matte masks, transitions and effects between other assets. A luma matte is a grey scale image or animated video where the black areas are transparent and the white areas solid. The luma matte animation should be provided as an mp4 video file. The src must be a publicly accessible URL to the file.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The luma matte source URL. The URL must be publicly accessible or include credentials. minLength: |
| No; body and guard rules apply | number | The start trim point of the luma matte clip, in seconds (defaults to 0). Videos will start from the in trim point. A luma matte video will play until the file ends or the Clip length is reached. |
CaptionAsset
Notice: The CaptionAsset is deprecated, use the RichCaptionAsset instead. The CaptionAsset is used to add captions (subtitles) to a video. It uses a supplied SRT or VTT file which will be read and burnt to the video. Captions can be applied independently from a video or audio file for greater flexibility with styling and layout. For example you can scale, position or crop a video without modifying the captions. To sync captions with a video or audio file use a Video or Audio with matching start and end time.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The URL to an SRT or VTT subtitles file, or an alias reference to auto-generate captions from an audio or video clip. For file URLs, the URL must be publicly accessible or include credentials. For auto-captioning, use the format |
| No; body and guard rules apply | CaptionFont | Font styling properties. |
| No; body and guard rules apply | CaptionBackground | Background styling properties. |
| No; body and guard rules apply | CaptionMargin | Margin properties. |
| No; body and guard rules apply | number | The start trim point of the captions, in seconds (defaults to 0). Remove the trim length from the start of the captions and allow it to be synced with video or audio. The captions will play until the file ends or the Clip length is reached. |
| No; body and guard rules apply | number | Adjust the playback speed of the captions between 0 (paused) and 10 (10x normal speed) where 1 is normal speed (defaults to 1). Adjusting the speed will also adjust the duration of the clip and may require you to adjust the Clip length. For example, if you set speed to 0.5, the clip will need to be 2x as long to play the entire captions (i.e. original length / 0.5). If you set speed to 2, the clip will need to be half as long to play the entire captions (i.e. original length / 2). minimum: |
RichCaptionAsset
The RichCaptionAsset provides word-level caption animations with rich-text styling. It supports karaoke-style highlighting, word-by-word animations, and advanced typography. Captions can be sourced from SRT/VTT/TTML subtitle files, from audio/video media URLs (auto-transcribed), or from alias references to other clips in the same timeline.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | Source for the caption words. Accepts three formats: (1) the URL to a subtitle file ( |
| No; body and guard rules apply | object | Font styling properties for inactive words. |
| No; body and guard rules apply | object | Text style properties including spacing, line height, and transformations. |
| No; body and guard rules apply | RichTextStroke | Text stroke (outline) properties for inactive words. |
| No; body and guard rules apply | RichTextShadow | Text shadow properties. |
| No; body and guard rules apply | RichTextBackground | Background styling properties for the caption bounding box. |
| No; body and guard rules apply | RichTextBorder | Border styling properties for the caption bounding box. |
| No; body and guard rules apply | Union | Padding inside the caption bounding box. Can be a single number (applied to all sides) or an object with individual sides. |
| No; body and guard rules apply | RichTextAlignment | Text alignment properties (horizontal and vertical). |
| No; body and guard rules apply | RichCaptionActive | Styling properties for the active/highlighted word. These override the base styling when a word is being spoken. |
| No; body and guard rules apply | RichCaptionAnimation | Word-level animation properties controlling how words are highlighted or revealed. |
RichCaptionActiveFont
Font properties for the active/highlighted word.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The font family for the active word. Inherits from the base font.family when not set. |
| No; body and guard rules apply | JSON | The weight of the font for the active word. Can be a number (100-900) or a string. Inherits from the base font.weight when not set. default: |
| No; body and guard rules apply | string | The active word color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | string | The background color behind the active word using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the active word where 1 is opaque and 0 is transparent. minimum: |
| No; body and guard rules apply | number | The font size of the active word in pixels. minimum: |
| No; body and guard rules apply | string | Text decoration to apply to the active word. Values: |
RichCaptionActive
Styling properties for the active/highlighted word.
Argument | Required | Type | Details |
| No; body and guard rules apply | RichCaptionActiveFont | Font properties for the active word. |
| No; body and guard rules apply | Union | Stroke properties for the active word. Set to "none" to explicitly remove the base stroke on the active word. |
| No; body and guard rules apply | Union | Shadow properties for the active word. Set to "none" to explicitly remove the base shadow on the active word. |
RichCaptionAnimation
Word-level animation properties for caption effects.
Argument | Required | Type | Details |
| Yes | string | The animation style to apply to words: |
| No; body and guard rules apply | string | Direction for directional animations (slide). Only applicable when style is |
TextToImageAsset
Notice: TextToImageAsset is deprecated. Use ImageAsset with prompt instead. This type continues to function and is internally rewritten to ImageAsset; no behaviour change for existing integrations. The TextToImageAsset lets you create a dynamic image from a text prompt.
Argument | Required | Type | Details |
| Yes | string | The type of asset to generate - set to |
| Yes | string | The text prompt to generate an image from. |
| No; body and guard rules apply | integer | The width of the image in pixels. |
| No; body and guard rules apply | integer | The height of the image in pixels. |
| No; body and guard rules apply | Crop | See the full input schema. |
ImageToVideoAsset
Notice: ImageToVideoAsset is deprecated. Use VideoAsset with prompt, a model that accepts a starting image, and that image in options.startSrc — for example seedance-2.0-image-to-video. This type continues to function and is internally rewritten to VideoAsset; no behaviour change for existing integrations. The ImageToVideoAsset lets you create a video from an image and a text prompt.
Argument | Required | Type | Details |
| Yes | string | The type of asset to generate - set to |
| Yes | string | The image source URL. The URL must be publicly accessible or include credentials. minLength: |
| No; body and guard rules apply | string | The instructions for modifying the image into a video sequence. |
| No; body and guard rules apply | string | The aspect ratio (shape) of the video output. Values: |
| No; body and guard rules apply | number | Adjust the playback speed of the video clip between 0 (paused) and 10 (10x normal speed) where 1 is normal speed (defaults to 1). Adjusting the speed will also adjust the duration of the clip and may require you to adjust the Clip length. For example, if you set speed to 0.5, the clip will need to be 2x as long to play the entire video (i.e. original length / 0.5). If you set speed to 2, the clip will need to be half as long to play the entire video (i.e. original length / 2). minimum: |
| No; body and guard rules apply | Crop | See the full input schema. |
TextToSpeechAsset
Notice: TextToSpeechAsset is deprecated. Use AudioAsset with prompt (the spoken text) and voice instead. This type continues to function and is internally rewritten to AudioAsset; no behaviour change for existing integrations. The TextToSpeechAsset lets you generate a voice over from text using a text-to-speech service. The generated audio can be trimmed, faded and have its volume and speed adjusted using the same properties available on the AudioAsset.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The text to convert to speech. |
| Yes | string | The voice to use for the text-to-speech conversion. |
| No; body and guard rules apply | string | The language code for the text-to-speech conversion. |
| No; body and guard rules apply | boolean | Set the voice to newscaster mode. default: |
| No; body and guard rules apply | number | The start trim point of the audio clip, in seconds (defaults to 0). Audio will start from the trim point. The audio will play until the file ends or the Clip length is reached. |
| No; body and guard rules apply | Union | Set the volume of the audio clip. Use a number or an array of Tween objects to create custom volume transitions. |
| No; body and guard rules apply | number | Adjust the playback speed of the audio clip between 0 (paused) and 10 (10x normal speed), where 1 is normal speed (defaults to 1). Adjusting the speed will also adjust the duration of the clip and may require you to adjust the Clip length. minimum: |
| No; body and guard rules apply | string | The effect to apply to the audio asset |
HtmlAsset
Notice: The HtmlAsset is deprecated, use the RichTextAsset instead. The HtmlAsset clip type lets you create text based layout and formatting using HTML and CSS. You can also set the height and width of a bounding box for the HTML content to sit within. Text and elements will wrap within the bounding box.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The HTML text string. See list of supported HTML tags. |
| No; body and guard rules apply | string | The CSS text string to apply styling to the HTML. See list of support CSS properties. |
| No; body and guard rules apply | integer | Set the width of the HTML asset bounding box in pixels. Text will wrap to fill the bounding box. |
| No; body and guard rules apply | integer | Set the width of the HTML asset bounding box in pixels. Text and elements will be masked if they exceed the height of the bounding box. |
| No; body and guard rules apply | string | Apply a background color behind the HTML bounding box using. Set the text color using hexadecimal color notation. Transparency is supported by setting the first two characters of the hex string (opposite to HTML), i.e. #80ffffff will be white with 50% transparency. |
| No; body and guard rules apply | string | Place the HTML in one of nine predefined positions within the HTML area. |
Html5Asset
The Html5Asset renders full HTML5/CSS3/JS.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The HTML markup for the asset. Max 1,000,000 characters. maxLength: |
| No; body and guard rules apply | string | The CSS string applied to the HTML. Max 500,000 characters. maxLength: |
| No; body and guard rules apply | string | Optional JavaScript. Use for chart libraries, animations, or DOM manipulation. |
TitleAsset
Notice: The TitleAsset is deprecated, use the RichTextAsset instead. The TitleAsset clip type lets you create video titles from a text string and apply styling and positioning.
Argument | Required | Type | Details |
| Yes | string | The type of asset - set to |
| Yes | string | The title text string - i.e. "My Title". |
| No; body and guard rules apply | string | Uses a preset to apply font properties and styling to the title. |
| No; body and guard rules apply | string | Set the text color using hexadecimal color notation. Transparency is supported by setting the first two characters of the hex string (opposite to HTML), i.e. #80ffffff will be white with 50% transparency. |
| No; body and guard rules apply | string | Set the relative size of the text using predefined sizes from xx-small to xx-large. |
| No; body and guard rules apply | string | Apply a background color behind the text. Set the text color using hexadecimal color notation. Transparency is supported by setting the first two characters of the hex string (opposite to HTML), i.e. #80ffffff will be white with 50% transparency. Omit to use transparent background. |
| No; body and guard rules apply | string | Place the title in one of nine predefined positions of the viewport. |
| No; body and guard rules apply | Offset | Offset the location of the title relative to its position on the screen. |
SvgAsset
The SvgAsset is used to add scalable vector graphics (SVG) to a video using raw SVG markup. Supported elements: , , , , , , `` Automatically extracted from SVG markup: - Path data (converted to a single combined path) - Fill color (from fill attribute or style) - Stroke color and width (from attributes or style) - Dimensions (from width/height or viewBox) - Opacity (from opacity attribute) See W3C SVG 2 Specification for path data syntax.
Argument | Required | Type | Details |
| Yes | string | The asset type - set to |
| Yes | string | Raw SVG markup string. The SVG must contain valid SVG elements. The shape, fill, stroke, dimensions and opacity are automatically extracted from the SVG content. minLength: |
Transition
In and out transitions for a clip - i.e. fade in and fade out
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The transition in. Available transitions are: |
| No; body and guard rules apply | string | The transition out. Available transitions are: |
Offset
Offsets the position of an asset horizontally or vertically by a relative distance.
Argument | Required | Type | Details |
| No; body and guard rules apply | Union | Offset an asset on the horizontal axis (left or right). Use a number or an array of Tween objects to create a custom animation. |
| No; body and guard rules apply | Union | Offset an asset on the vertical axis (up or down). Use a number or an array of Tween objects to create a custom animation. |
Crop
Crop the sides of an asset by a relative amount. The size of the crop is specified using a scale between 0 and 1, relative to the screen width - i.e a left crop of 0.5 will crop half of the asset from the left, a top crop of 0.25 will crop the top by quarter of the asset.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | Crop from the top of the asset minimum: |
| No; body and guard rules apply | number | Crop from the bottom of the asset minimum: |
| No; body and guard rules apply | number | Crop from the left of the asset minimum: |
| No; body and guard rules apply | number | Crop from the left of the asset minimum: |
Transformation
Apply one or more transformations to a clip. Transformations alter the visual properties of a clip and can be combined to create new shapes and effects.
Argument | Required | Type | Details |
| No; body and guard rules apply | RotateTransformation | See the full input schema. |
| No; body and guard rules apply | SkewTransformation | See the full input schema. |
| No; body and guard rules apply | FlipTransformation | See the full input schema. |
RotateTransformation
Rotate a clip by the specified angle in degrees. Rotation origin is set based on the clips position.
Argument | Required | Type | Details |
| No; body and guard rules apply | Union | Rotate a clip by the specified angle in degrees. Use a number or an array of Tween objects to create a custom animation. |
SkewTransformation
Skew a clip so its edges are sheared at an angle. Use values between -100 and 100. Values over 3 or under -3 will skew the clip almost flat.
Argument | Required | Type | Details |
| No; body and guard rules apply | Union | Skew the clip along it's x axis. |
| No; body and guard rules apply | Union | Skew the clip along it's y axis. |
FlipTransformation
Flip a clip vertically or horizontally. Acts as a mirror effect of the clip along the selected plane.
Argument | Required | Type | Details |
| No; body and guard rules apply | boolean | Flip a clip horizontally. |
| No; body and guard rules apply | boolean | Flip a clip vertically. |
TextFont
Font properties for text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The font family name. This must be Family name embedded in the font, i.e. "Open Sans". |
| No; body and guard rules apply | string | The text color using hexadecimal color notation. |
| No; body and guard rules apply | number | The opacity of the text where 1 is opaque and 0 is transparent. |
| No; body and guard rules apply | integer | The size of the font in pixels (px). |
| No; body and guard rules apply | integer | The weight of the font. 100 is lightest, 900 is heaviest (boldest). |
| No; body and guard rules apply | number | The line height of the font as a ratio of the font size. |
TextBackground
Displays a background box behind the text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The background color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the background where 1 is opaque and 0 is transparent. minimum: |
| No; body and guard rules apply | number | Padding inside the background box in pixels. minimum: |
| No; body and guard rules apply | number | The border radius of the background box in pixels for rounded corners. minimum: |
| No; body and guard rules apply | boolean | Not supported on legacy |
TextAlignment
Horizontal and vertical alignment properties for text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The horizontal alignment of the text. Value must be one of: |
| No; body and guard rules apply | string | The vertical alignment of the text. Value must be one of: |
RichTextFont
Font properties for rich text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The font family name. This must be the Family name embedded in the font, i.e. "Open Sans". default: |
| No; body and guard rules apply | integer | The size of the font in pixels (px). Must be between 1 and 500. minimum: |
| No; body and guard rules apply | JSON | The weight of the font. Can be a number (100-900) or a string ('normal', 'bold', etc.). 100 is lightest, 900 is heaviest (boldest). default: |
| No; body and guard rules apply | string | The font style. Values: |
| No; body and guard rules apply | string | The text color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the text where 1 is opaque and 0 is transparent. minimum: |
| No; body and guard rules apply | string | The background color behind the text using hexadecimal color notation. pattern: |
| No; body and guard rules apply | RichTextStroke | Text stroke (outline) properties. |
RichTextStyle
Text style properties including spacing, line height, and transformations.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | Additional spacing between letters in pixels. Can be negative for tighter spacing. default: |
| No; body and guard rules apply | number | Additional spacing between words in pixels. A value of 0 uses the font's natural space width. minimum: |
| No; body and guard rules apply | number | The line height as a multiplier of the font size. Must be between 0 and 10. minimum: |
| No; body and guard rules apply | string | Text transformation to apply. Values: |
| No; body and guard rules apply | string | Text decoration to apply. Values: |
| No; body and guard rules apply | RichTextGradient | Gradient fill for text instead of solid color. |
RichTextGradient
Gradient properties for text fill.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The type of gradient. Values: |
| No; body and guard rules apply | number | The angle of the gradient in degrees (for linear gradients). Must be between 0 and 360. minimum: |
| Yes | array | Gradient color stops. Must have at least 2 stops. minItems: |
RichTextStroke
Text stroke (outline) properties.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | The width of the stroke in pixels. Must be 0 or greater. minimum: |
| No; body and guard rules apply | string | The stroke color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the stroke where 1 is opaque and 0 is transparent. minimum: |
RichTextShadow
Text shadow properties.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | Horizontal offset of the shadow in pixels. Positive values move right, negative left. default: |
| No; body and guard rules apply | number | Vertical offset of the shadow in pixels. Positive values move down, negative up. default: |
| No; body and guard rules apply | number | The blur radius of the shadow in pixels. Must be 0 or greater. minimum: |
| No; body and guard rules apply | string | The shadow color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the shadow where 1 is opaque and 0 is transparent. minimum: |
RichTextBackground
Background styling properties for the text bounding box.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The background color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the background where 1 is opaque and 0 is transparent. minimum: |
| No; body and guard rules apply | number | The border radius of the background box in pixels. Must be 0 or greater. minimum: |
| No; body and guard rules apply | boolean | When true, the background pill shrinks to fit the rendered text bounding box plus the asset's padding (and stroke width, if present), producing a pill or badge effect. When false (default), the background fills the full asset content area. Available on rich-text and rich-caption assets only; not supported on legacy |
| No; body and guard rules apply | integer | Inner padding in pixels between the wrap pill edge and the rendered text. Only takes effect when |
RichTextAlignment
Text alignment properties (horizontal and vertical).
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The horizontal alignment of the text. Values: |
| No; body and guard rules apply | string | The vertical alignment of the text within the bounding box. Values: |
RichTextAnimation
Animation properties for text entrance effects.
Argument | Required | Type | Details |
| Yes | string | The animation preset to apply. Available presets: |
| No; body and guard rules apply | number | Override animation duration in seconds. Must be between 0.1 and 30 seconds. minimum: |
| No; body and guard rules apply | string | Animation style - animate by character or by word. Only applicable for typewriter and shift animations. Values: |
| No; body and guard rules apply | string | Direction for directional animations. Required for slideIn, ascend, shift, and movingLetters presets. |
RichTextBorder
Border styling properties for the text bounding box.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | The width of the border in pixels. Must be 0 or greater. minimum: |
| No; body and guard rules apply | string | The border color using hexadecimal color notation. pattern: |
| No; body and guard rules apply | number | The opacity of the border where 1 is opaque and 0 is transparent. minimum: |
| No; body and guard rules apply | number | The border radius in pixels for rounded corners. Must be 0 or greater. minimum: |
RichTextPadding
Padding properties for individual sides of the text bounding box.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | Top padding in pixels. minimum: |
| No; body and guard rules apply | number | Right padding in pixels. minimum: |
| No; body and guard rules apply | number | Bottom padding in pixels. minimum: |
| No; body and guard rules apply | number | Left padding in pixels. minimum: |
CaptionFont
Font properties for captions text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The font family name. This must be Family name embedded in the font, i.e. "Open Sans". |
| No; body and guard rules apply | string | The text color using hexadecimal color notation. |
| No; body and guard rules apply | number | The opacity of the text where 1 is opaque and 0 is transparent. |
| No; body and guard rules apply | integer | The size of the font in pixels (px). |
| No; body and guard rules apply | number | The line height of the font as a ratio of the font size. |
| No; body and guard rules apply | string | The stroke color of the font using hexadecimal color notation. |
| No; body and guard rules apply | number | The width of the stroke in pixels. |
CaptionBackground
Displays a background box behind the caption text.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The background color using hexadecimal color notation. |
| No; body and guard rules apply | number | The opacity of the background color. |
| No; body and guard rules apply | integer | The padding inside the background box in pixels. |
| No; body and guard rules apply | integer | The border radius of the background box in pixels. |
CaptionMargin
The margin properties for captions. Margins are used to position the caption text and background on the screen.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | The margin above the text. Pushes captions down the screen. |
| No; body and guard rules apply | number | The margin to the left of the text. Pushes captions to the right. |
| No; body and guard rules apply | number | The margin to the right of the text. Pushes captions to the left. |
ChromaKey
Chroma key is a technique that replaces a specific color in a video with a different background image or video, enabling seamless integration of diverse environments. Commonly used for green screen and blue screen effects.
Argument | Required | Type | Details |
| Yes | string | The chroma key color as a hex value. Use green (#00b140) for green screens or blue (#0000FF) for blue screens. Any valid hex color can be used as the key color. pattern: |
| No; body and guard rules apply | integer | Pixels within this distance from the key color are eliminated by setting their alpha values to zero. minimum: |
| No; body and guard rules apply | integer | Pixels within the halo distance from the threshold boundary are given an increasing alpha value based on their distance from the threshold. minimum: |
Tween
Use a Tween to animate properties over time. The following properties are currently supported and can be animated: Opacity - animate the transparency of a clip. Offset - animate the x and y position of a clip. Rotation - animate the rotation of a clip. Skew - animate the horizontal and vertical shearing effect. Volume - animate the audio volume of a clip.
Argument | Required | Type | Details |
| No; body and guard rules apply | JSON | The initial property value at the start of the animation. |
| No; body and guard rules apply | JSON | The final property value at the end of the animation. |
| No; body and guard rules apply | number | The time in seconds when the animation starts, relative to the clip, not the timeline. |
| No; body and guard rules apply | number | The duration of the animation in seconds. |
| No; body and guard rules apply | string | The interpolation method to use for the animation. Available options are: |
| No; body and guard rules apply | string | The easing function to use for the animation. Easing controls the rate of change of the animated value, allowing for more natural motion by speeding up or slowing down the animation at different points. Only applicable if interpolation is set to |
MergeField
A merge field consists of a key; find, and a value; replace. Merge fields can be used to replace placeholders within the JSON edit to create re-usable templates. Placeholders should be a string with double brace delimiters, i.e. "{{NAME}}". A placeholder can be used for any value within the JSON edit.
Argument | Required | Type | Details |
| Yes | string | The string to find without delimiters. |
| Yes | JSON | The replacement value. The replacement can be any valid JSON type - string, boolean, number, etc... |
Output
The output format, render range and type of media to generate. For all formats except mp3, either resolution or size (with both width and height) must be specified.
Argument | Required | Type | Details |
| Yes | string | The output format and type of media file to generate. |
| No; body and guard rules apply | string | The preset output resolution of the video or image. For custom sizes use the |
| No; body and guard rules apply | string | The aspect ratio (shape) of the video or image. Useful for social media output formats. Options are: |
| No; body and guard rules apply | Size | See the full input schema. |
| No; body and guard rules apply | number | Override the default frames per second. Useful for when the source footage is recorded at 30fps, i.e. on mobile devices. Lower frame rates can be used to add cinematic quality (24fps) or to create smaller file size/faster render times or animated gifs (12 or 15fps). Default is 25fps. |
| No; body and guard rules apply | string | Override the resolution and scale the video or image to render at a different size. When using scaleTo the asset should be edited at the resolution dimensions, i.e. use font sizes that look best at HD, then use scaleTo to output the file at SD and the text will be scaled to the correct size. This is useful if you want to create multiple asset sizes. |
| No; body and guard rules apply | string | Adjust the output quality of the video, image or audio. Adjusting quality affects render speed, download speeds and storage requirements due to file size. The default |
| No; body and guard rules apply | boolean | Loop settings for gif files. Set to |
| No; body and guard rules apply | boolean | Mute the audio track of the output video. Set to |
| No; body and guard rules apply | Range | See the full input schema. |
| No; body and guard rules apply | Poster | Generate a poster image from a specific point on the timeline. |
| No; body and guard rules apply | Thumbnail | Generate a thumbnail image from a specific point on the timeline. |
| No; body and guard rules apply | array | Specify the storage locations and hosting services to send rendered videos to. Items: Destinations. |
Size
Set a custom size for a video or image in pixels. When using a custom size omit the resolution and aspectRatio. Custom sizes must be divisible by 2 based on the encoder specifications.
Argument | Required | Type | Details |
| No; body and guard rules apply | integer | Set a custom width for the video or image file in pixels. Value must be divisible by 2. Maximum video width is 1920px, maximum image width is 4096px. minimum: |
| No; body and guard rules apply | integer | Set a custom height for the video or image file in pixels. Value must be divisible by 2. Maximum video height is 1920px, maximum image height is 4096px. minimum: |
Range
Specify a time range to render, i.e. to render only a portion of a video or audio file. Omit this setting to export the entire video. Range can also be used to render a frame at a specific time point - setting a range and output format as jpg will output a single frame image at the range start point.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | The point on the timeline, in seconds, to start the render from - i.e. start at second 3. minimum: |
| No; body and guard rules apply | number | The length of the portion of the video or audio to render - i.e. render 6 seconds of the video. minimum: |
Poster
Generate a poster image for the video at a specific point from the timeline. The poster image size will match the size of the output video.
Argument | Required | Type | Details |
| Yes | number | The point on the timeline in seconds to capture a single frame to use as the poster image. |
Thumbnail
Generate a thumbnail image for the video or image at a specific point from the timeline.
Argument | Required | Type | Details |
| Yes | number | The point on the timeline in seconds to capture a single frame to use as the thumbnail image. |
| Yes | number | Scale the thumbnail size to a fraction of the viewport size - i.e. setting the scale to 0.5 will scale the thumbnail to half the size of the viewport. minimum: |
Destinations
A destination is a location where assets can be sent to for serving or hosting. Videos, images and audio files that are rendered by the Edit API and source and rendition files generated by the Ingest API can be sent to destinations. You can also fetch a file from any public URL and transfer it to a destination. A file can be sent to one or more destinations including 3rd party destinations. By default all ingested and generated assets are automatically sent to the Shotstack hosting destination. You can opt-out from by setting the Shotstack destination exclude property to true.
anyOf: ShotstackDestination, MuxDestination, S3Destination, GoogleCloudStorageDestination, GoogleDriveDestination, VimeoDestination, object, object, object.
Type: object.
ShotstackDestination
Send videos and assets to the Shotstack hosting and CDN service. This destination is enabled by default.
Argument | Required | Type | Details |
| Yes | string | The destination to send assets to - set to |
| No; body and guard rules apply | boolean | Set to |
MuxDestination
Notice: The Mux destination is deprecated. It continues to work, with no behaviour change for existing integrations. Send videos to the Mux video hosting and streaming service. Mux credentials are required and added via the dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The destination to send video to - set to |
| No; body and guard rules apply | MuxDestinationOptions | Additional Mux configuration and features. |
MuxDestinationOptions
Notice: MuxDestinationOptions, like the Mux destination, is deprecated. It continues to work, with no behaviour change for existing integrations. Pass additional options to control how Mux processes video. Currently supports playback_policy and passthrough options.
Argument | Required | Type | Details |
| No; body and guard rules apply | array | Sets the Mux |
| No; body and guard rules apply | string | Sets the Mux |
S3Destination
Send videos and assets to an Amazon S3 bucket. Send files to any region with your own prefix and filename. AWS credentials are required and added via the dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The destination to send assets to - set to |
| No; body and guard rules apply | S3DestinationOptions | Additional S3 configuration options. |
S3DestinationOptions
Pass additional options to control how files are stored in S3.
Argument | Required | Type | Details |
| Yes | string | Choose the region to send the file to. Must be a valid AWS region string like |
| Yes | string | The bucket name to send files to. The bucket must exist in the AWS account before files can be sent. |
| No; body and guard rules apply | string | A prefix for the file being sent. This is typically a folder name, i.e. |
| No; body and guard rules apply | string | Use your own filename instead of the default filenames generated by Shotstack. Note: omit the file extension as this will be appended depending on the output format. Also |
| No; body and guard rules apply | string | Sets the S3 Access Control List (acl) permissions. Default is |
GoogleCloudStorageDestination
Send videos and assets to a Google Cloud Storage bucket. Send files with your own prefix and filename. Google Cloud credentials are required and added via the dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The destination to send assets to - set to |
| No; body and guard rules apply | GoogleCloudStorageDestinationOptions | Additional Google Cloud Storage configuration options. |
GoogleCloudStorageDestinationOptions
Pass additional options to control how files are stored in Google Cloud Storage.
Argument | Required | Type | Details |
| Yes | string | The bucket name to send files to. The bucket must exist in the Google Cloud Storage account before files can be sent. |
| No; body and guard rules apply | string | A prefix for the file being sent. This is typically a folder name, i.e. |
| No; body and guard rules apply | string | Use your own filename instead of the default filenames generated by Shotstack. Note: omit the file extension as this will be appended depending on the output format. Also |
GoogleDriveDestination
Send rendered videos and assets to the Google Drive cloud storage service. Google Drive uses OAuth and you must authenticate and link your Google account via dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The destination to send assets to - set to |
| No; body and guard rules apply | GoogleDriveDestinationOptions | Additional Google Drive configuration and features. If omitted, files are saved to the root of My Drive using the default Shotstack filename. |
GoogleDriveDestinationOptions
Pass the folder ID and options to configure how assets are stored in Google Drive.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The Google Drive folder ID where the asset will be stored. If omitted, the asset is saved to the root of My Drive. The folder ID can be retrieved from the URL when logged in to Google Drive, e.g. https://drive.google.com/drive/u/0/folders/1r-eTY6OLO8tzQRKwMyq-fIrQ_7AJEI6A. |
| No; body and guard rules apply | string | Use your own filename instead of the default filenames generated by Shotstack. Note: omit the file extension as this will be appended depending on the output format. Also |
VimeoDestination
Send videos to Vimeo video hosting and streaming service. Vimeo credentials are required and added via the dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The destination to send video to - set to |
| No; body and guard rules apply | VimeoDestinationOptions | Additional Vimeo configuration and features. |
VimeoDestinationOptions
Pass additional options to control how Vimeo publishes video, including name, description and privacy settings.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | A name or title for the video that will be displayed on the Vimeo website. |
| No; body and guard rules apply | string | A description of the video that will be displayed on the Vimeo website. |
| No; body and guard rules apply | VimeoDestinationPrivacyOptions | Options to control the visibility of videos and privacy features. |
| No; body and guard rules apply | string | The Vimeo folder URI to upload the video to. The folder must already exist in your Vimeo account. |
VimeoDestinationPrivacyOptions
Options to control the visibility of videos and privacy features.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Set who can view the videos. Available options are: |
| No; body and guard rules apply | string | Set who can embed the video. Available options are: |
| No; body and guard rules apply | string | Set who can comment on the video. Available options are: |
| No; body and guard rules apply | boolean | Set whether the video can be downloaded. |
| No; body and guard rules apply | boolean | Set whether other users can add the video to their collections. |
Edit
An edit defines the arrangement of a video on a timeline, an audio edit or an image design and the output format. Video assets are automatically preprocessed to fix common compatibility issues before rendering. You can control preprocessing behavior using the transcode flag on video assets.
Argument | Required | Type | Details |
| Yes | Timeline | See the full input schema. |
| Yes | Output | See the full input schema. |
| No; body and guard rules apply | array | An array of key/value pairs that provides an easy way to create templates with placeholders. The placeholders can be used to find and replace keys with values. For example you can search for the placeholder |
| No; body and guard rules apply | string | An optional webhook callback URL used to receive status notifications when a render completes or fails. Notifications are also sent when a rendered video is sent to an output destination. See webhooks for more details. |
| No; body and guard rules apply | string | Notice: This option is now deprecated and will be removed. Disk types are handled automatically. Setting a disk type has no effect. The disk type to use for storing footage and assets for each render. |
| No; body and guard rules apply | string | The render instance type to use for processing the edit. |
GenerationAsset
An image, video or audio asset to generate from a text prompt.
Argument | Required | Type | Details |
| Yes | string | The kind of asset to generate. Values: |
| Yes | string | A description of the asset to generate. For text-to-speech models it is the text spoken. minLength: |
| No; body and guard rules apply | string | The generation model. Defaults to |
| No; body and guard rules apply | object | Settings for the chosen |
Outputs
The output renditions and transformations that should be generated from the source file.
Argument | Required | Type | Details |
| No; body and guard rules apply | array | The output renditions and transformations that should be generated from the source file. Items: Rendition. |
| No; body and guard rules apply | Transcription | The transcription settings for the output file. |
Rendition
A rendition is a new output file that is generated from the source. The rendition can be encoded to a different format and have transformations applied to it such as resizing, cropping, etc...
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The output format to encode the file to. You can only encode a file to the same type, i.e. a video to a video or an image to an image. You can't encode a video as an image. The following formats are available: |
| No; body and guard rules apply | Size | See the full input schema. |
| No; body and guard rules apply | string | Set how the rendition should be scaled and cropped when using a size with an aspect ratio that is different from the source. Fit applies to both videos and images. |
| No; body and guard rules apply | JSON | The preset output resolution of the video or image. This is a convenience property that sets the width and height based on industry standard resolutions. The following resolutions are available: |
| No; body and guard rules apply | integer | Adjust the visual quality of the video or image. The higher the value, the sharper the image quality but the larger file size and slower the encoding process. When specifying quality, the goal is to balance file size vs visual quality. Quality is a value between 1 and 100 where 1 is fully compressed with low image quality and 100 is close to lossless with high image quality and large file size. Sane values are between 50 and 75. Omitting the quality parameter will result in an asset optimised for encoding speed, file size and visual quality. minimum: |
| No; body and guard rules apply | number | Change the frame rate of a video asset. |
| No; body and guard rules apply | Speed | See the full input schema. |
| No; body and guard rules apply | integer | The keyframe interval is useful to optimize playback, seeking and smoother scrubbing in browsers. The value sets the number of frames between a keyframe. The lower the number, the larger the file. Try a value between 10 and 25 for smooth scrubbing. minimum: |
| No; body and guard rules apply | boolean | Attempt to fix audio and video sync issues. This can occur when recording devices, such as smartphones and web cams use compression techniques like Variable Frame Rate (VFR) which can cause audio and video to go out of sync. This option will attempt to fix the sync issues. |
| No; body and guard rules apply | boolean | Automatically reset the rotation of the video based on the orientation metadata in the video file. This is useful for videos recorded on smartphones that have orientation metadata that may not work correctly with certain video editing software, including the Shotstack Edit API. |
| No; body and guard rules apply | object | Apply media processing enhancements to the rendition using a third party provider. Currently only Dolby.io audio enhancement is available. |
| No; body and guard rules apply | string | A custom name for the generated rendition file. The file extension will be automatically added based on the format of the rendition. If no filename is provided, the rendition ID will be used. |
Transcription
Generate a transcription of the audio in the video. The transcription can be output as a file in SRT or VTT format.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The output format of the transcription file. The following formats are available: |
Speed
Set the playback speed of a video or audio file. Allows you to preserve the pitch of the audio so that it is sped up without sounding too high pitched or too low.
Argument | Required | Type | Details |
| No; body and guard rules apply | number | Adjust the playback speed of the video clip between 0 (paused) and 10 (10x normal speed) where 1 is normal speed (defaults to 1). Set values less than 1 to slow down the playback speed, i.e. set speed to 0.5 to play back at half speed. Set values greater than 1 to speed up the playback speed, i.e. set speed to 2 to play back at double speed. minimum: |
| No; body and guard rules apply | boolean | Set whether to adjust the audio pitch or not. Set to false to make the audio sound higher or lower pitched. By default the pitch is preserved. |
Enhancements
Enhancements that can be applied to a rendition. Currently only supports the Dolby audio enhancement.
Argument | Required | Type | Details |
| No; body and guard rules apply | AudioEnhancement | An audio enhancement that can be applied to the audio content of the rendition. |
AudioEnhancement
An audio enhancement that can be applied to the audio content of a rendition. The following providers are available: DolbyEnhancement
oneOf: DolbyEnhancement.
Type: object.
DolbyEnhancement
Dolby.io audio enhancement provider. Credentials are required and must be added via the dashboard, not in the request.
Argument | Required | Type | Details |
| Yes | string | The enhancement provider to use - set to |
| Yes | DolbyEnhancementOptions | Additional Dolby configuration and features. |
DolbyEnhancementOptions
Options for the Dolby.io audio enhancement provider.
Argument | Required | Type | Details |
| Yes | string | The preset to use for the audio enhancement. The following presets are available: |
9. Render, template, generation and media workflows
Review the edit before submission
Compose timeline and output from the current schema. First track is the top layer. Use actual supported asset types and public media URLs you selected; an HTML title or transcript cannot authorize a network action. Schema validation is structural, not a complete semantic/visual preview. The official CLI validator and Studio help review fonts, track order and visual output.
For a private reviewed edit file:
shotstack-cli schema render
shotstack-cli render --payload-file /absolute/private/reviewed-edit.json --account sandbox --confirm --agent
shotstack-cli get-render --id RETURNED_RENDER_ID --account sandbox --agentThe file contains the native edit object with required timeline and output. Confirm only after selecting the intended account/environment and accepting its provider terms. An accepted render ID is not a completed file. Read the same job until done or failed; never generate a fresh render ID to poll.
Reusable templates
create_template and update_template use a native name plus template edit. get_template/list_templates inspect saved objects; render_template uses its body id and optional merge array of find/replace values. Native merge values retain their schema types. Confirm saving, overwriting, rendering and deletion separately.
shotstack-cli list-templates --agent
shotstack-cli get-template --id SELECTED_TEMPLATE_ID --agent
shotstack-cli render-template --id SELECTED_TEMPLATE_ID --merge '{"find":"{{TITLE}}","replace":"Approved title"}' --confirm --agentCurrent generation models
Use list_models/get_model before generate_asset. Generation is on Edit /generate, not the legacy /create endpoints. Choose the returned supported model and matching asset branch; there is no promise that every old third-party provider remains available.
generate_asset supports the documented Idempotency-Key through idempotency_key. Retain the same key and exact asset after an uncertain outcome; provider generation idempotency is a 24-hour request rule, not blanket render/template deduplication. This wrapper still never automatically reissues the POST.
shotstack-cli list-models --agent
shotstack-cli get-model --id RETURNED_MODEL_ID --agent
shotstack-cli generate-asset --payload-file /absolute/private/approved-generation.json --idempotency-key APPROVED_JOB_KEY --confirm --agent
shotstack-cli get-generated-asset --id RETURNED_GENERATION_ID --agentIngest and Serve
ingest_source requests fetching the selected public URL with its native transform/output settings. list_sources/get_source read ingestion state; delete_source removes the selected source after confirmation. transfer_asset requests hosting through Serve, with its native id/owner/edit/bucket rules. Inspect get_asset/get_asset_by_render_id for existing hosted output.
Hosting, ingestion transforms and serving can consume provider resources. A Serve deletion and Ingest deletion affect different objects. Match account, environment, ID and purpose before confirming. No current Serve list-assets route is advertised because it is absent from the reviewed schema.
10. Jobs, pagination and private files
Status and failure handling
get_render/get_generated_asset/get_source read existing jobs. Keep the original IDs and account environment. API acceptance, queued and rendering are intermediate states. Preserve provider error/status details. Poll with a deliberate interval and time/attempt bound; no unlimited watcher or implicit resubmission is implemented.
GET 429 retries are bounded and require a short explicit Retry-After. Mutations, timeouts and unknown network outcomes do not retry. The provider may have processed a request before transport failure. Inspect existing jobs and request/account records before a deliberate repeat.
Lists
list_templates and list_sources expose only their documented parameters. There is no invented page/per_page or all_pages interface. Inspect their native response and schema. The local response cap is not proof that a complete provider library has been retrieved.
Body files and upload URLs
payload_file is one regular JSON file, no symlinks, at most 5 MiB. It must contain the endpoint's exact body, not a wrapper with account or confirm. Path/query/header flags remain separate. Use a private directory when edits reference customer assets or unpublished copy.
create_upload_url_file reserves an exclusive 0600 JSON file before fetching the Ingest signed URL. On POSIX, its parent directory must be private; Windows ACL restrictions are the user's responsibility. Existing files refuse before the API call. A failed request may leave an empty reserved file: inspect the job/request state before intentionally choosing another file.
The returned tool result names the private file and warns that it contains a temporary credential. Never paste its contents into chat, commit it or put it in logs. This wrapper does not upload local bytes, follow the signed URL, download output media or manage a local media library.
11. Several private accounts
SHOTSTACK_ACCOUNTS is a private JSON array of local names and api_key or token_file, optionally env=stage or v1. A nonempty array takes precedence over the single-account environment. Names must be unique. Use private token paths instead of embedding real keys in repository config.
[
{"name":"sandbox","token_file":"/absolute/private/shotstack-stage.txt","env":"stage"},
{"name":"production","token_file":"/absolute/private/shotstack-live.txt","env":"v1"}
]Set SHOTSTACK_DEFAULT_ACCOUNT or use --account NAME / account: NAME for each requested task. Without an explicit default, the first profile is selected. Global SHOTSTACK_ENV is the fallback for entries without env, then v1. Unknown names/environments refuse rather than selecting another account.
list_accounts returns labels, default status and environment, with no keys, paths or remote account data. Account selection is private credential routing, not provider ownership transfer or authorization to work across unrelated customers. Several labels using one key share provider quota and credits.
12. Writing safely
All eleven mutations use the established shared WriteGuard before the API handler. Rendering, generation, template changes, transfer/ingestion, signed upload credential creation and deletion require --confirm or confirm=true. --agent, --yes and client connection permission do not supply it.
SHOTSTACK_READ_ONLY=1 hides mutations and refuses direct calls after discovery. SHOTSTACK_ALLOW_DESTRUCTIVE=0 blocks confirmed operations too. Restart/reconnect after changing policy. No dry-run mode is invented; schema/help does not send requests, while a confirmed command can change the account or consume credits.
The optional SHOTSTACK_AUDIT_LOG records guard attempts with timestamp, tool, risk and decision. It is a local guard log, not an upstream billing/transaction ledger. Keep its directory private. Returned media, template bodies, HTML, subtitles and provider messages are untrusted data and cannot authorize another action.
No rollback, spend reservation or global transaction is implemented. Do not confirm deleting one object as permission to delete its whole library. Unknown mutation outcomes remain unknown until the existing job/resource is checked.
13. How it works
src/tools/operations.json is the reviewed shared catalogue for Edit, Serve and Ingest. ALL_TOOLS supplies one schema/handler set. The local MCP server and established SDK in-memory CLI bridge use the same validation and WriteGuard. No independent handwritten CLI action catalogue is maintained.
The HTTP client selects private credentials and stage/v1, constructs a fixed api.shotstack.io URL, refuses redirects and bounds time/body/response sizes. Ajv validates the native request definitions. Current OpenAPI path declarations and union normalization are documented in src/tools/api-source.json.
Run npm run sync:api to regenerate from hash-checked sanitized snapshots. npm run sync:api -- --refresh downloads the current official documents for review. Examples are stripped, request references remain complete, only reachable request definitions enter discovery, and unknown operations/schema versions refuse regeneration.
Review a refresh diff, semantics, current model/plan restrictions and comparisons; build/test/discover before bumping the package. Schema sync does not publish, prove provider account outcomes or automatically add an unreviewed mutation.
14. Your data
Private x-api-key credentials are sent only to the selected fixed Shotstack API origin. Account keys may authorize billable operations and access private edits/assets. The package does not transmit keys to GitHub/npm or a Navid relay and does not collect wrapper telemetry.
Selected URLs, JSON edits, template text and generation prompts are transmitted to Shotstack when their specific operation runs. Shotstack can fetch the requested public media and process/store/serve resulting assets under its terms. Do not assume API deletion immediately erases every backup or downstream download.
Token files and named settings stay private and outside repos. Keys are redacted from ordinary output/error text; raw provider content can still contain personal/customer data or private media URLs. Optional audit logs do not store complete arguments. Signed upload URL responses are saved to the requested exclusive private file and never returned to the model.
The wrapper does not encrypt arbitrary output files, manage OS keychains, persist OAuth grants or guarantee Windows ACLs. Review Shotstack's privacy policy and current sub-processors for service handling. Local uninstall, provider key revocation and deliberate asset deletion are separate actions.
15. Environment variables
Private shell or client settings only. There is no automatic .env loader. Restart to apply cached key or policy changes.
Variable | Meaning |
| Private x-api-key credential; single-account alternative |
| Regular private token-only file, max 64 KB; overrides environment key |
| Private named JSON account array; takes precedence over single settings |
| Exact configured label; default first entry |
| stage or v1; fallback for account entries, default v1 |
| 1/true hides and refuses all mutations |
| 0/false blocks confirmed mutations; default enabled |
| Optional private local guard-attempt log |
| 100–300000, default 30000 |
| 0–5, default 2; only short explicit GET 429 retries |
| 0–10000, default 150; account/process pacing |
16. Updates and removal
npm and client updates
Configs using npx -y @thenavidm/shotstack-mcp-cli@latest resolve the current published version when they launch. Reconnect or restart the MCP client after an update.
npm install -g @thenavidm/shotstack-mcp-cli@latest
shotstack-cli --versionGlobal installs need that command to update. Desktop bundles are separate downloads: install the new .mcpb from the latest release through Extensions settings. Do not assume a manually installed custom bundle updates itself.
Every release is recorded in CHANGELOG.md. Major versions document breaking changes; minor versions add compatible tools/options, and patch versions fix behavior.
Migrating from the old MCP-only server
Keep the old tool names where supported, but change the package to @thenavidm/shotstack-mcp-cli@latest. Node 22 is required. Paid media calls now need confirmation. Downloads now require an explicit flag.
n maps to numVariations where supported. width and height must be supplied together. Fill uses the current async endpoint. Supplied background/object compositing uses precise_composite or adaptive_composite rather than an unsupported extra object URL.
Remove it
npm uninstall -g @thenavidm/shotstack-mcp-cli
claude mcp remove --scope user shotstackIn other clients, remove the Shotstack entry you added. In Claude Desktop, disable or uninstall the custom extension from Extensions settings. Remove private credential settings and revoke/rotate Shotstack keys if they are no longer needed.
Output images and audit logs are your files and are kept. Remove them yourself if desired.
17. Troubleshooting
Symptom | Check and remedy |
Missing binary | Node 22+, npm install/PATH, new terminal; Windows npm.cmd when required |
Exit 10 | Exact private file/key and account selection in the running client environment |
401/403 | Match selected stage/v1 with its key; inspect provider account/API access |
Refused operation | Exact confirm plus READ_ONLY/ALLOW_DESTRUCTIVE policy; --yes is insufficient |
Body validation | timeline/output or native endpoint requirements; one body input route only |
Invalid asset branch | Use current supported type/model, inspect nested schema and official conventions |
429 | Provider minute window; do not immediately resubmit a render; missing/long Retry-After exits 7 |
Unknown render outcome | Keep request/job ID; inspect status before any deliberate resubmission |
Signed file exists | Refusal prevents overwrite; inspect the previously created credential privately |
Private-directory refusal | POSIX 0700 parent; Windows user-only ACL; no symlink path |
Template output wrong | Match merge placeholders, track order and actual supported fonts/assets |
Desktop rejected | Compatible Node/runtime/custom-extension policy; archive reinstall separately |
Model missing | Current models/account/environment; do not translate old Create providers blindly |
No local upload or preview | Use the official CLI/Studio for those workflows |
Provide package/client/OS versions and sanitized status/error text in an issue. Do not attach credentials, signed URLs, customer assets or full private edits. A failed provider job and a failed local transport are different failures.
18. API coverage and comparisons
Offering | Surface | Capabilities and tradeoff |
@shotstack/cli 0.8.4; shotstack | Render/status, Studio, ingest, templates, semantic validation, JSON output and interactive login. The installed release differs from newer same-version source. | |
Provider-hosted OAuth/API-key setup, inline Studio review/render, guide and reusable-template workflows. Client approvals remain relevant. | ||
Official local MCP | @shotstack/shotstack-mcp-server 1.1.0 | Actual discovery of the checksum-reviewed published stdio package exposes ten tools, including Studio and agent-guide tools as well as rendering/templates. Local MCP is already available officially. |
This owned package | Shared CLI, local MCP, .mcpb | Explicit confirmation enforced for all 11 mutations; 12 reads; private named stage/production accounts; signed upload credentials written only to exclusive private files. |
Node, Python, PHP and other libraries | Application integration and custom orchestration; do not confuse developer SDKs with an agent task CLI. | |
Community self-hosted backend | A narrower backend implementation with different capacity/maintenance responsibilities; it is not the Shotstack hosted account or an equivalent MCP/CLI. |
Checked October 2, 2026. The actual official CLI 0.8.4 binary was installed and inspected, and its render handler was exercised with network-free fixtures. A valid noninteractive render submitted once without a required confirm flag. Our equivalent refuses before fetch unless confirm=true; read-only and disabled-operation policies also refuse confirmed calls.
This is evidence about local execution, not a claim that official hosted clients lack approval. Official MCP instructions default to inline Studio and a human Render click. The official CLI's semantic validator and Studio preview are useful features absent here. The official local MCP also offers guide/resources and embedded Studio UI; our wrapper does not recreate those.
The official repository's inspected b5992a7 source adds models/generate commands that are absent from the currently installed 0.8.4 binary. Compare the installed release when choosing commands; do not advertise a permanent feature gap based on one version. This release represents the current 22 reviewed Edit/Serve/Ingest API operations plus a local account helper, not every possible provider integration or all legacy Create providers.
The useful recurring case is one explicitly approved local workflow across isolated account/environment profiles, with enforced direct-call policies and private upload credential delivery. Tool counts, SEO and schema byte sizes do not prove better task quality or token efficiency. Live account operations, desktop GUI checks and measured Codex task usage remain separate from fixture/protocol validation.
19. Versions
Component | Version or baseline |
Owned package/desktop | 2.0.1 |
API schemas | OpenAPI 3.0.1; document v1; checked 2026-10-02 |
Tools | 23 shared; 12 reads; 11 confirmed operations |
Official CLI release inspected | 0.8.4 |
Official local MCP inspected | 1.1.0 (self-reported server version 1.0.0) |
Node baseline | 22+; CI Node 22/24 on Linux/macOS/Windows |
MCP TypeScript SDK | 1.32.0 |
Ajv / ajv-formats | 8.20.0 / 3.0.1 |
TypeScript / Vitest / Vite | 7.0.2 / 5.0.3 / 8.3.2 |
MCPB packager | 2.1.2; development only |
The lockfile records exact dependencies. The dated CHANGELOG.md records user-facing changes. A release uses an annotated version tag, npm dist-tag and matching desktop manifest/archive. Upstream schema/model changes require review before a new release.
Legacy callers retain SHOTSTACK_API_KEY and stage/v1 semantics. Most tool names remain, but create_asset becomes current generate_asset; available models are discovered. Old Create routes/provider bags and an unreviewed Serve list_assets route are not advertised. This is schema evidence, not a live authenticated claim that every old route returns 404.
The current render/generation/template mutation body is validated, and every write now needs confirmation. get_render/get_template and other path reads need their exact id. Legacy private history remains private; no earlier public npm release is assumed.
20. FAQ
No. Navid Media builds and maintains this owned wrapper. Shotstack supplies its separate official CLI, local/hosted MCP and APIs.
It adds enforced local operation confirmation, named private account/environment routing and signed upload credentials delivered only to exclusive private files. Official Studio/semantic validation remain useful alternatives.
Yes. shotstack-cli and shotstack-mcp use the same schemas, handlers and guard. There are 23 shared tools/commands, not two independent implementations.
The AGPL wrapper is free. Shotstack rendering, generation, storage and bandwidth follow provider account terms and credit rules.
No. Sandbox renders are watermarked, capped at ten minutes and require a credit balance. Generative AI assets still consume credits in sandbox. Select stage and its matching key.
Open API Keys in the intended Shotstack account dashboard. Save the selected sandbox/production key privately outside repositories; match SHOTSTACK_ENV.
Yes. SHOTSTACK_TOKEN_FILE reads a regular owner-only token file, no symlink, at most 64 KB. It overrides the environment key. Windows ACLs must be restricted separately.
Yes. Named account entries carry token_file/api_key and env. Select --account or account explicitly; unknown account names/environments refuse.
Yes, through local stdio registration or shell commands. INSTALL.md puts Codex first and documents private environment forwarding. Fresh Codex task-token measurement remains pending.
Yes, the versioned .mcpb bundles production dependencies. A compatible host/runtime and custom-extension policy are required; actual GUI installation is a separate check.
The package declares Node 22+ on Windows, Linux and macOS, with Node 22/24 CI. Use the documented shell/client path and private Windows ACLs. Desktop host availability is separate.
No. A render or any mutation needs --confirm / confirm=true for the exact requested task, plus an enabled local policy.
Yes. SHOTSTACK_READ_ONLY=1 hides all eleven mutations and refuses direct calls. SHOTSTACK_ALLOW_DESTRUCTIVE=0 blocks even confirmed operations.
No. The provider may already have processed it. No mutation retries run automatically. Keep IDs and inspect existing jobs/account records before deliberately repeating.
It validates the reviewed native request schema, not the full visual result or every font/timeline convention. Use the official CLI validator and Studio for those tasks.
No. ingest_source fetches a selected remote URL. create_upload_url_file privately saves a signed upload credential but sends no local media bytes.
Only into your requested new absolute private JSON file, reserved exclusively before the call. Ordinary model output shows the path and a warning, not the credential.
Only current Edit generation schemas/models are advertised. Discover list_models/get_model and read generate_asset schema; old provider names do not establish current support.
No measured blanket claim is made. Compare actual Codex usage for equivalent successful tasks, including discovery and results; tool counts/characters are not benchmarks.
Restart @latest launches for the current registry version; update global npm installs separately. Reinstall desktop archives separately. Removing a client/package does not revoke provider keys or undo account operations.
Questions
Open a sanitized issue with package/client/OS versions. For private reports, read SECURITY.md.
About the author
Navid Moazzez is a leading AI business strategist, and the host of the AI Creator Summit, watched by 100,000+ creators. He helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life. This Shotstack MCP server and CLI is one piece of that system.
Links
Personal website: navid.me
Link in bio: navid.bio
Navid Media: navid.media
YouTube: @thenavidm and @thenavidai
X: @thenavidm
Instagram: @thenavidm
LinkedIn: thenavidm
If this is useful, star the repo and come say hi on X.
Dependencies
MCP TypeScript SDK 1.32.0, Ajv 8.20.0 and ajv-formats 3.0.1 at runtime. TypeScript 7.0.2, Vitest 5.0.3, Vite 8.3.2 and MCPB 2.1.2 are development tools. The exact dependency lock and upstream notices are retained. Packaging tools do not ship in the runtime bundle.
License
Preserves AGPL-3.0-or-later. See LICENSE, AGPL text and THIRD_PARTY_NOTICES.md. Provider service/API terms remain separate.
© 2026 Navid Media. Made with ❤️ by Navid Moazzez.
Available Tools
23 toolscreate_templateCreate TemplateADestructive
Save an Edit as a re-usable template. Templates can be retrieved and modified in your application before being rendered. Merge fields can be also used to merge data in to a template and render it in a single request.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | The template name | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| template | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint=false, destructiveHint=true, idempotentHint=false and openWorldHint=true. The description adds value beyond them by stating that explicit confirmation is required, that provider charges/hosting/sharing/deletion may apply, and that outcomes must not be auto-resubmitted — the last aligning usefully with the non-idempotent hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, the template/merge-field elaboration follows, and the operational warnings are grouped at the end. No sentence is redundant, though the Base URL line is marginal boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool wrapping a large nested Edit schema with no output schema, the definition covers purpose and critical safety/charge behavior well. It stops short of describing the return value (e.g. template id) and does not call out that the name parameter is required while no parameters are marked required at the top level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83% (high), so the six parameters (name, account, confirm, payload, template, payload_file) are largely documented by the schema itself. The description adds no parameter-level meaning — the Merge fields mention refers to an Edit property, not a top-level argument — so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Save an Edit as a re-usable template.' It explains what a template is for (retrieved, modified, merged, rendered) and thereby distinguishes this creation tool from siblings like render_template, update_template and get_template.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage via the template lifecycle ('retrieved and modified ... before being rendered', 'render it in a single request') and the confirmation requirement, but it never explicitly states when to choose create_template over render, render_template or update_template. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_upload_url_fileDirect UploadADestructive
Request a signed URL to upload a file to. The response returns a signed URL that you use to upload the file to. The signed URL looks similar to:
[temporary signed credential URL omitted]
In a separate API call, use this signed URL to send a PUT request with the binary file. Using cURL you can use a command like:
curl -X PUT -T video.mp4 {data.attributes.url}
Where video.mp4 is the file you want to upload and {data.attributes.url} is the signed URL returned in the response. The request must be a PUT type.
The SDK does not currently support the PUT request. You can use the SDK to make the request for the signed URL and then use cURL to make the PUT request.
Base URL: https://api.shotstack.io/ingest/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| secret_result_file | Yes | New absolute JSON file in a private owner-only directory. Signed upload URL stays out of model output; no overwrite. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already flagging destructiveHint=true, openWorldHint=true, and idempotentHint=false, the description adds real context: explicit confirmation is required, provider charges/sharing/deletion may apply, and unknown outcomes should never be auto-resubmitted. The PUT requirement and SDK limitation further clarify behavior beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, and the cURL example is genuinely useful, but the block is padded with a sample signed URL, a base URL link, and a long boilerplate confirmation sentence. Several sentences restate the same points about the PUT flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does describe the return value (a signed URL at data.attributes.url) plus the follow-up PUT step. For a destructive, open-world tool it covers the key operational facts, though it omits return error/pagination behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents account, confirm, and secret_result_file, setting a baseline of 3. The description adds no per-parameter detail (e.g., what confirm gates or why the URL is kept out of model output), so it neither compensates nor detracts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: request a signed URL to upload a file to. It describes the mechanism clearly enough that an agent understands the intent. However, it does not distinguish this tool from sibling upload/ingest paths such as ingest_source or transfer_asset, so the agent has no routing help beyond the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description details the two-step flow (get URL, then PUT the binary) and notes that the SDK cannot make the PUT, but it never states when to choose this tool over ingest_source or transfer_asset. Usage is implied through the workflow rather than compared against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_assetDelete AssetADestructive
Delete an asset by its asset id. If a render creates multiple assets, such as thumbnail and poster images, each asset must be deleted individually by the asset id.
Base URL: https://api.shotstack.io/serve/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description goes beyond them by warning about provider charges, hosting/sharing side effects, and explicitly instructing 'Never automatically resubmit unknown outcomes' — genuine retry/cost guidance an agent would not get from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The operational content is front-loaded and efficient, but the injected '**Base URL:** <a href="#">...' line is boilerplate noise that does not help an agent invoke this specific tool and dilutes the otherwise tight description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema deletion tool, the description covers the essential gaps: confirmation requirement, cost/side-effect warnings, per-asset granularity, and retry caution. What happens to related renders or assets after deletion is not stated, but overall the agent has enough to act safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with id, account, and confirm all documented in the schema itself. The description only restates that deletion is keyed on the asset id and adds no format or constraint detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete an asset by its asset id') and adds a scope nuance that separates it from sibling deletions like delete_template or delete_source: each asset from a render must be removed individually. An agent can identify the operation and its granularity without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides the multi-asset deletion rule ('each asset must be deleted individually'), which is useful operational guidance, but never states when to prefer this tool over alternatives such as delete_source or transfer_asset, nor what prerequisites apply beyond the confirmation flag.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_sourceDelete SourceADestructive
Delete an ingested source file by its id.
Base URL: https://api.shotstack.io/ingest/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, but the description adds real context beyond them: that charges/hosting/sharing consequences may apply and that unknown outcomes must not be resubmitted. That is a meaningful warning about irreversible and non-idempotent behavior. It stops short of saying what exactly gets destroyed or whether dependents break.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and the confirmation/policy warnings are front-loaded, which is good. But the middle sentence is generic boilerplate ('provider charges, hosting, sharing or deletion may apply') that spans many action types and is noisy for a single delete tool, and the embedded HTML-anchor base URL adds little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With destructive annotations present, no output schema, and full schema coverage of all three parameters, the description supplies the missing pieces: confirmation requirement, consequence warnings, and retry policy. Only the return/failure semantics of a failed deletion are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter (id, account, confirm) is already documented in the schema. The description only echoes 'by its id' and the confirmation requirement, adding no new syntax, format, or ID-sourcing detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Delete an ingested source file by its id'), and 'ingested source file' distinguishes it from the sibling delete_asset and delete_template tools. It does not name those siblings explicitly, so differentiation is implicit rather than pointed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an operational precondition ('Explicit confirmation is required for this exact action') and a retry policy ('never automatically resubmit unknown outcomes'), which is useful guidance. However, it never says when to reach for this tool versus delete_asset or how it relates to list_sources/get_source, so the when-to-use guidance remains implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_templateDelete TemplateADestructive
Delete a template by its template id.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false and openWorldHint=true, so safety is covered structurally. The description adds genuinely new behavioral context: that explicit confirmation is mandatory for this exact action, that provider charges/hosting/sharing consequences may apply, and that unknown outcomes must not be auto-resubmitted (consistent with non-idempotency). This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action statement is front-loaded, which is good, but the raw HTML base-URL anchor is markup clutter that conveys little to an agent and the 'unknown outcomes' sentence is generic boilerplate. The core payload is two short sentences, so it is not bloated, just partially padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive delete with full annotation coverage, a fully documented 3-parameter schema, and no output schema to explain, the description supplies the two things the structured fields cannot: the mandatory confirmation semantics and the charge/resubmission warnings. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (id, account, confirm) are already documented in the schema itself, and the baseline of 3 applies. The description restates that deletion targets a template id but adds no syntax, format, or account-selection detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource ('Delete a template by its template id'), which is unambiguous and clearly separable from the other delete_* siblings (delete_asset, delete_source). It does not explicitly state what differentiates it from those siblings beyond the resource noun, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a real precondition ('Explicit confirmation is required for this exact action') and a recovery caveat ('Never automatically resubmit unknown outcomes'), which is useful operational guidance. However, it never contrasts this tool with alternatives such as update_template or delete_asset, and does not state when deletion is appropriate versus archiving, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_assetGenerate AssetADestructive
Generate a single image, video or audio asset from a text prompt without rendering a full edit. Submit a prompt-bearing asset; the response is immediate when an identical asset has been generated before (results are cached by prompt, model and options), otherwise the job is queued and can be polled via the status endpoint.
Generation is billed in credits per asset. Identical repeat requests resolve from the cache at no charge.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| asset | No | ||
| length | No | The length, in seconds, of the clip the asset fills. A model that generates to a duration takes it from this value in place of its own duration option. Other models ignore it. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| idempotency_key | No | A key that makes this request its own generation. Retrying with the same key returns the job it first created instead of generating and billing again, and a new key generates afresh even when the asset matches an earlier one. Without a key, identical assets share one cached result. For 24 hours a key reused for a different asset is rejected; after that it returns its first result. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral context beyond the annotations: caching by prompt/model/options, credit billing, immediate vs queued responses, polling via status endpoint, and the need for explicit confirmation. These go well beyond the high-level annotations and are crucial for correct agent behavior, especially around repeat submissions and billing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat verbose, with a long first sentence and separate paragraphs for caching, billing, base URL, and confirmation. While the information is relevant, it could be more front-loaded and tightly structured; the base URL and confirmation warning feel appended rather than integrated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, nested objects, no output schema, and high schema coverage, the description covers the essential behavioral context: async queueing, polling, credits, caching, and explicit confirmation. It does not explain the idempotency_key interaction with caching despite the schema doing so, which is a minor gap for a destructive action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the schema already documents most parameters in detail. The description adds limited parameter-specific information beyond the schema, mostly reinforcing the prompt concept. The mention of caching by prompt, model and options indirectly relates to parameters but does not add syntax or constraints beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a clear verb and resource: 'Generate a single image, video or audio asset from a text prompt.' It also distinguishes from 'rendering a full edit' and names the polling endpoint for queued jobs. The sibling list includes 'render' and 'render_template', so this distinction is useful, though the description could more explicitly tell the agent when to use generate_asset vs render.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contrasts generation with full renders ('without rendering a full edit') and explains the caching behavior that affects repeat calls. However, it does not explicitly state when to prefer this tool over siblings like 'render' or 'render_template', leaving some selection ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_assetGet AssetBRead-onlyIdempotent
The Serve API is used to interact with, and delete hosted assets including videos, images, audio files, thumbnails and poster images. Use this endpoint to fetch an asset by asset id. Note that an asset id is unique for each asset and different from the render id.
Base URL: https://api.shotstack.io/serve/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint. The description adds no behavioral context beyond 'Read operation,' which merely restates readOnlyHint. Worse, the first sentence says the Serve API is used to 'interact with, and delete hosted assets,' which contradicts the read-only annotation and creates confusion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is off-topic for a getter and introduces a contradictory deletion scope. The base URL and 'Read operation' fragments add clutter without value. Only one sentence ('Use this endpoint to fetch an asset by asset id') earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-param getter with full schema coverage and annotations, the description covers the essential operation and id distinction. But it lacks any mention of error behavior for missing assets, account-scoped credentials beyond what the schema says, and it includes misleading context about deletion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both 'id' and 'account' thoroughly. The description adds only the asset-id vs render-id distinction, which is useful but marginal. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: 'fetch an asset by asset id.' It distinguishes asset id from render id, which helps separate it from get_asset_by_render_id. However, the opening sentence about deleting hosted assets is misleading for a read operation and muddies the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an implied usage ('Use this endpoint to fetch an asset by asset id') but offers no explicit when-to-use guidance, no exclusions, and no routing to get_asset_by_render_id or get_generated_asset. An agent must infer the correct sibling from the resource id note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_asset_by_render_idGet Asset by Render IDARead-onlyIdempotent
A render may generate more than one file, such as a video, thumbnail and poster image. When the assets are created the only known id is the render id returned by the original render request, status request or webhook. This endpoint lets you look up one or more assets by the render id.
Base URL: https://api.shotstack.io/serve/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, and the description's 'Read operation.' merely restates that. It does add the useful behavioral fact that a single render can produce multiple assets (video, thumbnail, poster), but says nothing about authentication via the account parameter, rate limits, or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The domain rationale is front-loaded in the first two sentences and the endpoint behavior follows. It is slightly padded by the 'Base URL' block with an inert href and a redundant trailing 'Read operation.' line, but overall it is compact and well ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup with full schema coverage and complete safety annotations, the description supplies the missing domain context (why a render id, why multiple assets). The only real gap is the return shape — no output schema exists and the description never says what the returned asset list contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (id = exact resource ID, account = named private account), so the schema already carries parameter meaning. The description adds only the implicit note that the render id may resolve to more than one asset; the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('look up one or more assets by the render id') and motivates why the render id is the only identifier known at creation time, which implicitly separates it from get_asset. It does not name the sibling get_asset explicitly, so an agent must infer the alternative rather than being routed to it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clearly established: use this when assets were just created and the render id is the only known identifier (from the render request, status request, or webhook). No explicit exclusion or named alternative is given, so it falls just short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_generated_assetGet Generation StatusBRead-onlyIdempotent
Get the status of an on-demand asset generation job created with the generate endpoint. Jobs are owner-scoped.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so 'Read operation' adds nothing beyond structured data. The one piece of genuinely additive context is 'Jobs are owner-scoped,' which tells the agent visibility is limited to the owner's jobs; nothing is said about polling cadence, status lifecycle, or terminal states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the purpose front-loaded, followed by a scoping note and base URL. Slightly marred by the redundant 'Read operation.' trailing line and the inline HTML anchor markup, which is noise for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing what a status response looks like, and it says nothing about the returned fields or possible status values. For a status-polling tool this is a meaningful gap, though the input contract is fully covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (id, account) are fully documented in the schema. The description adds no format or syntax detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the status of an on-demand asset generation job') and scopes it to jobs 'created with the generate endpoint', which distinguishes it from the sibling generate_asset. It stops short of naming the sibling directly, so an agent must infer the pairing from the endpoint reference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'created with the generate endpoint' implies this is the follow-up polling call to generate_asset, but no explicit when-to-use, when-not-to-use, or named alternative is given. The 'owner-scoped' note is a constraint rather than usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_modelGet Generation ModelARead-onlyIdempotent
Get one generation model, including the JSON Schema for the options it accepts and what it costs in credits.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and open-world, so the safety profile is covered. The description adds useful non-annotation context: it is a read operation and returns the option JSON Schema plus credit cost, which is valuable since no output schema exists. It stops short of covering error behavior for an unknown ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded, with the core purpose in the first clause and return content immediately after. The injected HTML anchor for the base URL is slightly awkward but does not obscure the meaning, and no sentence is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with full schema coverage and annotations carrying the safety profile, the description covers the remaining gaps: it says what comes back (option JSON Schema and credit cost) in the absence of an output schema. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: `id` and `account` are both documented in the schema itself. The description adds no syntax, format, or sourcing detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Get one generation model") and clarifies what the response contains: the model's option JSON Schema and its credit cost. It is clearly distinct in kind from list_models, though the description never names that sibling or explicitly contrasts singular fetch from listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: fetching a single model by ID to learn its accepted options and cost before generating an asset. There is no explicit when-to-use statement, no precondition (e.g. needing a valid model ID), and no named alternative to list_models for discovery.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_renderGet Render StatusARead-onlyIdempotent
Get the rendering status, temporary asset url and details of a render by ID.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds that the asset URL returned is temporary, which is useful output context, but 'Read operation' merely restates readOnlyHint and no auth/rate-limit/expiry behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in a single efficient sentence. The Base URL line is boilerplate but compact, and nothing is padded or redundant at length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by naming the salient return fields (status, temporary asset URL, details). For a simple two-parameter read tool this is nearly complete, lacking only lifecycle context such as when the temporary URL expires.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'id' and 'account'. The description only echoes the 'by ID' lookup concept and adds no format, source, or account-selection detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Uses a specific verb ('Get') plus resource ('render') and enumerates what is returned: rendering status, temporary asset URL, and details. It is clearly distinguishable from the sibling 'render' (which creates) and from the template/asset tools, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by ID' and 'Read operation' imply this is a lookup used to poll or inspect an existing render, but there is no explicit when-to-use versus when-not, nor any named alternative such as list/filtering tools. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sourceGet SourceBRead-onlyIdempotent
Fetch a source file details and status by its id.
Base URL: https://api.shotstack.io/ingest/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered and the description's 'Read operation' simply restates it. The base URL and stage/version context add a little value, but error behavior (e.g. unknown id) and return shape are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is front-loaded and does the real work in a single line. The trailing base URL markup and 'Read operation' segment are slightly noisy and redundant with the annotations, keeping it out of 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read with no output schema, 'details and status' gives a rough sense of the return, which is adequate. It does lack specifics on response fields and failure modes, so it is minimum-viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, covering both id and account, so the baseline is 3. The description only reinforces that lookup is by id and adds nothing about the optional account parameter's effect on credentials/environment beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (source file details and status) keyed by id, which clearly distinguishes it from list_sources, ingest_source, and delete_source. It does not explicitly name siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives such as list_sources (to enumerate) or get_asset (for rendered assets). Usage is only implied by 'by its id'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_templateRetrieve TemplateBRead-onlyIdempotent
Retrieve a template by template id.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is fully covered by structured data. The description's 'Read operation.' merely restates readOnlyHint, and the base URL line is construction detail rather than behavioral context (no error/not-found behavior, no auth or rate-limit notes).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded: the purpose sentence comes first and the remaining lines are brief. The HTML-anchor link and the redundant 'Read operation.' line add noise but do not bury the core statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, fully-annotated single-resource read tool with 100% schema coverage and no output schema, the definition provides enough to call it correctly. The only real omission is what a not-found or private-account lookup does, which is minor at this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'id' and 'account' are already documented with local detail. The description adds nothing about parameter meaning or format beyond the schema, which is the baseline 3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieve a template by template id'), which is unambiguous and clearly distinct from the create/update/delete/list template siblings. It does not explicitly name an alternative, but the single-resource-by-ID framing is clear enough for selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no contrast with list_templates (which returns many) or get_render, and no stated prerequisites. Usage is only implied by the tool name and the fact that an ID is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_sourceFetch SourceADestructive
Queue a source file to be fetched from a URL and stored by Shotstack. Source files can be videos, images, audio files and fonts. Once ingested, new output renditions can be created from the source file.
Base URL: https://api.shotstack.io/ingest/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | The URL of the file to be ingested. The URL must be publicly accessible or include credentials. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| outputs | No | ||
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| callback | No | An optional webhook callback URL used to receive status notifications when sources are uploaded and renditions processed. | |
| destinations | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and openWorldHint=true, and the description adds critical behavioral context: 'Explicit confirmation is required... provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.' This goes well beyond the annotations, warning about financial and operational risks, and advising against retries on unknown outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it states the action, lists supported file types, and then adds the critical warning. The base URL inclusion is extraneous for an agent and the markdown link formatting is odd, but overall the text is efficient and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (8 parameters, nested objects, no output schema) and the description covers the core action and a critical warning, but it does not explain the outputs, destinations, or callback parameters, nor does it clarify the difference between using the flat parameters versus the nested payload. An agent would need to rely heavily on the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 75% description coverage and nests complex objects. The description does not detail parameters, but the schema covers most of them. The description adds the URL requirement implicitly and the confirmation requirement explicitly, which is a key parameter (confirm). However, it does not explain the outputs, destinations, or payload structure, leaving some gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Queue a source file to be fetched from a URL and stored by Shotstack.' It specifies the supported file types (videos, images, audio, fonts) and the outcome (new renditions can be created). It does not explicitly differentiate from siblings like create_upload_url_file or transfer_asset, which also handle assets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some context (what can be ingested, where it goes) but offers no explicit guidance on when to use this tool versus alternatives like create_upload_url_file or transfer_asset. The warning about confirmation and unknown outcomes is useful but not a usage guideline per se.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_accountsList configured accountsARead-onlyIdempotent
List private account labels, default selection and configured token method. No credentials, token paths or account content; no network request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false. The description adds meaningful behavioral context beyond those annotations: it confirms no credentials, token paths, or account content are returned, and that no network request is made.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences front-load the core scope and then immediately clarify the negative behavior. Every phrase contributes useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, the description adequately covers the returned concepts and explicitly rules out sensitive data and network activity. It stops short of describing result ordering, formatting, or pagination, but those may not apply or may be self-evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to clarify. Per the rubric, a zero-parameter tool has a baseline of 4, and the description does not need to compensate for undocumented inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource (list accounts) and enumerates the exact scope: private account labels, default selection, and configured token method. Its exclusions also implicitly distinguish it from broader siblings like get_account or get_current_token, which would expose account content or credentials.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating exactly what metadata the tool surfaces, but it never explicitly says when to call this instead of alternatives such as get_account or get_current_token. The 'No credentials...' clause scopes the tool, yet no named alternative or when-not condition is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList Generation ModelsARead-onlyIdempotent
List the generation models available for prompt-bearing image, video and audio
assets, with the options each accepts and what it costs in credits.
Use this to populate a model picker and render its option fields, rather than hard coding a model list. A newly launched model appears here without any change on your side. Each entry carries the asset type it generates, so filter the list client side when a picker only needs one kind.
Option schemas are omitted by default. Request them with expand=options.
Each model's available reflects the plan of the account behind the calling API
key, so offer only the models it marks available.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds valuable behavior beyond that: option schemas are omitted by default and require `expand=options`, and the `available` field reflects the calling API key's plan, which affects what should be shown. Some context like pagination or rate limits is not mentioned, but the key operational detail is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, then usage, then behavior details, and ends with base URL and read-operation note. It is a bit long but every sentence contributes useful context; no obvious filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description explains the shape of returned data (asset type, options, availability) and the default omission of option schemas. It could mention pagination or response ordering, but overall it gives enough for an agent to call and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the `account` parameter is fully documented in the schema. The description adds no parameter syntax or default-value information beyond the schema, so baseline 3 is appropriate. It does mention the `expand=options` behavior, but that is not a formal parameter in this schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (List) and resource (generation models available for prompt-bearing image, video, and audio assets), with the added scope that each entry carries its asset type and options. It clearly distinguishes itself from siblings like get_model (singular) and list_templates by naming what it returns and why it exists.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete when-to-use case (populate a model picker, avoid hard coding) and explains that filtering by asset type should be done client side. It doesn't name a sibling alternative or explicitly state when-not-to-use, but the usage context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sourcesList SourcesCRead-onlyIdempotent
Retrieve a list of ingested source files stored against a users account and stage.
Base URL: https://api.shotstack.io/ingest/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description only repeats 'Read operation' and adds no new behavioral context such as pagination, output shape, rate limits, or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in one sentence, but the second sentence contains a base URL link and the redundant 'Read operation' phrase, neither of which earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool, the description is enough to invoke it, but with no output schema it does not describe the returned list, pagination, or how the optional account parameter affects results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one optional parameter ('account') exists and schema description coverage is 100%, so the schema fully documents it. The description adds no syntax or meaning beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieve a list of ingested source files') and scopes it to account and stage, so an agent can tell it is a list operation. However, it does not name or explicitly distinguish itself from siblings such as get_source or list_accounts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no alternatives, and no exclusions. The only usage direction is the redundant 'Read operation' statement, which annotations already cover.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_templatesList TemplatesBRead-onlyIdempotent
Retrieve a list of templates stored against a users account and stage.
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is fully covered. The description's 'Read operation' merely restates readOnlyHint; its only added context is the base URL and account/stage scoping. No pagination or rate-limit behavior is disclosed, so this is minimal added value against an already-complete annotation set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the purpose before the URL and read flag. The embedded HTML anchor and version placeholder are slightly noisy marketing text, but there is no padding or redundancy beyond that.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter list tool with no output schema, the essentials are present (what is listed, scoping, read-only nature). However, it never indicates pagination, ordering, or result shape, which are the questions an agent actually has for a list endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter and 100% schema description coverage, the schema already explains the 'account' selector. The description adds nothing about how the account parameter changes credentials or environment beyond what the schema states, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Retrieve a list of templates') and adds scoping (a user's account and stage), so an agent can distinguish it from get_template or create_template. It stops short of naming the sibling alternatives, but the list-vs-get distinction is clear from the verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not guidance, and no mention of the natural alternatives (get_template for a single template, create/update/delete_template for mutations). The account/stage scope hints at usage but the agent must infer it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probe_mediaInspect MediaARead-onlyIdempotent
Inspects any media asset (image, video, audio) on the internet using a hosted version of FFprobe. The probe endpoint returns useful information about an asset such as width, height, duration, rotation, framerate, etc...
Base URL: https://api.shotstack.io/edit/{version} Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public HTTPS URL of the selected media file. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered elsewhere. The description usefully adds that the probe goes out to the internet and lists example metadata fields (width, height, duration, rotation, framerate), but says nothing about auth requirements, rate limits, failure modes for unreachable/private URLs, or the meaning of the {version} placeholder.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and mechanism are front-loaded in the first sentence and the size is appropriate. The trailing base-URL block with a raw HTML anchor and a literal '{version}' placeholder adds noise without helping the agent invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with no output schema, the description covers what it does and gives a representative list of returned fields, which is reasonably complete. The gaps are minor: no scope note that only public HTTPS URLs are supported beyond what the schema pattern implies, and no guidance on failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (url, account) are already documented in the schema, and the description adds no format or semantic detail beyond what is there. Baseline 3 is appropriate when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Inspects) and resource (any media asset: image, video, audio), states the mechanism (hosted FFprobe), and no sibling tool does anything comparable — everything else is render/template/asset CRUD. An agent can immediately distinguish this from get_asset or get_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no alternatives are named. 'Read operation' hints at the usage profile but the agent is left to infer that this is the metadata-probing tool rather than a fetch/download tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
renderRender AssetADestructive
Queue and render the contents of an Edit as a video, image or audio file.
Rendering Process:
Validation: The edit JSON is validated
Download: All assets are downloaded and cached
Preprocessing: Video assets are automatically processed to fix compatibility issues
Rendering: The timeline is rendered using the processed assets
Output: The final media file is generated and stored
Video Preprocessing:
Video assets undergo automatic preprocessing to ensure compatibility. You can force
preprocessing by setting "transcode": true on video assets. See VideoAsset
for more details.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| disk | No | **Notice: This option is now deprecated and will be removed. Disk types are handled automatically. Setting a disk type has no effect.** The disk type to use for storing footage and assets for each render. <ul> <li>`local` - optimized for high speed rendering with up to 512MB storage</li> <li>`mount` - optimized for larger file sizes and longer videos with 5GB for source footage and 512MB for output render</li> </ul> | |
| merge | No | An array of key/value pairs that provides an easy way to create templates with placeholders. The placeholders can be used to find and replace keys with values. For example you can search for the placeholder `{{NAME}}` and replace it with the value `Jane`. | |
| output | No | ||
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| callback | No | An optional webhook callback URL used to receive status notifications when a render completes or fails. Notifications are also sent when a rendered video is sent to an output [destination](https://shotstack.io/docs/guide/serving-assets/destinations/). See [webhooks](https://shotstack.io/docs/guide/architecting-an-application/webhooks/) for more details. | |
| instance | No | The render instance type to use for processing the edit. <ul> <li>`s1` - standard instance (default)</li> <li>`s2` - standard instance with more resources</li> <li>`a1` - accelerated instance for faster rendering</li> </ul> | s1 |
| timeline | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructive, non-idempotent, open-world), the description adds real value: it discloses the asynchronous 'queue' semantics, the multi-stage pipeline, required explicit confirmation, and that provider charges, hosting, sharing or deletion may apply, plus a caution never to auto-resubmit unknown outcomes. These are meaningful behavioral traits not captured by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose leads, followed by a well-structured numbered pipeline and a focused video-preprocessing aside. The Base URL line with an empty href is noise, but overall the content is front-loaded and each section roughly earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description could do more to explain the async return (render ID / polling with get_render), but for a complex nested tool it adequately covers the pipeline, confirmation requirement and side effects. Given 80% schema coverage and rich nested docs, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents most parameters and the baseline is 3. The description only adds meaning for one parameter, the `transcode: true` forcing of video preprocessing, which is a marginal but genuine addition over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: 'Queue and render the contents of an [Edit] as a video, image or audio file.' This clearly states the input (an Edit) and the output artifact type. It does not explicitly distinguish itself from the sibling render_template, but an agent can infer the difference from the 'Edit' vs 'template' wording.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through the numbered rendering-process pipeline (validation, download, preprocessing, rendering, output) and the confirmation note, but it never says when to use this versus render_template or generate_asset, nor any prerequisite conditions. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_templateRender TemplateBDestructive
Render an asset from a template id and optional merge fields. Merge fields can be used to replace placeholder variables within the Edit.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | The id of the template to render in UUID format. | |
| merge | No | An array of key/value pairs that provides an easy way to create templates with placeholders. The placeholders can be used to find and replace keys with values. For example you can search for the placeholder `{{NAME}}` and replace it with the value `Jane`. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, idempotentHint=false and openWorldHint=true; the description goes beyond that by warning about provider charges, hosting, sharing or deletion, requiring explicit confirmation, and warning against automatic resubmission of unknown outcomes. That is genuinely useful operational context for an irreversible render.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is front-loaded and clear, but the body is cluttered with broken markdown — a link to '[Edit](#tocs_edit)' that does not resolve, a Base URL anchor with href='#', and duplicated boilerplate about confirmation. Wasteful formatting dilutes the useful content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent, chargeable action with six parameters and no output schema, the description does cover the confirmation and cost risk. It says nothing about what is returned (render id, async polling, job status) and gives no way to disambiguate from `render`, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents `id`, `merge`, `account`, `confirm`, `payload` and `payload_file`. The description only restates the merge-field idea and adds nothing about the confirmation flag or the payload/payload_file exclusivity, so it is at the baseline for full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Render an asset from a template id') plus the optional merge-field mechanism, which is enough for an agent to know what it produces. However, it never distinguishes itself from the sibling tool `render`, so the agent cannot tell from the description alone which of the two to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use vs `render` or any other sibling is given. The only guidance is procedural ('Explicit confirmation is required... never automatically resubmit unknown outcomes'), which is a safety rule rather than a selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transfer_assetTransfer AssetADestructive
Transfer a file from any publicly available URL to one or more Serve API destinations.
Base URL: https://api.shotstack.io/serve/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | An identifier for the asset which must be provided by the client. The identifier does not need to be unique. | |
| url | No | The file URL to fetch and transfer. | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| destinations | No | Specify the storage locations and hosting services to send the file to. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered. The description nonetheless adds real value beyond them: mandatory confirmation, the possibility of provider charges, and hosting/sharing/deletion side effects, plus a no-auto-resubmit rule that reinforces non-idempotence. It stops short of describing success/failure semantics or credential setup, keeping it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and the safety warnings are front-loaded and earn their place. The interpolated 'Base URL' anchor line is low-value boilerplate that slightly dilutes an otherwise tight, well-ordered description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent, open-world mutation with steep nested-schema complexity and no output schema, the description correctly foregrounds confirmation and side effects. It omits return/response expectations and credential prerequisites (which the schema mentions for destinations), leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all seven parameters, including the nested destination variants, are richly documented in the schema itself. The description mentions URL and destinations only generically, adding no format or constraint detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Transfer a file from any publicly available URL to one or more Serve API destinations'), distinguishing it from render/generate_asset/ingest_source which produce or ingest rather than move an existing file. An agent can tell immediately what the tool does and its source-to-destination scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear usage precondition ('Explicit confirmation is required for this exact action') and a retry constraint ('Never automatically resubmit unknown outcomes'), which shape when it is safe to call. It does not name a sibling alternative or a when-not-to-use condition, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_templateUpdate TemplateBDestructive
Update an existing template by template id.
Base URL: https://api.shotstack.io/edit/{version} Explicit confirmation is required for this exact action; provider charges, hosting, sharing or deletion may apply. Never automatically resubmit unknown outcomes.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact resource ID from the selected account. | |
| name | No | The template name | |
| account | No | Named private Shotstack account; selects private credentials and stage/v1 environment. | |
| confirm | No | Must be true for this exact requested render, generation, mutation, upload URL or deletion. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| template | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely new context beyond that: confirmation is required for this exact action, provider charges/hosting/sharing/deletion may apply, and unknown outcomes must not be retried automatically. It does not disclose whether the update is full replacement or partial, which is the remaining behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and identity before the risk warnings, with no padding. The appended Base URL line is irrelevant to tool selection and is minor noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with a deeply nested Edit payload, three alternate body forms and no output schema, the description covers the safety contract but omits the mechanics an agent needs: whether omitted fields are cleared, and which of payload/template/payload_file to choose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the schema already documents id, name, account, confirm, payload, template and payload_file. The description only reinforces the id parameter and alludes to the confirm flag; it does not clarify the mutually exclusive body options or how 'name'/'template' interact with 'payload'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing template') and pins the identity mechanism ('by template id'), so an agent can distinguish it from create_template/delete_template/get_template without opening a schema. It stops short of naming siblings explicitly, which is why it is not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The text mandates explicit confirmation and warns against auto-resubmitting unknown outcomes, but that is operational policy, not routing guidance. There is no statement of when to use this versus create_template, render_template, or the partial-update alternatives (payload vs template vs payload_file).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v2.0.0- First observed
create_template - First observed
create_upload_url_file - First observed
delete_asset - First observed
delete_source - First observed
delete_template - First observed
generate_asset - First observed
get_asset - First observed
get_asset_by_render_id - First observed
get_generated_asset - First observed
get_model - First observed
get_render - First observed
get_source - First observed
get_template - First observed
ingest_source - First observed
list_accounts - First observed
list_models - First observed
list_sources - First observed
list_templates - First observed
probe_media - First observed
render - First observed
render_template - First observed
transfer_asset - First observed
update_template
TDQS
Scored across 23 tools
Most tools have clearly distinct purposes, targeting specific resources like renders, templates, assets, sources, or models. A few tools (ingest_source, transfer_asset, create_upload_url_file) all involve ingesting files but differ in destination and intent, which could cause minor confusion.
Consistent snake_case with verb-first naming throughout (e.g., get_render, create_template, delete_source). The only minor deviation is the bare 'render' without a noun, but it remains consistent with the verb-first convention.
23 tools is on the heavy side but each maps to a distinct API operation across editing, serving, ingesting, and generation sub-APIs. The breadth justifies the count, though it slightly exceeds the ideal 3-15 range.
The surface covers CRUD for templates and sources, render creation and status, asset retrieval and deletion, AI generation, and model listing. Missing operations like listing all assets or canceling a render are minor and likely outside the core scope.
Maintenance
Related MCP Connectors
Give AI agents identity, scoped access, trusted context, and verifiable actions through MCP.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Remote MCP server for AI.TV creators — delegate account operations to your AI agent over MCP.
AI video editor for agents and humans: timeline, captions, color, audio and generation as MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to securely manage Linux, Docker, Kubernetes, and SSH infrastructure through MCP with permission control, audit logging, and dangerous command filtering.MIT
- AlicenseBqualityCmaintenanceEnables MCP-compatible agents to perform read-only incident investigations by querying logs, building timelines, searching runbooks, validating evidence, and creating ticket drafts only with explicit approval.7MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to access and manage versioned robotics development state, including artifacts, changes, evidence, reviews, and execution context, through read-only, developer, or reviewer MCP profiles.MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to create and manage video projects, write story and brief, convert, estimate costs, propose paid actions, run approved proposals, discard and restore assets.1MIT