Skip to main content
Glama

ModelsLab

Motion Control

motion-control

Control motion in video generation. Uses an image and video to create motion-controlled output. Returns a request ID that can be used with fetch-video to retrieve results.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoProcessing mode: std (standard) or pro (professional).
promptNoOptional text prompt (max 2500 characters).
webhookNoURL to receive webhook notification when processing completes.
model_idYesThe model ID to use for motion control.
track_idNoCustom tracking ID for the request.
init_imageYesURL or base64 string of the input image (character/subject).
init_videoYesURL or base64 string of the video for motion reference.
keep_original_soundNoKeep original sound from the video.
character_orientationYesUse character orientation from image or video.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint declared, the description carries most of the behavioral burden. It usefully discloses the async nature (returns a request ID) and the fetch-video handoff, but says nothing about cost, generation latency, permissions, or what happens if the optional webhook is omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded and the return/consumption detail last. There is a small amount of wasted leading whitespace and no further padding, so it is efficient without being exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, 4-required generation tool with no output schema, the description covers the essential async return contract via the request ID and fetch-video pointer. It remains thin on the required-input expectations and mode/model_id selection, leaving gaps that annotations (openWorldHint only) cannot fill.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 — every parameter, including enums for mode, keep_original_sound, and character_orientation, is already documented in the schema. The description only restates the image/video roles, adding no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Control motion in video generation') and clarifies the input combination (image + video) that produces motion-controlled output. It does not explicitly distinguish itself from close siblings like video-to-video or image-to-video, so it lands at clear-but-undifferentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one useful workflow cue — the returned request ID is consumed by fetch-video — which implies the async two-step pattern. However, it gives no guidance on when to choose this over video-to-video, image-to-video, or lip-sync, nor any prerequisites for the required image/video inputs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources