wan-mcp
The wan-mcp server gives AI agents direct access to Wan AI models for image/video generation, editing, and animation via RunAPI. Key capabilities:
Authenticate (
login): Browser-based PKCE OAuth flow to save credentials locally; headless/CI environments can use theRUNAPI_API_KEYenv var instead.Animate images (
animate): Create animation tasks from source images using models likewan-2.2-animate-moveorwan-2.2-animate-replace.Edit videos (
edit_video): Apply prompt-driven edits to existing videos using models likewan-2.6-edit-videoorwan-2.7-edit-video, supporting resolutions up to 1080p.Image to video (
image_to_video): Generate videos from a first-frame image across 5 model variants, with control over duration, resolution, and optional audio.Speech/audio to video (
speech_to_video): Generate video from a source image and audio file.Text to image (
text_to_image): Create images from text prompts with aspect ratio options and resolutions up to 4K.Text to video (
text_to_video): Produce videos from text prompts with duration (2–10 seconds), aspect ratio, and resolution controls across 5 model variants.Poll task status (
get_task): Fetch status and output URLs for any previously created task by ID. Each creation tool can also wait for completion or return a task ID for later polling.Check pricing (
check_pricing): Look up current pricing for any Wan model and endpoint — no authentication required.
Supports 18 model variants across 6 endpoints (Wan 2.2–2.7) and integrates with MCP clients like Claude, Cursor, Windsurf, VS Code, and more.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@wan-mcpgenerate a video of a cat dancing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Why This Package?
@runapi.ai/wan-mcp is a focused Model Context Protocol server for the Wan model line on RunAPI.
It gives MCP-compatible assistants direct access to 6 endpoints and 18 model variants without loading the full RunAPI catalog.
Use this per-model server when an agent should stay scoped to Wan. Use @runapi.ai/mcp when one assistant should discover every RunAPI model line.
Related MCP server: @runapi.ai/gpt-4o-image-mcp
Install
Add it to Claude Code:
claude mcp add wan -s user -- npx -y @runapi.ai/wan-mcpUse project scope when the server should be shared with a repository:
claude mcp add wan -s project -- npx -y @runapi.ai/wan-mcpCodex, Cursor, Windsurf, VS Code, Roo Code, and other MCP hosts can use the same stdio command:
{
"mcpServers": {
"wan": {
"command": "npx",
"args": ["-y", "@runapi.ai/wan-mcp"]
}
}
}check_pricing works before sign-in. For task creation and status polling, ask your assistant to call the login tool. It opens a browser login and saves credentials to ~/.config/runapi/config.json, the same file used by runapi login.
Headless and CI hosts can still set RUNAPI_API_KEY before starting the MCP host.
Ready-made examples are in examples/ for Claude, Cursor, Windsurf, VS Code, and Roo Code.
Tools
Tool | Auth | Purpose |
| Yes | Create a Wan animate task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Create a Wan edit video task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Create a Wan image to video task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Create a Wan speech to video task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Create a Wan text to image task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Create a Wan text to video task and optionally wait for a terminal status. Returns the task id, status, and output URLs. |
| Yes | Fetch the current status and latest payload for an existing task. |
| No | Look up current pricing for a Wan model and endpoint. |
Models
Wan covers 18 model variants across 6 endpoints. Each tool accepts the models listed for it:
Tool | Models |
|
|
|
|
|
|
|
|
|
|
|
|
Model availability can change between releases. Use check_pricing or the Wan model page for the current catalog view.
Agent Prompts
Ask your assistant in natural language; it can inspect pricing, create the task, and return the task id plus output URLs.
Create a task
Run a Wan animate task with RunAPI.The assistant can call check_pricing, then animate, and return the task id, status, and output URLs.
Submit without waiting
Create the task but don't wait for it to finish.The assistant calls the create tool with wait: false and returns the task id. Check on it later with get_task.
Check pricing before creating
Check current Wan pricing, then create the task if it matches my request.The assistant calls check_pricing and can link to the Wan model page for the canonical catalog entry.
Configuration
The server resolves auth in this order:
RUNAPI_API_KEYenvironment variable, useful for headless and CI hosts~/.config/runapi/config.json, created by the MCPlogintool orrunapi loginNo key, which still allows
check_pricing
The config file is normally managed by login. A pre-provisioned headless config can use:
{
"apiKey": "your_runapi_key"
}Do not commit real API keys.
Links
Resource | URL |
Wan model page | |
npm package | |
GitHub repository | |
RunAPI MCP overview | |
RunAPI docs |
License
Licensed under the Apache License, Version 2.0.
Available Tools
9 toolsanimateC
Create a Wan task on RunAPI (animate). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Poll until the task reaches a terminal status. | |
| model | No | RunAPI model slug for this model line. | |
| timeout_ms | No | ||
| callback_url | No | Declared type: string. | |
| poll_interval_ms | No | ||
| source_image_url | Yes | Declared type: string. | |
| output_resolution | No | Declared type: string. Known values: "480p", "580p", "720p". | |
| reference_video_url | Yes | Declared type: string. | |
| enable_safety_checker | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the async task pattern and return shape (task id, status, output URLs). However it never mentions the long-running/polling nature implied by the 'wait' and 'poll_interval_ms' parameters, nor cost, auth, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and followed by the return shape; nothing is padded. It loses a point only because the terse phrasing leans on unexplained provider jargon rather than spending a few words on what the tool produces.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter generative media tool with no annotations and no output schema, the description is far too thin. It explains the return fields but not the async lifecycle, the roles of the image/video inputs, or how polling/timeout interact, leaving an agent under-equipped to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 78%, just under the high-coverage threshold, and several parameters (source_image_url, reference_video_url, callback_url, enable_safety_checker) carry only 'Declared type' placeholders. The description adds no parameter meaning beyond the schema, so a baseline-adjacent 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a verb and resource ('Create a Wan task on RunAPI (animate)'), so the general action is clear. But it does not differentiate from the close siblings image_to_video and text_to_video, and 'animate' plus 'Wan task' is provider jargon rather than a description of what the animation actually does with a source image and reference video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives, despite four visually similar siblings (image_to_video, text_to_video, edit_video, speech_to_video). An agent cannot infer from the text alone why it should pick 'animate' over image_to_video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_pricingB
Look up RunAPI pricing for the wan model line.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model slug. Defaults to the line's primary model. | |
| action | No | Endpoint name. Defaults to the endpoint that offers the model. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. 'Look up' implies a read-only operation, but there is no mention of auth requirements, whether calls are billable/rate-limited, or the format of the pricing data returned — all relevant for a pricing endpoint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the resource and scope come first. Nothing repeats the name or wastes tokens.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description could usefully sketch the shape of returned pricing (units, currency, per-endpoint granularity) but does not. Scope is adequately stated, but return-value expectations are left unset for a tool whose whole purpose is retrieving data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'model' and 'action' (including the action enum) are already documented in the schema. The description adds no syntax, defaulting, or relationship detail beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('look up') plus a bounded resource ('RunAPI pricing') and scope ('the wan model line'), so an agent knows exactly what the tool returns. It does not, however, distinguish itself from the sibling action tools or explain how the pricing scope relates to them, leaving that inference to the caller.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, no mention of the defaulting behavior for model/action, and no guidance on what to do with the result. Only the verb 'look up' implies usage, which is weak guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_videoC
Create a Wan task on RunAPI (edit video). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Declared type: integer. | |
| wait | No | Poll until the task reaches a terminal status. | |
| audio | No | Declared type: boolean. | |
| model | No | RunAPI model slug for this model line. | |
| prompt | No | Declared type: string. | |
| watermark | No | Declared type: boolean. | |
| timeout_ms | No | ||
| multi_shots | No | Controls whether the generated video uses multiple shots with transitions instead of one continuous shot. Declared type: boolean. | |
| aspect_ratio | No | Declared type: string. | |
| callback_url | No | Declared type: string. | |
| audio_setting | No | Declared type: string. | |
| negative_prompt | No | Declared type: string. | |
| duration_seconds | No | Declared type: integer. | |
| poll_interval_ms | No | ||
| source_video_url | No | Declared type: string. | |
| output_resolution | No | Declared type: string. | |
| source_video_urls | No | Declared type: array. | |
| reference_image_url | No | Declared type: string. | |
| enable_safety_checker | No | Declared type: boolean. | |
| enable_prompt_expansion | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It does disclose the async task pattern and the returned fields (task id, status, output URLs), which is useful, but it says nothing about required inputs (source_video_url?), whether wait=true blocks, cost/pricing implications, or failure behavior. For a 20-parameter generation tool with zero annotation coverage, that is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the operation stated first and the return contract second — no filler or repetition. It is appropriately sized, though the parenthetical "(edit video)" is doing less work than a real scope statement would.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 20-parameter, zero-required, no-annotation, no-output-schema tool needs more than two sentences. The description partially compensates for the missing output schema by naming the returned fields, but it leaves the agent unable to tell which of source_video_url / source_video_urls / reference_image_url are actually needed for an edit, nor how the wait/poll parameters interact.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 90%, so the schema already documents nearly all 20 parameters, and the description adds no parameter-level meaning beyond that. Baseline 3 is appropriate when the schema does the heavy lifting; the description contributes nothing extra here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It gives a verb and resource ("Create a Wan task on RunAPI (edit video)") and names the return shape, but "edit video" is never qualified — it does not say whether it restyles an existing source video, concatenates clips, or performs some other transform. Siblings like text_to_video and image_to_video imply this line is for editing a source video, but the description never makes that distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance at all. Nothing tells the agent how this differs from text_to_video, image_to_video, or animate, nor that get_task is the alternative for polling an existing task, even though this tool also exposes wait/poll_interval_ms.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_taskA
Fetch the current status and latest result payload for a wan task.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Asynchronous endpoint the task was created on. | |
| task_id | Yes | Task id returned when the task was created. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only fetch and names the returned content at a high level, but omits auth requirements, polling behavior, status value semantics, rate limits, and error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It states the action and returned payload immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the input schema is fully documented, but there is no output schema and no annotations. The description gives only a high-level view of the return payload and does not explain status values, polling expectations, or authentication needs, leaving gaps for a task-status tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema fully documents both parameters, including the action enum. The description adds no additional parameter meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') and resource ('current status and latest result payload for a wan task'), making it clearly distinct from the sibling creation and account tools. An agent can identify this as the task-status retrieval tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'current status' and 'latest result payload' for a task, suggesting it is used after an asynchronous task has been created. However, the description gives no explicit when-to-use guidance, prerequisites, or alternatives, leaving the agent to infer the polling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_to_videoB
Create a Wan task on RunAPI (image to video). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Declared type: integer. | |
| wait | No | Poll until the task reaches a terminal status. | |
| audio | No | Declared type: boolean. | |
| model | No | RunAPI model slug for this model line. | |
| ratio | No | Declared type: string. | |
| prompt | No | Declared type: string. | |
| watermark | No | Declared type: boolean. | |
| timeout_ms | No | ||
| multi_shots | No | Declared type: boolean. | |
| acceleration | No | Declared type: string. | |
| aspect_ratio | No | Declared type: string. | |
| callback_url | No | Declared type: string. | |
| negative_prompt | No | Declared type: string. | |
| duration_seconds | No | Declared type: integer. | |
| poll_interval_ms | No | ||
| source_video_url | No | Declared type: string. | |
| driving_audio_url | No | Declared type: string. | |
| output_resolution | No | Declared type: string. Known values: "480p", "720p", "1080p". | |
| background_audio_url | No | Declared type: string. | |
| last_frame_image_url | No | Declared type: string. | |
| enable_safety_checker | No | Declared type: boolean. | |
| first_frame_image_url | No | Declared type: string. | |
| enable_prompt_expansion | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does disclose that this is an asynchronous task-creation call returning a task id, status, and output URLs. It says nothing about authentication, cost, rate limits, failure behavior, or what the 'wait' polling default implies; for a 23-parameter generation tool with zero annotation coverage this is only partially adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the verb and resource front-loaded and the return shape compactly stated; nothing is wasted. It is arguably under-specified for the tool's complexity, but that is a completeness concern rather than a verbosity one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a task with 23 parameters, no required fields, no annotations, and no output schema, the description leaves major gaps: which image fields to supply, how model/ratio/duration interact, and what the polling behavior actually does. The return-value sentence is the only substantive addition, and it is not enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 91%, so the baseline is 3 even though the description adds no parameter-level meaning. In practice most schema entries are placeholders ('Declared type: integer'), so the description does nothing to explain key knobs like first_frame_image_url, last_frame_image_url, ratio, or wait.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (a Wan image-to-video task on RunAPI), and the parenthetical 'image to video' distinguishes it from text_to_video, edit_video, and speech_to_video siblings. However, it never clarifies which input fields constitute the 'image' input, which matters given the tool has zero required parameters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no conditions, and no mention of alternatives such as text_to_video or animate, despite several closely related siblings. The agent must infer selection purely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
loginA
Authenticate RunAPI by opening a browser PKCE login flow and saving the API key to ~/.config/runapi/config.json.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Re-run browser login when the current credential comes from the local config file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the interactive browser flow and the file write side effect (config.json). However, it does not mention that it may overwrite existing credentials or that it could block waiting for user input, though these are implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the action ('Authenticate RunAPI') and provides necessary details without extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple login tool with one optional parameter and no output schema, the description covers the core purpose and side effect. It lacks an explicit statement that this is a prerequisite for other tools, but that is implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (the only parameter 'force' has a description). The tool description adds no additional meaning about parameters beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Authenticate'), target resource ('RunAPI'), method ('browser PKCE login flow'), and side effect (saving to config.json). It is distinct from sibling tools, none of which relate to authentication.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (to authenticate RunAPI) but does not explicitly say when to run it (e.g., before other RunAPI tools) or when to use the 'force' parameter. Since there are no alternative auth tools among siblings, 'vs alternatives' is not applicable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speech_to_videoC
Create a Wan task on RunAPI (speech to video). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Declared type: integer. | |
| wait | No | Poll until the task reaches a terminal status. | |
| model | No | RunAPI model slug for this model line. | |
| shift | No | Declared type: number. | |
| prompt | Yes | Declared type: string. | |
| num_frames | No | Declared type: integer. | |
| timeout_ms | No | ||
| callback_url | No | Declared type: string. | |
| guidance_scale | No | Declared type: number. | |
| negative_prompt | No | Declared type: string. | |
| poll_interval_ms | No | ||
| source_audio_url | Yes | Declared type: string. | |
| source_image_url | Yes | Declared type: string. | |
| frames_per_second | No | Declared type: integer. | |
| output_resolution | No | Declared type: string. Known values: "480p", "580p", "720p". | |
| num_inference_steps | No | Declared type: integer. | |
| enable_safety_checker | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the return shape (task id, status, output URLs), which is useful, but says nothing about async/long-running generation, polling behavior implied by the 'wait' default, timeouts, cost, or auth requirements for the RunAPI call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and immediately followed by the return shape. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter media-generation tool with no annotations and no output schema, the description is too thin. It should cover the async task lifecycle, polling/timeout semantics, and the fact that generated video output is the payoff, none of which appear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema itself already documents the parameters and the baseline is 3. The description adds no parameter meaning at all – notably it doesn't mention that both a source image and a source audio URL are required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource – 'Create a Wan task ... (speech to video)' – and the parenthetical modality distinguishes it from image_to_video and text_to_video. However it never says what inputs drive the generation (image + audio), leaving the agent to infer the distinction from the schema rather than the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no reference to any alternative sibling tool. An agent cannot tell from the description why it would pick speech_to_video over animate or image_to_video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text_to_imageB
Create a Wan task on RunAPI (text to image). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Declared type: integer. | |
| wait | No | Poll until the task reaches a terminal status. | |
| model | No | RunAPI model slug for this model line. | |
| prompt | No | Declared type: string. | |
| bbox_list | No | Declared type: array. | |
| watermark | No | Declared type: boolean. | |
| timeout_ms | No | ||
| aspect_ratio | No | Declared type: string. Known values: "1:1", "16:9", "4:3", "21:9", "3:4", "9:16", "8:1", "1:8". | |
| callback_url | No | Declared type: string. | |
| output_count | No | Declared type: integer. | |
| color_palette | No | Declared type: array. | |
| thinking_mode | No | Declared type: boolean. | |
| poll_interval_ms | No | ||
| enable_sequential | No | Declared type: boolean. | |
| output_resolution | No | Declared type: string. Known values: "1k", "2k", "4k". | |
| source_image_urls | No | Declared type: array. | |
| enable_safety_checker | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden. It does disclose the return shape (task id, status, output URLs), which is valuable with no output schema, and implies an async task model. However, it never explains the polling behavior implied by 'wait', 'timeout_ms', and 'poll_interval_ms', nor cost, auth, or rate-limit considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler and the action front-loaded. The jargon ('Wan task on RunAPI') is slightly opaque but does not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter generative tool with zero annotations and no output schema, this is thin. Return values are mentioned, which helps, but there is nothing about async/wait semantics, valid model slugs, cost (despite a check_pricing sibling), or how to poll with get_task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema already documents the parameters, and the description adds no parameter-level meaning. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource: creating a text-to-image Wang task via RunAPI, and the parenthetical '(text to image)' distinguishes it from the sibling text_to_video. It stops short of naming alternatives explicitly, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no mention of the sibling tools an agent must choose between (text_to_video, edit_video, animate). Nothing tells the agent which model line or scenario this tool is appropriate for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text_to_videoB
Create a Wan task on RunAPI (text to video). Returns a task id, status, and output URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Declared type: integer. | |
| wait | No | Poll until the task reaches a terminal status. | |
| model | No | RunAPI model slug for this model line. | |
| ratio | No | Declared type: string. | |
| prompt | No | Declared type: string. | |
| watermark | No | Declared type: boolean. | |
| timeout_ms | No | ||
| multi_shots | No | Declared type: boolean. | |
| acceleration | No | Declared type: string. | |
| aspect_ratio | No | Declared type: string. | |
| callback_url | No | Declared type: string. | |
| negative_prompt | No | Declared type: string. | |
| duration_seconds | No | Declared type: integer. | |
| poll_interval_ms | No | ||
| output_resolution | No | Declared type: string. Known values: "480p", "580p", "720p", "1080p". | |
| reference_audio_url | No | Declared type: string. | |
| background_audio_url | No | Declared type: string. | |
| reference_image_urls | No | Declared type: array. | |
| reference_video_urls | No | Declared type: array. | |
| enable_safety_checker | No | Declared type: boolean. | |
| first_frame_image_url | No | Declared type: string. | |
| enable_prompt_expansion | No | Declared type: boolean. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the async task model ('Create a Wan task') and the return shape (task id, status, output URLs), which is genuinely useful. It omits cost implications, expected generation latency, auth requirements, and whether the task is billable or cancelable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero padding, front-loaded with the action and platform, then the return values. Nothing is wasted and nothing needs reordering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter, zero-required tool with no annotations and no output schema, the description is thin. Naming the returned fields partially compensates for the missing output schema, but it says nothing about defaults, polling behavior tied to the 'wait' flag, or which parameters matter most, leaving an agent to guess at invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 91%, so the schema nominally owns parameter documentation, which sets the baseline at 3. In practice those schema descriptions are largely filler ('Declared type: string'), and the tool description adds no parameter meaning at all — notably absent is any hint that prompt is the key input while the other 21 fields are optional knobs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a Wan task on RunAPI (text to video)'. The parenthetical modality cleanly separates it from image_to_video and speech_to_video siblings, though those siblings are never named directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the '(text to video)' qualifier tells the agent to pick this when the input is a prompt rather than an image, audio clip, or existing video. There is no explicit when-to-use/when-not statement, no mention of the login prerequisite, and no comparison against edit_video or animate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.2.0- Changed
animate10 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.2-animate-move", - "wan-2.2-animate-replace" -] - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string. Known values: \"480p\", \"580p\", \"720p\"." - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "480p", - "580p", - "720p" -] - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / reference_video_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / source_image_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991
- Changed
check_pricing2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.2-animate-move", - "wan-2.2-animate-replace", - "wan-2.6-edit-video", - "wan-2.6-flash-edit-video", - "wan-2.7-edit-video", - "wan-2.2-a14b-image-to-video-turbo", - "wan-2.5-image-to-video", - "wan-2.6-flash-image-to-video", - "wan-2.6-image-to-video", - "wan-2.7-image-to-video", - "wan-2.2-a14b-speech-to-video-turbo", - "wan-2.7-image", - "wan-2.7-image-pro", - "wan-2.2-a14b-text-to-video-turbo", - "wan-2.5-text-to-video", - "wan-2.6-text-to-video", - "wan-2.7-r2v", - "wan-2.7-text-to-video" -]
- Changed
edit_video23 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / aspect_ratio / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / audio / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / audio_setting / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / duration_seconds / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / duration_seconds / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / enable_prompt_expansion / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.6-edit-video", - "wan-2.6-flash-edit-video", - "wan-2.7-edit-video" -] - added
Input schema / properties / multi_shots / descriptionAdded value: +"Controls whether the generated video uses multiple shots with transitions instead of one continuous shot. Declared type: boolean." - added
Input schema / properties / negative_prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / reference_image_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / seed / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / seed / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / source_video_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / source_video_urls / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / watermark / descriptionAdded value: +"Declared type: boolean." - added
Input schema / requiredAdded value: +[]
- Changed
get_task1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
image_to_video27 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / acceleration / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / aspect_ratio / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / audio / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / background_audio_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / driving_audio_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / duration_seconds / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / duration_seconds / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / enable_prompt_expansion / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / first_frame_image_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / last_frame_image_url / descriptionAdded value: +"Declared type: string." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.2-a14b-image-to-video-turbo", - "wan-2.5-image-to-video", - "wan-2.6-flash-image-to-video", - "wan-2.6-image-to-video", - "wan-2.7-image-to-video" -] - added
Input schema / properties / multi_shots / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / negative_prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string. Known values: \"480p\", \"720p\", \"1080p\"." - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "480p", - "720p", - "1080p" -] - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / ratio / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / seed / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / seed / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / source_video_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / watermark / descriptionAdded value: +"Declared type: boolean." - added
Input schema / requiredAdded value: +[]
- Changed
speech_to_video22 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / frames_per_second / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / frames_per_second / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / guidance_scale / descriptionAdded value: +"Declared type: number." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.2-a14b-speech-to-video-turbo" -] - added
Input schema / properties / negative_prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / num_frames / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / num_frames / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / num_inference_steps / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / num_inference_steps / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string. Known values: \"480p\", \"580p\", \"720p\"." - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "480p", - "580p", - "720p" -] - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / seed / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / seed / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / shift / descriptionAdded value: +"Declared type: number." - added
Input schema / properties / source_audio_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / source_image_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991
- Changed
text_to_image22 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / aspect_ratio / descriptionAdded value: +"Declared type: string. Known values: \"1:1\", \"16:9\", \"4:3\", \"21:9\", \"3:4\", \"9:16\", \"8:1\", \"1:8\"." - removed
Input schema / properties / aspect_ratio / enumRemoved value: -[ - "1:1", - "16:9", - "4:3", - "21:9", - "3:4", - "9:16", - "8:1", - "1:8" -] - added
Input schema / properties / bbox_list / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / color_palette / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / enable_sequential / descriptionAdded value: +"Declared type: boolean." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.7-image", - "wan-2.7-image-pro" -] - added
Input schema / properties / output_count / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / output_count / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string. Known values: \"1k\", \"2k\", \"4k\"." - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "1k", - "2k", - "4k" -] - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / seed / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / seed / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / source_image_urls / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / thinking_mode / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / watermark / descriptionAdded value: +"Declared type: boolean." - added
Input schema / requiredAdded value: +[]
- Changed
text_to_video26 fields changed- changed
Input schema / additionalPropertiesPrevious value: -falseNew value: +{} - added
Input schema / properties / acceleration / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / aspect_ratio / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / background_audio_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / callback_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / duration_seconds / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / duration_seconds / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / enable_prompt_expansion / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / enable_safety_checker / descriptionAdded value: +"Declared type: boolean." - added
Input schema / properties / first_frame_image_url / descriptionAdded value: +"Declared type: string." - removed
Input schema / properties / model / enumRemoved value: -[ - "wan-2.2-a14b-text-to-video-turbo", - "wan-2.5-text-to-video", - "wan-2.6-text-to-video", - "wan-2.7-r2v", - "wan-2.7-text-to-video" -] - added
Input schema / properties / multi_shotsAdded value: +{ + "description": "Declared type: boolean.", + "type": "boolean" +} - added
Input schema / properties / negative_prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / output_resolution / descriptionAdded value: +"Declared type: string. Known values: \"480p\", \"580p\", \"720p\", \"1080p\"." - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "480p", - "580p", - "720p", - "1080p" -] - added
Input schema / properties / poll_interval_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / prompt / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / ratio / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / reference_audio_url / descriptionAdded value: +"Declared type: string." - added
Input schema / properties / reference_image_urls / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / reference_video_urls / descriptionAdded value: +"Declared type: array." - added
Input schema / properties / seed / descriptionAdded value: +"Declared type: integer." - changed
Input schema / properties / seed / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / timeout_ms / maximumAdded value: +9007199254740991 - added
Input schema / properties / watermark / descriptionAdded value: +"Declared type: boolean." - added
Input schema / requiredAdded value: +[]
7 tool updates
v0.1.9- Changed
animate5 fields changed- added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / reference_video_url / typeAdded value: +"string" - added
Input schema / properties / source_image_url / typeAdded value: +"string" - added
Input schema / requiredAdded value: +[ + "source_image_url", + "reference_video_url" +]
- Changed
edit_video17 fields changed- removed
Input schema / properties / aspect_ratio / enumRemoved value: -[ - "16:9", - "9:16", - "1:1", - "4:3", - "3:4" -] - added
Input schema / properties / audioAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / audio_settingAdded value: +{ + "type": "string" +} - added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / duration_secondsAdded value: +{ + "type": "number" +} - added
Input schema / properties / enable_prompt_expansionAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / multi_shotsAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / negative_promptAdded value: +{ + "type": "string" +} - removed
Input schema / properties / output_resolution / enumRemoved value: -[ - "720p", - "1080p" -] - added
Input schema / properties / prompt / typeAdded value: +"string" - added
Input schema / properties / reference_image_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / seedAdded value: +{ + "type": "number" +} - added
Input schema / properties / source_video_url / typeAdded value: +"string" - added
Input schema / properties / source_video_urls / itemsAdded value: +{} - added
Input schema / properties / source_video_urls / typeAdded value: +"array" - added
Input schema / properties / watermarkAdded value: +{ + "type": "boolean" +}
- Changed
get_task1 field changed- changed
Input schema / properties / action / descriptionPrevious value: -"Endpoint the task was created on."New value: +"Asynchronous endpoint the task was created on."
- Changed
image_to_video18 fields changed- added
Input schema / properties / accelerationAdded value: +{ + "type": "string" +} - added
Input schema / properties / aspect_ratioAdded value: +{ + "type": "string" +} - added
Input schema / properties / audio / typeAdded value: +"boolean" - added
Input schema / properties / background_audio_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / driving_audio_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / duration_seconds / typeAdded value: +"number" - added
Input schema / properties / enable_prompt_expansionAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / first_frame_image_url / typeAdded value: +"string" - added
Input schema / properties / last_frame_image_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / multi_shotsAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / negative_promptAdded value: +{ + "type": "string" +} - added
Input schema / properties / prompt / typeAdded value: +"string" - added
Input schema / properties / ratioAdded value: +{ + "type": "string" +} - added
Input schema / properties / seedAdded value: +{ + "type": "number" +} - added
Input schema / properties / source_video_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / watermarkAdded value: +{ + "type": "boolean" +}
- Changed
speech_to_video13 fields changed- added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / frames_per_secondAdded value: +{ + "type": "number" +} - added
Input schema / properties / guidance_scaleAdded value: +{ + "type": "number" +} - added
Input schema / properties / negative_promptAdded value: +{ + "type": "string" +} - added
Input schema / properties / num_framesAdded value: +{ + "type": "number" +} - added
Input schema / properties / num_inference_stepsAdded value: +{ + "type": "number" +} - added
Input schema / properties / prompt / typeAdded value: +"string" - added
Input schema / properties / seedAdded value: +{ + "type": "number" +} - added
Input schema / properties / shiftAdded value: +{ + "type": "number" +} - added
Input schema / properties / source_audio_url / typeAdded value: +"string" - added
Input schema / properties / source_image_url / typeAdded value: +"string" - added
Input schema / requiredAdded value: +[ + "source_image_url", + "source_audio_url", + "prompt" +]
- Changed
text_to_image11 fields changed- added
Input schema / properties / bbox_listAdded value: +{ + "items": {}, + "type": "array" +} - added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / color_paletteAdded value: +{ + "items": {}, + "type": "array" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / enable_sequentialAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / output_countAdded value: +{ + "type": "number" +} - added
Input schema / properties / promptAdded value: +{ + "type": "string" +} - added
Input schema / properties / seedAdded value: +{ + "type": "number" +} - added
Input schema / properties / source_image_urlsAdded value: +{ + "items": {}, + "type": "array" +} - added
Input schema / properties / thinking_modeAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / watermarkAdded value: +{ + "type": "boolean" +}
- Changed
text_to_video17 fields changed- added
Input schema / properties / accelerationAdded value: +{ + "type": "string" +} - removed
Input schema / properties / aspect_ratio / enumRemoved value: -[ - "16:9", - "9:16", - "1:1", - "4:3", - "3:4" -] - added
Input schema / properties / background_audio_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / callback_urlAdded value: +{ + "type": "string" +} - removed
Input schema / properties / duration_seconds / maximumRemoved value: -10 - removed
Input schema / properties / duration_seconds / minimumRemoved value: -2 - added
Input schema / properties / enable_prompt_expansionAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / enable_safety_checkerAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / first_frame_image_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / negative_promptAdded value: +{ + "type": "string" +} - added
Input schema / properties / promptAdded value: +{ + "type": "string" +} - added
Input schema / properties / ratioAdded value: +{ + "type": "string" +} - added
Input schema / properties / reference_audio_urlAdded value: +{ + "type": "string" +} - added
Input schema / properties / reference_image_urlsAdded value: +{ + "items": {}, + "type": "array" +} - added
Input schema / properties / reference_video_urlsAdded value: +{ + "items": {}, + "type": "array" +} - added
Input schema / properties / seedAdded value: +{ + "type": "number" +} - added
Input schema / properties / watermarkAdded value: +{ + "type": "boolean" +}
9 tool updates
v0.1.6- First observed
animate - First observed
check_pricing - First observed
edit_video - First observed
get_task - First observed
image_to_video - First observed
login - First observed
speech_to_video - First observed
text_to_image - First observed
text_to_video
TDQS
Scored across 9 tools
Most tools are clearly differentiated by input modality (text_to_image, image_to_video, speech_to_video, etc.), but 'animate' and 'edit_video' have overlapping video-manipulation purposes and vague descriptions, making selection uncertain. Overall, the set is mostly distinct with only minor ambiguity.
All names use snake_case, which is consistent, but the patterns vary widely: some follow input_to_output form (text_to_image), some are verb_noun (edit_video, get_task), and some are single verbs (animate, login). This mixed convention reduces predictability.
Nine tools is a well-scoped set for a generation API, covering multiple creation endpoints plus utility functions. No tool feels redundant.
The surface covers task creation for various modalities, task status retrieval, pricing, and authentication. However, it lacks task management operations like cancel, list, or delete, which could be needed for longer workflows.
Maintenance
Related MCP Connectors
MCP server for Wan AI video generation
MCP server for Pixapi: check live credit pricing and balance, then generate images and video.
MCP server for Hailuo (MiniMax) AI video generation
MCP server for Google Veo AI video generation
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP server for the Gemini Omni model line, enabling task creation (audio, character, text-to-video) and pricing checks through RunAPI.6289 npmApache 2.0
- AlicenseBqualityAmaintenanceMCP server for GPT-4o Image model line, enabling text-to-image task creation, status polling, and pricing checks via RunAPI.4273 npmApache 2.0
- AlicenseAqualityAmaintenanceMCP server for Ideogram V3 model line enabling image editing, reframing, remixing, and text-to-image generation with task polling and pricing checks through RunAPI.7251 npmApache 2.0
- AlicenseBqualityAmaintenanceMCP server for the InfiniteTalk model line, enabling audio-to-video task creation, status polling, and pricing checks via a RunAPI API key.4261 npm1Apache 2.0