AI-video-generator-MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AI-video-generator-MCPCreate a 10 second video of a futuristic city at night in 16:9."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI Video Generator
An MCP-based AI video generation server built with Python and FastMCP.
The project is being developed incrementally, starting with a local proof-of-concept and eventually evolving into a remotely accessible, containerized AI video-generation service.
The initial goal is to evaluate local LLMs such as Qwen3.5 27B and gpt-oss-20B as agents capable of controlling an AI video-generation workflow through MCP tools.
Project Goals
The project aims to provide an MCP interface for AI video generation, allowing an LLM to perform actions such as:
Create a video from a prompt
Check video-generation status
Retrieve generated videos
Cancel video-generation jobs
Eventually work with different video-generation backends
The MCP interface should remain independent from the underlying video-generation implementation.
This allows the project to evolve from a local GPU-based prototype into a remotely deployed service without redesigning the MCP tools.
Project setup
You should have installed python 3.14.7. You can download it from Python Install Manager here: https://www.python.org/downloads/
Install Bionic (the LM Studio agent app): https://lmstudio.ai/
Clone repository to your local machine
Open cloned repository folder with terminal and create local virtual environment:
py -m venv .venvRun from root folder:
uv syncCheck that the server starts:
uv run python -m video_mcp.serverIt should print the FastMCP banner and then sit waiting for input; stop it with Ctrl+C. This only proves the server starts. Do not leave it running: the transport is stdio, which means the MCP client launches its own copy of the server and talks to it over that process's stdin/stdout. A server started by hand in a terminal is not connected to anything.
Related MCP server: Seedance MCP
Connecting the server to Bionic
Add an MCP server in Bionic's settings with these values:
Field | Value |
Command |
|
Working directory | your cloned repository folder |
Env |
|
The command must run the server as a module (-m video_mcp.server) from the
project root. Pointing it at the script path instead (python video_mcp/server.py) puts the video_mcp/ folder on sys.path rather than the
project root, so from video_mcp.models import ... fails, the process dies on
startup, and the client reports MCP error -32000: Connection closed.
Bionic stores this as its own JSON (servers: [...]), not the mcpServers
shape used by most other MCP clients. mcp.json in this repository is the
portable version, for clients that read that format; adjust cwd to your own
folder.
The env values are optional and only control the mock timings. Without them
the job finishes about 10 seconds after it is created, which is too fast to
observe the running state.
To confirm it worked, ask the model in a new chat:
List every tool you have available, with their exact names.You should get create_video, get_video_status, get_video_result and
cancel_video. If the tools are missing, check whether the server process is
actually running while Bionic is open:
Get-CimInstance Win32_Process | Where-Object { $_.CommandLine -like "*video_mcp*" }Nothing listed means the client never started the server, or it crashed on startup.
Downloading a model
Change the model install folder to your bigger ssd
Search for qwen 3.5 27B GGUF and download the unsloth version
(Both screenshots are from LM Studio; Bionic's model settings are equivalent.)
Note on hardware: Qwen3.5 27B at Q4 is about 17 GB. On a card with less VRAM than that, most of the model runs on the CPU, and a single tool-calling turn can take minutes. Check that the GPU is actually being used:
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv -l 2Testing that the LLM can drive the tools
In chat write:
Create a 10 second video of a futuristic city at night in 16:9.It should call the create_video() tool.
The mock job does not finish instantly. With the env values above it stays
queued for 5 seconds and runs for 120, and only then reports completed. A
model that handles the workflow correctly will poll get_video_status and then
call get_video_result.
Raise AI_VIDEO_MOCK_RUNNING_SECONDS for scenarios that need a longer window,
such as cancelling a job while it is still running.
Running the tests
uv run pytest
uv run ruff check .Architecture
The project will be developed in several stages.
Current target architecture
Bionic
│
Local LLM model
┌─────────┴─────────┐
│ │
Qwen3.5 27B gpt-oss-20B
│ │
└─────────┬─────────┘
│
MCP Client
│
stdio
│
▼
FastMCP Server
│
▼
Video Service
│
▼
Local GPU / BackendThe LLM is responsible for understanding the user's request and deciding which MCP tools to use.
The actual video generation is performed by a separate video-generation backend.
Roadmap
Phase 1 — Local MVP / Proof of Concept
Status: In progress — mock backend and LLM evaluation done; real video generation not connected yet.
The first phase focuses entirely on proving that the concept works.
There will be:
No Docker
No HTTPS
No remote deployment
No Terraform
No Ansible
No Jenkins
No VM
The MCP server will run locally using stdio transport.
The video-generation workload will initially use the GPU and other resources available on the developer's PC.
Initial architecture
Bionic
│
│ Local LLM
▼
MCP Client
│
│ stdio
▼
FastMCP Server
│
▼
Video-generation tools
│
▼
Developer PC GPUInitial MCP tools
The first version will provide a minimal set of tools, for example:
create_video()
get_video_status()
get_video_result()
cancel_video()The first implementation may use a mock/fake video backend.
This is intentional.
Before connecting an actual video-generation model, the project should establish that the selected LLM can reliably:
Understand the available MCP tools
Select the correct tool
Generate valid tool arguments
Handle returned job IDs
Check job status
Handle errors
Complete a multi-step video-generation workflow
LLM evaluation
The initial models to evaluate are:
Qwen3.5 27B
gpt-oss-20B
Additional models may be tested later.
The models will be tested through Bionic using the same MCP server and the same tool definitions.
The goal is to determine which model provides the best combination of:
Tool-calling reliability
Reasoning
Parameter accuracy
Context handling
Speed
Resource consumption
Error recovery
Overall reliability as an MCP agent
Phase 2 — Remote MVP
Status: Planned
Once the local proof-of-concept works, the MCP server will be adapted for remote access.
The transport will move from:
stdioto:
Streamable HTTPThe service will eventually be exposed through:
HTTPSTarget architecture
Internet
│
HTTPS
│
▼
Streamable HTTP
│
▼
FastMCP Server
│
▼
Video Backend
│
▼
GPU / ComputeThis phase introduces concerns that are not necessary during local development, including:
HTTPS/TLS
Authentication
Authorization
Secrets management
Network security
Request validation
Logging
Error handling
Rate limiting
Remote configuration
The goal is to make the MCP server usable remotely while keeping the underlying MCP tool interface stable.
Phase 3 — Production MVP
Status: Planned
After the remote MVP has been validated, the application will be moved away from the developer's personal PC and deployed to dedicated infrastructure.
This phase introduces:
Docker
Virtual machines
Dedicated GPU compute
Persistent storage
Environment configuration
Service management
Target architecture
Internet
│
HTTPS
│
▼
┌───────────┐
│ VM │
│ │
│ FastMCP │
│ Server │
└─────┬─────┘
│
▼
Video Backend
│
▼
GPU ComputeDocker will package the application and its Python dependencies into a reproducible environment.
The development environment and production environment should remain consistent as much as practical.
Phase 4 — Infrastructure & Automation
Status: Planned
Once the application is running reliably on dedicated infrastructure, infrastructure automation and CI/CD will be introduced.
Potential technologies include:
Terraform — infrastructure provisioning
Ansible — server configuration and deployment
Jenkins — CI/CD and deployment automation
GitHub — source control and collaboration
Target workflow
Developer
│
▼
GitHub
│
▼
CI / Tests
│
▼
Jenkins
│
├── Terraform
│ │
│ ▼
│ Infrastructure
│
└── Ansible
│
▼
VM / Services
│
▼
Docker
│
▼
FastMCP Server
│
▼
Video Backend
│
▼
GPU WorkerThe purpose of this phase is to make the system reproducible, deployable, and maintainable rather than manually configured.
Development Strategy
The project intentionally follows an incremental approach.
We do not want to solve infrastructure problems before the core application is proven.
The progression is:
1. Prove the MCP concept
↓
2. Test local LLMs
↓
3. Connect real video generation
↓
4. Enable remote access
↓
5. Move to dedicated infrastructure
↓
6. Containerize
↓
7. Automate infrastructure and deploymentEach phase should produce a working system before the next layer of complexity is introduced.
Technology Stack
Initial
Python
FastMCP 4.x
MCP
Bionic (LM Studio agent app)
Qwen3.5 27B
gpt-oss-20B
Local GPU
Development
uvpyproject.tomlPython virtual environment
Git
GitHub
pytest
Ruff
Later
Streamable HTTP
HTTPS
Docker
Virtual machines
GPU infrastructure
Terraform
Ansible
Jenkins
The exact video-generation model and backend will be selected after the initial MCP/LLM proof-of-concept.
Project Structure
The project is expected to follow a structure similar to:
AI-video-generator/
│
├── video_mcp/
│ ├── __init__.py
│ ├── server.py MCP tool layer (thin)
│ ├── video.py video-generation backend
│ └── models.py request/response models
│
├── tests/
│ ├── conftest.py
│ ├── test_backend_config.py timings read from the environment
│ ├── test_models.py request validation and terminal states
│ ├── test_server.py tool layer, via an in-memory MCP client
│ ├── test_stdio.py real subprocess launch over stdio
│ └── test_video.py backend job lifecycle
│
├── mcp.json
├── pyproject.toml
├── uv.lock
├── .gitignore
└── readme.mdThe package is named video_mcp, not mcp. A local package called mcp shadows
the installed mcp SDK that FastMCP depends on, which breaks both the server and
the test suite when anything runs from the project root.
Subpackages (tools/, services/, models/) can be introduced later if the
flat modules grow. Docker-related files are not needed during Phase 1 and will
become relevant during the deployment phase.
Development Environment
A dedicated Python environment will be used during development.
The project should not install its dependencies globally into the developer's system Python installation.
uv will be used to manage the project environment and dependencies.
For example:
uv syncThis creates/updates the project's isolated environment based on pyproject.toml and uv.lock.
The virtual environment should not be committed to Git.
MCP Transport
Phase 1
stdiostdio is used because the MCP server is running locally and the main objective is rapid development and testing.
Phase 2+
Streamable HTTPStreamable HTTP will be introduced when the MCP server needs to be accessed remotely.
The MCP tools themselves should remain largely independent of the transport.
For example:
@mcp.tool
def create_video(
prompt: str,
duration: int = 5,
aspect_ratio: str = "16:9",
): ...The same logical tool should be usable regardless of whether the MCP server is accessed through stdio or Streamable HTTP.
Design Principles
1. Keep the MCP layer thin
MCP tools should expose a clean interface to the LLM.
Complex video-generation logic should live in services/backend components rather than directly inside the MCP tool implementation.
2. Keep the video backend replaceable
The MCP server should not be tightly coupled to one video-generation implementation.
Possible future backends include:
Local video model
ComfyUI
Remote video-generation API
Dedicated GPU worker3. Keep the LLM replaceable
The MCP server should not be designed around a specific LLM.
The same MCP tools should be testable with:
Qwen3.5 27B
gpt-oss-20B
Other local models
Future models4. Introduce infrastructure only when needed
The project will start locally and become progressively more production-oriented.
There is no need to introduce Docker, HTTPS, Terraform, Ansible, Jenkins, or cloud infrastructure before the core application has been validated.
Current Status
Phase 1 — Local MVP / Proof of Concept
The current focus is:
Create project structure
Configure Python environment
Configure
pyproject.tomlInstall FastMCP 4.x
Create basic stdio MCP server
Implement
create_video()mock toolImplement
get_video_status(),get_video_result(),cancel_video()Mock backend with real job state and error cases
Test suite covering the backend and the tool layer
Connect MCP client to Bionic
Test Qwen3.5 27B
Test gpt-oss-20B
Compare tool-calling performance
Select initial LLM
Select video-generation backend
Connect real video generation
Implemented tools
Tool | Returns | Errors |
| queued job with | invalid prompt, duration outside 1-60, unsupported aspect ratio |
| job status and progress | unknown |
| video path for a completed job | unknown |
| cancelled job | unknown |
Generation itself is still mocked: no file is written and video_path is a
placeholder. What is real is the job lifecycle — queued → running →
completed, with cancelled as a sticky terminal state — so the workflow and
the error paths can be evaluated before a backend exists.
LLM evaluation results
Each model was run through the same scenarios against the same mock backend. Both Qwen3.5 27B and gpt-oss-20B passed all eight.
# | Scenario | What it tests | Qwen3.5 27B | gpt-oss-20B |
1 | "Create a 10 second video of a futuristic city at night in 16:9." | tool selection, argument accuracy | pass | pass |
2 | Ask for the video immediately after creating it | does it poll, or give up on the error | pass | pass |
3 | Ask for a 5 minute video | recovery from the duration limit | pass | pass |
4 | Ask about a job id that was never created | recovery from an unknown id | pass | pass |
5 | Create a video, then cancel it while it is running | multi-step state handling | pass | pass |
6 | Create two videos, then ask about the second | keeping two job ids apart | pass | pass |
7 | Ask for an unsupported aspect ratio | recovery from an invalid enum | pass | pass |
8 | Cancel a job that has already completed | terminal-state error handling | pass | pass |
Because both models pass every scenario, capability is not what separates them. The useful comparison is in the softer measures, which should be recorded per run: wall time and number of turns to completion, whether the job_id was carried between calls without prompting, instruction adherence, memory use, and overclaiming.
Overclaiming is worth watching closely. The tools return only prompt,
duration, aspect_ratio, video_path and message — nothing describing the
imagery. A model that reports what the video looks like has invented it. In
testing, Qwen3.5 27B did exactly that on one run, describing snow-covered pines
and a frozen stream that no tool ever returned, and dropped the mock/placeholder
caveat it had correctly relayed on an earlier run.
Running the scenarios
Use a fresh chat for each scenario, or a previous job id in the context will be what you are testing. Repeat each scenario about three times; tool calling is nondeterministic and a single pass proves little.
Scenarios 2 and 5 need the mock job to still be running when you take your turn.
Raise AI_VIDEO_MOCK_RUNNING_SECONDS to 600 for those, and for scenario 5 tell
the model not to wait:
Create a 10 second video of a snowy forest. Do not wait for it to finish and do
not check its status — just give me the job id.Otherwise the model polls the job to completion inside its own turn, and there is never a running job left for you to cancel.
Long-Term Vision
The final system is intended to become a remotely accessible MCP-based AI video-generation service.
The long-term architecture may look like:
User / AI Agent
│
▼
MCP Client
│
HTTPS
│
▼
┌─────────────────┐
│ FastMCP API │
└────────┬────────┘
│
Job Management
│
┌────────┴────────┐
│ │
▼ ▼
Video Queue Other Services
│
▼
GPU Worker(s)
│
▼
Video Generation
│
▼
Storage / ResultThe exact architecture will evolve as the project progresses.
The primary objective is to keep the system modular, testable, replaceable, and deployable while avoiding unnecessary complexity during the early development stages.
Available Tools
4 toolscancel_videoCancel VideoA
Cancel a video-generation job that has not finished yet.
Jobs that are already completed or cancelled cannot be cancelled.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Identifier returned by create_video. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job_id | Yes | Identifier used to track this job. |
| prompt | Yes | Prompt the job was created from. |
| status | Yes | Current lifecycle state of the job. |
| message | Yes | Human-readable summary of the job state. |
| duration | Yes | Requested video length in seconds. |
| progress | Yes | Completion percentage, 0-100. |
| aspect_ratio | Yes | Requested aspect ratio. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses a key behavioral constraint: cancellation is only possible for jobs that have not finished, and completed or already-cancelled jobs cannot be cancelled. It doesn't describe failure modes or whether cancellation is asynchronous, but the most important behavior is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The main purpose is front-loaded, followed by the key limitation. Every part of the description earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description explains the core precondition and edge cases. It could add what happens on attempting to cancel an already completed job, but the provided information is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the job_id parameter is described as 'Identifier returned by create_video,' which is precise. The tool description adds no additional parameter meaning, but the baseline of 3 applies because the schema fully documents the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('cancel') and resource ('video-generation job'), and immediately clarifies the scope with 'that has not finished yet.' This clearly distinguishes it from sibling tools like create_video, get_video_status, and get_video_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when the tool is applicable (only for unfinished jobs) and explicitly excludes jobs that are completed or cancelled. While it doesn't name an alternative tool, the condition is clear enough that an agent knows not to call it for finished jobs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_videoCreate VideoA
Start generating a video from a text prompt.
Returns immediately with a job_id; the video is not ready yet. Poll get_video_status with that job_id until the status is 'completed', then call get_video_result to retrieve the video.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of the video to generate. | |
| duration | No | Length in seconds, from 1 to 60. | |
| aspect_ratio | No | One of '16:9', '9:16', '1:1', '4:3', '21:9'. | 16:9 |
Output Schema
| Name | Required | Description |
|---|---|---|
| job_id | Yes | Identifier used to track this job. |
| prompt | Yes | Prompt the job was created from. |
| status | Yes | Current lifecycle state of the job. |
| message | Yes | Human-readable summary of the job state. |
| duration | Yes | Requested video length in seconds. |
| progress | Yes | Completion percentage, 0-100. |
| aspect_ratio | Yes | Requested aspect ratio. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of explaining behavior. It clearly discloses the asynchronous nature ('Returns immediately with a job_id; the video is not ready yet') and directs the agent to poll for completion. It does not mention failure modes or cancellation, but the core non-obvious behavior is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero redundancy. It front-loads the purpose, then gives the async caveat and next steps in order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the async workflow is described, the description is complete. It tells the agent exactly what to do after calling this tool, including which siblings to invoke and when. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description mentions 'text prompt' but does not add meaning beyond what the schema already provides for the three parameters. It adds no new semantic context for duration or aspect_ratio.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Start generating a video from a text prompt.' It clearly distinguishes this tool from its siblings by framing it as the submission step, while the follow-up tools are explicitly named for later stages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit workflow: call this tool to start generation, then poll get_video_status until 'completed', then call get_video_result. This directly tells the agent how to use the tool versus the alternatives, leaving no ambiguity about the lifecycle.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_resultGet Video ResultA
Retrieve the finished video for a completed job.
Only works once get_video_status reports 'completed'. Calling it earlier is an error, not a wait.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Identifier returned by create_video. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job_id | Yes | Identifier of the completed job. |
| prompt | Yes | Prompt the video was generated from. |
| status | Yes | Always 'completed' for a result. |
| message | Yes | Human-readable summary of the result. |
| duration | Yes | Video length in seconds. |
| video_path | Yes | Path to the generated video file. |
| aspect_ratio | Yes | Aspect ratio of the video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It clearly states the precondition and the failure behavior when called too early. It does not discuss return format or side effects, but for a read-only retrieval tool this covers the most important non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two crisp sentences with no filler. The core purpose is front-loaded, and the critical usage caveat follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description covers the essential context: what it returns, when it is valid, and what happens if used prematurely. Sibling tools are available as context, and no additional prerequisities or side effects are needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage for job_id, including its source ('Identifier returned by create_video'). The description adds no new parameter-level meaning beyond tying the job to a completed video, which is implicit in the tool's purpose. Baseline of 3 is appropriate given the schema's thoroughness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve') and resource ('finished video for a completed job'), making the tool's function unambiguous. It also implies the distinction from get_video_status by focusing on the result rather than status. An agent can clearly tell this tool's purpose without looking at the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit precondition: only use after get_video_status reports 'completed'. It even warns that calling earlier is an error, not a wait, which prevents misuse. This is strong, actionable guidance for when to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_statusGet Video StatusA
Check the progress of a video-generation job.
Use the job_id returned by create_video. Statuses are 'queued', 'running', 'completed' and 'cancelled'. Keep polling while the job is queued or running.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Identifier returned by create_video. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job_id | Yes | Identifier used to track this job. |
| prompt | Yes | Prompt the job was created from. |
| status | Yes | Current lifecycle state of the job. |
| message | Yes | Human-readable summary of the job state. |
| duration | Yes | Requested video length in seconds. |
| progress | Yes | Completion percentage, 0-100. |
| aspect_ratio | Yes | Requested aspect ratio. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It discloses the polling nature of the tool and lists all possible statuses, which is meaningful behavioral context beyond the tool name. It does not mention error behavior or rate limits, but for a read-only status check this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and efficiently structured: purpose first, followed by job_id source, status vocabulary, and polling guidance. Every sentence earns its place and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with an output schema present, the description covers the essential behavior and polling loop. It would be slightly more complete if it explicitly mentioned using get_video_result once the status becomes 'completed', but this is reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully describes job_id with 100% coverage, so the baseline is 3. The description restates that job_id comes from create_video, matching the schema without adding new semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Check the progress'), targets a specific resource ('video-generation job'), and enumerates the statuses. It differentiates itself from create_video and cancel_video, though it does not explicitly contrast with get_video_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete usage guidance: use the job_id from create_video and keep polling while status is queued or running. It does not explicitly state when to switch to get_video_result, but the polling instruction makes the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
cancel_video - First observed
create_video - First observed
get_video_result - First observed
get_video_status
TDQS
Scored across 4 tools
Each tool maps to a distinct stage of the asynchronous video-generation workflow: create starts a job, status polls progress, result retrieves output, and cancel aborts a pending job. No two tools overlap in purpose or return type.
All tool names use a consistent verb_noun snake_case pattern: create_video, get_video_status, get_video_result, and cancel_video. The get_ prefix is consistently used for retrieval operations, making the API surface easy to predict.
Four tools is an appropriate size for this narrow domain; each tool has a clear, necessary role in the async job lifecycle. There is no redundancy or obvious missing basic operation.
The set supports the full create, poll, retrieve, and cancel workflow for video generation. The main gap is that get_video_status does not expose a 'failed' terminal state, so error handling requires a workaround such as a client-side timeout.
Maintenance
Related MCP Connectors
AI video editor for agents and humans: timeline, captions, color, audio and generation as MCP tools.
Create and manage AI image and video generations through Quriov's fixed public MCP tools.
Plan, compare, price, generate, and recover AI video from compatible MCP clients.
Build, run, schedule, and publish AI video pipelines to YouTube and TikTok from any MCP client.
Related MCP Servers
AlicenseNot gradedqualityFmaintenanceEnables video generation from text, images, and more through MCP-compatible apps like Claude and Cursor.52MIT- FlicenseAqualityDmaintenanceEnables video generation using the Seedance 2.0 model through MCP, supporting both OpenAI and Volcengine API formats with tools for creating, monitoring, and downloading videos.6-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to perform deterministic video editing operations like trim, resize, add text, and more using MCP tools.8 npm5MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to generate professional storyboards and videos from scripts or creative descriptions via MCP-compatible clients.21 npmMIT