AI Video Generator MCP Server
Enables configuration of API keys through environment variables in a .env file for secure credential management.
Provides package management for installing dependencies and running the server through npm commands.
Offers access to Luma Ray2 Flash, an AI model for image-to-video conversion, allowing for video generation from text prompts or images.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AI Video Generator MCP ServerCreate a 9-second video of a sunset over mountains with Kling model"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI Video Generator MCP Server
This MCP (Model Context Protocol) server provides tools to generate videos from text prompts and images using AI image generation models.
Compatible Models
Luma Ray2 Flash - Luma's cutting edge image to video conversion model
Kling v1.6 Pro - Kling's high quality image to video conversion model
Related MCP server: Video Edit MCP Server
function
Video generation from text prompts
Video generation with start and/or end images
Control video parameters (aspect ratio, resolution, duration, loop)
Check the generation status
Choosing which AI model to use
install
Clone this repository
Install the dependencies:
npm installCreate a
.envfile and set your FAL.AI API key:FAL_KEY=your_fal_key_hereYou can get the API key from FAL.AI
Building the Server
npm run buildRunning the Server
You can run the server directly:
npm startIntegration with Claude Desktop
To use this server with Claude Desktop, add the following to your claude_desktop_config.json file:
{
"mcpServers": {
"video-generator": {
"command": "node",
"args": ["your_install_path/fal-mcp-server/build/index.js"],
"env": {
"FAL_KEY": "your_fal_key_here"
}
}
}
}Available Tools
generate-video
It uses AI models to generate videos from text prompts and/or images.
Parameters:
prompt(required): A text description of the content of the video you want to generate.image_url(optional): The starting image URL for the video (URL or base64 data URI).end_image_url(optional): The end image URL for the video (URL or base64 data URI).aspect_ratio(default "16:9"): Video aspect ratio ("16:9", "9:16", "4:3", "3:4", "21:9", "9:21")resolution(default "540p"): Video resolution ("540p", "720p", "1080p")duration(default "5s"): video length ("5s", "9s")loop(default false): whether the video should loopmodel(default "luma"): AI model to use ("luma"=Ray2, "kling"=Kling v1.6 Pro)
check-video-status
Check the status of your video generation request.
Parameters:
request_id(required): The request ID to check.model(default "luma"): AI model used for the request ("luma"=Ray2, "kling"=Kling v1.6 Pro)
Claude usage example
猫が毛糸玉で遊んでいる動画を生成してください。縦向きモードでお願いします。Klingモデルを使用してください。Claude calls the generate-video tool with the appropriate parameters and provides the resulting video URL.
Compare Models
Luma Ray2 Flash : Excellent for smooth motion and realistic physics, producing natural results.
Kling v1.6 Pro : Excellent for detailed textures and special effects, producing stylized results.
Depending on the prompt and the desired outcome, different models may work best.
Limitations
Video generation may take some time (especially at higher resolutions)
A valid FAL.AI API key and sufficient credits are required
Higher resolution and longer videos cost more credits
Both models consume FAL.AI credits (prices may vary per model)
troubleshooting
API Key Error
Make sure the FAL_KEY environment variable is set correctly, or you can set it directly in the Claude Desktop configuration file.
Video Generation Error
If an error occurs during video generation, a detailed error message will be logged. Common issues are:
Invalid or expired API key
Insufficient credits on your account
Inappropriate prompts or images
A temporary server-side issue
If the error persists, wait a while and try again or try changing the prompt.
license
MIT
Available Tools
2 toolscheck-video-statusC
Check the status of a video generation request
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | AI model used for the request (luma=Ray2, kling=Kling) | luma |
| request_id | Yes | The request ID to check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states what the tool does but reveals nothing about response format, error conditions, rate limits, authentication requirements, or whether this is a read-only operation (though implied by 'check'). For a status-checking tool with zero annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core purpose without any wasted words. It's appropriately sized and front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a status-checking tool with no annotations and no output schema, the description is incomplete. It doesn't explain what status information will be returned, possible states (pending, completed, failed), or how to interpret results. The agent would need to guess about the tool's behavior and output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters. The description doesn't add any parameter semantics beyond what's in the schema (like explaining the relationship between request_id and previous video generation). Baseline 3 is appropriate when the schema does all the parameter documentation work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('check') and resource ('video generation request'), making the purpose immediately understandable. However, it doesn't differentiate this status-checking tool from its sibling 'generate-video' tool, which would be helpful for an agent choosing between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. There's no mention of prerequisites (like needing a request_id from a previous generation), nor does it explain the relationship with the sibling 'generate-video' tool, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate-videoB
Generate a video from text prompt and/or images using AI models (Luma or Kling)
| Name | Required | Description | Default |
|---|---|---|---|
| aspect_ratio | No | Aspect ratio of the video | 16:9 |
| duration | No | Duration of the video (9s costs 2x more) | 5s |
| end_image_url | No | Final image to end the video with (URL or base64 data URI) | |
| image_url | No | Initial image to start the video from (URL or base64 data URI) | |
| loop | No | Whether the video should loop (blend end with beginning) | |
| model | No | AI model to use (luma=Ray2, kling=Kling) | luma |
| prompt | Yes | Text description of the desired video content | |
| resolution | No | Resolution of the video (higher resolutions use more credits) | 540p |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions AI models but doesn't disclose important behavioral traits like: whether this is a synchronous or asynchronous operation, what permissions or authentication are needed, rate limits, credit costs (beyond the hint in the duration parameter schema), or what the output looks like. The description is minimal and lacks crucial operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core purpose without unnecessary words. It's appropriately sized and front-loaded with the essential information. Every word earns its place in this concise formulation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex video generation tool with 8 parameters and no annotations or output schema, the description is insufficient. It doesn't explain the operation's nature (async/sync), authentication requirements, cost implications beyond the duration hint, error conditions, or what happens after invocation. The combination of complexity and lack of structured metadata demands more comprehensive description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds no parameter-specific information beyond what's in the schema. The baseline score of 3 reflects adequate coverage through the schema alone, but the description doesn't enhance understanding of parameter usage or relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate a video from text prompt and/or images using AI models (Luma or Kling)'. It specifies the verb ('generate'), resource ('video'), and input sources ('text prompt and/or images'), but doesn't differentiate from its sibling tool 'check-video-status' beyond the obvious generation vs. status check distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by mentioning AI models (Luma or Kling), suggesting this is for AI-generated video creation. However, it doesn't provide explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites or exclusions. The sibling tool 'check-video-status' is clearly complementary rather than an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have completely distinct purposes with no overlap: one checks status of existing requests, the other creates new videos. An agent would never confuse these functions as they operate on different stages of the video generation workflow.
Both tools follow a consistent verb-object naming pattern with hyphen separation: 'check-video-status' and 'generate-video'. The naming is predictable and follows the same convention throughout the set.
With only 2 tools for a video generation server, the surface feels severely limited. While the basic create+status pair covers minimal functionality, a video generation domain typically requires more operations like listing videos, canceling generations, or managing templates. The count is too low for the apparent scope.
The toolset provides only generation initiation and status checking, creating significant gaps. Missing are operations like listing existing videos, canceling pending generations, retrieving generated content, managing templates/presets, or configuring generation parameters. Agents will hit dead ends when trying to manage the video lifecycle beyond initial creation.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Kling AI video generation
MCP server for Google Veo AI video generation
MCP server for Hailuo (MiniMax) AI video generation
MCP server for Wan AI video generation
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA server that provides Luma AI's video generation API as the Model Context Protocol (MCP)3
- AlicenseNot gradedqualityFmaintenanceA Model Context Protocol server that enables AI assistants to perform comprehensive video and audio editing operations including trimming, effects, overlays, audio processing, and YouTube downloads.25MIT
- AlicenseAqualityAmaintenanceFastMCP server for Google's gemini-omni-flash-preview video model, enabling text-to-video, image-to-video, and video editing with stateful interactions and batch generation.2MIT
- AlicenseNot gradedqualityCmaintenanceA Model Context Protocol server for AI image and video generation using Midjourney through the AceDataCloud API.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/el-el-san/fal-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server