Skip to main content
Glama
el-el-san

AI Video Generator MCP Server

by el-el-san

AI Video Generator MCP Server

This MCP (Model Context Protocol) server provides tools to generate videos from text prompts and images using AI image generation models.

Compatible Models

  • Luma Ray2 Flash - Luma's cutting edge image to video conversion model

  • Kling v1.6 Pro - Kling's high quality image to video conversion model

Related MCP server: Video Edit MCP Server

function

  • Video generation from text prompts

  • Video generation with start and/or end images

  • Control video parameters (aspect ratio, resolution, duration, loop)

  • Check the generation status

  • Choosing which AI model to use

install

  1. Clone this repository

  2. Install the dependencies:

    npm install
  3. Create a .env file and set your FAL.AI API key:

    FAL_KEY=your_fal_key_here

    You can get the API key from FAL.AI

Building the Server

npm run build

Running the Server

You can run the server directly:

npm start

Integration with Claude Desktop

To use this server with Claude Desktop, add the following to your claude_desktop_config.json file:

{
  "mcpServers": {
    "video-generator": {
      "command": "node",
      "args": ["your_install_path/fal-mcp-server/build/index.js"],
      "env": {
        "FAL_KEY": "your_fal_key_here"
      }
    }
  }
}

Available Tools

generate-video

It uses AI models to generate videos from text prompts and/or images.

Parameters:

  • prompt (required): A text description of the content of the video you want to generate.

  • image_url (optional): The starting image URL for the video (URL or base64 data URI).

  • end_image_url (optional): The end image URL for the video (URL or base64 data URI).

  • aspect_ratio (default "16:9"): Video aspect ratio ("16:9", "9:16", "4:3", "3:4", "21:9", "9:21")

  • resolution (default "540p"): Video resolution ("540p", "720p", "1080p")

  • duration (default "5s"): video length ("5s", "9s")

  • loop (default false): whether the video should loop

  • model (default "luma"): AI model to use ("luma"=Ray2, "kling"=Kling v1.6 Pro)

check-video-status

Check the status of your video generation request.

Parameters:

  • request_id (required): The request ID to check.

  • model (default "luma"): AI model used for the request ("luma"=Ray2, "kling"=Kling v1.6 Pro)

Claude usage example

猫が毛糸玉で遊んでいる動画を生成してください。縦向きモードでお願いします。Klingモデルを使用してください。

Claude calls the generate-video tool with the appropriate parameters and provides the resulting video URL.

Compare Models

  • Luma Ray2 Flash : Excellent for smooth motion and realistic physics, producing natural results.

  • Kling v1.6 Pro : Excellent for detailed textures and special effects, producing stylized results.

Depending on the prompt and the desired outcome, different models may work best.

Limitations

  • Video generation may take some time (especially at higher resolutions)

  • A valid FAL.AI API key and sufficient credits are required

  • Higher resolution and longer videos cost more credits

  • Both models consume FAL.AI credits (prices may vary per model)

troubleshooting

API Key Error

Make sure the FAL_KEY environment variable is set correctly, or you can set it directly in the Claude Desktop configuration file.

Video Generation Error

If an error occurs during video generation, a detailed error message will be logged. Common issues are:

  • Invalid or expired API key

  • Insufficient credits on your account

  • Inappropriate prompts or images

  • A temporary server-side issue

If the error persists, wait a while and try again or try changing the prompt.

license

MIT

Available Tools

2 tools
check-video-statusC

Check the status of a video generation request

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoAI model used for the request (luma=Ray2, kling=Kling)luma
request_idYesThe request ID to check

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states what the tool does but reveals nothing about response format, error conditions, rate limits, authentication requirements, or whether this is a read-only operation (though implied by 'check'). For a status-checking tool with zero annotation coverage, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that states the core purpose without any wasted words. It's appropriately sized and front-loaded with the essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a status-checking tool with no annotations and no output schema, the description is incomplete. It doesn't explain what status information will be returned, possible states (pending, completed, failed), or how to interpret results. The agent would need to guess about the tool's behavior and output format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents both parameters. The description doesn't add any parameter semantics beyond what's in the schema (like explaining the relationship between request_id and previous video generation). Baseline 3 is appropriate when the schema does all the parameter documentation work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('check') and resource ('video generation request'), making the purpose immediately understandable. However, it doesn't differentiate this status-checking tool from its sibling 'generate-video' tool, which would be helpful for an agent choosing between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There's no mention of prerequisites (like needing a request_id from a previous generation), nor does it explain the relationship with the sibling 'generate-video' tool, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate-videoB

Generate a video from text prompt and/or images using AI models (Luma or Kling)

ParametersJSON Schema
NameRequiredDescriptionDefault
aspect_ratioNoAspect ratio of the video16:9
durationNoDuration of the video (9s costs 2x more)5s
end_image_urlNoFinal image to end the video with (URL or base64 data URI)
image_urlNoInitial image to start the video from (URL or base64 data URI)
loopNoWhether the video should loop (blend end with beginning)
modelNoAI model to use (luma=Ray2, kling=Kling)luma
promptYesText description of the desired video content
resolutionNoResolution of the video (higher resolutions use more credits)540p

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions AI models but doesn't disclose important behavioral traits like: whether this is a synchronous or asynchronous operation, what permissions or authentication are needed, rate limits, credit costs (beyond the hint in the duration parameter schema), or what the output looks like. The description is minimal and lacks crucial operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that states the core purpose without unnecessary words. It's appropriately sized and front-loaded with the essential information. Every word earns its place in this concise formulation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex video generation tool with 8 parameters and no annotations or output schema, the description is insufficient. It doesn't explain the operation's nature (async/sync), authentication requirements, cost implications beyond the duration hint, error conditions, or what happens after invocation. The combination of complexity and lack of structured metadata demands more comprehensive description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds no parameter-specific information beyond what's in the schema. The baseline score of 3 reflects adequate coverage through the schema alone, but the description doesn't enhance understanding of parameter usage or relationships.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate a video from text prompt and/or images using AI models (Luma or Kling)'. It specifies the verb ('generate'), resource ('video'), and input sources ('text prompt and/or images'), but doesn't differentiate from its sibling tool 'check-video-status' beyond the obvious generation vs. status check distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by mentioning AI models (Luma or Kling), suggesting this is for AI-generated video creation. However, it doesn't provide explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites or exclusions. The sibling tool 'check-video-status' is clearly complementary rather than an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.1/5.0
Disambiguation5/5

The two tools have completely distinct purposes with no overlap: one checks status of existing requests, the other creates new videos. An agent would never confuse these functions as they operate on different stages of the video generation workflow.

Naming Consistency5/5

Both tools follow a consistent verb-object naming pattern with hyphen separation: 'check-video-status' and 'generate-video'. The naming is predictable and follows the same convention throughout the set.

Tool Count2/5

With only 2 tools for a video generation server, the surface feels severely limited. While the basic create+status pair covers minimal functionality, a video generation domain typically requires more operations like listing videos, canceling generations, or managing templates. The count is too low for the apparent scope.

Completeness2/5

The toolset provides only generation initiation and status checking, creating significant gaps. Missing are operations like listing existing videos, canceling pending generations, retrieving generated content, managing templates/presets, or configuring generation parameters. Agents will hit dead ends when trying to manage the video lifecycle beyond initial creation.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/el-el-san/fal-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server