MCP Video Generation with Veo2
Use this MCP server to generate Google Veo videos and configured image-model images, then list or retrieve saved media through MCP clients over stdio, HTTP, or SSE.
Generate videos from text prompts.
Generate videos from image inputs: base64, MCP image object, allowed HTTPS URL, or allowed local file path.
Generate images from text prompts.
Generate a video from a generated image in one step.
List generated videos and images.
Retrieve a saved image by UUID.
Receive metadata plus
videos://<uuid>orimages://<uuid>resource links; optionally include full media data when within size limits.Configure video/image settings such as aspect ratio, duration, negative prompt, person-generation mode, and inclusion of full data.
Run over MCP stdio, Streamable HTTP, or legacy SSE; HTTP/SSE require
MCP_AUTH_TOKEN.Use discovery, templates, and saved-media access without a Google API key; generation requires one.
Configure models, storage directory, host/port, allowed hosts/origins, image URL allowlist, timeouts, and media-size limits via environment variables.
Integrates with Google's Veo2 video generation capabilities, allowing generation of videos from text prompts or images with various configuration options such as aspect ratio, duration, and person generation settings.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Video Generation with Veo2generate a cinematic video of a sunset over mountains with golden hour lighting"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Google Veo and Gemini image MCP server
Generate videos with Google Veo and images with Gemini. The package name remains
mcp-video-generation-veo2 for existing installations; version 2 uses maintained,
configurable models and MCP SDK 2.
Install and run
Use Node.js 22 or 24. From this repository:
npm ci
npm run build
node dist/index.jsFor an MCP host, launch node directly with the absolute dist/index.js path.
Do not launch npm start as the stdio child: npm's own banners can corrupt the
JSON-RPC stream. The installed package also provides mcp-video-generation-veo2.
{
"mcpServers": {
"veo": {
"command": "node",
"args": ["/absolute/path/mcp-veo2/dist/index.js"],
"env": {
"GOOGLE_API_KEY": "YOUR_GOOGLE_API_KEY",
"VIDEO_MODEL": "veo-3.1-generate-preview",
"IMAGE_MODEL": "gemini-3.1-flash-image"
}
}
}
}An API key is needed only for generation; discovery, templates, and saved-media access work without one. Google model access, billing and quota must be enabled. Generation may incur charges even if the MCP request is cancelled or times out. The server never automatically retries a paid creation request.
Related MCP server: hyper-video-service
HTTP and existing SSE clients
Set MCP_AUTH_TOKEN to a random value of at least 24 characters, then run:
node dist/index.js httpOpen http://127.0.0.1:3000/ for the included browser connection check. It asks for
the MCP token, keeps it in memory, and lists saved media using real MCP requests.
/mcp: MCP 2026-07-28 and stateless legacy Streamable HTTP./sseplus/messages?sessionId=...: legacy SSE compatibility for existing clients such as LibreChat. Each connection owns its own MCP server; request bodies are parsed once. SSE sessions expire after 10 minutes without a POST.node dist/index.js sseremains an alias for the same HTTP listener.Send
Authorization: Bearer <MCP_AUTH_TOKEN>on every MCP request, including the initial SSE GET. The Google API key is a separate, server-only credential.
The default bind address is loopback. To expose the service, configure HOST,
ALLOWED_HOSTS and ALLOWED_ORIGINS for the proxy deployment and terminate TLS at
that proxy. Only exact configured Host values and browser origins are trusted.
Each deployment and storage directory belongs to one operator; do not share its
bearer token/storage across untrusted tenants. This is static bearer deployment
authentication, not a complete OAuth authorization server.
Issue #5's file:// frontend and /sse/tool/listGeneratedVideos URL were never an
MCP REST API. Use the included page or an MCP client. Opaque Origin: null and
hostile origins are rejected; no wildcard CORS is enabled. The old public
/videos directory server has been removed.
Tools and media
All seven tool names are retained:
Tool | Inputs |
|
|
|
|
|
|
|
|
| No inputs |
| No inputs |
| Saved image UUID, optional |
Settings are flat tool arguments, for example:
{
"prompt": "A slow camera pan through a sunlit forest",
"aspectRatio": "16:9",
"durationSeconds": 8,
"resolution": "720p"
}Veo 3.1 accepts durations 4, 6, or 8 seconds; 1080p/4k requires 8 seconds. Text
generation uses personGeneration=allow_all; image animation uses
allow_adult, subject to Google's regional restrictions and filtering. One
video/image is generated per call. Explicit unsupported settings fail before
provider generation.
An image input can be a base64 string, an MCP image object
{ "type": "image", "mimeType": "image/png", "data": "..." }, an allowed HTTPS URL,
or an absolute file path beneath IMAGE_INPUT_DIR. PNG, JPEG and WebP are checked
by MIME type and file signature. Input images are limited to 10 MiB. URL imports
require an explicit IMAGE_URL_HOSTS allowlist controlled by the operator; redirects
must remain in that allowlist. Do not allow hosts you do not trust.
Outputs contain metadata and videos://<uuid> or images://<uuid> resource links.
Resources return actual media blobs; video data is never mislabeled as an MCP
image. includeFullData=true additionally embeds media in the tool result, limited
to 8 MiB. Resources are likewise limited to 8 MiB for interoperability. Larger
assets remain in STORAGE_DIR for the operator to retrieve on the server.
Video downloads always run on the server. API keys are sent in headers, never in returned or saved URLs. Redirects are bounded; cross-host redirects do not receive the Google key. Downloaded bytes, metadata and API responses have size limits. Provider error bodies and raw request arguments are not logged. Startup rewrites valid historical metadata through an allowlist, dropping old key-bearing URLs and absolute paths while preserving media files. Rotate keys previously exposed by an older release and remove historical copies from any downstream logs/backups.
Configuration
Copy .env.example or set environment variables:
Variable | Default / meaning |
| Provider key; |
|
|
|
|
|
|
|
|
| Required for HTTP/SSE, at least 24 characters |
| Additional comma-separated exact |
| Additional comma-separated exact HTTP(S) origins |
|
|
| Empty; comma-separated trusted HTTPS source hosts |
| 600000, maximum 1800000 |
| 10000 |
| 67108864; maximum 134217728 |
Veo 2, Imagen 3 and Imagen 4 are not suitable defaults for the end of 2026.
Google's Veo guide documents the current
video interface; image generation
lists current Gemini image models. This implementation uses the supported
generateContent image API; Google's
Interactions overview
confirms it remains supported. Recheck model availability and
deprecations before deployment.
Configurable model IDs permit compatible replacements; they do not guarantee that
a future model accepts the same parameters. Veo 3.1's current identifier is a
preview, so end-of-2026 availability cannot be promised.
Version 2 migration
MCP stdio and HTTP serve revision 2026-07-28 through SDK2's actual
serveStdio/createMcpHandlerentries. Legacy stdio/HTTP/SSE remains tested.Duration 5/7,
numberOfVideos=2,numberOfImages>1, unsupported person modes,enhancePrompt=true, andautoDownload=falsenow fail explicitly.Media always downloads locally, and public storage directory access is removed.
Use direct Node/package binary for MCP stdio. Missing generation credentials no longer prevent listing tools.
Unsupported resource subscriptions are not advertised.
Verification
npm run checkTests use credential-free provider fixtures and actual stdio, Streamable HTTP, legacy SSE and production-only packed installs. They cover schema errors, cancellation, key handling, redirects, size limits, media types, path boundaries, connection isolation, Host/Origin/auth checks, and clean EOF. CI runs on Node 22 and 24. No paid generation is performed in automated tests; a live account test is still required to establish provider billing, regional access and visual quality.
MIT license. Original examples remain in example-files/.
Available Tools
7 toolsgenerateImageC
Generate an image from a text prompt using Google Imagen
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| numberOfImages | No | ||
| includeFullData | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It does not mention side effects, authentication needs, rate limits, or whether the image is stored or returned directly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise. However, it sacrifices valuable information for brevity. Not every sentence earns its place when it omits critical details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters, no output schema, and no annotations, the description is severely incomplete. It does not explain return values or parameter behavior, leaving the agent without enough context to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description adds no meaning beyond the parameter names and types in the input schema. It fails to explain prompt constraints, the number of images, or the includeFullData field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Generate'), the resource ('an image'), and the input ('from a text prompt using Google Imagen'). It effectively distinguishes from sibling tools focused on video generation or listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like getImage or listGeneratedImages. The description does not mention prerequisites, exclusions, or context for invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateVideoFromGeneratedImageC
Generate a video from a generated image (one-step process)
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| aspectRatio | No | 16:9 | |
| videoPrompt | No | ||
| autoDownload | No | ||
| enhancePrompt | No | ||
| negativePrompt | No | ||
| numberOfImages | No | ||
| numberOfVideos | No | ||
| durationSeconds | No | ||
| includeFullData | No | ||
| personGeneration | No | dont_allow |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, and the description only mentions 'one-step process'. Important behavioral traits such as cost, async behavior, output format, and side effects are not disclosed. The description fails to compensate for the lack of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is too brief given the tool's complexity (11 parameters). It is under-specified and does not effectively convey necessary information despite being concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal given the large parameter count and lack of output schema. It does not explain the overall workflow (e.g., whether the image is generated first or input is needed) or how the tool fits with siblings. Important context like return values and constraints is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters but does not. None of the 11 parameters are described in the description, and the schema itself lacks descriptions. The agent cannot understand what parameters like 'autoDownload' or 'personGeneration' do.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action (generate video) and the source (generated image). However, it does not differentiate from the sibling tool 'generateVideoFromImage', which could lead to confusion about when to use this tool vs that one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like generateVideoFromImage or generateVideoFromText. There is no explanation of prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateVideoFromImageC
Generate a video from an image
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | ||
| prompt | No | Generate a video from this image | |
| aspectRatio | No | 16:9 | |
| autoDownload | No | ||
| enhancePrompt | No | ||
| negativePrompt | No | ||
| numberOfVideos | No | ||
| durationSeconds | No | ||
| includeFullData | No | ||
| personGeneration | No | dont_allow |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only states the basic operation. It does not disclose any behavioral traits such as processing time, image format requirements, or potential limitations, leaving the agent without important context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is under-specified for a tool with 10 parameters. It lacks structure and important details, making it minimally informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (10 parameters, no output schema, no annotations), the description is grossly incomplete. It fails to provide context for the output, parameter behavior, or usage scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero description coverage for parameters, and the description adds no meaning about any of the 10 parameters. Parameters like 'prompt', 'aspectRatio', 'durationSeconds' are not explained, so the agent cannot infer their purpose from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates a video from an image, using a specific verb and resource. However, it does not differentiate from sibling tool 'generateVideoFromGeneratedImage', which has a similar purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you have an image to convert to a video, but provides no explicit guidance on when to use this tool versus alternatives like 'generateVideoFromText' or 'generateVideoFromGeneratedImage'. No exclusions or context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateVideoFromTextC
Generate a video from a text prompt
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| aspectRatio | No | 16:9 | |
| autoDownload | No | ||
| enhancePrompt | No | ||
| negativePrompt | No | ||
| numberOfVideos | No | ||
| durationSeconds | No | ||
| includeFullData | No | ||
| personGeneration | No | dont_allow |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full burden for behavioral transparency. It only states the basic function without disclosing any behavioral traits such as generation time, cost, success/failure handling, or return format. This is insufficient for responsible agent use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise at one sentence, but given the tool's complexity (9 parameters), it is too brief. While front-loaded with the core action, it sacrifices necessary detail for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Considering the tool's complexity (9 parameters, no output schema, no annotations), the description is critically incomplete. It lacks details on output, behavior, parameter effects, and usage context, making it inadequate for correct agent selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description adds no value by explaining parameters like enhancePrompt, autoDownload, includeFullData, or personGeneration. Agents must infer meaning from names alone, risking misuse. At minimum, key parameters should be explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Generate a video from a text prompt' clearly states the verb (Generate) and resource (video from text), distinguishing it from siblings like generateImage (image from text) and generateVideoFromImage (video from existing image). It's specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool vs alternatives. It doesn't mention scenarios, prerequisites, or exclusions, leaving the agent without context for appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
getImageB
Get a specific image by ID
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| includeFullData | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose behavioral traits beyond the basic operation. It lacks details about error handling, response format, or side effects. With no annotations, the description carries the full burden but provides minimal insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence. Every word is necessary and there is no wasted information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, no parameter descriptions in the schema, and no annotations, the description is insufficient. It does not explain return values, error states, or the effect of optional parameters, leaving significant gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description only hints at the 'id' parameter through the phrase 'by ID', but does not explain the 'includeFullData' parameter. Schema description coverage is 0%, so the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Get a specific image by ID' clearly states the verb (Get), resource (specific image), and method (by ID). It effectively distinguishes from sibling tools like generateImage, which are generative in nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For example, it does not clarify that this tool is for retrieving existing images, while siblings handle generation or listing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
listGeneratedImagesB
List all generated images
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description does not disclose read-only nature, auth requirements, or any side effects. Simply states 'list all', offering no behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise at one sentence, no wasted words. Could be slightly improved by front-loading key info, but currently adequate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple list tool with no params or output schema; however, lacks any mention of pagination or ordering, which may be necessary for completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has zero parameters; schema coverage is 100%. Description adds no parameter info, but none is needed. Baseline 4 is appropriate for no-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states verb (list) and resource (generated images), distinguishing from siblings like getImage (single) and generateImage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives; no mention of filtering, pagination, or context such as if the list is all images or scoped to a user.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
listGeneratedVideosB
List all generated videos
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description merely repeats the tool name without disclosing any behavioral traits (e.g., pagination, scope of 'all', or read-only nature).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, efficient sentence with no extraneous words. Appropriate for a simple list tool with no parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool, the description is acceptable but lacks details like result format, sorting, or limits. Could be improved with a note on scope or output structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist (schema is empty), so the description has no parameter semantics to add. Baseline 4 applies as the schema already covers 100% of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'List all generated videos' clearly states the verb (List) and resource (generated videos), distinguishing it from siblings like generateImage, generateVideoFromText, and listGeneratedImages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as getImage or generateVideoFromText. The description does not specify context, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- First observed
generateImage - First observed
generateVideoFromGeneratedImage - First observed
generateVideoFromImage - First observed
generateVideoFromText - First observed
getImage - First observed
listGeneratedImages - First observed
listGeneratedVideos
TDQS
Scored across 7 tools
Most tools have distinct purposes, but generateVideoFromGeneratedImage and generateVideoFromImage overlap in functionality; one is described as 'one-step process' but the distinction is not immediately clear from names alone.
Tool names follow a verb_noun pattern, but there is inconsistency: 'generateImage' uses a direct object while video tools use 'from' prepositional phrases. Also, 'getImage' differs from 'listGeneratedImages/listGeneratedVideos' in tense.
With 7 tools, the set is well-scoped for an image/video generation server, covering creation and listing without being overly numerous.
The tool surface includes generation and listing but lacks retrieval of individual videos (no getVideo), update, and delete operations, leaving notable gaps in lifecycle management.
Maintenance
Related MCP Connectors
MCP server for Google Veo AI video generation
MCP server for Kling AI video generation
MCP server for Hailuo (MiniMax) AI video generation
MCP server for Wan AI video generation
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn async video generation MCP server with multi-provider support. Currently in skeleton phase with stub implementations, it will eventually enable video generation through providers like Veo 3.1, Grok Imagine Video, and Sora 2 Pro.MIT
- FlicenseNot gradedqualityDmaintenanceMCP server for programmatic video generation. Send a prompt, get an MP4.-
- AlicenseNot gradedqualityAmaintenanceMCP server for generating images and videos using Google Gemini and VEO models, with support for multiple AI models and credential modes.1Apache 2.0
- FlicenseBqualityDmaintenanceA production-ready MCP server that enables Claude and other LLMs to generate images and videos using Google's Gemini AI models (Gemini 2.0 Flash and Veo 2.0).32-