image-video-mcp
This server provides tools for generating AI images and videos using OpenAI and Google models, saving outputs locally.
OpenAI Image Generation (
openai_generate_image): Generate 1–10 images from a text prompt using OpenAI Images API (default model:gpt-image-2). Customizesize(e.g.,1024x1024,1536x1024; dimensions divisible by 16),quality(low,medium,high), number of images, optional outputfilename, and model override via parameter or environment variable. Images are saved toassets/images/.Google Gemini Image Generation (
gemini_generate_image): Generate a single image from a text prompt using Gemini (default model:gemini-2.5-flash-image). Supports optionalfilenameand model override. Images saved toassets/images/.Google Veo Video Generation (
gemini_generate_video): Generate a video from a text prompt using Veo via the Gemini API (default model:veo-3.1-generate-001). Supports optionalnegative_prompt,aspect_ratio(16:9or9:16),filename, model override (e.g., faster preview model), and automatic asynchronous polling until completion. Videos saved toassets/videos/.Requirements: Requires user-provided OpenAI and Google Gemini API keys.
Provides image generation via Gemini's Nano Banana model and video generation via Veo through the Gemini API.
Provides image generation through OpenAI's Images API, supporting models such as gpt-image-2.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@image-video-mcpgenerate an image of a cat wearing a spacesuit"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
image-video-mcp
One MCP server that adds OpenAI image generation and Google Gemini image + video generation to Claude Code — and any MCP client. Bring your own API keys.
Tool | Provider | Default model |
| OpenAI Images API |
|
| Gemini (Nano Banana) |
|
| Veo via Gemini API |
|
Images save to assets/images/, videos to assets/videos/. Every model is
overridable with an env var, so it keeps working as models roll over.
Quick install (one command)
Run this — it installs everything, asks for your two API keys, and registers the tools in both Claude Code and the Claude desktop app:
bash <(curl -fsSL https://raw.githubusercontent.com/eracom-technologies/image-video-mcp/main/setup.sh)That's it. Restart the Claude desktop app when it finishes, and the image/video
tools are ready in both places. (Prefer to set keys inline? OPENAI_API_KEY=sk-... GEMINI_API_KEY=... bash <(curl -fsSL .../setup.sh).)
The manual steps below are only if you'd rather do it yourself.
Related MCP server: mcp-media-engine
Manual install
One-line, straight from git:
pip install "git+https://github.com/eracom-technologies/image-video-mcp.git"Or from a local clone:
git clone https://github.com/eracom-technologies/image-video-mcp.git
cd image-video-mcp && pip install .Both create the image-video-mcp command. Prefer isolation? pipx install .
or uv tool install ..
Register in Claude Code
claude mcp add-json image-video '{
"command": "image-video-mcp",
"args": [],
"env": { "OPENAI_API_KEY": "sk-...", "GEMINI_API_KEY": "..." }
}'
claude mcp list # confirm it shows "image-video"Or drop the included .mcp.json into a project folder (Claude Code auto-detects
it). Get keys at platform.openai.com and
aistudio.google.com/apikey.
Use
Ask Claude Code:
"Use gemini_generate_video: a slow drone shot over a misty forest, 16:9."
Verify
PYTHONPATH=src python scripts/smoke_test.py # offline: tools register
PYTHONPATH=src python scripts/smoke_test.py --live # real generations (uses credits)Full docs
See README-MCP.md for OS-specific config locations, all environment variables, model/deprecation notes, sharing, and troubleshooting.
License
MIT — see LICENSE.
Available Tools
3 toolsgemini_generate_imageA
Generate an image with Google Gemini (Nano Banana) and save to assets/images.
Args: prompt: Text description of the image to generate. model: Override the model id (default from GEMINI_IMAGE_MODEL / gemini-2.5-flash-image). Newer option: a Nano Banana 2 / Gemini 3.x Flash Image id if enabled on your key. filename: Optional base filename (without extension).
Returns: A human-readable summary listing the saved file path(s).
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| filename | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden. It discloses that images are saved to assets/images and returns a human-readable summary listing file paths, which is useful. However, it omits details about overwrite behavior, file format, error handling, or authentication requirements, leaving gaps in behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main purpose and structured into Args/Returns sections. It is scannable and includes only necessary detail about model override, though the model note adds a bit of length. Overall, it is appropriately sized and well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no visible output schema, the description covers the primary action and return value but misses contextual details like failure handling, file extension behavior, and permission or API key prerequisites. It is adequate for the main happy path but not fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description effectively compensates by explaining all three parameters: prompt as text description, model with default from environment and newer option, and filename as optional base filename without extension. This adds meaningful context beyond the schema's bare types and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Generate an image with Google Gemini (Nano Banana) and save to assets/images', which is a specific verb+resource+destination. The mention of 'Google Gemini (Nano Banana)' clearly differentiates it from the sibling openai_generate_image, and 'image' distinguishes it from gemini_generate_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this tool is for Gemini-based image generation, but it does not explicitly name alternatives or provide when-not-to-use conditions. It does include usage guidance via the model override option and default from environment variables, giving clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_generate_videoA
Generate a video with Google Veo (via Gemini API) and save to assets/videos.
This is an async operation: the tool submits the job, polls until done, then downloads the result. It can take a few minutes.
Args: prompt: Text description of the video to generate. model: Override the model id (default from GEMINI_VIDEO_MODEL / veo-3.1-generate-001). Use veo-3.1-fast-generate-preview for speed. negative_prompt: Optional description of what to avoid. aspect_ratio: "16:9" or "9:16". filename: Optional base filename (without extension).
Returns: A human-readable summary with the saved video path.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| filename | No | ||
| aspect_ratio | No | 16:9 | |
| negative_prompt | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for disclosing behavior. It clearly explains the async lifecycle: submits the job, polls until done, downloads the result, and can take a few minutes. It also specifies the save location, which is valuable context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear opening sentence, a short async-behavior note, and an Args list. Every sentence provides needed information without padding; no redundancy with schema fields, and the structure makes it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex async tool with five parameters and no annotations, the description is remarkably complete. It covers input parameters, behavior, timing, save destination, and return value. The stated return summary is sufficient especially given that an output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates by explaining every parameter: prompt, model with default details, negative_prompt, aspect_ratio allowed values, and filename semantics. This adds meaning well beyond the schema's bare type/default information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Generate a video with Google Veo (via Gemini API) and save to assets/videos', naming the specific verb, resource, and destination. This distinguishes it from sibling image-generation tools by focusing explicitly on video generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for use, including that it is async and can take minutes, and offers model guidance ('Use veo-3.1-fast-generate-preview for speed'). It does not explicitly mention alternatives such as the image tools, but the video-specific purpose makes the primary use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
openai_generate_imageA
Generate image(s) with OpenAI's Images API and save them to assets/images.
Args: prompt: Text description of the image to generate. size: WIDTHxHEIGHT (e.g. "1024x1024", "1536x1024"). gpt-image-2 supports arbitrary sizes with width/height divisible by 16. quality: "low", "medium", or "high". n: Number of images (1-10 for gpt-image-2). model: Override the model id (default from OPENAI_IMAGE_MODEL / gpt-image-2). filename: Optional base filename (without extension) for a single image.
Returns: A human-readable summary listing the saved file path(s).
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| size | No | 1024x1024 | |
| model | No | ||
| prompt | Yes | ||
| quality | No | high | |
| filename | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently states the side effect of saving images to assets/images, the return summary format, and model-specific constraints (e.g., 'gpt-image-2 supports arbitrary sizes with width/height divisible by 16' and 'Number of images (1-10 for gpt-image-2)'). This goes beyond a basic summary, though it omits auth requirements or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a one-sentence summary, a bulleted Args list, and a Returns line. Every sentence provides useful information, and there is no filler. The format is front-loaded and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all six parameters and the return format, making it self-sufficient for invocation. Since an output schema exists, the return description is a bonus. It lacks only a note on prerequisites like API key setup, but that is not critical for the tool's core operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must fully compensate. It does so by explaining each parameter: prompt, size with examples and constraints, quality values, n range, model default, and filename usage. This adds essential meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with 'Generate image(s) with OpenAI's Images API and save them to assets/images,' which specifies the verb (generate), resource (images via OpenAI), and location (assets/images). This clearly distinguishes the tool from its Gemini siblings by naming OpenAI explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for OpenAI image generation but provides no explicit when-to-use or when-not-to-use guidance. It doesn't mention alternatives like 'use gemini_generate_image for Google's model' or exclusions for video generation. The context is implied through the OpenAI-specific details but lacks direct comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two image generation tools (openai_generate_image and gemini_generate_image) overlap in purpose, both generating an image from a prompt. The provider prefixes help distinguish them, but an agent may still be uncertain which to choose without explicit context. The video tool is clearly distinct.
All tool names follow the same [provider]_[verb]_[object] pattern with snake_case, making the set predictable and easy to navigate. The consistent use of 'generate' as the verb reinforces a clear convention.
Three tools is well-scoped for a media generation server covering image and video output. Each tool serves a distinct provider or modality, and the count is within the ideal range for a focused MCP.
The toolset covers image generation via OpenAI and Gemini, and video generation via Gemini, satisfying the core 'image-video' purpose. Minor gaps exist, such as no OpenAI video generation or image editing capabilities, but these are not essential given the apparent scope.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
MCP server for Google Veo AI video generation
MCP server for OpenAI Sora AI video generation
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Related MCP Servers
- AlicenseAqualityDmaintenanceAn MCP server that enables Claude Code to generate images and short videos using Google's Gemini API (Imagen 4 for stills, Veo 3.1 for video). It integrates directly into development workflows, allowing AI assistants to create visual assets and reference them in code without context switching.61MIT
- AlicenseAqualityBmaintenanceMCP server for AI-powered image, audio, and video generation, enabling media creation directly from Claude, Cursor, and other MCP clients.1164MIT
- AlicenseBqualityDmaintenanceMCP server for generating and editing images using OpenAI, and creating videos using OpenAI Sora and Google Veo. Enables fetching media from URLs or disk with smart output placement.14289MIT
- FlicenseBqualityDmaintenanceA production-ready MCP server that enables Claude and other LLMs to generate images and videos using Google's Gemini AI models (Gemini 2.0 Flash and Veo 2.0).32
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/eracom-technologies/image-video-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server