image-video-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GEMINI_API_KEY | Yes | Your Google Gemini API key | |
| OPENAI_API_KEY | Yes | Your OpenAI API key |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| openai_generate_imageA | Generate image(s) with OpenAI's Images API and save them to assets/images. Args: prompt: Text description of the image to generate. size: WIDTHxHEIGHT (e.g. "1024x1024", "1536x1024"). gpt-image-2 supports arbitrary sizes with width/height divisible by 16. quality: "low", "medium", or "high". n: Number of images (1-10 for gpt-image-2). model: Override the model id (default from OPENAI_IMAGE_MODEL / gpt-image-2). filename: Optional base filename (without extension) for a single image. Returns: A human-readable summary listing the saved file path(s). |
| gemini_generate_imageA | Generate an image with Google Gemini (Nano Banana) and save to assets/images. Args: prompt: Text description of the image to generate. model: Override the model id (default from GEMINI_IMAGE_MODEL / gemini-2.5-flash-image). Newer option: a Nano Banana 2 / Gemini 3.x Flash Image id if enabled on your key. filename: Optional base filename (without extension). Returns: A human-readable summary listing the saved file path(s). |
| gemini_generate_videoA | Generate a video with Google Veo (via Gemini API) and save to assets/videos. This is an async operation: the tool submits the job, polls until done, then downloads the result. It can take a few minutes. Args: prompt: Text description of the video to generate. model: Override the model id (default from GEMINI_VIDEO_MODEL / veo-3.1-generate-001). Use veo-3.1-fast-generate-preview for speed. negative_prompt: Optional description of what to avoid. aspect_ratio: "16:9" or "9:16". filename: Optional base filename (without extension). Returns: A human-readable summary with the saved video path. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The two image generation tools (openai_generate_image and gemini_generate_image) overlap in purpose, both generating an image from a prompt. The provider prefixes help distinguish them, but an agent may still be uncertain which to choose without explicit context. The video tool is clearly distinct.
All tool names follow the same [provider]_[verb]_[object] pattern with snake_case, making the set predictable and easy to navigate. The consistent use of 'generate' as the verb reinforces a clear convention.
Three tools is well-scoped for a media generation server covering image and video output. Each tool serves a distinct provider or modality, and the count is within the ideal range for a focused MCP.
The toolset covers image generation via OpenAI and Gemini, and video generation via Gemini, satisfying the core 'image-video' purpose. Minor gaps exist, such as no OpenAI video generation or image editing capabilities, but these are not essential given the apparent scope.