Gemini Imagen 3.0 MCP Server
Provides access to Google's state-of-the-art Imagen 3.0 model for high-quality image generation from text prompts with automated file management.
Uses the Gemini API to harness the Imagen 3.0 model for generating up to four professional images per request and creating HTML previews for local viewing.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini Imagen 3.0 MCP Servergenerate 2 high-quality images of a cozy cabin in the woods during winter"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini Imagen 3.0 MCP Server
A professional Model Context Protocol (MCP) server implementation that harnesses Google's Imagen 3.0 model through the Gemini API for high-quality image generation. Built with TypeScript and designed for seamless integration with Claude Desktop and other MCP-compatible hosts.
🌟 Features
Leverage Google's state-of-the-art Imagen 3.0 model via Gemini API
Generate up to 4 high-quality images per request
Automatic file management with intelligent naming
HTML preview generation with file:// protocol support
Built on MCP protocol for AI agent compatibility
TypeScript implementation with robust error handling
Related MCP server: KOF Nano Banana MCP Server
🚀 Quick Start
Prerequisites
Node.js 18 or higher
Google Gemini API key
Claude Desktop or another MCP-compatible host
Installation
Clone the repository:
git clone https://github.com/yourusername/gemini-imagen-mcp-server.git
cd gemini-imagen-mcp-serverInstall dependencies:
npm installBuild the TypeScript code:
npm run build⚙️ Configuration
Configure Claude Desktop by adding to
claude_desktop_config.json:
{
"mcpServers": {
"gemini-image-gen": {
"command": "node",
"args": ["./build/index.js"],
"cwd": "<path-to-project-directory>",
"env": {
"GEMINI_API_KEY": "your-gemini-api-key"
}
}
}
}Replace placeholders:
<path-to-project-directory>: Your project pathyour-gemini-api-key: Your Gemini API key
🛠️ Available Tools
1. generate_images
Generates images using Google's Imagen 3.0 model.
Parameters:
prompt(required): Text description of the image to generatenumberOfImages(optional): Number of images (1-4, default: 1)
File Management:
Images are automatically saved in
G:\image-gen3-google-mcp-server\imagesFilenames follow the pattern:
{sanitized-prompt}-{timestamp}-{index}.pngTimestamps ensure unique filenames
Prompts are sanitized for safe filesystem usage
Example:
Generate an image of a futuristic city at night2. create_image_html
Creates HTML preview tags for generated images.
Parameters:
imagePaths(required): Array of image file pathswidth(optional): Image width in pixels (default: 512)height(optional): Image height in pixels (default: 512)
Returns HTML tags with absolute file:// URLs for local viewing.
Example:
Create HTML tags for the generated images with width=400🔧 Development
# Install dependencies
npm install
# Build TypeScript
npm run build
# Run tests (when available)
npm test🤝 Contributing
Contributions are welcome! Please feel free to submit a Pull Request. For major changes:
Fork the repository
Create your feature branch (
git checkout -b feature/AmazingFeature)Commit your changes (
git commit -m 'Add some AmazingFeature')Push to the branch (
git push origin feature/AmazingFeature)Open a Pull Request
📝 Error Handling
The server implements two main error codes:
tool_not_found(1): When the requested tool is not availableexecution_error(2): When image generation or HTML creation fails
📄 License
MIT License - see the LICENSE file for details.
✨ Author
Falah G. Salieh
Copyright © 2025
GitHub: @yourgithubhandle
Email: your.email@example.com
🙏 Acknowledgments
Google Gemini API and Imagen 3.0 model
Model Context Protocol (MCP) by Anthropic
Claude Desktop team for MCP host implementation
📌 Tags
#MCP #Gemini #Imagen3 #AI #ImageGeneration #TypeScript #NodeJS #GoogleAI #ClaudeDesktop
Made with ❤️ by Falah G. Salieh
Available Tools
2 toolscreate_image_htmlB
Create HTML img tags from image file paths with gallery view
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Image width in pixels | |
| height | No | Image height in pixels | |
| gallery | No | Whether to create a gallery view with CSS | |
| imagePaths | Yes | Array of image file paths |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only mentions 'gallery view' but does not specify the output format (e.g., HTML string, file write), side effects, or any prerequisites, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, succinct sentence that directly conveys the core function without unnecessary words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple and the schema covers all parameters, but the description omits the return type or whether it returns an HTML fragment or full document. Without an output schema, this missing context leaves the agent unsure about the tool's exact behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are documented in the schema. The description adds minimal extra meaning beyond the 'gallery' parameter by referencing 'gallery view,' but does not elaborate on details like the exact CSS behavior or default values. This aligns with the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates HTML img tags from image file paths, using a specific verb and resource. It also distinguishes itself from the sibling tool generate_images, which presumably generates images rather than HTML.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given on when to use this tool versus alternatives. The sibling tool generate_images is present but never mentioned in the description, so the agent has no basis for choosing between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imagesC
Generate images using Google Gemini AI
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate | |
| category | No | Optional category folder for organizing images | |
| numberOfImages | No | Number of images to generate (1-4) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It only states the action without explaining side effects, response format, rate limits, costs, or whether images are saved or returned. This is insufficient for a generation tool with side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no redundant phrasing. It is front-loaded and easy to parse, though it is possibly too brief for a tool that would benefit from usage guidance. Still, it earns its place as a clear one-line summary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (3 params, no output schema, no annotations), the description is too sparse. It does not clarify what the generated images look like, how they are returned, or how the category and numberOfImages parameters affect behavior. The sibling tool's existence also suggests more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters with descriptions, so the baseline is 3. The description does not add extra semantic meaning beyond the schema; it only mentions the model provider. Since schema coverage is 100%, no significant gap exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates images using Google Gemini AI, which is a specific verb-resource combination. However, it does not differentiate from the sibling tool create_image_html, so some ambiguity remains about when to choose one over the other.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus create_image_html or any alternatives. There is no mention of prerequisites, exclusions, or contexts where sibling tools would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
create_image_html - First observed
generate_images
TDQS
The two tools have completely distinct purposes: generate_images creates new images via AI, while create_image_html produces HTML img tags from existing file paths. There is no overlap or ambiguity between them.
Both tool names follow a verb_noun pattern (generate_images and create_image_html). The verbs 'generate' and 'create' are similar but not identical, and the nouns differ in structure (plural vs. compound), but the overall pattern is consistent and readable.
With only 2 tools, the server feels somewhat thin. While the scope is narrow (image generation and HTML formatting), this is borderline on the low end; a small utility set would benefit from at least one additional tool for managing or inspecting generated images.
The core workflow of generating images and then creating HTML for viewing is covered, but there are notable gaps: no way to list, delete, or manage previously generated images, and no tool to adjust model parameters beyond what might be embedded in generate_images. This limits the server to a single-generation flow.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Create and manage AI image and video generations through Quriov's fixed public MCP tools.
Generate images with any major model — one API key, one prepaid balance, one MCP.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceGenerate high-quality images from text descriptions using Google's Imagen 4.0 models with multiple quality variants, flexible aspect ratios, and local file storage.3-
- AlicenseAqualityCmaintenanceEnables image generation using Gemini native models, supporting both single prompts and batch processing via a file-based queue. It allows for detailed configuration of aspect ratios and models using YAML frontmatter across various MCP-enabled clients.313MIT
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- FlicenseBqualityDmaintenanceEnables high-quality image generation using Google's Imagen 3.0 model via the Gemini API, with support for multiple images per request and automatic file management and HTML preview generation.2-
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/falahgs/imagen-3.0-generate-google-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server