image-gen3-google-mcp-server
Generates high-quality images using Google's Imagen 3.0 model via the Gemini API, with configurable number of images and automatic file management.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@image-gen3-google-mcp-serverGenerate an image of a futuristic city at night"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini Imagen 3.0 MCP Server
A professional Model Context Protocol (MCP) server implementation that harnesses Google's Imagen 3.0 model through the Gemini API for high-quality image generation. Built with TypeScript and designed for seamless integration with Claude Desktop and other MCP-compatible hosts.
🌟 Features
Leverage Google's state-of-the-art Imagen 3.0 model via Gemini API
Generate up to 4 high-quality images per request
Automatic file management with intelligent naming
HTML preview generation with file:// protocol support
Built on MCP protocol for AI agent compatibility
TypeScript implementation with robust error handling
Related MCP server: Nano Banana MCP Server
🚀 Quick Start
Prerequisites
Node.js 18 or higher
Google Gemini API key
Claude Desktop or another MCP-compatible host
Installation
Clone the repository:
git clone https://github.com/yourusername/gemini-imagen-mcp-server.git
cd gemini-imagen-mcp-serverInstall dependencies:
npm installBuild the TypeScript code:
npm run build⚙️ Configuration
Configure Claude Desktop by adding to
claude_desktop_config.json:
{
"mcpServers": {
"gemini-image-gen": {
"command": "node",
"args": ["./build/index.js"],
"cwd": "<path-to-project-directory>",
"env": {
"GEMINI_API_KEY": "your-gemini-api-key"
}
}
}
}Replace placeholders:
<path-to-project-directory>: Your project pathyour-gemini-api-key: Your Gemini API key
🛠️ Available Tools
1. generate_images
Generates images using Google's Imagen 3.0 model.
Parameters:
prompt(required): Text description of the image to generatenumberOfImages(optional): Number of images (1-4, default: 1)
File Management:
Images are automatically saved in
G:\image-gen3-google-mcp-server\imagesFilenames follow the pattern:
{sanitized-prompt}-{timestamp}-{index}.pngTimestamps ensure unique filenames
Prompts are sanitized for safe filesystem usage
Example:
Generate an image of a futuristic city at night2. create_image_html
Creates HTML preview tags for generated images.
Parameters:
imagePaths(required): Array of image file pathswidth(optional): Image width in pixels (default: 512)height(optional): Image height in pixels (default: 512)
Returns HTML tags with absolute file:// URLs for local viewing.
Example:
Create HTML tags for the generated images with width=400🔧 Development
# Install dependencies
npm install
# Build TypeScript
npm run build
# Run tests (when available)
npm test🤝 Contributing
Contributions are welcome! Please feel free to submit a Pull Request. For major changes:
Fork the repository
Create your feature branch (
git checkout -b feature/AmazingFeature)Commit your changes (
git commit -m 'Add some AmazingFeature')Push to the branch (
git push origin feature/AmazingFeature)Open a Pull Request
📝 Error Handling
The server implements two main error codes:
tool_not_found(1): When the requested tool is not availableexecution_error(2): When image generation or HTML creation fails
📄 License
MIT License - see the LICENSE file for details.
✨ Author
Falah G. Salieh
Copyright © 2025
GitHub: @yourgithubhandle
Email: your.email@example.com
🙏 Acknowledgments
Google Gemini API and Imagen 3.0 model
Model Context Protocol (MCP) by Anthropic
Claude Desktop team for MCP host implementation
📌 Tags
#MCP #Gemini #Imagen3 #AI #ImageGeneration #TypeScript #NodeJS #GoogleAI #ClaudeDesktop
Made with ❤️ by Falah G. Salieh
Available Tools
2 toolscreate_image_htmlB
Create HTML img tags from image file paths
| Name | Required | Description | Default |
|---|---|---|---|
| imagePaths | Yes | Array of image file paths | |
| width | No | Image width in pixels | |
| height | No | Image height in pixels |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should disclose behavioral traits but does not. It omits details about error handling, output format, or side effects, leaving the agent uninformed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no waste. It is concise but could be slightly more informative without sacrificing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 3 parameters and no output schema, the description gives the basic idea but lacks details such as tag format, handling of missing files, or default attributes. It is adequate but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds no extra parameter information beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it creates HTML img tags from file paths. It distinguishes from sibling tool generate_images by focusing on HTML output rather than image generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling generate_images. No context on prerequisites or appropriate scenarios is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imagesB
Generate images using Google Gemini AI
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate | |
| numberOfImages | No | Number of images to generate (1-4) | |
| outputDir | No | Directory to save generated images | G:\image-gen3-google-mcp-server\images |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It only states that images are generated, but omits key traits such as whether the tool saves files to disk (though implied by 'outputDir' in schema), cost implications, rate limits, or error behavior. The agent learns little beyond the basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence achieving conciseness, but it lacks structure and front-loading of critical constraints. It does not highlight the image count range (1-4) or output directory default, which are important for usage but are buried in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain what the tool returns (e.g., file paths, base64 data). It does not. Given the sibling tool suggesting alternative image-related functionality, more context on output format and typical use cases would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters. The description adds the context of using 'Google Gemini AI', which is not in the schema. However, it does not clarify or enrich the parameter meanings beyond the schema descriptions. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('generate images') and specifies the service ('Google Gemini AI'). It effectively distinguishes from the sibling 'create_image_html' tool by indicating this generates actual images rather than HTML.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling 'create_image_html'. There are no conditions, prerequisites, or exclusions mentioned, leaving the agent to infer context from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
create_image_html - First observed
generate_images
TDQS
Each tool targets a completely different operation: generate_images creates new images via AI, while create_image_html outputs HTML for existing image files. There is no ambiguity or overlap between them.
Both tools follow a clear verb_noun pattern in snake_case. 'generate_images' and 'create_image_html' use consistent styling and predictable naming conventions.
Only 2 tools for an image generation service is borderline. While the core functionality (generation and HTML output) is covered, the small number suggests a limited scope that may require additional tools for a complete workflow.
The tool set is notably incomplete: it lacks tools for listing, deleting, or managing generated images, and there is no way to configure generation parameters beyond what might be in the description. Users are left with a generation-and-output loop without lifecycle management.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
AI image generation across 5 quality tiers (SDXL to Gemini 3 Pro), 50 free credits on signup.
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
LLM chat, text tools, image generation, editing and batch image jobs
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceGenerate high-quality images from text descriptions using Google's Imagen 4.0 models with multiple quality variants, flexible aspect ratios, and local file storage.3-
- AlicenseNot gradedqualityDmaintenanceEnables generating, editing, and manipulating images using Google Gemini Flash 2.5 through natural language prompts. Supports text-to-image generation, image editing, multi-image composition, and batch processing with direct file management.724MIT
- FlicenseBqualityDmaintenanceGenerates high-quality images using Google's Imagen 3.0 model via the Gemini API with support for up to four images per request. It provides automated file management and creates HTML previews for seamless image viewing within MCP-compatible hosts.24-
- AlicenseBqualityDmaintenanceEnables image generation using Google Gemini models like Gemini 2.0 Flash and Imagen 3.0 with support for custom aspect ratios and negative prompts. It also allows users to list and manage generated images stored in local directories.225MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/falahgs/image-gen3-google-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server