gemini-imggen
Provides token-optimized image generation and transformation using Google Gemini's API, supporting text-to-image and image-to-image modes, and returning file paths to avoid token limits.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-imggenGenerate a cute cat illustration in flat design style"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini Image Generation MCP Server
A token-optimized MCP server that enables Gemini image generation in MCP clients by returning file paths instead of base64 data.
Why This Exists
Existing Gemini image generation MCP servers fail in Claude Code with MCP tool response exceeded token limit errors. They return base64-encoded image data (~2.4M tokens per image), exceeding Claude Code's 25,000 token limit.
This implementation solves the problem by saving images to disk and returning only file paths (~20 tokens) — a 120,000× reduction in token usage.
Implementation | Response | Tokens | Result |
Existing servers | Base64 data | 2.4M | ❌ Error |
This server | File path | ~20 | ✅ Works |
Related MCP server: gemini-image-generator
Features
Token-optimized: Returns file paths only (~20 tokens vs 2.4M)
Two generation modes: Text-to-image and image-to-image transformation
Claude Code compatible: Works within 25,000 token limit
ISO 8601 UTC timestamps: Globally sortable filenames (
YYYYMMDDTHHMMSSZ.png)Lightweight: Minimal dependencies
Fast: uv-powered startup
Simple: No build step required
Requirements
Python 3.10+
uv - Modern Python package manager (10-100× faster than pip)
Gemini API key from Google AI Studio
Install uv
# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
# Homebrew
brew install uv
# Verify installation
uv --versionQuick Start
# 1. Clone and navigate
git clone https://github.com/couhie/mcp-gemini-imggen.git
cd mcp-gemini-imggen
# 2. Configure settings
cp .env.example .env
# Edit .env and set:
# GEMINI_API_KEY - Your API key from Google AI Studio
# OUTPUT_DIR - Directory for generated images (e.g., ~/Pictures/ai)
# Directory will be created automatically if it doesn't exist
# 3. Add to Claude Code
claude mcp add -s user gemini-imggen uv -- --directory $(pwd) run mcp-gemini-imggenConfiguration
Claude Code CLI (Recommended)
claude mcp add -s user gemini-imggen uv -- --directory /absolute/path/to/mcp-gemini-imggen run mcp-gemini-imggenManual Setup
Add to ~/.claude.json:
{
"mcpServers": {
"gemini-imggen": {
"type": "stdio",
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/mcp-gemini-imggen",
"run",
"mcp-gemini-imggen"
],
"env": {}
}
}
}Note: Use absolute paths, not ~ (e.g., /Users/yourname/dev/mcp-gemini-imggen)
Usage
Once configured, use the MCP tools in Claude Code:
Text-to-Image Generation
Generate a flat design style cute cat illustrationImage-to-Image Transformation
Transform /Users/name/Pictures/ai/20251015T120000Z.png: make the background blueNote: You must provide the file path to an existing image. Common use cases:
Modify previously generated images
Transform images already saved on your system
Chain transformations: generate → transform → transform again
The server will:
Generate/transform the image using Gemini 2.5 Flash
Save it to
$OUTPUT_DIR/YYYYMMDDTHHMMSSZ.png(ISO 8601 UTC format)Return only the file path (~20 tokens)
Claude Code will automatically display the generated image.
Technical Details
Token Optimization
Base64-encoded responses cause token explosion:
1536×1536 PNG ≈ 1.4MB → Base64 ≈ 1.9MB (33% overhead)
Token conversion: 1.9MB ÷ 4 chars/token ≈ 475,000 tokens
Multiple images (4×): ~1,900,000 tokens
JSON wrapper: +500,000 tokens
Total: ~2,400,000 tokens (exceeds 25,000 limit)
Solution: Return file path instead of data
# ❌ Existing: 2.4M tokens
{"type": "image", "data": "iVBORw0KGgo...", "mimeType": "image/png"}
# ✅ This server: ~20 tokens
[{"type": "text", "text": "/Users/name/Pictures/ai/20251015T120000Z.png"}]Troubleshooting
"uv: command not found"
Install uv first:
curl -LsSf https://astral.sh/uv/install.sh | sh"GEMINI_API_KEY environment variable is required"
Get your API key from Google AI Studio and add to .env
"OUTPUT_DIR environment variable is required"
Set your desired output directory in .env (e.g., OUTPUT_DIR=~/Pictures/ai). The directory will be created automatically if it doesn't exist.
Images not generating
Verify API key is valid at Google AI Studio
Check API quota limits
Verify OUTPUT_DIR path is valid (parent directories must be writable)
Contributing
Contributions are welcome! Please submit a Pull Request.
License
MIT License - see LICENSE for details.
Links
Available Tools
2 toolsgenerate_image_from_imageA
Transform or edit an existing image using Gemini 2.5 Flash. Returns only the file path.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Instruction for how to transform or edit the image | |
| input_image_path | Yes | Path to the input image file (supports ~/ expansion) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full disclosure burden. It usefully states that the tool 'returns only the file path', giving a clear output expectation. However, it does not clarify whether the input image is overwritten or a new file is created, nor does it mention input format or size constraints, leaving important behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the purpose and then the return type. Every word is useful, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two parameters and no output schema, the description covers the essential aspects: what it does and what it returns. Some behavioral details (e.g., side effects on the input file) are missing, but the tool is simple enough that the description is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage for both parameters, each having clear descriptions. The tool description adds no additional parameter-specific information, so it does not improve upon what the schema already provides. The baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('Transform or edit') and clearly identifies the resource ('existing image'), while also naming the model (Gemini 2.5 Flash). This distinguishes it from the sibling generate_image_from_text tool by emphasizing editing existing images rather than generating from text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'existing image' provides clear context for when to use this tool, implying it is for editing images you already have. However, it does not explicitly mention alternatives (e.g., generate_image_from_text) or state when not to use it, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_from_textA
Generate an image from a text prompt using Gemini 2.5 Flash. Returns only the file path.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the burden of behavioral disclosure. It does add a useful constraint: 'Returns only the file path,' which clarifies the output shape and avoids the assumption of returning inline image data. However, it does not disclose other traits such as persistence of the file, token/cost implications, or any rate limits, leaving gaps in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loads the primary action, and adds a critical return-value note. Every word earns its place—there is no padding or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description covers the essential operational facts: what it does, what model it uses, and what it returns. It does not explicitly contrast with the sibling, but the naming convention makes that clear. The only minor gap is that it does not mention whether the returned file path is temporary or persistent, but this is not critical for basic usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter fully with a description ('Text description of the image to generate'), so schema coverage is 100%. The tool description adds no further parameter-specific details beyond restating the idea of a text prompt, thus the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Generate'), resource ('an image'), and the required input modality ('from a text prompt'). It also names the exact model ('Gemini 2.5 Flash'), which adds useful specificity. The sibling tool is clearly differentiated by the explicit 'from text' versus the sibling's 'from image' orientation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for text-to-image generation but does not explicitly discuss when to choose it over generate_image_from_image. No alternatives or exclusions are mentioned, so usage context is only implied by the tool name and the phrase 'from a text prompt.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have clearly distinct purposes: one generates from a text prompt, the other transforms an existing image. There is no ambiguity or overlap in their functionality.
Both tool names follow the exact same pattern: 'generate_image_from_' followed by the source type ('text' or 'image'). This is perfectly consistent and predictable.
With only 2 tools, the server feels minimal, but the scope is narrowly defined as image generation, so the count is borderline appropriate. A few more tools (e.g., for variations or parameter presets) could enhance the set, but it's not excessive.
For an image generation server, the two tools cover the primary modes: text-to-image and image-to-image (editing/transformation). No obvious dead ends or missing core operations exist within this narrow domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Image processing for AI agents: resize, convert, compress, crop, and web-ready AI-generated images.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceGenerates images using Google's Gemini model, with options to save to file or return as base64 data URL.
- AlicenseNot gradedqualityDmaintenanceEnables image generation and transformation using Google Gemini AI, with automatic English translation for multilingual prompts and local saving.MIT
- FlicenseNot gradedqualityDmaintenanceEnables Claude to generate and edit images using Google Gemini AI.
- AlicenseAqualityDmaintenanceProvides image generation, modification, and analysis capabilities using Google's Gemini API, enabling AI-powered image operations through natural language.511MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/couhie/mcp-gemini-imggen'
If you have feedback or need assistance with the MCP directory API, please join our Discord server