Skip to main content
Glama

Gemini Image Generation MCP Server

A token-optimized MCP server that enables Gemini image generation in MCP clients by returning file paths instead of base64 data.

Python uv License: MIT

Why This Exists

Existing Gemini image generation MCP servers fail in Claude Code with MCP tool response exceeded token limit errors. They return base64-encoded image data (~2.4M tokens per image), exceeding Claude Code's 25,000 token limit.

This implementation solves the problem by saving images to disk and returning only file paths (~20 tokens) — a 120,000× reduction in token usage.

Implementation

Response

Tokens

Result

Existing servers

Base64 data

2.4M

❌ Error

This server

File path

~20

✅ Works

Related MCP server: gemini-image-generator

Features

  • Token-optimized: Returns file paths only (~20 tokens vs 2.4M)

  • Two generation modes: Text-to-image and image-to-image transformation

  • Claude Code compatible: Works within 25,000 token limit

  • ISO 8601 UTC timestamps: Globally sortable filenames (YYYYMMDDTHHMMSSZ.png)

  • Lightweight: Minimal dependencies

  • Fast: uv-powered startup

  • Simple: No build step required

Requirements

  • Python 3.10+

  • uv - Modern Python package manager (10-100× faster than pip)

  • Gemini API key from Google AI Studio

Install uv

# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"

# Homebrew
brew install uv

# Verify installation
uv --version

Quick Start

# 1. Clone and navigate
git clone https://github.com/couhie/mcp-gemini-imggen.git
cd mcp-gemini-imggen

# 2. Configure settings
cp .env.example .env
# Edit .env and set:
#   GEMINI_API_KEY - Your API key from Google AI Studio
#   OUTPUT_DIR - Directory for generated images (e.g., ~/Pictures/ai)
#                Directory will be created automatically if it doesn't exist

# 3. Add to Claude Code
claude mcp add -s user gemini-imggen uv -- --directory $(pwd) run mcp-gemini-imggen

Configuration

claude mcp add -s user gemini-imggen uv -- --directory /absolute/path/to/mcp-gemini-imggen run mcp-gemini-imggen

Manual Setup

Add to ~/.claude.json:

{
  "mcpServers": {
    "gemini-imggen": {
      "type": "stdio",
      "command": "uv",
      "args": [
        "--directory",
        "/absolute/path/to/mcp-gemini-imggen",
        "run",
        "mcp-gemini-imggen"
      ],
      "env": {}
    }
  }
}

Note: Use absolute paths, not ~ (e.g., /Users/yourname/dev/mcp-gemini-imggen)

Usage

Once configured, use the MCP tools in Claude Code:

Text-to-Image Generation

Generate a flat design style cute cat illustration

Image-to-Image Transformation

Transform /Users/name/Pictures/ai/20251015T120000Z.png: make the background blue

Note: You must provide the file path to an existing image. Common use cases:

  • Modify previously generated images

  • Transform images already saved on your system

  • Chain transformations: generate → transform → transform again

The server will:

  1. Generate/transform the image using Gemini 2.5 Flash

  2. Save it to $OUTPUT_DIR/YYYYMMDDTHHMMSSZ.png (ISO 8601 UTC format)

  3. Return only the file path (~20 tokens)

Claude Code will automatically display the generated image.

Technical Details

Token Optimization

Base64-encoded responses cause token explosion:

  1. 1536×1536 PNG ≈ 1.4MB → Base64 ≈ 1.9MB (33% overhead)

  2. Token conversion: 1.9MB ÷ 4 chars/token ≈ 475,000 tokens

  3. Multiple images (4×): ~1,900,000 tokens

  4. JSON wrapper: +500,000 tokens

  5. Total: ~2,400,000 tokens (exceeds 25,000 limit)

Solution: Return file path instead of data

# ❌ Existing: 2.4M tokens
{"type": "image", "data": "iVBORw0KGgo...", "mimeType": "image/png"}

# ✅ This server: ~20 tokens
[{"type": "text", "text": "/Users/name/Pictures/ai/20251015T120000Z.png"}]

Troubleshooting

"uv: command not found"

Install uv first:

curl -LsSf https://astral.sh/uv/install.sh | sh

"GEMINI_API_KEY environment variable is required"

Get your API key from Google AI Studio and add to .env

"OUTPUT_DIR environment variable is required"

Set your desired output directory in .env (e.g., OUTPUT_DIR=~/Pictures/ai). The directory will be created automatically if it doesn't exist.

Images not generating

  • Verify API key is valid at Google AI Studio

  • Check API quota limits

  • Verify OUTPUT_DIR path is valid (parent directories must be writable)

Contributing

Contributions are welcome! Please submit a Pull Request.

License

MIT License - see LICENSE for details.

Available Tools

2 tools
generate_image_from_imageA

Transform or edit an existing image using Gemini 2.5 Flash. Returns only the file path.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesInstruction for how to transform or edit the image
input_image_pathYesPath to the input image file (supports ~/ expansion)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full disclosure burden. It usefully states that the tool 'returns only the file path', giving a clear output expectation. However, it does not clarify whether the input image is overwritten or a new file is created, nor does it mention input format or size constraints, leaving important behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the purpose and then the return type. Every word is useful, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only two parameters and no output schema, the description covers the essential aspects: what it does and what it returns. Some behavioral details (e.g., side effects on the input file) are missing, but the tool is simple enough that the description is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage for both parameters, each having clear descriptions. The tool description adds no additional parameter-specific information, so it does not improve upon what the schema already provides. The baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Transform or edit') and clearly identifies the resource ('existing image'), while also naming the model (Gemini 2.5 Flash). This distinguishes it from the sibling generate_image_from_text tool by emphasizing editing existing images rather than generating from text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing image' provides clear context for when to use this tool, implying it is for editing images you already have. However, it does not explicitly mention alternatives (e.g., generate_image_from_text) or state when not to use it, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_image_from_textA

Generate an image from a text prompt using Gemini 2.5 Flash. Returns only the file path.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText description of the image to generate

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the burden of behavioral disclosure. It does add a useful constraint: 'Returns only the file path,' which clarifies the output shape and avoids the assumption of returning inline image data. However, it does not disclose other traits such as persistence of the file, token/cost implications, or any rate limits, leaving gaps in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loads the primary action, and adds a critical return-value note. Every word earns its place—there is no padding or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description covers the essential operational facts: what it does, what model it uses, and what it returns. It does not explicitly contrast with the sibling, but the naming convention makes that clear. The only minor gap is that it does not mention whether the returned file path is temporary or persistent, but this is not critical for basic usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single parameter fully with a description ('Text description of the image to generate'), so schema coverage is 100%. The tool description adds no further parameter-specific details beyond restating the idea of a text prompt, thus the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('Generate'), resource ('an image'), and the required input modality ('from a text prompt'). It also names the exact model ('Gemini 2.5 Flash'), which adds useful specificity. The sibling tool is clearly differentiated by the explicit 'from text' versus the sibling's 'from image' orientation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for text-to-image generation but does not explicitly discuss when to choose it over generate_image_from_image. No alternatives or exclusions are mentioned, so usage context is only implied by the tool name and the phrase 'from a text prompt.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: one generates from a text prompt, the other transforms an existing image. There is no ambiguity or overlap in their functionality.

Naming Consistency5/5

Both tool names follow the exact same pattern: 'generate_image_from_' followed by the source type ('text' or 'image'). This is perfectly consistent and predictable.

Tool Count3/5

With only 2 tools, the server feels minimal, but the scope is narrowly defined as image generation, so the count is borderline appropriate. A few more tools (e.g., for variations or parameter presets) could enhance the set, but it's not excessive.

Completeness5/5

For an image generation server, the two tools cover the primary modes: text-to-image and image-to-image (editing/transformation). No obvious dead ends or missing core operations exist within this narrow domain.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/couhie/mcp-gemini-imggen'

If you have feedback or need assistance with the MCP directory API, please join our Discord server