Skip to main content
Glama
jonchun

Gemini Image Generator MCP Server

by jonchun

Cover

Gemini Image Generator MCP Server

Generate and transform images using Google's Gemini AI through the Model Context Protocol (MCP).

MIT licensed Python 3.11+

Features

  • Text-to-Image Generation - Create images from natural language prompts

  • Image Transformation - Modify existing images with text descriptions

  • Automatic Filename Generation - Smart naming based on prompts

  • Multi-Language Support - Automatic prompt translation to English

Related MCP server: Gemini Image MCP

Installation

Get a free API key from Google AI Studio.

git clone https://github.com/jonchun/gemini-image-mcp.git
cd gemini-image-mcp

# Using uv (recommended)
uv venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
uv pip install -e .

Configuration

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "gemini-image-mcp": {
      "command": "uv",
      "args": [
        "--directory",
        "/ABSOLUTE/PATH/TO/gemini-image-mcp",
        "run",
        "gemini-image-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "your-api-key-here",
        "DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
      }
    }
  }
}
{
  "gemini-image-mcp": {
    "command": "uv",
    "args": [
      "--directory",
      "/ABSOLUTE/PATH/TO/gemini-image-mcp",
      "run",
      "gemini-image-mcp"
    ],
    "env": {
      "GEMINI_API_KEY": "your-api-key-here",
      "DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
    }
  }
}

Install from smithery.ai - search for "gemini-image-mcp".

Usage

Generate Images

Generate a photorealistic sunset over mountains with purple sky

Sunset example

Create a British Shorthair silver tabby kitten playing with a ball of yarn

Kitten example

Transform Images

Prompt: Add beautiful vibrant aurora borealis (northern lights) dancing across the sky with green, purple, and blue colors

Sunset with aurora


Prompt: Add soft natural sunlight streaming through a window, creating beautiful warm light rays and gentle shadows

Kitten with sunbeams

Available Tools

generate_image_from_text

Creates an image from a text description.

Parameters:

  • prompt (required): Text description of the image

  • output_dir (optional): Directory to save the image

  • model (optional): Gemini model to use (defaults to GEMINI_MODEL environment variable)

Returns: Path to the saved image file

transform_image_from_file

Transforms an existing image based on a text prompt.

Parameters:

  • image_file_path (required): Path to the source image

  • prompt (required): Description of the transformation

  • output_dir (optional): Directory to save the image

  • model (optional): Gemini model to use (defaults to GEMINI_MODEL environment variable)

Returns: Path to the transformed image file

transform_image_from_encoded

Transforms a base64-encoded image.

Parameters:

  • encoded_image (required): Base64 data URL (data:image/[format];base64,[data])

  • prompt (required): Description of the transformation

  • output_dir (optional): Directory to save the image

  • model (optional): Gemini model to use (defaults to GEMINI_MODEL environment variable)

Returns: Path to the transformed image file

Configuration

Variable

Required

Default

Description

GEMINI_API_KEY

Yes

-

Your Gemini API key

DEFAULT_OUTPUT_IMAGE_PATH

No

Current directory

Default save location

GEMINI_MODEL

No

gemini-2.5-flash-image

Model to use

GEMINI_BASE_URL

No

https://generativelanguage.googleapis.com

API base URL

Development

Test the server locally:

fastmcp dev src/gemini_image_mcp/server.py

Opens MCP Inspector at http://localhost:5173/

License

MIT

Available Tools

3 tools
generate_image_from_textA

Generate an image from a text prompt using Gemini.

Args: prompt: Text description of the desired image. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.

Returns: List containing ImageContent with the generated image, and optionally TextContent with file path if saved.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
promptYes
output_dirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses that the image is returned in the response and only saved to disk if output_dir is given, and that model falls back to the GEMINI_MODEL environment variable. It stops short of covering auth requirements, rate limits, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose, then organized Args/Returns sections. Efficient overall, though the Returns block partially duplicates the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All three parameters are explained, the optional persistence behavior is confirmed, and an output schema exists so return details need not be restated. Missing only selection guidance relative to the transform siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters and it largely does: prompt, output_dir (with its disk-vs-response consequence), and model (with its env-var fallback) each get meaningful explanation beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Generate an image from a text prompt using Gemini.' This clearly distinguishes it from the sibling transform_image_* tools, which transform existing images rather than generating new ones, though it never explicitly names those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is offered. The agent must infer that this is the generation path versus the transformation siblings purely from the verb. The output_dir default behavior is described, but that is parameter semantics rather than selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_encodedA

Transform a base64-encoded image using Gemini.

Args: encoded_image: Base64 data URL (data:image/[format];base64,[data]). prompt: Text description of desired transformation. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.

Returns: List containing ImageContent with the transformed image, and optionally TextContent with file path if saved.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
promptYes
output_dirNo
encoded_imageYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose some behavior: the output is returned in the response and optionally saved to disk depending on output_dir, and it reveals the GEMINI_MODEL env fallback for the model. However, it omits whether the call consumes Gemini quota, what happens on invalid base64, image size/format limits, and whether disk writes create or overwrite files. The return-format disclosure partially compensates but key mutation/side-effect info is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The Args/Returns structure is front-loaded with the purpose sentence, and each parameter line adds real information. It is somewhat verbose due to the Args/Returns scaffolding, and the model/ctx lines could be tighter, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not detail return values, though it still helpfully notes the return shape. For a 4-param tool with no annotations it covers purpose and all parameters well, but it leaves out API-key/auth requirements and failure/limit behavior that an agent might need to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate, and it does: it documents all four parameters. It specifies encoded_image as a data URL with the exact format, prompt as the desired transformation text, output_dir's conditional save behavior when omitted, and model's GEMINI_MODEL env fallback. This is exactly the compensation the low schema coverage demands.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Transform a base64-encoded image using Gemini.' It clearly distinguishes from the file-based sibling 'transform_image_from_file' via 'from_encoded'/base64 and from 'generate_image_from_text' by being a transformation of an existing image rather than generation. It doesn't explicitly name the siblings, but the base64 encoding distinction is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'from_encoded' scope (use it with base64 data URLs rather than local files), but the description never states when to prefer this over transform_image_from_file or generate_image_from_text. The data URL format note gives partial context. No explicit when-not or alternative tool routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_fileA

Transform an image file using Gemini.

Args: image_file_path: Path to the source image file. prompt: Text description of desired transformation. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.

Returns: List containing ImageContent with the transformed image, and optionally TextContent with file path if saved.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
promptYes
output_dirNo
image_file_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the image is only returned in the response unless output_dir is supplied, and the model falls back to the GEMINI_MODEL environment variable. It omits error behavior, cost/latency, and whether an existing file at the target path is overwritten, but this is solid disclosure for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The docstring is front-loaded with the purpose in one line, then structured Args/Returns sections. Slightly verbose in formatting but every parameter line carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and 0% schema coverage, the description covers inputs and return shape; an output schema exists so return details are partly redundant but helpful. The main gaps, sibling routing guidance and error/format handling, are the only omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all four parameters with semantics: image_file_path as source path, prompt as transformation text, output_dir's save-vs-return consequence, and model's env-var fallback. It does not specify accepted image formats, path constraints, or whether output_dir must exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transform an image file using Gemini') that clearly distinguishes it from transform_image_from_encoded (which takes encoded input) and generate_image_from_text (which creates rather than transforms). Differentiation comes implicitly from the name and the input-concept rather than an explicit sibling callout, so it stops just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over its siblings, e.g. 'use this when you have a path on disk; use transform_image_from_encoded when you have base64 data.' The agent must infer the routing decision entirely from the title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedgenerate_image_from_text
    • First observedtransform_image_from_encoded
    • First observedtransform_image_from_file

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: text-to-image generation versus image transformation from two different input sources (encoded data or file path). The input types prevent confusion, and descriptions explicitly state the expected arguments.

Naming Consistency5/5

All tool names follow the same snake_case verb_noun_preposition pattern (generate_image_from_text, transform_image_from_encoded, transform_image_from_file). The convention is consistent and predictable.

Tool Count5/5

Three tools are well-scoped for an image generation and transformation server; each tool covers a distinct input method without redundancy. No tool feels superfluous or missing from a minimal set.

Completeness4/5

The server covers the core lifecycle: generating an image from text and transforming existing images from both encoded data and file paths. Minor gaps exist (e.g., no batch generation or targeted editing), but the primary workflows are supported.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers