Gemini Image Generator MCP Server

# Gemini Image Generator MCP Server
Generate and transform images using Google's Gemini AI through the Model Context Protocol (MCP).
[](./LICENSE)
[](https://www.python.org/downloads/)
## Features
- **Text-to-Image Generation** - Create images from natural language prompts
- **Image Transformation** - Modify existing images with text descriptions
- **Automatic Filename Generation** - Smart naming based on prompts
- **Multi-Language Support** - Automatic prompt translation to English
## Installation
Get a free API key from [Google AI Studio](https://aistudio.google.com/apikey).
```bash
git clone https://github.com/jonchun/gemini-image-mcp.git
cd gemini-image-mcp
# Using uv (recommended)
uv venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
uv pip install -e .
```
### Configuration
<details>
<summary><b>Claude Desktop</b></summary>
Add to `claude_desktop_config.json`:
```json
{
"mcpServers": {
"gemini-image-mcp": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/gemini-image-mcp",
"run",
"gemini-image-mcp"
],
"env": {
"GEMINI_API_KEY": "your-api-key-here",
"DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
}
}
}
}
```
</details>
<details>
<summary><b>OpenCode</b></summary>
```json
{
"gemini-image-mcp": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/gemini-image-mcp",
"run",
"gemini-image-mcp"
],
"env": {
"GEMINI_API_KEY": "your-api-key-here",
"DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
}
}
}
```
</details>
<details>
<summary><b>Smithery</b></summary>
Install from [smithery.ai](https://smithery.ai) - search for "gemini-image-mcp".
</details>
## Usage
### Generate Images
```
Generate a photorealistic sunset over mountains with purple sky
```

```
Create a British Shorthair silver tabby kitten playing with a ball of yarn
```

### Transform Images
**Prompt:** `Add beautiful vibrant aurora borealis (northern lights) dancing across the sky with green, purple, and blue colors`

---
**Prompt:** `Add soft natural sunlight streaming through a window, creating beautiful warm light rays and gentle shadows`

## Available Tools
### `generate_image_from_text`
Creates an image from a text description.
**Parameters:**
- `prompt` (required): Text description of the image
- `output_dir` (optional): Directory to save the image
- `model` (optional): Gemini model to use (defaults to `GEMINI_MODEL` environment variable)
**Returns:** Path to the saved image file
### `transform_image_from_file`
Transforms an existing image based on a text prompt.
**Parameters:**
- `image_file_path` (required): Path to the source image
- `prompt` (required): Description of the transformation
- `output_dir` (optional): Directory to save the image
- `model` (optional): Gemini model to use (defaults to `GEMINI_MODEL` environment variable)
**Returns:** Path to the transformed image file
### `transform_image_from_encoded`
Transforms a base64-encoded image.
**Parameters:**
- `encoded_image` (required): Base64 data URL (`data:image/[format];base64,[data]`)
- `prompt` (required): Description of the transformation
- `output_dir` (optional): Directory to save the image
- `model` (optional): Gemini model to use (defaults to `GEMINI_MODEL` environment variable)
**Returns:** Path to the transformed image file
## Configuration
| Variable | Required | Default | Description |
| --------------------------- | -------- | ------------------------------------------- | --------------------- |
| `GEMINI_API_KEY` | Yes | - | Your Gemini API key |
| `DEFAULT_OUTPUT_IMAGE_PATH` | No | Current directory | Default save location |
| `GEMINI_MODEL` | No | `gemini-2.5-flash-image` | Model to use |
| `GEMINI_BASE_URL` | No | `https://generativelanguage.googleapis.com` | API base URL |
## Development
Test the server locally:
```bash
fastmcp dev src/gemini_image_mcp/server.py
```
Opens MCP Inspector at `http://localhost:5173/`
## License
MIT
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: text-to-image generation versus image transformation from two different input sources (encoded data or file path). The input types prevent confusion, and descriptions explicitly state the expected arguments.
All tool names follow the same snake_case verb_noun_preposition pattern (generate_image_from_text, transform_image_from_encoded, transform_image_from_file). The convention is consistent and predictable.
Three tools are well-scoped for an image generation and transformation server; each tool covers a distinct input method without redundancy. No tool feels superfluous or missing from a minimal set.
The server covers the core lifecycle: generating an image from text and transforming existing images from both encoded data and file paths. Minor gaps exist (e.g., no batch generation or targeted editing), but the primary workflows are supported.