Skip to main content
Glama

GenImgMCP ๐ŸŽจ

Local MCP (Model Context Protocol) server for image generation and editing using the google/gemini-3.1-flash-image model via OpenRouter.

Designed specifically for coding agents (such as Antigravity, Claude Code, Cursor, and other MCP clients), saving images directly to the project's local file system and returning the file path and structured metadata in a lightweight manner (without cluttering the context window with Base64 strings).


โœจ Features

  • ๐Ÿ–ผ๏ธ Image Generation (generate_image): Create images from detailed text prompts.

  • ๐Ÿ–Œ๏ธ Image Editing and Variations (edit_image): Transform, refine, or add elements to existing images on disk.

  • ๐Ÿ’พ Automatic Local Saving: Saves image files directly to the project directory requested by the agent (or in ./generated_images/).

  • โšก Lightweight & Efficient Response: Returns only the absolute file path and metadata (size, format, aspect ratio), preserving the caller LLM's context tokens.

  • ๐Ÿ“ Aspect Ratio & Format Control: Supports multiple aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9) and file formats (png, webp, jpeg).

  • ๐Ÿ”„ Model Flexibility: Defaults to google/gemini-3.1-flash-image, with support for any multimodal model available on OpenRouter.


Related MCP server: openrouter-imgen-mcp

๐Ÿš€ Installation and Build

1. Prerequisites

2. Install Dependencies and Build

# In the project root:
npm install
npm run build

โš™๏ธ Configuration

You can configure the API key in one of three ways:

  1. Environment Variable: Set OPENROUTER_API_KEY in your environment.

  2. .env File: Create a .env file in the root of the project:

    OPENROUTER_API_KEY=sk-or-v1-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
    DEFAULT_IMAGE_MODEL=google/gemini-3.1-flash-image
    DEFAULT_OUTPUT_DIR=./generated_images
  3. Command Line Argument: Pass --api-key <your_key> in arguments when starting the server.


๐Ÿ› ๏ธ Available Tools

1. generate_image

Generates a new image based on a text prompt.

Parameter

Type

Required

Description

prompt

string

Yes

Detailed description of the image to generate.

output_path

string

No

File path where the image will be saved (e.g., ./assets/banner.png).

aspect_ratio

string

No

Aspect ratio of the image (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9). Default: 1:1.

output_format

string

No

File format (png, webp, jpeg). Default: png.

model

string

No

Model on OpenRouter (Default: google/gemini-3.1-flash-image).

2. edit_image

Modifies an existing image provided by local file path.

Parameter

Type

Required

Description

image_path

string

Yes

Source image path (e.g., ./assets/logo.png) or Data URL.

prompt

string

Yes

Modification instructions or elements to add.

output_path

string

No

Destination path for the edited image.

aspect_ratio

string

No

Desired aspect ratio for the resulting image.

output_format

string

No

Output file format (png, webp, jpeg).

model

string

No

Model on OpenRouter (Default: google/gemini-3.1-flash-image).


๐Ÿ”Œ How to Integrate with MCP Clients

Configuration in Claude Desktop (claude_desktop_config.json)

{
  "mcpServers": {
    "genimg": {
      "command": "node",
      "args": [
        "D:/MyProjs.Github/GenImgMCP/dist/index.js"
      ],
      "env": {
        "OPENROUTER_API_KEY": "sk-or-v1-your-key-here"
      }
    }
  }
}

Configuration in Antigravity / Cursor / MCP Config (mcp_config.json)

{
  "mcpServers": {
    "genimg": {
      "command": "node",
      "args": [
        "D:/MyProjs.Github/GenImgMCP/dist/index.js"
      ],
      "env": {
        "OPENROUTER_API_KEY": "sk-or-v1-your-key-here"
      }
    }
  }
}

๐Ÿงช Testing the Server Locally

You can run the server directly in the terminal to verify startup:

node dist/index.js

The server will start and wait for JSON-RPC messages on the stdio channel.


๐Ÿ“„ License

MIT

Available Tools

2 tools
edit_imageA

Edits or transforms an existing image from a local file and natural language instructions (add elements, change backgrounds, alter style, etc.).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOpenRouter model to use. Default: 'google/gemini-3.1-flash-image'.
promptYesInstructions on what to modify or add to the reference image.
image_pathYesLocal path to the base image (e.g., './assets/logo.png') or base64 Data URL.
output_pathNoFile path to save the resulting image. If omitted, automatically saves to the default folder.
aspect_ratioNoDesired aspect ratio for the edited image.
output_formatNoSaved file format ('png', 'webp', 'jpeg').

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does not mention that the tool invokes an external model (OpenRouter), that it produces and saves a new image file, or whether the original file is left untouched. These are relevant side effects for an image-editing tool and are not disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler; the operation, resource, and input type are front-loaded, and the examples are compactly contained in parentheses. Every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is adequately covered for a straightforward call because the schema documents all parameters. However, without annotations or an output schema, the description could usefully disclose that the operation uses a remote model and saves the resulting image to output_path or a default folder.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all six parameters are already documented in the schema. The description adds only illustrative examples of prompt instructions (add elements, change backgrounds, alter style), which adds minor context but does not go beyond the schema's parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('edits or transforms') and a clear resource ('existing image from a local file'), making the tool's function unmistakable. The phrase 'existing image' also implicitly distinguishes it from the sibling generate_image, which creates images rather than modifying an existing one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this tool is for modifying an existing image using natural-language instructions, which is enough to guide selection versus generate_image. It does not, however, explicitly name the sibling or state when to prefer one over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generates an image from a detailed text prompt using the Gemini 3.1 Flash Image model via OpenRouter. Saves the image locally and returns the absolute path and metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOpenRouter model to use. Default: 'google/gemini-3.1-flash-image'.
promptYesDetailed description of the image to create (subject, setting, style, lighting, composition).
output_pathNoLocal path (relative or absolute) where the generated image file should be saved (e.g., './assets/hero.png'). If omitted, automatically saves to the default folder.
aspect_ratioNoAspect ratio of the generated image. Default: '1:1'.
output_formatNoFormat of the saved image file. Default: 'png'.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose two key behaviors: it saves the image locally and returns the absolute path and metadata. It does not mention potential overwrite behavior, default folder location, network/API costs, authentication requirements, or rate limits, leaving some gaps for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short, information-dense clauses with no filler. It front-loads the core action, then immediately conveys the outcome and return value, which is exactly what an agent needs to know at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one required parameter, complete schema coverage, and no output schema, the description covers the essential operational facts: what it does, how it saves output, and what it returns. It is missing only an explicit differentiation from edit_image and finer detail about the returned metadata, but neither prevents correct selection or invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all five parameters, including enums for aspect_ratio and output_format. The description adds little parameter-level meaning beyond reinforcing that the prompt should be detailed and that the model is provided via OpenRouter, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Generates an image from a detailed text prompt'), identifies the model and service, and explains the resulting side effect (saving locally) and return value. It is easily distinguishable from the sibling edit_image by the word 'generates,' though it does not explicitly contrast itself with that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

When to use the tool is implied: call it when you need to create a new image from a text prompt. However, there is no explicit mention of the sibling edit_image, no when-not-to-use guidance, and no mention of prerequisites or cases where another tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observededit_image
    • First observedgenerate_image

TDQS

A3.9/5.0

Scored across 2 tools

Disambiguation5/5

generate_image and edit_image have clearly distinct purposes: one creates new images from a prompt, while the other transforms an existing local image. There is no meaningful overlap or ambiguity between the two tools.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern: generate_image and edit_image. The naming is predictable, clear, and uniform.

Tool Count3/5

Two tools is on the thin side for an image generation/editing server, even though both are core operations. The count is borderline but not unreasonable for a minimal focused toolset.

Completeness4/5

The server covers the two primary image operations: generation and editing. It lacks supporting operations like listing locally saved images or deleting them, but these are minor gaps that agents can work around using file paths.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers