GenImgMCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@GenImgMCPGenerate a 16:9 hero image of a futuristic city and save it to ./assets/hero.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
GenImgMCP ๐จ
Local MCP (Model Context Protocol) server for image generation and editing using the google/gemini-3.1-flash-image model via OpenRouter.
Designed specifically for coding agents (such as Antigravity, Claude Code, Cursor, and other MCP clients), saving images directly to the project's local file system and returning the file path and structured metadata in a lightweight manner (without cluttering the context window with Base64 strings).
โจ Features
๐ผ๏ธ Image Generation (
generate_image): Create images from detailed text prompts.๐๏ธ Image Editing and Variations (
edit_image): Transform, refine, or add elements to existing images on disk.๐พ Automatic Local Saving: Saves image files directly to the project directory requested by the agent (or in
./generated_images/).โก Lightweight & Efficient Response: Returns only the absolute file path and metadata (size, format, aspect ratio), preserving the caller LLM's context tokens.
๐ Aspect Ratio & Format Control: Supports multiple aspect ratios (
1:1,16:9,9:16,4:3,3:4,3:2,2:3,21:9) and file formats (png,webp,jpeg).๐ Model Flexibility: Defaults to
google/gemini-3.1-flash-image, with support for any multimodal model available on OpenRouter.
Related MCP server: openrouter-imgen-mcp
๐ Installation and Build
1. Prerequisites
Node.js 18 or higher
OpenRouter API Key
2. Install Dependencies and Build
# In the project root:
npm install
npm run buildโ๏ธ Configuration
You can configure the API key in one of three ways:
Environment Variable: Set
OPENROUTER_API_KEYin your environment..envFile: Create a.envfile in the root of the project:OPENROUTER_API_KEY=sk-or-v1-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx DEFAULT_IMAGE_MODEL=google/gemini-3.1-flash-image DEFAULT_OUTPUT_DIR=./generated_imagesCommand Line Argument: Pass
--api-key <your_key>in arguments when starting the server.
๐ ๏ธ Available Tools
1. generate_image
Generates a new image based on a text prompt.
Parameter | Type | Required | Description |
|
| Yes | Detailed description of the image to generate. |
|
| No | File path where the image will be saved (e.g., |
|
| No | Aspect ratio of the image ( |
|
| No | File format ( |
|
| No | Model on OpenRouter (Default: |
2. edit_image
Modifies an existing image provided by local file path.
Parameter | Type | Required | Description |
|
| Yes | Source image path (e.g., |
|
| Yes | Modification instructions or elements to add. |
|
| No | Destination path for the edited image. |
|
| No | Desired aspect ratio for the resulting image. |
|
| No | Output file format ( |
|
| No | Model on OpenRouter (Default: |
๐ How to Integrate with MCP Clients
Configuration in Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"genimg": {
"command": "node",
"args": [
"D:/MyProjs.Github/GenImgMCP/dist/index.js"
],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-your-key-here"
}
}
}
}Configuration in Antigravity / Cursor / MCP Config (mcp_config.json)
{
"mcpServers": {
"genimg": {
"command": "node",
"args": [
"D:/MyProjs.Github/GenImgMCP/dist/index.js"
],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-your-key-here"
}
}
}
}๐งช Testing the Server Locally
You can run the server directly in the terminal to verify startup:
node dist/index.jsThe server will start and wait for JSON-RPC messages on the stdio channel.
๐ License
MIT
Available Tools
2 toolsedit_imageA
Edits or transforms an existing image from a local file and natural language instructions (add elements, change backgrounds, alter style, etc.).
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | OpenRouter model to use. Default: 'google/gemini-3.1-flash-image'. | |
| prompt | Yes | Instructions on what to modify or add to the reference image. | |
| image_path | Yes | Local path to the base image (e.g., './assets/logo.png') or base64 Data URL. | |
| output_path | No | File path to save the resulting image. If omitted, automatically saves to the default folder. | |
| aspect_ratio | No | Desired aspect ratio for the edited image. | |
| output_format | No | Saved file format ('png', 'webp', 'jpeg'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does not mention that the tool invokes an external model (OpenRouter), that it produces and saves a new image file, or whether the original file is left untouched. These are relevant side effects for an image-editing tool and are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler; the operation, resource, and input type are front-loaded, and the examples are compactly contained in parentheses. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is adequately covered for a straightforward call because the schema documents all parameters. However, without annotations or an output schema, the description could usefully disclose that the operation uses a remote model and saves the resulting image to output_path or a default folder.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all six parameters are already documented in the schema. The description adds only illustrative examples of prompt instructions (add elements, change backgrounds, alter style), which adds minor context but does not go beyond the schema's parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('edits or transforms') and a clear resource ('existing image from a local file'), making the tool's function unmistakable. The phrase 'existing image' also implicitly distinguishes it from the sibling generate_image, which creates images rather than modifying an existing one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this tool is for modifying an existing image using natural-language instructions, which is enough to guide selection versus generate_image. It does not, however, explicitly name the sibling or state when to prefer one over the other.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generates an image from a detailed text prompt using the Gemini 3.1 Flash Image model via OpenRouter. Saves the image locally and returns the absolute path and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | OpenRouter model to use. Default: 'google/gemini-3.1-flash-image'. | |
| prompt | Yes | Detailed description of the image to create (subject, setting, style, lighting, composition). | |
| output_path | No | Local path (relative or absolute) where the generated image file should be saved (e.g., './assets/hero.png'). If omitted, automatically saves to the default folder. | |
| aspect_ratio | No | Aspect ratio of the generated image. Default: '1:1'. | |
| output_format | No | Format of the saved image file. Default: 'png'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose two key behaviors: it saves the image locally and returns the absolute path and metadata. It does not mention potential overwrite behavior, default folder location, network/API costs, authentication requirements, or rate limits, leaving some gaps for a tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, information-dense clauses with no filler. It front-loads the core action, then immediately conveys the outcome and return value, which is exactly what an agent needs to know at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one required parameter, complete schema coverage, and no output schema, the description covers the essential operational facts: what it does, how it saves output, and what it returns. It is missing only an explicit differentiation from edit_image and finer detail about the returned metadata, but neither prevents correct selection or invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all five parameters, including enums for aspect_ratio and output_format. The description adds little parameter-level meaning beyond reinforcing that the prompt should be detailed and that the model is provided via OpenRouter, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Generates an image from a detailed text prompt'), identifies the model and service, and explains the resulting side effect (saving locally) and return value. It is easily distinguishable from the sibling edit_image by the word 'generates,' though it does not explicitly contrast itself with that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
When to use the tool is implied: call it when you need to create a new image from a text prompt. However, there is no explicit mention of the sibling edit_image, no when-not-to-use guidance, and no mention of prerequisites or cases where another tool would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- First observed
edit_image - First observed
generate_image
TDQS
Scored across 2 tools
generate_image and edit_image have clearly distinct purposes: one creates new images from a prompt, while the other transforms an existing local image. There is no meaningful overlap or ambiguity between the two tools.
Both tools follow a consistent verb_noun pattern: generate_image and edit_image. The naming is predictable, clear, and uniform.
Two tools is on the thin side for an image generation/editing server, even though both are core operations. The count is borderline but not unreasonable for a minimal focused toolset.
The server covers the two primary image operations: generation and editing. It lacks supporting operations like listing locally saved images or deleting them, but these are minor gaps that agents can work around using file paths.
Maintenance
Related MCP Connectors
AI image + video generation for agents: --flag prompt DSL, async generate/poll, x402 pay-per-use.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables image generation via OpenRouter API, supporting models like Gemini 2.5 Flash Image Preview with options to save files locally.21Do What The F*ck You Want To Public
- AlicenseAqualityCmaintenanceGives AI assistants image generation and editing capabilities through OpenRouter, supporting multiple models, style presets, variations, and batch operations.546 npm2MIT
- AlicenseAqualityAmaintenanceEnables coding agents to generate and edit images using Gemini and OpenAI image models, saving files directly into the project with configurable providers, models, and security restrictions.345 npm1MIT
- AlicenseNot gradedqualityAmaintenanceEnables image generation and editing using OpenAI's GPT Image 2 model within a Claude Desktop Code project workspace, with security boundaries and no-overwrite file handling.MIT