image-gen-mcp
Generates images using OpenAI's gpt-image-2 model, allowing text-to-image generation and editing with reference images and masks, saving directly to specified paths.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@image-gen-mcpgenerate a 1024x1024 banner with a sunset and mountains, save as public/banner.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
image-gen-mcp
An MCP server that lets Claude Code generate real images to use in your projects (banners, icons, backgrounds, textures, etc.) with OpenAI gpt-image-2, and save them straight to a path you specify. It is meant for producing production assets — not for mocking up UI designs.
ภาษาไทย: ดู README.th.md
Tools
generate_image
Generate a new image from a text prompt and write it to a file.
param | values | default |
| text description of the image | — |
| where to save, e.g. | — |
|
|
|
|
|
|
|
| inferred from |
|
|
|
| 0-100 (jpeg/webp only) | — |
| 1-10 (if >1, saves |
|
| override the model |
|
edit_image
Edit / build on existing images. Pass 1-16 reference images, plus an optional mask for inpainting.
param | values |
| what to edit / create |
| 1-16 source (reference) image paths |
| where to save the result |
| (optional) mask for inpainting — transparent areas of the mask are what gets edited; must match the first image's dimensions |
+ |
Both tools return only { path, bytes, size, quality, ... } — never the raw image bytes (to avoid flooding the agent's context). A relative output_path is resolved against the calling project's CWD, and parent directories are created automatically.
Related MCP server: ImageForge MCP
Authentication
Two modes. IMAGE_GEN_AUTH_MODE picks between them (auto by default: an API key in the environment wins, otherwise a stored login is used).
api_key — OPENAI_API_KEY (recommended)
Calls api.openai.com/v1. Every parameter above works. Billed to your API account.
chatgpt — browser login
npx image-gen-mcp login # opens a browser; also `status` and `logout`Signs in with your ChatGPT account and generates on your ChatGPT plan quota instead of API credit. If you call a tool with no credential configured, the server asks the MCP client to open the same login URL (URL-mode elicitation) and retries the call once you're done — Claude Code supports this; claude -p does not, so use the CLI there.
Tokens are stored in ~/.image-gen-mcp/auth.json (0600) and refreshed automatically. Codex's own ~/.codex/auth.json is never read or written — sharing one rotating refresh token between two tools logs both out.
This mode is a reduced feature set. It goes through chatgpt.com/backend-api/codex, which accepts only prompt, model, quality, background, always returns PNG, and silently ignores everything else. Rather than hand back an asset that quietly isn't what was asked for, the tools reject the unsupported parameters:
|
| |
| yes | rejected — dimensions come from the prompt; say "landscape"/"portrait"/"square" in it |
| yes | rejected — one image per call |
| png / jpeg / webp | png only |
| yes | rejected |
| yes | rejected |
| yes | rejected |
| yes | yes |
| yes | accepted, but not reliably honored — the response's actual |
Caveats. OpenAI publishes no OAuth flow for third-party apps: this reuses the login its own Codex CLI performs, including that client's ID and its two allow-listed loopback ports (1455/1457). It is undocumented, unsupported, and can break without notice. The resulting token is not accepted by the Platform API (GET /v1/models with one answers 403 Missing scopes: api.model.read), which is why the two modes talk to different hosts. api_key remains the supported path.
Install
cd image-gen-mcp
npm installConnect to Claude Code
For api_key mode, provide the key via the server's env block in your MCP config (no .env file required) — in a project .mcp.json or a user-scoped config. For chatgpt mode, drop the env block and run npx image-gen-mcp login instead:
{
"mcpServers": {
"image-gen": {
"command": "node",
"args": ["/Users/palm/Desktop/projects/custom-mcp/image-gen-mcp/src/index.mjs"],
"env": { "OPENAI_API_KEY": "sk-..." }
}
}
}Or add it via the CLI:
claude mcp add image-gen -e OPENAI_API_KEY=sk-... -- node /Users/palm/Desktop/projects/custom-mcp/image-gen-mcp/src/index.mjsTest
npm run selftest # real in-memory MCP round-trip with a fake client — no key/network needed
OPENAI_API_KEY=sk-... npm run smoke # one real API call (quality low) that verifies a valid PNGPricing (approx., gpt-image-2, 1024x1024)
low ~$0.006 · medium ~$0.053 · high ~$0.211 per image — start at medium, use high only for final assets.
Available Tools
2 toolsedit_imageA
แก้ไข/ต่อยอดรูปที่มีอยู่ด้วย gpt-image-2: ใส่รูปต้นทางได้ 1-16 ไฟล์เป็น reference (เช่น แก้รูปเดิม, รวมสไตล์จากหลายรูป, ใส่โลโก้), และใส่ mask (mask_path) เพื่อ inpaint เฉพาะบางส่วนได้. ผลลัพธ์เซฟเป็นไฟล์ที่ output_path. เหมาะกับการปรับ asset ที่เจนไว้แล้วหรือทำให้ตรงแบรนด์.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | จำนวนรูป (default 1). ถ้า >1 จะเซฟเป็น name-1.ext, name-2.ext, ... | |
| size | No | auto | 1024x1024 (จตุรัส) | 1536x1024 (แนวนอน) | 1024x1536 (แนวตั้ง) | 2048x2048 | WxH (gpt-image-2 สูงสุด 3840px, ทวีคูณของ 16) | 1024x1024 |
| model | No | โมเดล (default gpt-image-2). override ได้ เช่น gpt-image-1.5 | gpt-image-2 |
| format | No | นามสกุลไฟล์ผลลัพธ์ (ไม่ใส่ = เดาจาก output_path, ไม่รู้ = png) | |
| prompt | Yes | บอกว่าจะให้แก้/สร้างอะไรจากรูปต้นทาง | |
| quality | No | คุณภาพ/ราคา: low (ร่าง, ถูกสุด) | medium (ค่าเริ่มต้น) | high (งาน final, แพงสุด) | auto | medium |
| mask_path | No | (ไม่บังคับ) path ไฟล์ mask สำหรับ inpaint — บริเวณโปร่งใสใน mask คือส่วนที่จะถูกแก้ ต้องขนาดเท่ารูปแรก | |
| background | No | transparent = พื้นหลังโปร่งใส (ใช้ได้เฉพาะ png/webp) | opaque | auto | auto |
| moderation | No | ระดับ moderation (default auto) | auto |
| compression | No | ระดับบีบอัด 0-100 (เฉพาะ jpeg/webp) | |
| image_paths | Yes | path รูปต้นทาง 1-16 ไฟล์ (reference images) | |
| output_path | Yes | ที่จะเซฟไฟล์ผลลัพธ์ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It explains the process (reference images, mask, output path) and model, but does not disclose potential destructive actions, auth requirements, rate limits, or side effects. Adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph with three sentences, reasonably concise. It front-loads the purpose. Could be slightly more terse but is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 12 parameters and no output schema. The description covers the main use case and some parameter details (mask size, number of images), but does not explain return values or behavior. Adequate for an experienced agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description paraphrases some parameters (reference images, mask) but adds little beyond the schema. No significant additional semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool edits/enhances existing images using gpt-image-2, with support for multiple reference images and mask inpainting. It distinguishes from the sibling generate_image tool, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes it's suitable for adjusting already-generated assets or making them on-brand, providing context for when to use. However, it does not explicitly state when not to use or mention alternatives, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
เจนรูปใหม่จากข้อความ (prompt) ด้วย OpenAI gpt-image-2 แล้วเซฟเป็นไฟล์ที่ output_path ที่สั่ง (สร้างโฟลเดอร์ให้อัตโนมัติ, path relative จะอิงโฟลเดอร์โปรเจกต์ปัจจุบัน). ใช้สำหรับผลิต asset จริงเพื่อเอาไปใช้ในงาน เช่น banner เว็บ, ไอคอน, พื้นหลัง, texture. เลือกคุณภาพได้ (quality) — low ร่างถูกสุด, high งาน final. คืนแค่ path + metadata ไม่คืนข้อมูลรูปดิบ.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | จำนวนรูป (default 1). ถ้า >1 จะเซฟเป็น name-1.ext, name-2.ext, ... | |
| size | No | auto | 1024x1024 (จตุรัส) | 1536x1024 (แนวนอน) | 1024x1536 (แนวตั้ง) | 2048x2048 | WxH (gpt-image-2 สูงสุด 3840px, ทวีคูณของ 16) | 1024x1024 |
| model | No | โมเดล (default gpt-image-2). override ได้ เช่น gpt-image-1.5 | gpt-image-2 |
| format | No | นามสกุลไฟล์ผลลัพธ์ (ไม่ใส่ = เดาจาก output_path, ไม่รู้ = png) | |
| prompt | Yes | คำอธิบายรูปที่ต้องการ (อังกฤษได้ผลดีสุด, ไทยก็ได้) | |
| quality | No | คุณภาพ/ราคา: low (ร่าง, ถูกสุด) | medium (ค่าเริ่มต้น) | high (งาน final, แพงสุด) | auto | medium |
| background | No | transparent = พื้นหลังโปร่งใส (ใช้ได้เฉพาะ png/webp) | opaque | auto | auto |
| moderation | No | ระดับ moderation (default auto) | auto |
| compression | No | ระดับบีบอัด 0-100 (เฉพาะ jpeg/webp) | |
| output_path | Yes | ที่จะเซฟไฟล์ เช่น public/images/banner.png หรือ /abs/path/icon.webp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that it returns only path+metadata, not raw data, and mentions auto-creation of folders. However, it does not address overwrite behavior or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, well-structured, and front-loaded with the primary purpose. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (10 parameters, no output schema), the description covers key behaviors: return format, parameter interactions, quality differences, and file handling. It provides sufficient context for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds significant value beyond the schema: explains how 'n' affects filenames, size options including auto, quality levels, format inference, background restrictions, and compression applicability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (generate), resource (images from prompt using gpt-image-2), and provides specific use cases (banners, icons, textures). It distinguishes from the sibling tool edit_image by focusing on creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context for when to use this tool (producing actual assets) but does not explicitly exclude alternatives or mention when not to use it. The sibling differentiation is implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- First observed
edit_image - First observed
generate_image
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: one for generating new images from prompts, the other for editing existing images with references and masks. No overlap in functionality.
Both tools follow the same verb_noun pattern (edit_image, generate_image), creating a clear and predictable naming convention.
With only 2 tools, the server feels minimal; typically 3-15 tools are expected for a well-scoped server, making this borderline underpowered.
The server covers the core workflows of image generation and editing. Minor gaps exist (e.g., no tool to list or delete images), but these are not essential for the stated purpose.
Maintenance
Related MCP Connectors
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
LLM chat, text tools, image generation, editing and batch image jobs
Generate and edit images and videos with imageat.
Image processing for AI agents: resize, convert, compress, crop, and web-ready AI-generated images.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceGenerates and edits images using OpenAI image models via MCP tools.MIT
- AlicenseAqualityBmaintenanceEnables image generation and editing through OpenAI's gpt-image-2 model, supporting text-to-image, reference-guided generation, and editing with local files or URLs.23882MIT
- AlicenseNot gradedqualityAmaintenanceEnables image generation and editing using OpenAI's GPT Image 2 model within a Claude Desktop Code project workspace, with security boundaries and no-overwrite file handling.MIT
- AlicenseAqualityBmaintenanceGenerates and edits images using OpenAI GPT Image or Google Gemini models, saving every result to disk and returning local file paths so AI assistants can continue working with the images. It enables prompt-based image creation, editing, inpainting, multi-image composition, and model listing through MCP tools.131Apache 2.0