gpt-image-2-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OPENAI_ORG_ID | No | Forwarded as the organization header. | |
| OPENAI_API_KEY | Yes | Your OpenAI API key required for authentication. | |
| OPENAI_BASE_URL | No | Override base URL for proxies or enterprise routes. | |
| OPENAI_PROJECT_ID | No | Forwarded as the project header. | |
| GPT_IMAGE_2_JOB_MAX | No | LRU cap on finished image jobs kept for polling (0 = no cap). | 20 |
| GPT_IMAGE_2_MCP_DEBUG | No | Set to '1' to emit verbose debug logs on stderr. | 0 |
| GPT_IMAGE_2_JOB_TTL_MS | No | Age before a finished job becomes un-pollable (0 = never expire). Running jobs are never expired. | 1800000 |
| GPT_IMAGE_2_OUTPUT_DIR | No | Global default directory where images are saved. Absolute paths used as-is, relative resolved from CWD. | |
| GPT_IMAGE_2_SESSION_MAX | No | Max concurrent in-memory edit sessions, LRU-evicted beyond this (0 = no cap). | 20 |
| GPT_IMAGE_2_ASYNC_AFTER_MS | No | Sync window before generate_image / edit_image background themselves (<=0 disables backgrounding entirely). | 20000 |
| GPT_IMAGE_2_SESSION_TTL_MS | No | Idle TTL before an edit session is swept (0 = never expire). | 3600000 |
| OPENAI_RESPONSES_EDIT_MODEL | No | Host model used by the Responses-API fallback edit route. | gpt-4.1-mini |
| OPENAI_FORCE_RESPONSES_EDITS | No | Set to '1' to pin edits to the Responses-API fallback route instead of /v1/images/edits. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_imageA | Generate an image from a text prompt using OpenAI's image model family (models: "gpt-image-2.5-sunburst" (default), "gpt-image-2.5-flare", "gpt-image-2"). The image is written to disk and also returned inline so you can see it. These models handle photoreal, illustrations, infographics, multilingual text (incl. CJK), and complex structured visuals. For a transparent logo pass background="transparent" with output_format="png" — verified against this family; the origin decides, so check applied.background. Sizes accept presets or any custom "WxH" where edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta. Image work is slow, so this call hands off immediately: the first response carries a job_id — poll get_image_job until it reports state "completed" (the result arrives verbatim, images included). Set GPT_IMAGE_2_ASYNC_AFTER_MS to wait inline instead (below your host's tool-call timeout). |
| edit_imageA | Edit or compose images with OpenAI's image model family (models: "gpt-image-2.5-sunburst" (default), "gpt-image-2.5-flare", "gpt-image-2"). Give 1–8 input images plus a text prompt; optionally include a PNG mask whose transparent regions mark what to change (mask applies to the first image). Great for: swap backgrounds, retouch products, combine multiple reference images into one composition, maintain a character across scenes. These models always process inputs at high fidelity (no input_fidelity knob needed). The edited image is saved to disk and returned inline. Image work is slow, so this call hands off immediately: the first response carries a job_id — poll get_image_job until it reports state "completed". Set GPT_IMAGE_2_ASYNC_AFTER_MS to wait inline instead (below your host's tool-call timeout). |
| get_image_jobA | Poll a backgrounded gpt-image job: every image tool (generate_image, edit_image, start_edit_session, continue_edit_session) hands off immediately and returns a job_id rather than blocking past MCP client timeouts. Call this with that job_id until state becomes "completed" (files are already written to disk and returned inline, with the model, requested/applied settings, usage, and cost) or "failed". While still "running", wait a few seconds between polls. |
| list_image_jobsA | List the image jobs this server still remembers (newest first), with their state, prompt preview, and elapsed time. Useful after a reconnect or a context reset to recover a job_id instead of re-running — and paying for — a generation you already started. Running jobs are never expired; finished ones are kept up to GPT_IMAGE_2_JOB_MAX / GPT_IMAGE_2_JOB_TTL_MS and are lost on server restart. |
| start_edit_sessionA | Begin a stateful multi-turn edit session. Returns a session_id you then pass to continue_edit_session to iteratively refine the image (each turn uses the previous turn's output as the input). Use end_edit_session when done. The first turn hands off to a background job like every other image call: poll get_image_job for it, and the session_id arrives with its "completed" result. |
| continue_edit_sessionA | Apply another edit turn to an existing session. The previous turn's output image is used as the input. Use short, focused prompts like "make the sky more orange" or "add a small boat on the horizon"; include "keep everything else the same" to limit drift. Returns the new image and the updated session. Omitting size, quality, background, output_format, or model inherits what the session already uses (list_edit_sessions reports those settings). The turn hands off to a background job (poll get_image_job); while one is still running the session stays busy and further turns are refused until it lands. |
| end_edit_sessionA | Free an iterative-edit session. Safe to skip — sessions are in-memory only and are discarded on server restart — but calling this frees memory sooner and keeps list_edit_sessions tidy. |
| list_edit_sessionsA | List active iterative-edit sessions (in-memory only, discarded on server restart). Useful to recover a session_id after a client reconnect; each entry also reports whether a turn is still running in the background and which job to poll for it. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Each tool targets a distinct concern: one-shot generation, one-shot editing, multi-turn session lifecycle, and job polling/listing. The overlap between edit_image and continue_edit_session is clarified by the former being a single edit and the latter being an iterative turn within a session.
All tool names follow a consistent snake_case verb_noun pattern: generate_image, edit_image, get_image_job, list_image_jobs, start_edit_session, continue_edit_session, end_edit_session, list_edit_sessions. The verb choices clearly reflect the action for each resource.
Eight tools is well-scoped for the server's purpose: image generation and editing plus the supporting async-job and edit-session infrastructure. Each tool earns its place and there is no redundant surface.
The tool set covers the full workflow: generate and edit images, poll and list background jobs, and start/continue/list/end iterative edit sessions. No critical dead ends remain for the stated domain.