Skip to main content
Glama

gpt-image-2-mcp

An MCP server that exposes OpenAI's gpt-image-2 family (default model gpt-image-2.5-sunburst) to any MCP client — Claude Desktop, Claude Code, Cursor, MCP Inspector, etc.

Eight tools:

Tool

What it does

generate_image

text → image

edit_image

1–8 reference images (+ optional mask) → image

get_image_job

poll a backgrounded generate/edit job by job_id

list_image_jobs

recover a job_id after a reconnect or context reset

start_edit_session

begin an iterative multi-turn edit

continue_edit_session

apply another refinement turn — previous output becomes the new input

end_edit_session

release a session

list_edit_sessions

show active sessions

Every generated image is saved to disk and returned inline so the calling model sees it.

Image work is slow — tens of seconds, sometimes minutes — so every image tool returns { job_id, state: "running" } straight away and you collect the result from get_image_job. See Background jobs.

Batch work, or clients that cannot run a stdio server? The same repository ships a queue-backed remote service with an HTTP API and a remote MCP endpoint (/mcp), built for thousands of images at a time — see Remote service.

Requirements

  • Node.js ≥ 20 for the stdio server (the published gpt-image-2-mcp package)

  • Node.js ≥ 24 for the remote service and for this repository's test suite: the SQLite driver is Node's built-in node:sqlite, which older versions do not ship (the service image is node:24 for the same reason)

  • An OpenAI API key on an org with gpt-image-2 access (Organization Verification may be required)

Related MCP server: gpt-image-2-mcp

Install

Nothing to build. The server runs straight from npm over stdio — there is no port to open, no URL to host, nothing to keep running in the background. The client you configure below starts it on demand. Node.js ≥ 20 has to be available to that client.

Add this to your MCP client's config (claude mcp add, Claude Desktop, Cursor, VS Code, … — per-client details below):

{
  "mcpServers": {
    "gpt-image-2": {
      "command": "npx",
      "args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "OPENAI_BASE_URL": "https://your-gateway.example/v1"
      }
    }
  }
}
  • OPENAI_BASE_URL is optional — only for a proxy/enterprise route. Drop the line when you talk to api.openai.com directly.

  • Put the key in this env block. Whether the server would also see variables exported in your shell depends on the client: a GUI client (Claude Desktop) does not, a client you launched from a shell (Claude Code, MCP Inspector) usually does. The env block works in both cases.

  • On Windows, some clients need "command": "cmd", "args": ["/c", "npx", "-y", "@speed-fullbar/gpt-image-2-mcp"] instead (npx is a shell script there).

  • The package is scoped (@speed-fullbar/gpt-image-2-mcp) because the bare name on npm belongs to the upstream project this one is forked from (see Credits); the executable it installs is gpt-image-2-mcp.

Check the install before wiring it in. This starts the server and sends one MCP initialize:

printf '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"check","version":"0"}}}\n' \
  | npx -y @speed-fullbar/gpt-image-2-mcp

Expect one JSON line containing "serverInfo":{"name":"gpt-image-2-mcp","version":"…"} (the version you just installed). That proves the process starts and speaks MCP — it does not prove the API key works: for that, call generate_image and poll get_image_job (see the worked example below).

From source (for development, or a pinned local checkout):

pnpm install
pnpm run build

This produces build/index.js, the server entry point. Point the client at "command": "node", "args": ["/absolute/path/to/gpt-image-2-mcp/build/index.js"].

Configure a client

Every client needs the same shape — a command that starts the server (the block above) and an env block holding OPENAI_API_KEY. The sections below only show where that block goes in each client; use npx -y @speed-fullbar/gpt-image-2-mcp (no checkout needed) or the node …/build/index.js form from a pinned checkout.

Claude Code

claude mcp add gpt-image-2 --env OPENAI_API_KEY=sk-... -- npx -y @speed-fullbar/gpt-image-2-mcp

Or add it to ~/.claude.json (or a project .mcp.json) with the JSON shape below.

Claude Desktop

Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "gpt-image-2": {
      "command": "npx",
      "args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Cursor and other stdio clients

Same shape — mcpServers (Cursor, Claude, most wrappers) or mcp.servers (VS Code):

{
  "mcpServers": {
    "gpt-image-2": {
      "command": "npx",
      "args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
      "env": { "OPENAI_API_KEY": "sk-..." }
    }
  }
}

Local checkout instead of npx

{
  "mcpServers": {
    "gpt-image-2": {
      "command": "node",
      "args": ["/absolute/path/to/gpt-image-2-mcp/build/index.js"],
      "env": { "OPENAI_API_KEY": "sk-..." }
    }
  }
}

MCP Inspector (interactive testing)

pnpm run inspect

Launches the official inspector UI pointed at your local build.

Environment variables

Var

Required

Purpose

OPENAI_API_KEY

✅

Auth

OPENAI_BASE_URL

Override for proxies / enterprise routes

OPENAI_ORG_ID

Forwarded as organization

OPENAI_PROJECT_ID

Forwarded as project

OPENAI_TIMEOUT_MS

Per-request timeout (default 600000 = 10 min, generous because image calls are slow). Raise it for a very slow proxy.

OPENAI_MAX_RETRIES

SDK retries for 429/5xx (default 2).

GPT_IMAGE_2_OUTPUT_DIR

Global default for where images are saved. Absolute paths used as-is, relative resolved from the server's working directory (process.cwd()), not yours.

GPT_IMAGE_2_ALLOW_UNSAFE_OUTPUT_DIR

Set to 1 to let output_dir / GPT_IMAGE_2_OUTPUT_DIR point into OS-sensitive directories (/etc, ~/.ssh, …), which are refused by default.

GPT_IMAGE_2_MCP_DEBUG

Set to 1 to emit verbose debug logs on stderr.

GPT_IMAGE_2_SESSION_MAX

Max concurrent in-memory edit sessions, LRU-evicted beyond this (default 20; 0 = no cap).

GPT_IMAGE_2_SESSION_TTL_MS

Idle TTL before an edit session is swept (default 3600000 = 1h; 0 = never expire).

GPT_IMAGE_2_ASYNC_AFTER_MS

How long generate_image / edit_image / the edit-session tools may block before handing off to a pollable job (default 1 = hand off immediately; a larger value waits inline, keep it below your MCP host's tool-call timeout; <= 0 never backgrounds). See Background jobs below.

GPT_IMAGE_2_JOB_MAX

LRU cap on finished image jobs kept for polling (default 20; 0 = no cap).

GPT_IMAGE_2_JOB_TTL_MS

Age before a finished job becomes un-pollable (default 1800000 = 30 min; 0 = never expire). Running jobs are never expired or evicted.

GPT_IMAGE_2_SANITIZATION

Cleaning applied to every returned image before it is written or sent inline: metadata-v2 (default — container cleaning plus a re-encode), metadata-v1 (container cleaning only, provider's pixels kept), or off. See What an image carries.

OPENAI_FORCE_RESPONSES_EDITS

Set to 1 to pin edits to the Responses-API fallback route instead of /v1/images/edits. See Edit routing below.

OPENAI_RESPONSES_EDIT_MODEL

Host model used by the Responses-API fallback edit route (default gpt-4.1-mini). See Edit routing below.

Where images go

Unless overridden, each tool writes to:

<OS config dir>/gpt-image-2-mcp/output/<project-name>-<hash>/
  • macOS/Linux: ~/.config/gpt-image-2-mcp/output/<project>-<hash>/

  • Windows: %APPDATA%\gpt-image-2-mcp\output\<project>-<hash>\

<project>-<hash> is derived from the git root (if any) or the server process's current working directory — each project gets its own folder so generations don't collide.

⚠️ When a client launches the server (npx, Claude Desktop, …) that working directory is the client's, not your shell's, so the folder name is not predictable. If you care where images land, set GPT_IMAGE_2_OUTPUT_DIR (or pass output_dir per call) to an absolute path.

Per-call override: pass output_dir: "/some/path" to any tool.

Filenames look like image-20260422-150301-a1b2c3.png. If you pass filename_prefix: "hero-banner", it becomes image-20260422-150301-a1b2c3-hero-banner.png.

What an image carries

Every image this server returns — the file it writes and the inline copy in the tool result — is cleaned first. Image origins attach provenance metadata to what they return: measured against two gateways in production, every output PNG carried a caBX chunk, a ~22 KB C2PA/Content Credentials manifest naming the tool that made it. Cleaning removes it before the bytes reach you:

Profile

What it does

Pixels

metadata-v2 (default)

container rewrite and a re-encode: the image is decoded and written again with this server's own deflate

byte-identical

metadata-v1

container rewrite only: provider metadata is dropped, the provider's compressed stream is kept verbatim

untouched

The container rewrite keeps only what a decoder needs — PNG IHDR/PLTE/tRNS/IDAT/IEND, JPEG coding segments, WebP VP8/VP8L/ALPH/VP8X — and drops everything else: C2PA/Content Credentials, EXIF, XMP, ICC, text chunks, and any bytes appended after the end of the file. metadata-v2 goes further and re-encodes, which also removes the encoder fingerprint that survives in the compressed bytes; on a real 1.5 MP output that produced a 6-9% smaller file in ~1 s with AE = 0 against the original, i.e. not one pixel differs.

A shape that cannot be re-encoded faithfully (16-bit, indexed, interlaced, above the pixel ceiling) is delivered under metadata-v1 and says so. Each entry of structuredContent.images[] carries a sanitization receipt — { status, policy, source_bytes, removed, reencoded } — plus a note in the summary when something could not be cleaned:

{ "file_path": "…/image-….png", "sanitization": { "status": "clean", "policy": "metadata-v2",
  "source_bytes": 1595439, "removed": ["caBX"], "reencoded": true } }

GPT_IMAGE_2_SANITIZATION=off writes the provider's bytes verbatim and records status: "disabled". An unrecognised value is not guessed at: the image is still written (it was already paid for) and the result says cleaning did not run. Two limits worth stating plainly — this is a container-level guarantee (it removes what an image says about itself and, under v2, the encoder fingerprint, but makes no claim about watermarks inside the pixels), and a file the profile cannot parse at all is still delivered, flagged skipped with the reason rather than thrown away.

The queue-backed service applies the same profiles to what it stores; see What a downloaded image carries in the service half of this README.

What the tools return

Image tools hand off first. generate_image, edit_image, start_edit_session, and continue_edit_session answer with a job hand-off:

{ "job_id": "img-1761149123-a1b2c3d4", "state": "running", "poll_hint": "…", … }

Poll get_image_job with that job_id until it reports state: "completed". The completed poll is the result those tools would have returned inline:

  1. An inline ImageContent block per generated image (so the LLM sees the image)

  2. A text summary: applied settings, file path, token usage, estimated cost

  3. structuredContent for programmatic consumers:

{
  "model": "gpt-image-2.5-sunburst",
  "prompt": "…",
  "requested": { "size": "auto", "quality": "auto", "n": 1, "format": "png" },
  "applied":   { "size": "1024x1024", "quality": "high", "background": "opaque", "output_format": "png" },
  "images": [ { "file_path": "…", "filename": "…", "size_bytes": 123456, "mime_type": "image/png" } ],
  "usage":   { "input_tokens": …, "output_tokens": …, "total_tokens": …, "input_tokens_details": { … } },
  "cost_usd_estimated": 0.2112,
  "notes": []
}

requested is what you asked for, applied is what the origin actually used, and notes (always present, [] when empty) explains any difference. A job poll additionally carries job_id, state, tool, started_at, completed_at, elapsed_ms, error, and poll_hint (while running). Session tools also return session_id and turn.

Worked example: generate, poll, then edit the result

generate_image      prompt: "a red fox in a snowy pine forest, photorealistic"
  → { job_id: "img-…", state: "running" }

get_image_job       job_id: "img-…"
  → state: "running" (elapsed 6s)                       # poll again in a few seconds
get_image_job       job_id: "img-…"
  → state: "completed"
    images: [ { file_path: "/home/me/.config/gpt-image-2-mcp/output/my-app-1a2b3c/image-20260921-101500-a1b2c3.png" } ]

edit_image          prompt: "give the fox a small gold crown, keep everything else identical"
                    images: ["/home/me/.config/…/image-20260921-101500-a1b2c3.png"]
  → { job_id: "img-…", state: "running" }               # feed file_path straight back in

get_image_job       job_id: "img-…"
  → state: "completed", images: [ { file_path: "…-edited.png" } ]

Background jobs (the normal path)

Image work is slow — tens of seconds through a fast route, sometimes minutes — past the tool-call timeout of many MCP hosts. Every image tool (generate_image, edit_image, start_edit_session, continue_edit_session) therefore hands off immediately: the first response is { job_id, state: "running" } and the request continues in-process. Keep calling get_image_job with that job_id every few seconds:

  • state: "running" → keep polling

  • state: "completed" → content is identical to a synchronous success: inline images, summary text with file paths / usage / cost, and structuredContent carrying the model, requested/applied settings, usage, cost estimate, and route the call would have returned

  • state: "failed" → error carries the reason; content mirrors the error

GPT_IMAGE_2_ASYNC_AFTER_MS controls how long a call may block before that hand-off: 1 (the default, milliseconds) hands off immediately; a larger value waits inline for that long — keep it below your MCP host's tool-call timeout; 0 never backgrounds, so every call blocks until it finishes.

Job state lives in memory: finished jobs are kept up to GPT_IMAGE_2_JOB_MAX and expire after GPT_IMAGE_2_JOB_TTL_MS; a server restart drops all jobs. If you lose a job_id (reconnect, context reset), call list_image_jobs and pick the running/completed job back up instead of re-running it and paying twice.

Models

generate_image, edit_image, start_edit_session, and continue_edit_session accept an optional model argument:

  • gpt-image-2.5-sunburst — the default

  • gpt-image-2.5-flare

  • gpt-image-2

The 2.5 variants accept the same parameters (sizes, quality, formats). Omitting model keeps using gpt-image-2.5-sunburst; in an edit session, continue_edit_session inherits the model the session was started with unless you override it per turn. Token/cost estimates assume gpt-image-2 pricing.

Sizes

Default is auto (the model picks). You can pass:

  • A preset: 1024x1024, 1536x1024, 1024x1536

  • Any custom WxH where:

    • Both edges are multiples of 16

    • Max edge ≤ 3840px (outputs above 2K are beta)

    • Aspect ratio within 1:3 and 3:1

    • Total pixels between 655,360 and 8,294,400

Invalid sizes fail before the API call with a clear error — no wasted requests.

What actually comes back is the origin's call, not yours. Some upstreams normalize the controls: a gateway on a ChatGPT subscription (for example sub2api routing to its OAuth/Codex backend) ignores the requested size, quality, and format, and answers with its own — commonly 1254×1254 PNG for the 2.5 models, whatever you asked for. The tool reports what came back in applied, and adds a notes entry when it differs from requested. To have size/format honored, point OPENAI_BASE_URL at an API-key upstream (sub2api relays the fields unmodified on that path).

Transparent PNGs work via background: "transparent" (pair it with output_format: "png"). When the origin honors it, the response reports background: "transparent" and the file carries a real alpha channel (only the artwork is opaque).

The origin decides, though, and this gateway picks transparency from the prompt/content rather than your argument — verified at byte level on the same origin: a vector-logo prompt came back with 68% fully transparent pixels even when background: "opaque" was requested, and a scene prompt came back fully opaque (0 transparent pixels) even when "transparent" was requested. Always check applied.background and the file's alpha channel; a notes entry explains any disagreement. (For the other controls on this gateway, pnpm run smoke:matrix reports what your own origin honors — on the sub2api ChatGPT path size/quality/format/n are normalized away.)

Iterative editing example

start_edit_session    prompt: "A coastal lighthouse at dawn, photorealistic", images: ["./sketch.png"]
  → session_id: edit-1761149123-a1b2c3d4, turn 1, saved to …/session-…-turn1.png

continue_edit_session session_id: "edit-…-a1b2c3d4", prompt: "Make the sky more orange. Keep everything else the same."
  → turn 2

continue_edit_session session_id: "edit-…-a1b2c3d4", prompt: "Add a small boat on the horizon."
  → turn 3

end_edit_session      session_id: "edit-…-a1b2c3d4"

A turn hands off like any other slow call: start_edit_session and continue_edit_session answer with { job_id, state: "running" } and the turn finishes in the background. The session_id arrives with the job's "completed" result (for a start, the session only exists once its first turn has landed), and the session stays busy until then — continuing it early returns an error, and list_edit_sessions reports state: "running" plus the pending_job_id to poll. A turn is claimed the moment it starts, so two overlapping continue_edit_session calls can't edit the same input image twice.

Every follow-up turn inherits what the session already uses — model, size, quality, background, output format — unless that turn overrides it, and list_edit_sessions reports those settings for each session.

Sessions are in-memory only and discarded on server restart — this is intentional (keeps the server stateless on the wire) and mirrors the Gemini MCP pattern.

Image inputs for edit_image and start_edit_session

Accepts any mix of:

  • Absolute path: /Users/me/photo.png

  • Relative path: ./photo.png (resolved from CWD)

  • file:///Users/me/photo.png

  • https://example.com/photo.png (downloaded, size-capped)

  • data:image/png;base64,iVBOR…

Up to 8 images per call. Each ≤ 50MB. PNG/WEBP/JPG supported.

Pass the file_path from any earlier result straight back in images — that is how you iterate on an image, and how you combine several references into one composition. More than 8 images is refused before anything is uploaded (verified live: 8 accepted, 9 refused). continue_edit_session takes no images; it reuses the session's last output. The same summary is in the server's own initialize instructions, so a client sees it without reading this file.

Cost guardrails

The server ships no hard spending limits — you should watch your OpenAI usage dashboard. Each tool result includes an estimated cost in USD computed from the token usage returned by the API, plus an approximate pre-flight estimate logged to stderr.

Rough per-image cost at common sizes:

Quality

1024×1024

1024×1536 / 1536×1024

low

~$0.006

~$0.005

medium

~$0.053

~$0.041

high

~$0.211

~$0.165

Custom sizes scale with pixel count. Edit calls additionally tokenize input images at high fidelity — large reference images are expensive.

Edit routing

edit_image, start_edit_session, and continue_edit_session call POST /v1/images/edits directly. This is the canonical endpoint: it supports n > 1, masks, and returns accurate per-call token usage for cost estimation.

History: at launch (2026-04-21) the endpoint rejected gpt-image-2 (and gpt-image-1.5) with 400 Invalid value: 'gpt-image-2'. Value must be 'dall-e-2'. — an OpenAI-side bug. Versions ≤ 0.2.0 of this server therefore routed edits through the Responses API by default. OpenAI fixed the endpoint silently in early May 2026 (verified live 2026-06-11), and since 0.3.0 the direct endpoint is the default again.

The Responses-API workaround is kept as a fallback (src/utils/edit-via-responses.ts):

  • It engages automatically if the direct endpoint ever returns the launch-era 400 again (matched narrowly; the rejection is remembered for 10 minutes so only the first call in that window pays the failed attempt, then the direct endpoint is re-probed).

  • Set OPENAI_FORCE_RESPONSES_EDITS=1 to pin it explicitly.

  • The legacy OPENAI_USE_DIRECT_EDITS toggle from 0.2.0 is deprecated and ignored (its only meaningful setting was 1 — opt into the direct endpoint, which is now the default).

Fallback mechanics: input images are uploaded via the Files API (purpose: "vision"), a cheap host model (default gpt-4.1-mini, override with OPENAI_RESPONSES_EDIT_MODEL) is forced to invoke the image_generation tool, the base64 result is extracted, and uploaded files are deleted afterwards.

Fallback trade-offs versus the direct endpoint (only apply when the fallback is active — the tool result carries route: "responses" and a note when they do):

  • n > 1 is not supported — the Responses path returns one image per call.

  • Cost accounting undercounts — usage only reports the host chat model's text tokens; the image tool is billed separately (~$0.04–0.05 extra for a 1024×1536 medium edit).

  • Masks still work — uploaded and referenced via input_image_mask.file_id.

Remote service (HTTP API + remote MCP)

The same repository also ships a queue-backed service for batch work: it accepts generation/editing requests over HTTP and MCP, stores them in SQLite, and processes them with background workers against one or more image providers. The stdio server above is untouched — this is a second entry point (pnpm run service).

cp .env.example .env      # set SERVICE_TOKENS, KRILL_TOKEN, RUSTFS_* secrets
docker compose up -d --build
curl localhost:8787/v1/ready

Three containers: the service, Redis (BullMQ transport) and RustFS (S3-compatible storage for uploaded inputs and generated outputs). The service's SQLite database lives on the app-data volume; nothing else is required. All service variables are documented in .env.example.

Deploying on a single VM

docker-compose.prod.yml is the overlay for a machine that also runs other things (a local gateway stack, a reverse proxy on :80/:443). It publishes the app and rustfs on 127.0.0.1 only, mounts an image directory for input_paths, joins the provider stack's docker network, and caps memory and log growth.

git clone https://github.com/FullBars/gpt-image-2-mcp.git /root/gpt-image-service
cd /root/gpt-image-service
cp .env.example .env && chmod 600 .env     # then set the values below
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d --build
curl localhost:8788/v1/ready               # the app listens on 127.0.0.1:8788

Settings that matter on a shared host (all in .env):

Variable

Why

SERVICE_TOKENS

bearer token(s) for MCP and REST — long random, e.g. ops:$(openssl rand -hex 32)

SERVICE_S3_PUBLIC_ENDPOINT

where clients reach rustfs (the proxy's TLS URL, or http://<host>:9002); the signature covers the host, so it must match what clients send

SERVICE_IMAGES_DIR

host directory mounted at /data/images for input_paths (a 3000-image manifest of paths)

SERVICE_PROVIDER_NETWORK

docker network of a local gateway stack, so SERVICE_SUB2API_BASE_URL=http://sub2api:8080/v1 stays on the host

SERVICE_PUBLIC_URL

the address clients use (e.g. https://images.example.com/images), so MCP instructions and tool results carry absolute URLs instead of paths the agent must prefix

SERVICE_WORKER_CONCURRENCY

worker slots; keep it near the vCPU count on a small box (a synchronous provider holds a slot for a whole generation)

Reverse proxy (Caddy, host network — the app under a path prefix so an existing service can keep /v1/*, and storage on its own port with the path and Host header forwarded untouched so S3 signatures stay valid):

images.example.com {
	# MCP at /images/mcp, REST at /images/v1/...
	handle_path /images/* {
		# Compress the JSON/NDJSON API responses — a batch manifest is ~9x smaller
		# (27 KB -> 3 KB for 24 items) and it is the one thing here that is text.
		# Caddy's default list covers text/* and application/json but not NDJSON, so
		# the content type is matched explicitly. The event stream is excluded: it
		# must flush as it happens, never through a compressor.
		@compressible not path /v1/events*
		encode @compressible zstd gzip {
			match {
				header Content-Type *json*
			}
			minimum_length 1024
		}
		reverse_proxy 127.0.0.1:8788
	}
}

https://images.example.com:9000 {
	reverse_proxy 127.0.0.1:9002
}

Then SERVICE_S3_PUBLIC_ENDPOINT=https://images.example.com:9000, and open exactly two inbound ports in the firewall/security group: the proxy's HTTPS port and the storage port (clients upload and download straight from storage, so it cannot be loopback-only). Upgrades are git pull && docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d --build; the SQLite record lives on the app-data volume, so back that up if the job history matters.

Retention: stored images are deleted once they are older than SERVICE_RETENTION_DAYS (default 3 days); the task records stay. The sweep deletes

  • the output of every job that succeeded more than the window ago, and

  • input images older than the window that nothing still needs — an input is kept while any job that references it is unfinished (queued/submitting/accepted/downloading) or was created inside the window, however old that job is.

A purged job keeps its request, result metadata, provider and attempt history, moves to the expired phase and stops being handed to a worker — it is still listed and exported, just without a download URL. Re-uploading input content that was purged re-opens the asset with a fresh lifetime. Retrying a job whose inputs are gone is refused with 409 (and a job that discovers a missing input at submit time fails immediately with input_missing, before any charge) instead of waiting forever. Set SERVICE_RETENTION_DAYS=0 to keep everything forever.

Endpoints

Method

Path

Purpose

POST

/v1/assets/presign

sign a direct upload (content addressed by sha256)

POST

/v1/assets/:id/complete

confirm the object exists, making the asset usable

GET

/v1/assets/:id/download

signed download URL

POST

/v1/jobs

submit one job (edit when input_asset_ids or input_paths is given)

GET

/v1/jobs, /v1/jobs/:id

list / inspect (add ?attempts=1 for provider receipts, ?resolutions=1 for the operator audit trail)

POST

/v1/jobs/:id/cancel

stop before it finishes (a submission already in flight cannot be recalled: its attempt is recorded as ambiguous and may still be billed)

POST

/v1/jobs/:id/resolve

handle a needs_attention/failed/canceled job: retry (needs acknowledge_duplicate_charge: true, grants exactly one extra submission) or fail; both close an attempt that was still in flight as ambiguous

POST

/v1/batches, /v1/batches/import

batch create / NDJSON import (accepts request_key, or an Idempotency-Key header, to make a retried batch a no-op)

GET

/v1/batches/:id, /v1/batches/:id/export

progress / per-item reconciliation as NDJSON (add ?sign=1 to get a download_url per line, instead of calling /v1/jobs/:id for each result)

GET

/v1/providers

provider health and cooldown

GET

/v1/health, /v1/ready

liveness / readiness (public; everything else needs a bearer token)

GET

/v1/events

live notifications (SSE): job events as jobs change, resumable via Last-Event-ID, filterable with ?batch_id= / ?job_id=

ALL

/mcp

remote MCP (streamable HTTP), same bearer token

Images never travel through the service: clients PUT straight to a presigned URL and download through a presigned URL, so a 3000-image batch does not stream through Node.

Instead of polling

Two ways to stop sleeping between polls, both of which react the moment a job settles:

  • Blocking reads. GET /v1/jobs/:id?wait_ms=25000 and GET /v1/batches/:id?wait_ms=25000 (and wait_ms on the MCP tools get_image_job / get_image_batch) return as soon as the job is terminal or the batch has nothing pending — or at the deadline, with the ordinary current state, in which case the caller simply asks again. MCP clients have no push channel, so this is the one to use from an agent.

  • A live stream. GET /v1/events is server-sent events: each job event carries the job, its batch and item index, and the phase it just moved into. Subscribe with ?batch_id= (or ?job_id=) to see only your own work, and reconnect with the standard Last-Event-ID header to replay what you missed — the underlying journal is append-only, so nothing is skipped and nothing is delivered twice.

# Follow one batch's progress; stop when nothing is pending (Ctrl-C, or a counter).
curl -N "http://localhost:8787/v1/events?batch_id=$BATCH_ID" \
  -H "Authorization: Bearer $TOKEN" -H 'accept: text/event-stream'
#   event: ready            <- the cursor you are starting from (fetch the snapshot now)
#   data: {"cursor":8123}
#
#   event: job
#   id: 8124
#   data: {"job_id":"job_…","batch_id":"batch_…","item_index":7,"phase":"succeeded","occurred_at":…,"finished_at":…,"purged_at":null}

An event says that something changed, not what the new state is: read GET /v1/jobs/:id (or the batch export) for the result, the download URL and the provider's applied parameters. The journal starts with migration 0007, so history predating it is only visible through the ordinary read routes.

Files in, files out

The whole interaction is file-oriented, at both ends.

Upload (client → service). Either the file is already on the service's machine and you pass its path, or you upload bytes straight to storage with a presigned URL:

# 1. ask for a slot (MCP: presign_image_asset returns the same values)
curl -sS -X POST http://localhost:8787/v1/assets/presign \
  -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d "$(jq -nc --arg sha "$(sha256sum fox.png | cut -d' ' -f1)" \
        '{sha256:$sha, extension:"png", content_type:"image/png"}')"
# 2. PUT the file itself — bytes never pass through a prompt or a JSON payload
curl -sS -T fox.png -H 'content-type: image/png' "$UPLOAD_URL"
# 3. confirm, then use the asset id in input_asset_ids
curl -sS -X POST "http://localhost:8787/v1/assets/$ASSET_ID/complete" -H "Authorization: Bearer $TOKEN"

Download (service → client). get_image_job returns a download URL (valid for six hours by default) and the MCP reply also carries it as a resource link, so the client saves it as a file:

curl -sSL "$DOWNLOAD_URL" -o result.png
# bulk: one manifest for the whole batch, one line per item, each with its own URL
curl -sS "http://localhost:8787/v1/batches/$BATCH_ID/export?sign=1" -H "Authorization: Bearer $TOKEN"

Presigned URLs are signed for SERVICE_S3_PUBLIC_ENDPOINT, so set it to the address your clients reach rustfs on (for a LAN deployment, e.g. http://10.0.0.5:9000) — localhost:9000 only works for a client on the same machine as the stack.

What a downloaded image carries

Generated images arrive from the providers with provenance metadata attached. Measured on the production bucket, every sampled output PNG carried a caBX chunk — a ~22 KB C2PA/Content Credentials manifest naming the tool that made it (c2pa.claim, OpenAI, gpt-image). Publication therefore cleans every image before storing it, under one of two versioned profiles:

Profile

What it does

Pixels

metadata-v2 (default)

container rewrite and a re-encode: the image is decoded and written again with this service's own deflate

byte-identical

metadata-v1

container rewrite only: provider metadata is dropped, the provider's compressed stream is kept verbatim

untouched

The container rewrite keeps only what a decoder needs, and drops everything a decoder does not:

Container

Kept

Dropped

PNG

IHDR, PLTE, tRNS, IDAT, IEND

caBX (C2PA), tEXt, iTXt, zTXt, eXIf, iCCP, gAMA, sRGB, pHYs, … and any bytes after IEND

JPEG

coding segments (DQT, DHT, SOFn, DRI, SOS + scans, EOI)

every APPn (EXIF/XMP/ICC/Adobe) and COM

WebP

VP8/VP8L, ALPH, VP8X (metadata flags cleared)

ICCP, EXIF, XMP , and the RIFF trailer

metadata-v1 copies the kept bytes verbatim, so the compressed pixel stream and its CRCs are untouched. metadata-v2 goes further: it inflates the IDAT stream, un-filters the scanlines, re-filters them (Paeth) and re-deflates with fixed parameters (level 9, Z_FILTERED) — which removes the encoder fingerprint that survives in the compressed bytes, at the cost of holding one decoded image in memory. On a 1.5 MP output (2,231,063 bytes) that produced 2,056,971 bytes (−8%) in ~2.3 s with AE = 0 against the original, i.e. not one pixel differs; the result is deterministic and re-running it on its own output is a no-op. Both are covered by tests that decode the result with an independent decoder (tests/helpers/png-fixtures.ts).

A shape metadata-v2 cannot re-encode faithfully is delivered under metadata-v1 and records that it was: 16-bit, indexed (palette), interlaced, above SERVICE_REENCODE_MAX_PIXELS (default 16,777,216 = 4096²), or a decode/encode failure. JPEG and WebP get the container profile only, with reencode_reason: "unsupported_format" — re-encoding them would mean a second encoder (and, for JPEG, generation loss).

The result reports what happened: sanitization.status is clean (with the policy actually applied, reencoded, and the list of removed structures), skipped (with a reason, for a file the profile cannot parse — it is still delivered and should be treated as carrying metadata), or disabled (SERVICE_OUTPUT_SANITIZATION=off). A crash-recovery adoption refuses an object that carries no receipt rather than adopting it as this attempt's output. Re-encoding runs at most SERVICE_REENCODE_MAX_CONCURRENT (3) images at a time per process; the work itself is in libuv's thread pool, so it does not block the event loop.

Two limits are worth stating plainly. This is a container-level guarantee: it removes what an image says about itself and, under v2, the encoding fingerprint — it makes no claim about watermarks inside the pixels, and no re-encode can rule out steganographic content. And a skip is not an error: the image was already generated and paid for, so it is delivered with the reason recorded rather than thrown away.

Existing objects are brought forward with scripts/clean-stored-outputs.mjs, which skips an object whose recorded policy already matches SERVICE_OUTPUT_SANITIZATION and re-cleans one that does not (so raising the profile re-encodes the store on the next run).

The test suite verifies this with parsers written independently of the rewriter (see tests/helpers/png-fixtures.ts and tests/helpers/image-fixtures.ts) and with real encoder output committed as fixtures (baseline/progressive/greyscale/CMYK JPEG, lossy/ lossless/animated WebP).

Downloading a batch quickly

Results are plain PNG/JPEG/WebP files served by rustfs through the proxy. Two facts decide how fast they arrive:

  • Images are already compressed. gzip on a 1.4 MB PNG saved 0.04% (1.43 MB → 1.42 MB) because PNG's own deflate has nothing left to give. Size comes from the output format and dimensions, not from transport compression: output_format: "webp" and a smaller size are the only real levers on the bytes.

  • A single connection is the slow unit. Measured on the production host: 20 MB/s locally through the proxy, ~0.75 MB/s for one connection from the public internet, and 5.4 MB/s with eight parallel connections — the limit is per flow, so parallelism scales roughly 7x until the host's total egress (~11 MB/s) is hit. If your client is far away or lossy, download items concurrently rather than segments of one file.

# One manifest for the whole batch, then pull it with 8 connections.
curl -sS "https://<host>/v1/batches/$BATCH_ID/export?sign=1" -H "Authorization: Bearer $TOKEN" \
  | jq -r 'select(.download_url) | .download_url' > urls.txt
xargs -P 8 -n 1 curl -sS -C - -O < urls.txt          # -C - resumes a partial file
# or, if aria2 is available:
aria2c -x8 -j8 -i urls.txt -d results/

Range requests work (206), so interrupted downloads resume rather than restart. The manifest itself is gzipped by the proxy when the client asks for it (26.9 KB → 2.9 KB for 24 items, ~360 KB for 3000), which matters most on a slow link.

Input images: files, not base64

Two ways to hand over a reference image, neither of which puts image bytes in a payload:

  • input_paths (with mask_path) — absolute paths on the machine running the service. The service reads, hashes and stores each file itself, so a manifest of 3000 paths is a small NDJSON file. This is the one to use for a bulk edit run.

  • input_asset_ids — for images the service cannot read (a client on another machine): sign with POST /v1/assets/presign, PUT the file to the URL, POST /v1/assets/:id/complete, then reference the asset id.

Both are content addressed, so re-listing the same file is free. A path that does not exist, is relative, or is not a PNG/JPEG/WebP is rejected with a 400 naming the problem, and a batch import reports it against that line without dropping the rest.

# One JSONL line per input image; the service reads the files (mount them into the
# container, e.g. `-v /data/images:/data/images:ro`).
find /data/images -name '*.png' | jq -c '{prompt:"clean up the background", input_paths:[.]}' > batch.jsonl
curl -sS -X POST http://localhost:8787/v1/batches/import \
  -H "Authorization: Bearer $TOKEN" -H 'Idempotency-Key: cleanup-2026-09' \
  --data-binary @batch.jsonl

Remote MCP

{
  "mcpServers": {
    "gpt-image-service": {
      "type": "http",
      "url": "https://images.internal.example/mcp",
      "headers": { "Authorization": "Bearer <SERVICE_TOKENS value>" }
    }
  }
}

Eleven job-oriented tools (submit_image_job, presign_image_asset, confirm_image_asset, get_image_job, list_image_jobs, cancel_image_job, resolve_image_job, submit_image_batch, import_image_batch_jsonl, get_image_batch, list_image_providers). They return job/batch references and signed links rather than inline bytes, so the same tools work for one image and for a batch of thousands. submit_image_batch / import_image_batch_jsonl take a request_key, so a retried call returns the batch it already created.

Generation: submit_image_job {prompt, size?, quality?, output_format?} → get_image_job {job_id} (returns object_key and a download_url that stays valid for hours — SERVICE_PRESIGN_TTL_SEC, six by default — and attaches the image as a resource link so the client fetches it as a file). The reply also carries applied_diffs: whatever the provider changed on the way through. Gateways are not exact — one may answer a 1024x1024 request with 1374x1145, or quietly raise low to medium — and a silently different image is worth knowing about before a 3000-image run is signed off.

Editing (prompt + reference image): submit_image_job {prompt, input_paths: ["/abs/path/fox.png"]} when the file is on the service's machine — the service reads it, so nothing about the image enters the conversation. When it is not (an MCP client on a laptop, say): presign_image_asset {sha256, extension, content_type} → PUT the file to the returned upload_url → confirm_image_asset {asset_id} → submit_image_job {prompt, input_asset_ids: [asset_id]}. Up to 8 references, with an optional mask_asset_id / mask_path.

Batch from files: import_image_batch_jsonl with one line per image, e.g. {"prompt":"…","input_paths":["/data/images/0001.png"]} — a 3000-image manifest of paths, no base64 anywhere.

MCP has no server→client push, so the instructions tell an agent with many jobs to open the event stream itself: curl -sN '<service>/v1/events?batch_id=…' with the same bearer token, resuming with Last-Event-ID. With SERVICE_PUBLIC_URL set the instructions (and get_image_batch's results_path) spell out the full address; without it they say how to derive it from the /mcp URL you already have. Single jobs need no shell at all — wait_ms blocks the tool call until the job settles.

Tools for agents, REST for scripts. The tool names, arguments and descriptions are written for a model to choose from; they are the interactive surface, and they are free to change. Anything a program drives — a script, a cron job, a CI step, a 3000-image ingest — should call the REST API directly with the same bearer token: every tool is a thin wrapper over those routes, so nothing is missing, and the paths above are the supported client interface. The MCP instructions say the same thing to the agent, so one that is asked to "write a script" reaches for POST /v1/batches/import instead of mimicking tool calls.

To collect a finished batch, GET /v1/batches/:id/export?sign=1 returns one NDJSON line per item with a fresh download_url (also reported by get_image_batch as results_path), so 3000 results do not need 3000 per-job calls — just one manifest to loop over with the client's own downloader.

A quick check:

curl -s -X POST localhost:8787/mcp \
  -H "Authorization: Bearer <token>" -H 'content-type: application/json' \
  -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'

Live smoke test

pnpm run smoke:service -- --yes --count=3 --providers=krill,sub2api runs the real pipeline (queue-shaped submission → executor → provider → object storage) against real accounts for a few images, then reads each published object back and checks it is a real image. It refuses to start without --yes.

How the service spends money safely

  • Five billed submissions per image, across providers. Polling, downloading, waiting for capacity and switching provider do not consume the budget. The wall-clock SERVICE_JOB_DEADLINE_MS budget starts with the first authorised submission — never while waiting for capacity — so a saturated upstream cannot make a 3000-item batch fail items that were never submitted.

  • The authorisation is persisted before the request is sent, and a job whose outcome is unknown (a lost response, a crash mid-submission) becomes needs_attention instead of being retried automatically — a duplicate submission costs real money, so the service parks it for a human.

  • Idempotency: POST /v1/jobs defaults to a request key derived from the request content, so a retried HTTP call returns the existing job. Batch items are keyed by position instead, so identical prompts in one batch stay separate jobs.

  • A batch is one billed decision covering thousands of items, so it takes a client key (request_key in the body, or an Idempotency-Key header). Replaying the same key with the same payload returns the original batch and resumes only the items that never made it into the database (created reports how many were new, replayed that the batch already existed); the same key with a different payload is answered with 409 rather than silently reusing another batch's id. Without a key, every call is a new batch.

  • At-least-once dispatch, at-most-once billing: wake-ups are stored in a transactional outbox, each under its own durable intent id (which is also the transport id), and re-created by a reconciler if Redis loses them — a new intent is always delivered, even though the transport still retains the delivery it replaced. The exclusive claim and lease fencing in SQLite keep one job to one worker, and the lease is renewed for as long as an activation runs, so a slow synchronous provider cannot be reclaimed and worked twice.

  • Free retries are retried, paid ones are not: publication to object storage is retried inside the activation (it costs nothing), and output keys are per attempt, so a worker that lost its claim cannot overwrite the winner's image. If storage stays down, a job whose provider still holds the task goes back to downloading; a job whose bytes arrived inline — and therefore cannot be fetched again — is parked as needs_attention instead of being regenerated and billed. A crash between the write and the row update is recovered by adopting the object that attempt already published.

  • The ceiling is enforced where the money is spent: per-provider in-flight limits are re-checked inside the transaction that authorises a submission, so two workers that both read "one slot free" cannot both submit — the loser waits in queued without spending an attempt. A provider-specific rejection (permanent: bad parameters, moderation) is recorded but does not count against provider health, so a healthy cheap tier is not taken out of rotation by one bad prompt. A rejection that is really about the provider — a missing endpoint or model, which a gateway answers with 404/405/501 (provider_unavailable) — is treated the other way round: the provider steps aside for the hard-failure window, and the retry of that same job lands on the next tier instead of dying with the request.

  • Decisions are auditable: resolve records the action, the reason, the free-text operator and the authenticated token that made the request (actor, which the client cannot set) in an append-only trail (GET /v1/jobs/:id?resolutions=1, MCP get_image_job{include_resolutions:true}, newest-last and stable within the same millisecond). It survives the job succeeding and clearing its current error — and max_attempts counts the extra submissions an operator authorised, so a job never looks as if it ran past its own limit.

  • Explicit resolution: a job that ends with an unknown outcome parks as needs_attention and is never retried on its own. POST /v1/jobs/:id/resolve (or the resolve_image_job MCP tool) is the only way past that: action: retry requires acknowledge_duplicate_charge: true, grants exactly one extra billed submission and re-queues the job immediately; action: fail closes it out with no spend. Both record the operator and the note on the job.

  • Provider economics: providers sit in preference tiers (SERVICE_SUB2API_PRIORITY 0, SERVICE_KRILL_PRIORITY 10) — the subscription-backed provider is filled first and a per-image billed provider only absorbs the overflow once the cheaper tier is at its ceiling or out of rotation. A single transient failure does not spill (SERVICE_PROVIDER_FAILURE_THRESHOLD, default 2), because retrying the cheap provider is cheaper than paying another one. A deliberate mix is available too: SERVICE_KRILL_SPILL_RATIO=0.2 sends ~20% of requests to krill even while sub2api has room.

  • Provider ceilings: SERVICE_KRILL_MAX_IN_FLIGHT (15) / SERVICE_SUB2API_MAX_IN_FLIGHT (8) cap how many jobs each provider holds at once — that capacity is what triggers the spill to the next tier. Occupancy is derived from live job rows, so a restart reconstructs it instead of oversubscribing the upstream; jobs over the ceiling wait in queued and spend nothing.

Troubleshooting

  • "OPENAI_API_KEY is not set" — add it to the env block of your MCP config.

  • A tool result is a job_id, not an image — that is the normal path. Poll get_image_job every few seconds until state is completed or failed. (Set GPT_IMAGE_2_ASYNC_AFTER_MS=0 to make calls block instead.)

  • "Unknown or expired image job" — the server restarted, or the job aged past GPT_IMAGE_2_JOB_TTL_MS. A finished result cannot be recovered; call list_image_jobs to see what is still held, and re-run only if the image is really gone (images already written to disk stay on disk).

  • continue_edit_session says the session is busy — its previous turn is still running in the background. Poll the pending_job_id from list_edit_sessions, then continue. Sessions are also lost on server restart.

  • 403 / organization verification — gpt-image-2 may require Organization Verification on your OpenAI org. Check the dashboard.

  • 429 — you hit the IPM (images per minute) cap for your tier. Lower n, or wait.

  • Size / quality / format came back different — the origin normalized them (a ChatGPT-subscription gateway does this: see Sizes). applied and notes report what happened; point OPENAI_BASE_URL at an API-key upstream to have your settings honored.

  • Image doesn't appear in the client — check the file path in the text block; the image is saved regardless of inline display.

  • Images land in an unexpected directory — the default is keyed on the server's working directory; set GPT_IMAGE_2_OUTPUT_DIR to an absolute path.

  • Client times out anyway — raise OPENAI_TIMEOUT_MS (if the proxy is slow) or lower GPT_IMAGE_2_ASYNC_AFTER_MS (if your host's tool timeout is short).

  • Protocol disconnects silently — something printed to stdout. Check src/**/*.ts — all logs must use utils/logger.ts (stderr). This is the single biggest MCP footgun.

Getting help

Smoke test against the real API

tests/smoke/features/*.feature describe the capability in Gherkin and run it against the real endpoint — generation, editing, multi-turn sessions, and the background-job queue:

pnpm run smoke              # every feature (~4 min, a few cents of images)
pnpm run smoke -- editing   # one feature, or scenarios whose name matches

Each step reports pass/fail on its own line; artifacts (the generated images) land in a per-run temp directory printed at the top — set SMOKE_OUT_DIR to put them somewhere stable. Requires OPENAI_API_KEY (and OPENAI_BASE_URL if you go through a proxy).

Two deliberate choices: the suite runs queue-first (the default window), so every call exercises the background-job path — set GPT_IMAGE_2_ASYNC_AFTER_MS to a larger value to run the inline path instead (the queue feature is then skipped); and it asserts what the upstream actually returned rather than assuming — a proxy that ignores output_format or size shows up as a reported observation, not a false failure. These tests are not part of pnpm test or CI, which run without credentials.

SMOKE_TRANSPORT=stdio pnpm run smoke drives the built server (pnpm run build first) over stdio — a real MCP client, the same way your host calls it — instead of the default in-process transport.

Which parameters the origin honors

pnpm run smoke:matrix runs a one-parameter-at-a-time matrix — a baseline plus model, size, quality, background, output_format, and n, then the same for edit_image — and compares what the API reports (applied) with what the bytes are (container, dimensions, alpha). A control only counts as honored when the response agrees and the bytes differ from the baseline row: a gateway can echo a field back and still ignore it.

It writes outcomes.json (re-render with --report <file>) and parameter-matrix.md into the artifacts directory. It is a study, not a gate: it takes minutes and costs a few cents. Filter rows with an argument, e.g. pnpm run smoke:matrix -- size.

Development

pnpm run dev         # tsx watch
pnpm run typecheck   # tsc --noEmit for src, plus tsconfig.test.json for tests
pnpm run test        # offline unit + integration tests (node:test, no credentials)
pnpm run build       # compile to build/
pnpm run inspect     # launch MCP Inspector

pnpm run scale (optionally -- --jobs=10000) seeds a throwaway SQLite file with thousands of jobs and prints per-query timings plus EXPLAIN QUERY PLAN output, so a missing index shows up as a scan-and-sort before it shows up in production.

pnpm test is the fast, offline suite — every test runs against a local HTTP mock, and CI runs it on Node 24 (the version the service and its node:sqlite driver require) with a Redis service container, so the queue integration tests run too. pnpm run smoke (below) is the live suite and spends money.

Releasing

Bump version in package.json, add a CHANGELOG section for the new version, commit, then tag it:

git tag v0.5.1 && git push origin v0.5.1   # must match package.json

The Release workflow verifies the tag matches package.json, runs typecheck + tests + build, publishes with --provenance, and opens the GitHub release. It needs an NPM_TOKEN repository secret (npm automation token for the @speed-fullbar scope). Running the workflow manually defaults to a --dry-run publish, and tagging a version that is already on npm skips the publish (with a notice) instead of failing — useful when a release was published by hand.

Credits

This is a fork of Borys520/gpt-image-2-mcp by Borys Kusmirek (MIT) — the original server, tool surface, and research brief are theirs, and the copyright notice in LICENSE is retained. This fork adds the background job queue (get_image_job, list_image_jobs), edit sessions that inherit their settings, the transparent-PNG and parameter-matrix work, the BDD smoke suite, and the packaging/docs here.

License

MIT

Available Tools

8 tools
continue_edit_sessionContinue Edit SessionA

Apply another edit turn to an existing session. The previous turn's output image is used as the input. Use short, focused prompts like "make the sky more orange" or "add a small boat on the horizon"; include "keep everything else the same" to limit drift. Returns the new image and the updated session. Omitting size, quality, background, output_format, or model inherits what the session already uses (list_edit_sessions reports those settings). The turn hands off to a background job (poll get_image_job); while one is still running the session stays busy and further turns are refused until it lands.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoOutput dimensions. "auto" (default), one of the presets "1024x1024", "1536x1024", "1024x1536", or a custom "WxH" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta. Omit to keep the size the session is already using.
userNoOptional end-user identifier forwarded to OpenAI for abuse monitoring. Pass a stable hashed user ID, not PII.
modelNoModel to use. One of "gpt-image-2.5-sunburst", "gpt-image-2.5-flare", "gpt-image-2"; defaults to "gpt-image-2.5-sunburst". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing.
promptYesImage description. gpt-image-2 handles very detailed prompts; use ALL CAPS or quote literal text you want rendered verbatim.
qualityNoEdit quality — same levels as generate. Omit to keep the quality the session is already using.
backgroundNoBackground behavior. "opaque" forces a filled background; "transparent" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; "auto" lets the model pick. Omit to keep the background the session is already using.
session_idYesThe session id returned by start_edit_session.
output_formatNoFile format. "png" (default, lossless), "jpeg" (smaller, lossy), "webp" (best compression). When omitted on continue_edit_session, the session's current format is kept.
filename_prefixNoShort label appended to the generated filename so you can find it later (e.g. "hero-banner"). Letters/digits/hyphens only; auto-sanitized.
output_compressionNoCompression level 0–100 for jpeg/webp outputs. Ignored for png. Defaults to 100 (minimal compression).

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolNo
turnNo
modelNo
notesNo
routeNo
stateNo
usageNo
imagesNoWritten image files — present once the job completed successfully.
job_idNoPresent on background hand-off — pass to get_image_job.
promptNo
appliedNo
poll_hintNo
requestedNo
session_idNoIdentifies the session for later continue_edit_session calls; present once a turn has landed.
started_atNo
async_after_msNo
prompt_previewNo
cost_usd_estimatedNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses asynchronous execution ('hands off to a background job (poll get_image_job)') and session-level locking ('while one is still running the session stays busy and further turns are refused until it lands'), neither of which is present in the annotations. It also explains parameter inheritance behavior, adding substantial value beyond the structured hints. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: purpose, chaining, prompt tips, inheritance, and async behavior. It is front-loaded with the core action and ends with the operational caveat. Slightly dense but not padded; a minor trim would push it to a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description correctly focuses on behavior and usage. It covers prerequisites (existing session), invocation conditions (not busy), the polling workflow, and inheritance semantics, and it references sibling tools (list_edit_sessions, get_image_job) for supporting context. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% parameter coverage, including per-parameter 'omit to keep' rules, so the schema already carries the semantic load. The description reiterates the inheritance concept at a high level but does not add detail beyond what the schema states. This meets the baseline of 3 for fully documented schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Apply another edit turn' to 'an existing session') and immediately clarifies the chaining behavior ('previous turn's output image is used as the input'). This distinguishes it clearly from start_edit_session (new session) and generate_image (standalone generation), and the session context is reinforced by the reference to list_edit_sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides concrete usage guidance: short focused prompts, including 'keep everything else the same' to limit drift, omitting parameters to inherit session settings, and the busy condition that refuses further turns until the background job lands. It does not explicitly name edit_image as an alternative, but the session-flow context and references to start_edit_session and get_image_job make the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_imageEdit ImageA

Edit or compose images with OpenAI's image model family (models: "gpt-image-2.5-sunburst" (default), "gpt-image-2.5-flare", "gpt-image-2"). Give 1–8 input images plus a text prompt; optionally include a PNG mask whose transparent regions mark what to change (mask applies to the first image). Great for: swap backgrounds, retouch products, combine multiple reference images into one composition, maintain a character across scenes. These models always process inputs at high fidelity (no input_fidelity knob needed). The edited image is saved to disk and returned inline. Image work is slow, so this call hands off immediately: the first response carries a job_id — poll get_image_job until it reports state "completed". Set GPT_IMAGE_2_ASYNC_AFTER_MS to wait inline instead (below your host's tool-call timeout).

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoHow many images to generate (1–10). Each counts toward rate limits and cost.
maskNoOptional PNG mask — fully transparent pixels mark the editable region. Must match the first input image's dimensions and be <4MB. Accepts the same source types as `images`.
sizeNoOutput dimensions. "auto" (default), one of the presets "1024x1024", "1536x1024", "1024x1536", or a custom "WxH" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta.auto
userNoOptional end-user identifier forwarded to OpenAI for abuse monitoring. Pass a stable hashed user ID, not PII.
modelNoModel to use. One of "gpt-image-2.5-sunburst", "gpt-image-2.5-flare", "gpt-image-2"; defaults to "gpt-image-2.5-sunburst". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing.
imagesYesInput images. Each entry can be: an absolute file path, a relative path (resolved from CWD), a file:// URL, an http(s):// URL, or a data:image/...;base64,... URL. PNG/WEBP/JPG, up to 50MB each.
promptYesImage description. gpt-image-2 handles very detailed prompts; use ALL CAPS or quote literal text you want rendered verbatim.
qualityNoEdit quality — same levels as generate.auto
backgroundNoBackground behavior. "opaque" forces a filled background; "transparent" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; "auto" lets the model pick.auto
output_dirNoAbsolute or relative directory where generated images should be written. Defaults to $GPT_IMAGE_2_OUTPUT_DIR or a per-project subfolder under the OS config dir. The directory is created if missing.
output_formatNoFile format. "png" (default, lossless), "jpeg" (smaller, lossy), "webp" (best compression). When omitted on continue_edit_session, the session's current format is kept.
filename_prefixNoShort label appended to the generated filename so you can find it later (e.g. "hero-banner"). Letters/digits/hyphens only; auto-sanitized.
output_compressionNoCompression level 0–100 for jpeg/webp outputs. Ignored for png. Defaults to 100 (minimal compression).

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolNo
modelNo
notesNo
routeNo
stateNo
usageNo
imagesNoWritten image files — present once the job completed successfully.
job_idNoPresent on background hand-off — pass to get_image_job.
promptNo
appliedNo
poll_hintNo
requestedNo
started_atNo
async_after_msNo
prompt_previewNo
cost_usd_estimatedNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing that the call hands off immediately, that the first response carries a job_id, that the caller should poll get_image_job until state 'completed', and that GPT_IMAGE_2_ASYNC_AFTER_MS can make it wait inline. It also notes that inputs are always processed at high fidelity and that outputs are saved to disk and returned inline. These are nontrivial behaviors not captured by readOnlyHint, openWorldHint, idempotentHint, or destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-organized paragraph that front-loads the core action, then covers inputs, use cases, and behavior. Every sentence adds useful information, and the async-handoff detail is placed exactly where it is needed. There is no redundancy with the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (13 parameters, async behavior, optional mask, output writing), the description covers all the essential workflow knowledge an agent needs: inputs, mask semantics, model family, job_id polling, disk output, and the inline-wait alternative. The comprehensive input schema handles parameter-level details, and the output schema handles return values, so nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 100% of the 13 parameters with detailed descriptions, so the baseline is 3. The description adds cross-cutting semantic value by explaining that the PNG mask's transparent regions mark the editable region, that the mask applies to the first image, and that 1–8 input images can be combined with a text prompt. This contextual mapping of parameters to the editing workflow elevates the score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb-resource pair ('Edit or compose images') and immediately names the model family, inputs, and typical use cases. It clearly differentiates itself from sibling tools like generate_image (edit vs. generate) and get_image_job (the tool's async counterpart). An agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong context by listing concrete use cases and explicitly instructing the agent to poll get_image_job after receiving a job_id. It does not explicitly state when to use generate_image or continue_edit_session instead, but the phrase 'Edit or compose' plus the extensive job-handling guidance makes the boundary clear enough. No exclusions or when-not-to-use caveats are given, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

end_edit_sessionEnd Edit SessionA
DestructiveIdempotent

Free an iterative-edit session. Safe to skip — sessions are in-memory only and are discarded on server restart — but calling this frees memory sooner and keeps list_edit_sessions tidy.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe session id to end.

Output Schema

ParametersJSON Schema
NameRequiredDescription
endedYes
session_idYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true. The description adds valuable context: sessions are in-memory only, discarded on server restart, and calling the tool frees memory sooner. No contradictions found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all essential. Front-loaded with the main action, then adds usage context and benefits. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (1 param, high schema coverage, output schema exists), the description covers purpose, usage, and side effects. Could mention what happens if session_id is invalid, but the idempotentHint implies safe handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%. The parameter 'session_id' is well-documented in the schema with minLength and description. The tool description does not add further parameter details, which is acceptable given the schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool ends an iterative-edit session and distinguishes from siblings like start_edit_session by noting it is optional cleanup. The verb 'Free' and resource 'iterative-edit session' are specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (to free memory sooner, keep list tidy) and when not to use (safe to skip because sessions are discarded on restart). Provides alternative: not calling the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageGenerate ImageA

Generate an image from a text prompt using OpenAI's image model family (models: "gpt-image-2.5-sunburst" (default), "gpt-image-2.5-flare", "gpt-image-2"). The image is written to disk and also returned inline so you can see it. These models handle photoreal, illustrations, infographics, multilingual text (incl. CJK), and complex structured visuals. For a transparent logo pass background="transparent" with output_format="png" — verified against this family; the origin decides, so check applied.background. Sizes accept presets or any custom "WxH" where edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta. Image work is slow, so this call hands off immediately: the first response carries a job_id — poll get_image_job until it reports state "completed" (the result arrives verbatim, images included). Set GPT_IMAGE_2_ASYNC_AFTER_MS to wait inline instead (below your host's tool-call timeout).

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoHow many images to generate (1–10). Each counts toward rate limits and cost.
sizeNoOutput dimensions. "auto" (default), one of the presets "1024x1024", "1536x1024", "1024x1536", or a custom "WxH" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta.auto
userNoOptional end-user identifier forwarded to OpenAI for abuse monitoring. Pass a stable hashed user ID, not PII.
modelNoModel to use. One of "gpt-image-2.5-sunburst", "gpt-image-2.5-flare", "gpt-image-2"; defaults to "gpt-image-2.5-sunburst". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing.
promptYesImage description. gpt-image-2 handles very detailed prompts; use ALL CAPS or quote literal text you want rendered verbatim.
qualityNoGeneration quality. "low" for fast drafts, "medium" balanced (default when model picks), "high" for dense layouts and text, "auto" lets the model choose.auto
backgroundNoBackground behavior. "opaque" forces a filled background; "transparent" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; "auto" lets the model pick.auto
moderationNoModeration strictness. "auto" (default) applies standard safety filtering; "low" is less restrictive (still subject to OpenAI policy).auto
output_dirNoAbsolute or relative directory where generated images should be written. Defaults to $GPT_IMAGE_2_OUTPUT_DIR or a per-project subfolder under the OS config dir. The directory is created if missing.
output_formatNoFile format. "png" (default, lossless), "jpeg" (smaller, lossy), "webp" (best compression). When omitted on continue_edit_session, the session's current format is kept.
filename_prefixNoShort label appended to the generated filename so you can find it later (e.g. "hero-banner"). Letters/digits/hyphens only; auto-sanitized.
output_compressionNoCompression level 0–100 for jpeg/webp outputs. Ignored for png. Defaults to 100 (minimal compression).

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolNo
modelNo
notesNo
routeNo
stateNo
usageNo
imagesNoWritten image files — present once the job completed successfully.
job_idNoPresent on background hand-off — pass to get_image_job.
promptNo
appliedNo
poll_hintNo
requestedNo
started_atNo
async_after_msNo
prompt_previewNo
cost_usd_estimatedNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide flags, but the description discloses meaningful side effects: the image is written to disk and also returned inline, the call hands off immediately with a job_id, and the result must be polled through get_image_job. It also surfaces non-obvious caveats such as the 2K beta status and the fact that the origin still decides whether transparency is honored.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well front-loaded and logically organized: purpose, capabilities, transparency recipe, size rules, then async behavior. It is not a perfect 5 because several sentences duplicate schema text (size math, 2K beta, model names) and the model-capabilities sentence, while useful, is optional context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter generator with full schema coverage, an output schema, and a slow async job, the description covers the critical unknown an agent needs: how results arrive and how to monitor completion. The transparent-background caveat, size limits, and disk-write behavior are all present, so nothing required to invoke or follow up on the call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description earns an extra point by pairing background="transparent" with output_format="png" as a verified practical recipe and by consolidating size constraints and model names for fast grounding. It adds little for quality, moderation, or compression parameters, which remain schema-only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific action (generate), a specific resource (image from a text prompt), and the model family, making the core purpose immediately clear. It is also unmistakably distinct from the edit_image and get_image_job siblings: generation versus editing versus polling, so an agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear generation context, a concrete transparent-logo recipe, and operational guidance for the async completion flow via get_image_job, including the env-var alternative. It stops short of explicitly saying when not to use this tool or when to prefer edit_image/start_edit_session, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_image_jobGet Image JobA
Read-onlyIdempotent

Poll a backgrounded gpt-image job: every image tool (generate_image, edit_image, start_edit_session, continue_edit_session) hands off immediately and returns a job_id rather than blocking past MCP client timeouts. Call this with that job_id until state becomes "completed" (files are already written to disk and returned inline, with the model, requested/applied settings, usage, and cost) or "failed". While still "running", wait a few seconds between polls.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by generate_image / edit_image when they moved the work to the background.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
turnNo
errorYes
modelNo
notesNo
routeNo
stateYes
usageNo
imagesNoWritten image files — present once the job completed successfully.
job_idYes
promptNo
appliedNo
poll_hintNoPresent while a job is still running.
requestedNo
elapsed_msYes
session_idNo
started_atYes
completed_atYes
prompt_previewYes
cost_usd_estimatedNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds valuable traits beyond annotations: it explains the non-blocking handoff, the possible states ('running', 'completed', 'failed'), and that completed jobs have files written to disk and returned inline with model, settings, usage, and cost. This is exactly the behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact but complete: it front-loads the core polling action, explains why it exists, states the terminal states, and gives the polling cadence in three sentences. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter polling tool with a rich description and output schema, nothing essential is missing. It explains the backgrounding context, how to poll, what terminal states mean, what completed results contain, and how long to wait between polls. The presence of an output schema also relieves the description of needing to document return structure in detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by noting that job_id is returned by every image tool, including start_edit_session and continue_edit_session, whereas the schema only mentions generate_image and edit_image. This extends the agent's understanding of valid job_id provenance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Poll a backgrounded gpt-image job') and clearly distinguishes this tool from the image-generation tools that create the job_id. It also names the exact sibling tools that return job_ids, so the agent can select this polling tool without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: call this with a job_id after any image tool hands off, poll until 'completed' or 'failed', and wait a few seconds between polls. It does not explicitly contrast with list_image_jobs or state when not to use it, but the polling context is clear enough to guide correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_edit_sessionsList Edit SessionsA
Read-onlyIdempotent

List active iterative-edit sessions (in-memory only, discarded on server restart). Useful to recover a session_id after a client reconnect; each entry also reports whether a turn is still running in the background and which job to poll for it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
sessionsYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only and non-destructive, and the description adds important behavioral context: sessions are in-memory only, lost on restart, and entries indicate whether a turn is running and which job to poll. This goes beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, and no filler. Every clause adds useful information about scope, lifecycle, or return contents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only list tool with an output schema already present, the description fully covers purpose, transient behavior, and the practical recovery scenario. Nothing necessary for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema coverage is 100%, so there is nothing for the description to add about parameters. The baseline for no-parameter tools applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List active iterative-edit sessions'. It clearly distinguishes this tool from image-related siblings by focusing on edit sessions and noting they are in-memory only.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states a clear use case: recovering a session_id after a client reconnect. It does not explicitly contrast with alternatives like start_edit_session or list_image_jobs, but the context is sufficient for an agent to know when this listing tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_image_jobsList Image JobsA
Read-onlyIdempotent

List the image jobs this server still remembers (newest first), with their state, prompt preview, and elapsed time. Useful after a reconnect or a context reset to recover a job_id instead of re-running — and paying for — a generation you already started. Running jobs are never expired; finished ones are kept up to GPT_IMAGE_2_JOB_MAX / GPT_IMAGE_2_JOB_TTL_MS and are lost on server restart.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
jobsYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true. The description adds valuable behavior beyond annotations: 'Running jobs are never expired; finished ones are kept up to GPT_IMAGE_2_JOB_MAX / GPT_IMAGE_2_JOB_TTL_MS and are lost on server restart.' This discloses retention and ephemerality, which the agent would not know from the schema or annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences. The first front-loads the primary action and result fields; the second adds usage guidance and retention policy. No filler or redundancy; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (not shown but indicated by 'Has output schema: true'), the description need not detail return types. It covers scope, ordering, fields included, retention, and the key use case. For a zero-parameter, read-only list tool, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description has nothing to explain about inputs. Schema coverage is 100% (trivially). Baseline for 0 parameters is 4, and the description adds no parameter-related details, which is fine because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List'), the resource ('image jobs'), and adds specific detail: 'newest first', with 'state, prompt preview, and elapsed time'. It differentiates from siblings like generate_image (creation) and get_image_job (single retrieval) by scope, even without naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use scenario: 'Useful after a reconnect or a context reset to recover a job_id instead of re-running — and paying for — a generation you already started.' It does not explicitly state when not to use it or name alternative tools, but the context strongly implies that get_image_job is for a known ID. Lacks explicit exclusion, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_edit_sessionStart Iterative Edit SessionA

Begin a stateful multi-turn edit session. Returns a session_id you then pass to continue_edit_session to iteratively refine the image (each turn uses the previous turn's output as the input). Use end_edit_session when done. The first turn hands off to a background job like every other image call: poll get_image_job for it, and the session_id arrives with its "completed" result.

ParametersJSON Schema
NameRequiredDescriptionDefault
maskNoOptional PNG mask — fully transparent pixels mark the editable region. Must match the first input image's dimensions and be <4MB. Accepts the same source types as `images`.
sizeNoOutput dimensions. "auto" (default), one of the presets "1024x1024", "1536x1024", "1024x1536", or a custom "WxH" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta.auto
userNoOptional end-user identifier forwarded to OpenAI for abuse monitoring. Pass a stable hashed user ID, not PII.
modelNoModel to use. One of "gpt-image-2.5-sunburst", "gpt-image-2.5-flare", "gpt-image-2"; defaults to "gpt-image-2.5-sunburst". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing.
imagesYes1–8 input images to seed the session (same source formats as edit_image).
promptYesImage description. gpt-image-2 handles very detailed prompts; use ALL CAPS or quote literal text you want rendered verbatim.
qualityNoEdit quality — same levels as generate.auto
backgroundNoBackground behavior. "opaque" forces a filled background; "transparent" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; "auto" lets the model pick.auto
output_dirNoAbsolute or relative directory where generated images should be written. Defaults to $GPT_IMAGE_2_OUTPUT_DIR or a per-project subfolder under the OS config dir. The directory is created if missing.
output_formatNoFile format. "png" (default, lossless), "jpeg" (smaller, lossy), "webp" (best compression). When omitted on continue_edit_session, the session's current format is kept.
filename_prefixNoShort label appended to the generated filename so you can find it later (e.g. "hero-banner"). Letters/digits/hyphens only; auto-sanitized.
output_compressionNoCompression level 0–100 for jpeg/webp outputs. Ignored for png. Defaults to 100 (minimal compression).

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolNo
turnNo
modelNo
notesNo
routeNo
stateNo
usageNo
imagesNoWritten image files — present once the job completed successfully.
job_idNoPresent on background hand-off — pass to get_image_job.
promptNo
appliedNo
poll_hintNo
requestedNo
session_idNoIdentifies the session for later continue_edit_session calls; present once a turn has landed.
started_atNo
async_after_msNo
prompt_previewNo
cost_usd_estimatedNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and destructive hints. The description adds meaningful behavioral context beyond those: the session is stateful, each turn consumes the previous output, the first turn is asynchronous, and the session_id arrives with the completed job result. This is useful and not contradicted by any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The description front-loads the core stateful behavior, then explains the session lifecycle and async job mechanics. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema and the existence of an output schema, the description sufficiently covers what an agent needs to invoke this tool correctly: session lifecycle, sibling tool routing, and the background-job polling behavior. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and every parameter already has a detailed schema description. The tool description itself adds no parameter-level detail, but none is needed since the schema carries the full burden. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool begins a stateful multi-turn edit session and names the sibling tools it connects to (continue_edit_session, end_edit_session). This makes it immediately distinguishable from one-off image tools like edit_image and generate_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly lays out the intended workflow: pass the returned session_id to continue_edit_session, and call end_edit_session when done. It does not explicitly contrast start_edit_session with one-off edit_image/generate_image, but the lifecycle guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.5.5
    • Changedcontinue_edit_session22 fields changed
      • removedInput schema / properties / background / default
        Removed value: -"auto"
      • changedInput schema / properties / background / description
        Previous value: -"Background behavior. \"opaque\" forces a filled background; \"auto\" lets the model pick. gpt-image-2 does NOT support transparent backgrounds — use a different model for that."New value: +"Background behavior. \"opaque\" forces a filled background; \"transparent\" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; \"auto\" lets the model pick. Omit to keep the background the session is already using."
      • changedInput schema / properties / background / enum
        Previous value: -[
        -  "auto",
        -  "opaque"
        -]New value: +[
        +  "auto",
        +  "opaque",
        +  "transparent"
        +]
      • changedInput schema / properties / model / description
        Previous value: -"Model to use. One of \"gpt-image-2\", \"gpt-image-2.5-flare\", \"gpt-image-2.5-sunburst\"; defaults to \"gpt-image-2\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."New value: +"Model to use. One of \"gpt-image-2.5-sunburst\", \"gpt-image-2.5-flare\", \"gpt-image-2\"; defaults to \"gpt-image-2.5-sunburst\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."
      • changedInput schema / properties / model / enum
        Previous value: -[
        -  "gpt-image-2",
        -  "gpt-image-2.5-flare",
        -  "gpt-image-2.5-sunburst"
        -]New value: +[
        +  "gpt-image-2.5-sunburst",
        +  "gpt-image-2.5-flare",
        +  "gpt-image-2"
        +]
      • removedInput schema / properties / quality / default
        Removed value: -"auto"
      • changedInput schema / properties / quality / description
        Previous value: -"Edit quality — same levels as generate."New value: +"Edit quality — same levels as generate. Omit to keep the quality the session is already using."
      • removedInput schema / properties / size / default
        Removed value: -"auto"
      • changedInput schema / properties / size / description
        Previous value: -"Output dimensions. \"auto\" (default), one of the presets \"1024x1024\", \"1536x1024\", \"1024x1536\", or a custom \"WxH\" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta."New value: +"Output dimensions. \"auto\" (default), one of the presets \"1024x1024\", \"1536x1024\", \"1024x1536\", or a custom \"WxH\" where both edges are multiples of 16, max edge ≤ 3840px, aspect ratio within 1:3–3:1, and total pixels 655,360–8,294,400. Outputs above 2K are beta. Omit to keep the size the session is already using."
      • addedOutput schema / properties / async_after_ms
        Added value: +{
        +  "exclusiveMinimum": 0,
        +  "type": "integer"
        +}
      • addedOutput schema / properties / images / description
        Added value: +"Written image files — present once the job completed successfully."
      • addedOutput schema / properties / job_id
        Added value: +{
        +  "description": "Present on background hand-off — pass to get_image_job.",
        +  "type": "string"
        +}
      • removedOutput schema / properties / notes / description
        Removed value: -"Caveats about how the request was served."
      • addedOutput schema / properties / poll_hint
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / prompt_preview
        Added value: +{
        +  "type": "string"
        +}
      • removedOutput schema / properties / route / description
        Removed value: -"Which API route served the request (edit tools only): \"direct\" = /v1/images/edits, \"responses\" = Responses-API fallback (one image per call, undercounted cost)."
      • addedOutput schema / properties / session_id / description
        Added value: +"Identifies the session for later continue_edit_session calls; present once a turn has landed."
      • addedOutput schema / properties / started_at
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / state
        Added value: +{
        +  "enum": [
        +    "running",
        +    "completed",
        +    "failed"
        +  ],
        +  "type": "string"
        +}
      • addedOutput schema / properties / tool
        Added value: +{
        +  "type": "string"
        +}
      • changedOutput schema / properties / usage / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "input_tokens": {
        -        "type": "number"
        -      },
        -      "input_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "output_tokens": {
        -        "type": "number"
        -      },
        -      "output_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "total_tokens": {
        -        "type": "number"
        -      }
        -    },
        -    "required": [
        -      "input_tokens",
        -      "output_tokens",
        -      "total_tokens"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "input_tokens": {
        +        "type": "number"
        +      },
        +      "input_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "output_tokens": {
        +        "type": "number"
        +      },
        +      "output_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "total_tokens": {
        +        "type": "number"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • removedOutput schema / required
        Removed value: -[
        -  "model",
        -  "prompt",
        -  "requested",
        -  "applied",
        -  "images",
        -  "usage",
        -  "cost_usd_estimated",
        -  "session_id",
        -  "turn"
        -]
    • Changededit_image6 fields changed
      • changedInput schema / properties / background / description
        Previous value: -"Background behavior. \"opaque\" forces a filled background; \"auto\" lets the model pick. gpt-image-2 does NOT support transparent backgrounds — use a different model for that."New value: +"Background behavior. \"opaque\" forces a filled background; \"transparent\" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; \"auto\" lets the model pick."
      • changedInput schema / properties / background / enum
        Previous value: -[
        -  "auto",
        -  "opaque"
        -]New value: +[
        +  "auto",
        +  "opaque",
        +  "transparent"
        +]
      • changedInput schema / properties / model / description
        Previous value: -"Model to use. One of \"gpt-image-2\", \"gpt-image-2.5-flare\", \"gpt-image-2.5-sunburst\"; defaults to \"gpt-image-2\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."New value: +"Model to use. One of \"gpt-image-2.5-sunburst\", \"gpt-image-2.5-flare\", \"gpt-image-2\"; defaults to \"gpt-image-2.5-sunburst\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."
      • changedInput schema / properties / model / enum
        Previous value: -[
        -  "gpt-image-2",
        -  "gpt-image-2.5-flare",
        -  "gpt-image-2.5-sunburst"
        -]New value: +[
        +  "gpt-image-2.5-sunburst",
        +  "gpt-image-2.5-flare",
        +  "gpt-image-2"
        +]
      • addedOutput schema / properties / images / description
        Added value: +"Written image files — present once the job completed successfully."
      • changedOutput schema / properties / usage / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "input_tokens": {
        -        "type": "number"
        -      },
        -      "input_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "output_tokens": {
        -        "type": "number"
        -      },
        -      "output_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "total_tokens": {
        -        "type": "number"
        -      }
        -    },
        -    "required": [
        -      "input_tokens",
        -      "output_tokens",
        -      "total_tokens"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "input_tokens": {
        +        "type": "number"
        +      },
        +      "input_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "output_tokens": {
        +        "type": "number"
        +      },
        +      "output_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "total_tokens": {
        +        "type": "number"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
    • Changedgenerate_image6 fields changed
      • changedInput schema / properties / background / description
        Previous value: -"Background behavior. \"opaque\" forces a filled background; \"auto\" lets the model pick. gpt-image-2 does NOT support transparent backgrounds — use a different model for that."New value: +"Background behavior. \"opaque\" forces a filled background; \"transparent\" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; \"auto\" lets the model pick."
      • changedInput schema / properties / background / enum
        Previous value: -[
        -  "auto",
        -  "opaque"
        -]New value: +[
        +  "auto",
        +  "opaque",
        +  "transparent"
        +]
      • changedInput schema / properties / model / description
        Previous value: -"Model to use. One of \"gpt-image-2\", \"gpt-image-2.5-flare\", \"gpt-image-2.5-sunburst\"; defaults to \"gpt-image-2\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."New value: +"Model to use. One of \"gpt-image-2.5-sunburst\", \"gpt-image-2.5-flare\", \"gpt-image-2\"; defaults to \"gpt-image-2.5-sunburst\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."
      • changedInput schema / properties / model / enum
        Previous value: -[
        -  "gpt-image-2",
        -  "gpt-image-2.5-flare",
        -  "gpt-image-2.5-sunburst"
        -]New value: +[
        +  "gpt-image-2.5-sunburst",
        +  "gpt-image-2.5-flare",
        +  "gpt-image-2"
        +]
      • addedOutput schema / properties / images / description
        Added value: +"Written image files — present once the job completed successfully."
      • changedOutput schema / properties / usage / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "input_tokens": {
        -        "type": "number"
        -      },
        -      "input_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "output_tokens": {
        -        "type": "number"
        -      },
        -      "output_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "total_tokens": {
        -        "type": "number"
        -      }
        -    },
        -    "required": [
        -      "input_tokens",
        -      "output_tokens",
        -      "total_tokens"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "input_tokens": {
        +        "type": "number"
        +      },
        +      "input_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "output_tokens": {
        +        "type": "number"
        +      },
        +      "output_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "total_tokens": {
        +        "type": "number"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
    • Changedget_image_job11 fields changed
      • addedOutput schema / properties / applied
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "background": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "output_format": {
        +      "type": "string"
        +    },
        +    "quality": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "size": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    }
        +  },
        +  "required": [
        +    "output_format"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / cost_usd_estimated
        Added value: +{
        +  "type": [
        +    "number",
        +    "null"
        +  ]
        +}
      • addedOutput schema / properties / model
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / notes
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedOutput schema / properties / poll_hint
        Added value: +{
        +  "description": "Present while a job is still running.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / prompt
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / requested
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "format": {
        +      "type": "string"
        +    },
        +    "n": {
        +      "type": "number"
        +    },
        +    "quality": {
        +      "type": "string"
        +    },
        +    "size": {
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "size",
        +    "quality",
        +    "n",
        +    "format"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / route
        Added value: +{
        +  "enum": [
        +    "direct",
        +    "responses"
        +  ],
        +  "type": "string"
        +}
      • addedOutput schema / properties / session_id
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / turn
        Added value: +{
        +  "type": "number"
        +}
      • addedOutput schema / properties / usage
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": false,
        +      "properties": {
        +        "input_tokens": {
        +          "type": "number"
        +        },
        +        "input_tokens_details": {
        +          "additionalProperties": false,
        +          "properties": {
        +            "image_tokens": {
        +              "type": "number"
        +            },
        +            "text_tokens": {
        +              "type": "number"
        +            }
        +          },
        +          "type": "object"
        +        },
        +        "output_tokens": {
        +          "type": "number"
        +        },
        +        "output_tokens_details": {
        +          "additionalProperties": false,
        +          "properties": {
        +            "image_tokens": {
        +              "type": "number"
        +            },
        +            "text_tokens": {
        +              "type": "number"
        +            }
        +          },
        +          "type": "object"
        +        },
        +        "total_tokens": {
        +          "type": "number"
        +        }
        +      },
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
    • Changedlist_edit_sessions8 fields changed
      • addedOutput schema / properties / sessions / items / properties / background
        Added value: +{
        +  "description": "Requested background; continue turns inherit it when omitted.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / sessions / items / properties / model
        Added value: +{
        +  "description": "Model the session uses; continue turns inherit it.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / sessions / items / properties / output_format
        Added value: +{
        +  "description": "Output format; continue turns inherit it when omitted.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / sessions / items / properties / pending_job_id
        Added value: +{
        +  "description": "Job to poll with get_image_job while state is \"running\".",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • addedOutput schema / properties / sessions / items / properties / quality
        Added value: +{
        +  "description": "Requested quality; continue turns inherit it when omitted.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / sessions / items / properties / size
        Added value: +{
        +  "description": "Requested size; continue turns inherit it when omitted.",
        +  "type": "string"
        +}
      • addedOutput schema / properties / sessions / items / properties / state
        Added value: +{
        +  "description": "\"running\" while a turn is in flight in the background; continue_edit_session refuses until it lands.",
        +  "enum": [
        +    "idle",
        +    "running"
        +  ],
        +  "type": "string"
        +}
      • changedOutput schema / properties / sessions / items / required
        Previous value: -[
        -  "session_id",
        -  "created_at",
        -  "updated_at",
        -  "turns",
        -  "last_prompt",
        -  "last_image_path"
        -]New value: +[
        +  "session_id",
        +  "created_at",
        +  "updated_at",
        +  "turns",
        +  "last_prompt",
        +  "last_image_path",
        +  "model",
        +  "size",
        +  "quality",
        +  "background",
        +  "output_format",
        +  "state",
        +  "pending_job_id"
        +]
    • Addedlist_image_jobs
    • Changedstart_edit_session17 fields changed
      • changedInput schema / properties / background / description
        Previous value: -"Background behavior. \"opaque\" forces a filled background; \"auto\" lets the model pick. gpt-image-2 does NOT support transparent backgrounds — use a different model for that."New value: +"Background behavior. \"opaque\" forces a filled background; \"transparent\" asks for alpha (PNG) — verified working for the gpt-image-2 family, and the origin still decides, so check applied.background; \"auto\" lets the model pick."
      • changedInput schema / properties / background / enum
        Previous value: -[
        -  "auto",
        -  "opaque"
        -]New value: +[
        +  "auto",
        +  "opaque",
        +  "transparent"
        +]
      • changedInput schema / properties / model / description
        Previous value: -"Model to use. One of \"gpt-image-2\", \"gpt-image-2.5-flare\", \"gpt-image-2.5-sunburst\"; defaults to \"gpt-image-2\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."New value: +"Model to use. One of \"gpt-image-2.5-sunburst\", \"gpt-image-2.5-flare\", \"gpt-image-2\"; defaults to \"gpt-image-2.5-sunburst\". The 2.5 variants accept the same parameters. Cost/token estimates assume gpt-image-2 pricing."
      • changedInput schema / properties / model / enum
        Previous value: -[
        -  "gpt-image-2",
        -  "gpt-image-2.5-flare",
        -  "gpt-image-2.5-sunburst"
        -]New value: +[
        +  "gpt-image-2.5-sunburst",
        +  "gpt-image-2.5-flare",
        +  "gpt-image-2"
        +]
      • addedOutput schema / properties / async_after_ms
        Added value: +{
        +  "exclusiveMinimum": 0,
        +  "type": "integer"
        +}
      • addedOutput schema / properties / images / description
        Added value: +"Written image files — present once the job completed successfully."
      • addedOutput schema / properties / job_id
        Added value: +{
        +  "description": "Present on background hand-off — pass to get_image_job.",
        +  "type": "string"
        +}
      • removedOutput schema / properties / notes / description
        Removed value: -"Caveats about how the request was served."
      • addedOutput schema / properties / poll_hint
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / prompt_preview
        Added value: +{
        +  "type": "string"
        +}
      • removedOutput schema / properties / route / description
        Removed value: -"Which API route served the request (edit tools only): \"direct\" = /v1/images/edits, \"responses\" = Responses-API fallback (one image per call, undercounted cost)."
      • addedOutput schema / properties / session_id / description
        Added value: +"Identifies the session for later continue_edit_session calls; present once a turn has landed."
      • addedOutput schema / properties / started_at
        Added value: +{
        +  "type": "string"
        +}
      • addedOutput schema / properties / state
        Added value: +{
        +  "enum": [
        +    "running",
        +    "completed",
        +    "failed"
        +  ],
        +  "type": "string"
        +}
      • addedOutput schema / properties / tool
        Added value: +{
        +  "type": "string"
        +}
      • changedOutput schema / properties / usage / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "input_tokens": {
        -        "type": "number"
        -      },
        -      "input_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "output_tokens": {
        -        "type": "number"
        -      },
        -      "output_tokens_details": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "image_tokens": {
        -            "type": "number"
        -          },
        -          "text_tokens": {
        -            "type": "number"
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "total_tokens": {
        -        "type": "number"
        -      }
        -    },
        -    "required": [
        -      "input_tokens",
        -      "output_tokens",
        -      "total_tokens"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "input_tokens": {
        +        "type": "number"
        +      },
        +      "input_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "output_tokens": {
        +        "type": "number"
        +      },
        +      "output_tokens_details": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "image_tokens": {
        +            "type": "number"
        +          },
        +          "text_tokens": {
        +            "type": "number"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "total_tokens": {
        +        "type": "number"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • removedOutput schema / required
        Removed value: -[
        -  "model",
        -  "prompt",
        -  "requested",
        -  "applied",
        -  "images",
        -  "usage",
        -  "cost_usd_estimated",
        -  "session_id",
        -  "turn"
        -]
  2. 7 tool updatesv0.3.0
    • First observedcontinue_edit_session
    • First observededit_image
    • First observedend_edit_session
    • First observedgenerate_image
    • First observedget_image_job
    • First observedlist_edit_sessions
    • First observedstart_edit_session

TDQS

A4.6/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct concern: one-shot generation, one-shot editing, multi-turn session lifecycle, and job polling/listing. The overlap between edit_image and continue_edit_session is clarified by the former being a single edit and the latter being an iterative turn within a session.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern: generate_image, edit_image, get_image_job, list_image_jobs, start_edit_session, continue_edit_session, end_edit_session, list_edit_sessions. The verb choices clearly reflect the action for each resource.

Tool Count5/5

Eight tools is well-scoped for the server's purpose: image generation and editing plus the supporting async-job and edit-session infrastructure. Each tool earns its place and there is no redundant surface.

Completeness5/5

The tool set covers the full workflow: generate and edit images, poll and list background jobs, and start/continue/list/end iterative edit sessions. No critical dead ends remain for the stated domain.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers