gpt-image-2-mcp
# gpt-image-2-mcp
An MCP server that exposes OpenAI's **gpt-image-2** family (default model `gpt-image-2.5-sunburst`) to any MCP client — Claude Desktop, Claude Code, Cursor, MCP Inspector, etc.
Eight tools:
| Tool | What it does |
|---|---|
| `generate_image` | text → image |
| `edit_image` | 1–8 reference images (+ optional mask) → image |
| `get_image_job` | poll a backgrounded generate/edit job by `job_id` |
| `list_image_jobs` | recover a `job_id` after a reconnect or context reset |
| `start_edit_session` | begin an iterative multi-turn edit |
| `continue_edit_session` | apply another refinement turn — previous output becomes the new input |
| `end_edit_session` | release a session |
| `list_edit_sessions` | show active sessions |
Every generated image is **saved to disk** and **returned inline** so the calling model sees it.
Image work is slow — tens of seconds, sometimes minutes — so every image tool
returns `{ job_id, state: "running" }` straight away and you collect the result
from `get_image_job`. See [Background jobs](#background-jobs-the-normal-path).
> **Batch work, or clients that cannot run a stdio server?** The same repository
> ships a queue-backed **remote service** with an HTTP API and a remote MCP endpoint
> (`/mcp`), built for thousands of images at a time — see
> [Remote service](#remote-service-http-api--remote-mcp).
## Requirements
- Node.js ≥ 20 for the **stdio server** (the published `gpt-image-2-mcp` package)
- Node.js ≥ 24 for the **remote service** and for this repository's test suite: the
SQLite driver is Node's built-in `node:sqlite`, which older versions do not ship
(the service image is `node:24` for the same reason)
- An OpenAI API key on an org with gpt-image-2 access (Organization Verification may be required)
## Install
**Nothing to build.** The server runs straight from npm over **stdio** — there is
no port to open, no URL to host, nothing to keep running in the background. The
client you configure below starts it on demand. Node.js ≥ 20 has to be available
to that client.
Add this to your MCP client's config (`claude mcp add`, Claude Desktop, Cursor,
VS Code, … — per-client details below):
```json
{
"mcpServers": {
"gpt-image-2": {
"command": "npx",
"args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
"env": {
"OPENAI_API_KEY": "sk-...",
"OPENAI_BASE_URL": "https://your-gateway.example/v1"
}
}
}
}
```
- `OPENAI_BASE_URL` is optional — only for a proxy/enterprise route. Drop the line
when you talk to api.openai.com directly.
- Put the key in this `env` block. Whether the server would also see variables
exported in your shell depends on the client: a GUI client (Claude Desktop) does
not, a client you launched from a shell (Claude Code, MCP Inspector) usually
does. The `env` block works in both cases.
- On Windows, some clients need `"command": "cmd", "args": ["/c", "npx", "-y", "@speed-fullbar/gpt-image-2-mcp"]`
instead (npx is a shell script there).
- The package is scoped (`@speed-fullbar/gpt-image-2-mcp`) because the bare name on
npm belongs to the upstream project this one is forked from (see
[Credits](#credits)); the executable it installs is `gpt-image-2-mcp`.
**Check the install before wiring it in.** This starts the server and sends one
MCP `initialize`:
```bash
printf '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"check","version":"0"}}}\n' \
| npx -y @speed-fullbar/gpt-image-2-mcp
```
Expect one JSON line containing `"serverInfo":{"name":"gpt-image-2-mcp","version":"…"}`
(the version you just installed). That proves the process starts and speaks MCP — it does not
prove the API key works: for that, call `generate_image` and poll `get_image_job` (see the
worked example below).
**From source** (for development, or a pinned local checkout):
```bash
pnpm install
pnpm run build
```
This produces `build/index.js`, the server entry point. Point the client at
`"command": "node", "args": ["/absolute/path/to/gpt-image-2-mcp/build/index.js"]`.
## Configure a client
Every client needs the same shape — a command that starts the server (the block
above) and an `env` block holding `OPENAI_API_KEY`. The sections below only show
where that block goes in each client; use `npx -y @speed-fullbar/gpt-image-2-mcp`
(no checkout needed) or the `node …/build/index.js` form from a pinned checkout.
### Claude Code
```bash
claude mcp add gpt-image-2 --env OPENAI_API_KEY=sk-... -- npx -y @speed-fullbar/gpt-image-2-mcp
```
Or add it to `~/.claude.json` (or a project `.mcp.json`) with the JSON shape below.
### Claude Desktop
Edit `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or `%APPDATA%\Claude\claude_desktop_config.json` (Windows):
```json
{
"mcpServers": {
"gpt-image-2": {
"command": "npx",
"args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
"env": {
"OPENAI_API_KEY": "sk-..."
}
}
}
}
```
### Cursor and other stdio clients
Same shape — `mcpServers` (Cursor, Claude, most wrappers) or `mcp.servers`
(VS Code):
```json
{
"mcpServers": {
"gpt-image-2": {
"command": "npx",
"args": ["-y", "@speed-fullbar/gpt-image-2-mcp"],
"env": { "OPENAI_API_KEY": "sk-..." }
}
}
}
```
### Local checkout instead of npx
```json
{
"mcpServers": {
"gpt-image-2": {
"command": "node",
"args": ["/absolute/path/to/gpt-image-2-mcp/build/index.js"],
"env": { "OPENAI_API_KEY": "sk-..." }
}
}
}
```
### MCP Inspector (interactive testing)
```bash
pnpm run inspect
```
Launches the official inspector UI pointed at your local build.
## Environment variables
| Var | Required | Purpose |
|---|---|---|
| `OPENAI_API_KEY` | ✅ | Auth |
| `OPENAI_BASE_URL` | | Override for proxies / enterprise routes |
| `OPENAI_ORG_ID` | | Forwarded as `organization` |
| `OPENAI_PROJECT_ID` | | Forwarded as `project` |
| `OPENAI_TIMEOUT_MS` | | Per-request timeout (default 600000 = 10 min, generous because image calls are slow). Raise it for a very slow proxy. |
| `OPENAI_MAX_RETRIES` | | SDK retries for 429/5xx (default 2). |
| `GPT_IMAGE_2_OUTPUT_DIR` | | Global default for where images are saved. Absolute paths used as-is, relative resolved from the **server's** working directory (`process.cwd()`), not yours. |
| `GPT_IMAGE_2_ALLOW_UNSAFE_OUTPUT_DIR` | | Set to `1` to let `output_dir` / `GPT_IMAGE_2_OUTPUT_DIR` point into OS-sensitive directories (`/etc`, `~/.ssh`, …), which are refused by default. |
| `GPT_IMAGE_2_MCP_DEBUG` | | Set to `1` to emit verbose debug logs on stderr. |
| `GPT_IMAGE_2_SESSION_MAX` | | Max concurrent in-memory edit sessions, LRU-evicted beyond this (default 20; `0` = no cap). |
| `GPT_IMAGE_2_SESSION_TTL_MS` | | Idle TTL before an edit session is swept (default 3600000 = 1h; `0` = never expire). |
| `GPT_IMAGE_2_ASYNC_AFTER_MS` | | How long `generate_image` / `edit_image` / the edit-session tools may block before handing off to a pollable job (default 1 = hand off immediately; a larger value waits inline, keep it below your MCP host's tool-call timeout; `<= 0` never backgrounds). See **Background jobs** below. |
| `GPT_IMAGE_2_JOB_MAX` | | LRU cap on finished image jobs kept for polling (default 20; `0` = no cap). |
| `GPT_IMAGE_2_JOB_TTL_MS` | | Age before a *finished* job becomes un-pollable (default 1800000 = 30 min; `0` = never expire). Running jobs are never expired or evicted. |
| `GPT_IMAGE_2_SANITIZATION` | | Cleaning applied to every returned image before it is written or sent inline: `metadata-v2` (default — container cleaning plus a re-encode), `metadata-v1` (container cleaning only, provider's pixels kept), or `off`. See **What an image carries**. |
| `OPENAI_FORCE_RESPONSES_EDITS` | | Set to `1` to pin edits to the Responses-API fallback route instead of `/v1/images/edits`. See **Edit routing** below. |
| `OPENAI_RESPONSES_EDIT_MODEL` | | Host model used by the Responses-API **fallback** edit route (default `gpt-4.1-mini`). See **Edit routing** below. |
## Where images go
Unless overridden, each tool writes to:
```
<OS config dir>/gpt-image-2-mcp/output/<project-name>-<hash>/
```
- macOS/Linux: `~/.config/gpt-image-2-mcp/output/<project>-<hash>/`
- Windows: `%APPDATA%\gpt-image-2-mcp\output\<project>-<hash>\`
`<project>-<hash>` is derived from the git root (if any) or the **server process's** current working directory — each project gets its own folder so generations don't collide.
⚠️ When a client launches the server (npx, Claude Desktop, …) that working directory is the client's, not your shell's, so the folder name is not predictable. If you care where images land, set `GPT_IMAGE_2_OUTPUT_DIR` (or pass `output_dir` per call) to an absolute path.
**Per-call override:** pass `output_dir: "/some/path"` to any tool.
Filenames look like `image-20260422-150301-a1b2c3.png`. If you pass `filename_prefix: "hero-banner"`, it becomes `image-20260422-150301-a1b2c3-hero-banner.png`.
## What an image carries
Every image this server returns — the file it writes *and* the inline copy in the tool result — is cleaned first.
Image origins attach provenance metadata to what they return: measured against two gateways in production,
every output PNG carried a `caBX` chunk, a ~22 KB C2PA/Content Credentials manifest naming the tool that made
it. Cleaning removes it before the bytes reach you:
| Profile | What it does | Pixels |
|---|---|---|
| `metadata-v2` (default) | container rewrite **and** a re-encode: the image is decoded and written again with this server's own deflate | byte-identical |
| `metadata-v1` | container rewrite only: provider metadata is dropped, the provider's compressed stream is kept verbatim | untouched |
The container rewrite keeps only what a decoder needs — PNG `IHDR/PLTE/tRNS/IDAT/IEND`, JPEG coding segments,
WebP `VP8`/`VP8L`/`ALPH`/`VP8X` — and drops everything else: C2PA/Content Credentials, EXIF, XMP, ICC, text
chunks, and any bytes appended after the end of the file. `metadata-v2` goes further and re-encodes, which also
removes the encoder fingerprint that survives in the compressed bytes; on a real 1.5 MP output that produced a
6-9% smaller file in ~1 s with `AE = 0` against the original, i.e. not one pixel differs.
A shape that cannot be re-encoded faithfully (16-bit, indexed, interlaced, above the pixel ceiling) is
delivered under `metadata-v1` and says so. Each entry of `structuredContent.images[]` carries a `sanitization`
receipt — `{ status, policy, source_bytes, removed, reencoded }` — plus a note in the summary when something
could not be cleaned:
```json
{ "file_path": "…/image-….png", "sanitization": { "status": "clean", "policy": "metadata-v2",
"source_bytes": 1595439, "removed": ["caBX"], "reencoded": true } }
```
`GPT_IMAGE_2_SANITIZATION=off` writes the provider's bytes verbatim and records `status: "disabled"`. An
unrecognised value is not guessed at: the image is still written (it was already paid for) and the result says
cleaning did not run. Two limits worth stating plainly — this is a **container-level** guarantee (it removes
what an image says about itself and, under v2, the encoder fingerprint, but makes no claim about watermarks
inside the pixels), and a file the profile cannot parse at all is still delivered, flagged `skipped` with the
reason rather than thrown away.
The queue-backed service applies the same profiles to what it stores; see **What a downloaded image carries**
in the service half of this README.
## What the tools return
**Image tools hand off first.** `generate_image`, `edit_image`,
`start_edit_session`, and `continue_edit_session` answer with a job hand-off:
```json
{ "job_id": "img-1761149123-a1b2c3d4", "state": "running", "poll_hint": "…", … }
```
Poll `get_image_job` with that `job_id` until it reports `state: "completed"`.
The completed poll is the result those tools *would* have returned inline:
1. An inline `ImageContent` block per generated image (so the LLM sees the image)
2. A text summary: applied settings, file path, token usage, estimated cost
3. `structuredContent` for programmatic consumers:
```json
{
"model": "gpt-image-2.5-sunburst",
"prompt": "…",
"requested": { "size": "auto", "quality": "auto", "n": 1, "format": "png" },
"applied": { "size": "1024x1024", "quality": "high", "background": "opaque", "output_format": "png" },
"images": [ { "file_path": "…", "filename": "…", "size_bytes": 123456, "mime_type": "image/png" } ],
"usage": { "input_tokens": …, "output_tokens": …, "total_tokens": …, "input_tokens_details": { … } },
"cost_usd_estimated": 0.2112,
"notes": []
}
```
`requested` is what you asked for, `applied` is what the origin actually used,
and `notes` (always present, `[]` when empty) explains any difference. A job
poll additionally carries `job_id`, `state`, `tool`, `started_at`,
`completed_at`, `elapsed_ms`, `error`, and `poll_hint` (while running). Session
tools also return `session_id` and `turn`.
### Worked example: generate, poll, then edit the result
```
generate_image prompt: "a red fox in a snowy pine forest, photorealistic"
→ { job_id: "img-…", state: "running" }
get_image_job job_id: "img-…"
→ state: "running" (elapsed 6s) # poll again in a few seconds
get_image_job job_id: "img-…"
→ state: "completed"
images: [ { file_path: "/home/me/.config/gpt-image-2-mcp/output/my-app-1a2b3c/image-20260921-101500-a1b2c3.png" } ]
edit_image prompt: "give the fox a small gold crown, keep everything else identical"
images: ["/home/me/.config/…/image-20260921-101500-a1b2c3.png"]
→ { job_id: "img-…", state: "running" } # feed file_path straight back in
get_image_job job_id: "img-…"
→ state: "completed", images: [ { file_path: "…-edited.png" } ]
```
## Background jobs (the normal path)
Image work is slow — tens of seconds through a fast route, sometimes minutes —
past the tool-call timeout of many MCP hosts. Every image tool
(`generate_image`, `edit_image`, `start_edit_session`, `continue_edit_session`)
therefore hands off immediately: the first response is
`{ job_id, state: "running" }` and the request continues in-process. Keep
calling `get_image_job` with that `job_id` every few seconds:
- `state: "running"` → keep polling
- `state: "completed"` → content is identical to a synchronous success:
inline images, summary text with file paths / usage / cost, and
`structuredContent` carrying the model, requested/applied settings, usage,
cost estimate, and route the call would have returned
- `state: "failed"` → `error` carries the reason; `content` mirrors the error
`GPT_IMAGE_2_ASYNC_AFTER_MS` controls how long a call may block before that
hand-off: `1` (the default, milliseconds) hands off immediately; a larger value
waits inline for that long — keep it below your MCP host's tool-call timeout;
`0` never backgrounds, so every call blocks until it finishes.
Job state lives in memory: finished jobs are kept up to `GPT_IMAGE_2_JOB_MAX`
and expire after `GPT_IMAGE_2_JOB_TTL_MS`; a server restart drops all jobs. If
you lose a `job_id` (reconnect, context reset), call `list_image_jobs` and pick
the running/completed job back up instead of re-running it and paying twice.
## Models
`generate_image`, `edit_image`, `start_edit_session`, and `continue_edit_session`
accept an optional `model` argument:
- `gpt-image-2.5-sunburst` — the default
- `gpt-image-2.5-flare`
- `gpt-image-2`
The 2.5 variants accept the same parameters (sizes, quality, formats). Omitting
`model` keeps using `gpt-image-2.5-sunburst`; in an edit session,
`continue_edit_session` inherits the model the session was started with unless
you override it per turn. Token/cost estimates assume gpt-image-2 pricing.
## Sizes
Default is `auto` (the model picks). You can pass:
- A preset: `1024x1024`, `1536x1024`, `1024x1536`
- Any custom `WxH` where:
- Both edges are multiples of 16
- Max edge ≤ 3840px (outputs above 2K are beta)
- Aspect ratio within 1:3 and 3:1
- Total pixels between 655,360 and 8,294,400
Invalid sizes fail **before** the API call with a clear error — no wasted requests.
**What actually comes back is the origin's call, not yours.** Some upstreams
normalize the controls: a gateway on a ChatGPT subscription (for example
[sub2api](https://github.com/Wei-Shaw/sub2api) routing to its OAuth/Codex
backend) ignores the requested size, quality, and format, and answers with its
own — commonly 1254×1254 PNG for the 2.5 models, whatever you asked for. The
tool reports what came back in `applied`, and adds a `notes` entry when it
differs from `requested`. To have size/format honored, point `OPENAI_BASE_URL`
at an API-key upstream (sub2api relays the fields unmodified on that path).
**Transparent PNGs work via `background: "transparent"`** (pair it with `output_format: "png"`). When the origin honors it, the response reports `background: "transparent"` and the file carries a real alpha channel (only the artwork is opaque).
The origin decides, though, and this gateway picks transparency from the **prompt/content** rather than your argument — verified at byte level on the same origin: a vector-logo prompt came back with 68% fully transparent pixels even when `background: "opaque"` was requested, and a scene prompt came back fully opaque (0 transparent pixels) even when `"transparent"` was requested. Always check `applied.background` and the file's alpha channel; a `notes` entry explains any disagreement. (For the other controls on this gateway, `pnpm run smoke:matrix` reports what your own origin honors — on the sub2api ChatGPT path `size`/`quality`/`format`/`n` are normalized away.)
## Iterative editing example
```
start_edit_session prompt: "A coastal lighthouse at dawn, photorealistic", images: ["./sketch.png"]
→ session_id: edit-1761149123-a1b2c3d4, turn 1, saved to …/session-…-turn1.png
continue_edit_session session_id: "edit-…-a1b2c3d4", prompt: "Make the sky more orange. Keep everything else the same."
→ turn 2
continue_edit_session session_id: "edit-…-a1b2c3d4", prompt: "Add a small boat on the horizon."
→ turn 3
end_edit_session session_id: "edit-…-a1b2c3d4"
```
A turn hands off like any other slow call: `start_edit_session` and
`continue_edit_session` answer with `{ job_id, state: "running" }` and the turn
finishes in the background. The `session_id` arrives with the job's
`"completed"` result (for a start, the session only exists once its first turn
has landed), and the session stays **busy** until then — continuing it early
returns an error, and `list_edit_sessions` reports `state: "running"` plus the
`pending_job_id` to poll. A turn is claimed the moment it starts, so two
overlapping `continue_edit_session` calls can't edit the same input image twice.
Every follow-up turn inherits what the session already uses — model, size,
quality, background, output format — unless that turn overrides it, and
`list_edit_sessions` reports those settings for each session.
Sessions are **in-memory only** and discarded on server restart — this is intentional (keeps the server stateless on the wire) and mirrors the Gemini MCP pattern.
## Image inputs for `edit_image` and `start_edit_session`
Accepts any mix of:
- Absolute path: `/Users/me/photo.png`
- Relative path: `./photo.png` (resolved from CWD)
- `file:///Users/me/photo.png`
- `https://example.com/photo.png` (downloaded, size-capped)
- `data:image/png;base64,iVBOR…`
Up to 8 images per call. Each ≤ 50MB. PNG/WEBP/JPG supported.
Pass the `file_path` from any earlier result straight back in `images` — that is how you iterate on an image, and how you combine several references into one composition. More than 8 images is refused before anything is uploaded (verified live: 8 accepted, 9 refused). `continue_edit_session` takes no images; it reuses the session's last output. The same summary is in the server's own `initialize` instructions, so a client sees it without reading this file.
## Cost guardrails
The server ships **no hard spending limits** — you should watch your OpenAI usage dashboard. Each tool result includes an estimated cost in USD computed from the token usage returned by the API, plus an approximate pre-flight estimate logged to stderr.
Rough per-image cost at common sizes:
| Quality | 1024×1024 | 1024×1536 / 1536×1024 |
|---|---|---|
| low | ~$0.006 | ~$0.005 |
| medium | ~$0.053 | ~$0.041 |
| high | ~$0.211 | ~$0.165 |
Custom sizes scale with pixel count. Edit calls additionally tokenize input images at high fidelity — large reference images are expensive.
## Edit routing
`edit_image`, `start_edit_session`, and `continue_edit_session` call `POST /v1/images/edits` directly. This is the canonical endpoint: it supports `n > 1`, masks, and returns accurate per-call token usage for cost estimation.
> **History:** at launch (2026-04-21) the endpoint rejected `gpt-image-2` (and `gpt-image-1.5`) with `400 Invalid value: 'gpt-image-2'. Value must be 'dall-e-2'.` — an OpenAI-side bug. Versions ≤ 0.2.0 of this server therefore routed edits through the Responses API by default. OpenAI fixed the endpoint silently in early May 2026 (verified live 2026-06-11), and since 0.3.0 the direct endpoint is the default again.
The Responses-API workaround is kept as a **fallback** (`src/utils/edit-via-responses.ts`):
- It engages automatically if the direct endpoint ever returns the launch-era 400 again (matched narrowly; the rejection is remembered for 10 minutes so only the first call in that window pays the failed attempt, then the direct endpoint is re-probed).
- Set `OPENAI_FORCE_RESPONSES_EDITS=1` to pin it explicitly.
- The legacy `OPENAI_USE_DIRECT_EDITS` toggle from 0.2.0 is deprecated and ignored (its only meaningful setting was `1` — opt into the direct endpoint, which is now the default).
Fallback mechanics: input images are uploaded via the Files API (`purpose: "vision"`), a cheap host model (default `gpt-4.1-mini`, override with `OPENAI_RESPONSES_EDIT_MODEL`) is forced to invoke the `image_generation` tool, the base64 result is extracted, and uploaded files are deleted afterwards.
**Fallback trade-offs versus the direct endpoint** (only apply when the fallback is active — the tool result carries `route: "responses"` and a note when they do):
- `n > 1` is not supported — the Responses path returns one image per call.
- Cost accounting undercounts — `usage` only reports the host chat model's text tokens; the image tool is billed separately (~$0.04–0.05 extra for a 1024×1536 medium edit).
- Masks still work — uploaded and referenced via `input_image_mask.file_id`.
## Remote service (HTTP API + remote MCP)
The same repository also ships a **queue-backed service** for batch work: it accepts
generation/editing requests over HTTP and MCP, stores them in SQLite, and processes
them with background workers against one or more image providers. The stdio server
above is untouched — this is a second entry point (`pnpm run service`).
```bash
cp .env.example .env # set SERVICE_TOKENS, KRILL_TOKEN, RUSTFS_* secrets
docker compose up -d --build
curl localhost:8787/v1/ready
```
Three containers: the service, Redis (BullMQ transport) and RustFS (S3-compatible
storage for uploaded inputs and generated outputs). The service's SQLite database
lives on the `app-data` volume; nothing else is required. All service variables are
documented in [`.env.example`](.env.example).
### Deploying on a single VM
`docker-compose.prod.yml` is the overlay for a machine that also runs other things
(a local gateway stack, a reverse proxy on `:80/:443`). It publishes the app and
rustfs on `127.0.0.1` only, mounts an image directory for `input_paths`, joins the
provider stack's docker network, and caps memory and log growth.
```bash
git clone https://github.com/FullBars/gpt-image-2-mcp.git /root/gpt-image-service
cd /root/gpt-image-service
cp .env.example .env && chmod 600 .env # then set the values below
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d --build
curl localhost:8788/v1/ready # the app listens on 127.0.0.1:8788
```
Settings that matter on a shared host (all in `.env`):
| Variable | Why |
|---|---|
| `SERVICE_TOKENS` | bearer token(s) for MCP and REST — long random, e.g. `ops:$(openssl rand -hex 32)` |
| `SERVICE_S3_PUBLIC_ENDPOINT` | where *clients* reach rustfs (the proxy's TLS URL, or `http://<host>:9002`); the signature covers the host, so it must match what clients send |
| `SERVICE_IMAGES_DIR` | host directory mounted at `/data/images` for `input_paths` (a 3000-image manifest of paths) |
| `SERVICE_PROVIDER_NETWORK` | docker network of a local gateway stack, so `SERVICE_SUB2API_BASE_URL=http://sub2api:8080/v1` stays on the host |
| `SERVICE_PUBLIC_URL` | the address clients use (e.g. `https://images.example.com/images`), so MCP instructions and tool results carry absolute URLs instead of paths the agent must prefix |
| `SERVICE_WORKER_CONCURRENCY` | worker slots; keep it near the vCPU count on a small box (a synchronous provider holds a slot for a whole generation) |
Reverse proxy (Caddy, host network — the app under a path prefix so an existing
service can keep `/v1/*`, and storage on its own port with the path and Host header
forwarded untouched so S3 signatures stay valid):
```caddyfile
images.example.com {
# MCP at /images/mcp, REST at /images/v1/...
handle_path /images/* {
# Compress the JSON/NDJSON API responses — a batch manifest is ~9x smaller
# (27 KB -> 3 KB for 24 items) and it is the one thing here that is text.
# Caddy's default list covers text/* and application/json but not NDJSON, so
# the content type is matched explicitly. The event stream is excluded: it
# must flush as it happens, never through a compressor.
@compressible not path /v1/events*
encode @compressible zstd gzip {
match {
header Content-Type *json*
}
minimum_length 1024
}
reverse_proxy 127.0.0.1:8788
}
}
https://images.example.com:9000 {
reverse_proxy 127.0.0.1:9002
}
```
Then `SERVICE_S3_PUBLIC_ENDPOINT=https://images.example.com:9000`, and open exactly
two inbound ports in the firewall/security group: the proxy's HTTPS port and the
storage port (clients upload and download straight from storage, so it cannot be
loopback-only). Upgrades are `git pull && docker compose -f docker-compose.yml -f
docker-compose.prod.yml up -d --build`; the SQLite record lives on the `app-data`
volume, so back that up if the job history matters.
**Retention**: stored images are deleted once they are older than
`SERVICE_RETENTION_DAYS` (default **3 days**); the task records stay. The sweep deletes
- the output of every job that succeeded more than the window ago, and
- input images older than the window that **nothing still needs** — an input is kept
while any job that references it is unfinished (queued/submitting/accepted/downloading)
or was created inside the window, however old that job is.
A purged job keeps its request, result metadata, provider and attempt history, moves to
the `expired` phase and stops being handed to a worker — it is still listed and
exported, just without a download URL. Re-uploading input content that was purged
re-opens the asset with a fresh lifetime. Retrying a job whose inputs are gone is
refused with `409` (and a job that discovers a missing input at submit time fails
immediately with `input_missing`, before any charge) instead of waiting forever. Set
`SERVICE_RETENTION_DAYS=0` to keep everything forever.
### Endpoints
| Method | Path | Purpose |
|---|---|---|
| POST | `/v1/assets/presign` | sign a direct upload (content addressed by `sha256`) |
| POST | `/v1/assets/:id/complete` | confirm the object exists, making the asset usable |
| GET | `/v1/assets/:id/download` | signed download URL |
| POST | `/v1/jobs` | submit one job (edit when `input_asset_ids` or `input_paths` is given) |
| GET | `/v1/jobs`, `/v1/jobs/:id` | list / inspect (add `?attempts=1` for provider receipts, `?resolutions=1` for the operator audit trail) |
| POST | `/v1/jobs/:id/cancel` | stop before it finishes (a submission already in flight cannot be recalled: its attempt is recorded as `ambiguous` and may still be billed) |
| POST | `/v1/jobs/:id/resolve` | handle a `needs_attention`/`failed`/`canceled` job: `retry` (needs `acknowledge_duplicate_charge: true`, grants exactly one extra submission) or `fail`; both close an attempt that was still in flight as `ambiguous` |
| POST | `/v1/batches`, `/v1/batches/import` | batch create / NDJSON import (accepts `request_key`, or an `Idempotency-Key` header, to make a retried batch a no-op) |
| GET | `/v1/batches/:id`, `/v1/batches/:id/export` | progress / per-item reconciliation as NDJSON (add `?sign=1` to get a `download_url` per line, instead of calling `/v1/jobs/:id` for each result) |
| GET | `/v1/providers` | provider health and cooldown |
| GET | `/v1/health`, `/v1/ready` | liveness / readiness (public; everything else needs a bearer token) |
| GET | `/v1/events` | live notifications (SSE): `job` events as jobs change, resumable via `Last-Event-ID`, filterable with `?batch_id=` / `?job_id=` |
| ALL | `/mcp` | remote MCP (streamable HTTP), same bearer token |
Images never travel through the service: clients PUT straight to a presigned URL and
download through a presigned URL, so a 3000-image batch does not stream through Node.
### Instead of polling
Two ways to stop sleeping between polls, both of which react the moment a job
settles:
- **Blocking reads.** `GET /v1/jobs/:id?wait_ms=25000` and
`GET /v1/batches/:id?wait_ms=25000` (and `wait_ms` on the MCP tools
`get_image_job` / `get_image_batch`) return as soon as the job is terminal or the
batch has nothing pending — or at the deadline, with the ordinary current state,
in which case the caller simply asks again. MCP clients have no push channel, so
this is the one to use from an agent.
- **A live stream.** `GET /v1/events` is server-sent events: each `job` event carries
the job, its batch and item index, and the phase it just moved into. Subscribe with
`?batch_id=` (or `?job_id=`) to see only your own work, and reconnect with the
standard `Last-Event-ID` header to replay what you missed — the underlying journal
is append-only, so nothing is skipped and nothing is delivered twice.
```bash
# Follow one batch's progress; stop when nothing is pending (Ctrl-C, or a counter).
curl -N "http://localhost:8787/v1/events?batch_id=$BATCH_ID" \
-H "Authorization: Bearer $TOKEN" -H 'accept: text/event-stream'
# event: ready <- the cursor you are starting from (fetch the snapshot now)
# data: {"cursor":8123}
#
# event: job
# id: 8124
# data: {"job_id":"job_…","batch_id":"batch_…","item_index":7,"phase":"succeeded","occurred_at":…,"finished_at":…,"purged_at":null}
```
An event says *that something changed*, not what the new state is: read
`GET /v1/jobs/:id` (or the batch export) for the result, the download URL and the
provider's applied parameters. The journal starts with migration `0007`, so history
predating it is only visible through the ordinary read routes.
### Files in, files out
The whole interaction is file-oriented, at both ends.
**Upload (client → service).** Either the file is already on the service's machine
and you pass its path, or you upload bytes straight to storage with a presigned URL:
```bash
# 1. ask for a slot (MCP: presign_image_asset returns the same values)
curl -sS -X POST http://localhost:8787/v1/assets/presign \
-H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d "$(jq -nc --arg sha "$(sha256sum fox.png | cut -d' ' -f1)" \
'{sha256:$sha, extension:"png", content_type:"image/png"}')"
# 2. PUT the file itself — bytes never pass through a prompt or a JSON payload
curl -sS -T fox.png -H 'content-type: image/png' "$UPLOAD_URL"
# 3. confirm, then use the asset id in input_asset_ids
curl -sS -X POST "http://localhost:8787/v1/assets/$ASSET_ID/complete" -H "Authorization: Bearer $TOKEN"
```
**Download (service → client).** `get_image_job` returns a download URL (valid for six
hours by default) and the
MCP reply also carries it as a resource link, so the client saves it as a file:
```bash
curl -sSL "$DOWNLOAD_URL" -o result.png
# bulk: one manifest for the whole batch, one line per item, each with its own URL
curl -sS "http://localhost:8787/v1/batches/$BATCH_ID/export?sign=1" -H "Authorization: Bearer $TOKEN"
```
Presigned URLs are signed for `SERVICE_S3_PUBLIC_ENDPOINT`, so set it to the address
*your clients* reach rustfs on (for a LAN deployment, e.g. `http://10.0.0.5:9000`) —
`localhost:9000` only works for a client on the same machine as the stack.
### What a downloaded image carries
Generated images arrive from the providers with provenance metadata attached. Measured on
the production bucket, every sampled output PNG carried a `caBX` chunk — a ~22 KB
C2PA/Content Credentials manifest naming the tool that made it (`c2pa.claim`, `OpenAI`,
`gpt-image`). Publication therefore cleans every image before storing it, under one of two
versioned profiles:
| Profile | What it does | Pixels |
|---|---|---|
| `metadata-v2` (default) | container rewrite **and** a re-encode: the image is decoded and written again with this service's own deflate | byte-identical |
| `metadata-v1` | container rewrite only: provider metadata is dropped, the provider's compressed stream is kept verbatim | untouched |
The container rewrite keeps only what a decoder needs, and drops everything a decoder does
not:
| Container | Kept | Dropped |
|---|---|---|
| PNG | `IHDR`, `PLTE`, `tRNS`, `IDAT`, `IEND` | `caBX` (C2PA), `tEXt`, `iTXt`, `zTXt`, `eXIf`, `iCCP`, `gAMA`, `sRGB`, `pHYs`, … and any bytes after `IEND` |
| JPEG | coding segments (`DQT`, `DHT`, `SOFn`, `DRI`, `SOS` + scans, `EOI`) | every `APPn` (EXIF/XMP/ICC/Adobe) and `COM` |
| WebP | `VP8`/`VP8L`, `ALPH`, `VP8X` (metadata flags cleared) | `ICCP`, `EXIF`, `XMP `, and the RIFF trailer |
`metadata-v1` copies the kept bytes verbatim, so the compressed pixel stream and its CRCs
are untouched. `metadata-v2` goes further: it inflates the IDAT stream, un-filters the
scanlines, re-filters them (Paeth) and re-deflates with fixed parameters (level 9,
`Z_FILTERED`) — which removes the encoder fingerprint that survives in the compressed
bytes, at the cost of holding one decoded image in memory. On a 1.5 MP output (2,231,063
bytes) that produced 2,056,971 bytes (**−8%**) in ~2.3 s with `AE = 0` against the
original, i.e. **not one pixel differs**; the result is deterministic and re-running it on
its own output is a no-op. Both are covered by tests that decode the result with an
independent decoder (`tests/helpers/png-fixtures.ts`).
A shape `metadata-v2` cannot re-encode faithfully is delivered under `metadata-v1` and
*records that it was*: 16-bit, indexed (palette), interlaced, above
`SERVICE_REENCODE_MAX_PIXELS` (default 16,777,216 = 4096²), or a decode/encode failure.
JPEG and WebP get the container profile only, with `reencode_reason:
"unsupported_format"` — re-encoding them would mean a second encoder (and, for JPEG,
generation loss).
The result reports what happened: `sanitization.status` is `clean` (with the `policy`
actually applied, `reencoded`, and the list of `removed` structures), `skipped` (with a
`reason`, for a file the profile cannot parse — it is still delivered and should be
treated as carrying metadata), or `disabled` (`SERVICE_OUTPUT_SANITIZATION=off`). A
crash-recovery adoption refuses an object that carries no receipt rather than adopting it
as this attempt's output. Re-encoding runs at most `SERVICE_REENCODE_MAX_CONCURRENT`
(3) images at a time per process; the work itself is in libuv's thread pool, so it does
not block the event loop.
Two limits are worth stating plainly. This is a **container-level** guarantee: it removes
what an image says about itself and, under v2, the encoding fingerprint — it makes no
claim about watermarks inside the pixels, and no re-encode can rule out steganographic
content. And a skip is not an error: the image was already generated and paid for, so it
is delivered with the reason recorded rather than thrown away.
Existing objects are brought forward with `scripts/clean-stored-outputs.mjs`, which skips
an object whose recorded `policy` already matches `SERVICE_OUTPUT_SANITIZATION` and
re-cleans one that does not (so raising the profile re-encodes the store on the next run).
The test suite verifies this with parsers written independently of the rewriter (see
`tests/helpers/png-fixtures.ts` and `tests/helpers/image-fixtures.ts`) and with real
encoder output committed as fixtures (baseline/progressive/greyscale/CMYK JPEG, lossy/
lossless/animated WebP).
### Downloading a batch quickly
Results are plain PNG/JPEG/WebP files served by rustfs through the proxy. Two facts
decide how fast they arrive:
- **Images are already compressed.** gzip on a 1.4 MB PNG saved 0.04% (1.43 MB →
1.42 MB) because PNG's own deflate has nothing left to give. Size comes from the
*output format and dimensions*, not from transport compression: `output_format:
"webp"` and a smaller `size` are the only real levers on the bytes.
- **A single connection is the slow unit.** Measured on the production host:
20 MB/s locally through the proxy, ~0.75 MB/s for one connection from the public
internet, and **5.4 MB/s with eight parallel connections** — the limit is per flow,
so parallelism scales roughly 7x until the host's total egress (~11 MB/s) is hit.
If your client is far away or lossy, download *items* concurrently rather than
segments of one file.
```bash
# One manifest for the whole batch, then pull it with 8 connections.
curl -sS "https://<host>/v1/batches/$BATCH_ID/export?sign=1" -H "Authorization: Bearer $TOKEN" \
| jq -r 'select(.download_url) | .download_url' > urls.txt
xargs -P 8 -n 1 curl -sS -C - -O < urls.txt # -C - resumes a partial file
# or, if aria2 is available:
aria2c -x8 -j8 -i urls.txt -d results/
```
Range requests work (`206`), so interrupted downloads resume rather than restart.
The manifest itself is gzipped by the proxy when the client asks for it (26.9 KB →
2.9 KB for 24 items, ~360 KB for 3000), which matters most on a slow link.
### Input images: files, not base64
Two ways to hand over a reference image, neither of which puts image bytes in a
payload:
- **`input_paths`** (with `mask_path`) — absolute paths on the machine running the
service. The service reads, hashes and stores each file itself, so a manifest of
3000 paths is a small NDJSON file. This is the one to use for a bulk edit run.
- **`input_asset_ids`** — for images the service cannot read (a client on another
machine): sign with `POST /v1/assets/presign`, PUT the file to the URL, POST
`/v1/assets/:id/complete`, then reference the asset id.
Both are content addressed, so re-listing the same file is free. A path that does not
exist, is relative, or is not a PNG/JPEG/WebP is rejected with a `400` naming the
problem, and a batch import reports it against that line without dropping the rest.
```bash
# One JSONL line per input image; the service reads the files (mount them into the
# container, e.g. `-v /data/images:/data/images:ro`).
find /data/images -name '*.png' | jq -c '{prompt:"clean up the background", input_paths:[.]}' > batch.jsonl
curl -sS -X POST http://localhost:8787/v1/batches/import \
-H "Authorization: Bearer $TOKEN" -H 'Idempotency-Key: cleanup-2026-09' \
--data-binary @batch.jsonl
```
### Remote MCP
```json
{
"mcpServers": {
"gpt-image-service": {
"type": "http",
"url": "https://images.internal.example/mcp",
"headers": { "Authorization": "Bearer <SERVICE_TOKENS value>" }
}
}
}
```
Eleven job-oriented tools (`submit_image_job`, `presign_image_asset`,
`confirm_image_asset`, `get_image_job`, `list_image_jobs`, `cancel_image_job`,
`resolve_image_job`, `submit_image_batch`, `import_image_batch_jsonl`,
`get_image_batch`, `list_image_providers`). They return job/batch references and
signed links rather than inline bytes, so the same tools work for one image and for a
batch of thousands. `submit_image_batch` / `import_image_batch_jsonl` take a
`request_key`, so a retried call returns the batch it already created.
**Generation:** `submit_image_job {prompt, size?, quality?, output_format?}` →
`get_image_job {job_id}` (returns `object_key` and a `download_url` that stays valid
for hours — `SERVICE_PRESIGN_TTL_SEC`, six by default — and
attaches the image as a resource link so the client fetches it as a file). The reply
also carries `applied_diffs`: whatever the provider changed on the way through.
Gateways are not exact — one may answer a `1024x1024` request with `1374x1145`, or
quietly raise `low` to `medium` — and a *silently* different image is worth knowing
about before a 3000-image run is signed off.
**Editing (prompt + reference image):** `submit_image_job {prompt, input_paths:
["/abs/path/fox.png"]}` when the file is on the service's machine — the service reads
it, so nothing about the image enters the conversation. When it is not (an MCP client
on a laptop, say): `presign_image_asset {sha256, extension, content_type}` → PUT the
file to the returned `upload_url` → `confirm_image_asset {asset_id}` →
`submit_image_job {prompt, input_asset_ids: [asset_id]}`. Up to 8 references, with an
optional `mask_asset_id` / `mask_path`.
**Batch from files:** `import_image_batch_jsonl` with one line per image, e.g.
`{"prompt":"…","input_paths":["/data/images/0001.png"]}` — a 3000-image manifest of
paths, no base64 anywhere.
MCP has no server→client push, so the `instructions` tell an agent with many jobs to
open the event stream itself: `curl -sN '<service>/v1/events?batch_id=…'` with the same
bearer token, resuming with `Last-Event-ID`. With `SERVICE_PUBLIC_URL` set the
instructions (and `get_image_batch`'s `results_path`) spell out the full address; without
it they say how to derive it from the `/mcp` URL you already have. Single jobs need no shell at all — `wait_ms` blocks the tool call
until the job settles.
**Tools for agents, REST for scripts.** The tool names, arguments and descriptions are
written for a model to choose from; they are the interactive surface, and they are free
to change. Anything a program drives — a script, a cron job, a CI step, a 3000-image
ingest — should call the REST API directly with the same bearer token: every tool is a
thin wrapper over those routes, so nothing is missing, and the paths above are the
supported client interface. The MCP `instructions` say the same thing to the agent, so
one that is asked to "write a script" reaches for `POST /v1/batches/import` instead of
mimicking tool calls.
To collect a finished batch, `GET /v1/batches/:id/export?sign=1` returns one NDJSON
line per item with a fresh `download_url` (also reported by `get_image_batch` as
`results_path`), so 3000 results do not need 3000 per-job calls — just one manifest to
loop over with the client's own downloader.
A quick check:
```bash
curl -s -X POST localhost:8787/mcp \
-H "Authorization: Bearer <token>" -H 'content-type: application/json' \
-H 'accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
```
### Live smoke test
`pnpm run smoke:service -- --yes --count=3 --providers=krill,sub2api` runs the real
pipeline (queue-shaped submission → executor → provider → object storage) against
real accounts for a few images, then reads each published object back and checks it
is a real image. It refuses to start without `--yes`.
### How the service spends money safely
- **Five billed submissions per image, across providers.** Polling, downloading,
waiting for capacity and switching provider do not consume the budget. The
wall-clock `SERVICE_JOB_DEADLINE_MS` budget starts with the first *authorised*
submission — never while waiting for capacity — so a saturated upstream cannot make a
3000-item batch fail items that were never submitted.
- **The authorisation is persisted before the request is sent**, and a job whose
outcome is unknown (a lost response, a crash mid-submission) becomes
`needs_attention` instead of being retried automatically — a duplicate submission
costs real money, so the service parks it for a human.
- **Idempotency**: `POST /v1/jobs` defaults to a request key derived from the request
content, so a retried HTTP call returns the existing job. Batch items are keyed by
position instead, so identical prompts in one batch stay separate jobs.
- **A batch is one billed decision covering thousands of items**, so it takes a client
key (`request_key` in the body, or an `Idempotency-Key` header). Replaying the same
key with the same payload returns the original batch and resumes only the items that
never made it into the database (`created` reports how many were new, `replayed`
that the batch already existed); the same key with a *different* payload is answered
with `409` rather than silently reusing another batch's id. Without a key, every call
is a new batch.
- **At-least-once dispatch, at-most-once billing**: wake-ups are stored in a
transactional outbox, each under its own durable intent id (which is also the
transport id), and re-created by a reconciler if Redis loses them — a *new* intent
is always delivered, even though the transport still retains the delivery it
replaced. The exclusive claim and lease fencing in SQLite keep one job to one
worker, and the lease is renewed for as long as an activation runs, so a slow
synchronous provider cannot be reclaimed and worked twice.
- **Free retries are retried, paid ones are not**: publication to object storage is
retried inside the activation (it costs nothing), and output keys are per attempt, so
a worker that lost its claim cannot overwrite the winner's image. If storage stays
down, a job whose provider still holds the task goes back to `downloading`; a job
whose bytes arrived inline — and therefore cannot be fetched again — is parked as
`needs_attention` instead of being regenerated and billed. A crash between the write
and the row update is recovered by adopting the object that attempt already
published.
- **The ceiling is enforced where the money is spent**: per-provider in-flight limits
are re-checked inside the transaction that authorises a submission, so two workers
that both read "one slot free" cannot both submit — the loser waits in `queued`
without spending an attempt. A provider-specific rejection (`permanent`: bad
parameters, moderation) is recorded but does not count against provider health, so a
healthy cheap tier is not taken out of rotation by one bad prompt. A rejection that is
really about the *provider* — a missing endpoint or model, which a gateway answers
with 404/405/501 (`provider_unavailable`) — is treated the other way round: the
provider steps aside for the hard-failure window, and the retry of that same job
lands on the next tier instead of dying with the request.
- **Decisions are auditable**: `resolve` records the action, the reason, the free-text
`operator` *and* the authenticated token that made the request (`actor`, which the
client cannot set) in an append-only trail (`GET /v1/jobs/:id?resolutions=1`, MCP
`get_image_job{include_resolutions:true}`, newest-last and stable within the same
millisecond). It survives the job succeeding and clearing its current error — and
`max_attempts` counts the extra submissions an operator authorised, so a job never
looks as if it ran past its own limit.
- **Explicit resolution**: a job that ends with an unknown outcome parks as
`needs_attention` and is never retried on its own. `POST /v1/jobs/:id/resolve`
(or the `resolve_image_job` MCP tool) is the only way past that: `action: retry`
requires `acknowledge_duplicate_charge: true`, grants exactly **one** extra billed
submission and re-queues the job immediately; `action: fail` closes it out with no
spend. Both record the operator and the note on the job.
- **Provider economics**: providers sit in preference tiers
(`SERVICE_SUB2API_PRIORITY` 0, `SERVICE_KRILL_PRIORITY` 10) — the subscription-backed
provider is filled first and a per-image billed provider only absorbs the overflow
once the cheaper tier is at its ceiling or out of rotation. A single transient
failure does not spill (`SERVICE_PROVIDER_FAILURE_THRESHOLD`, default 2), because
retrying the cheap provider is cheaper than paying another one. A deliberate mix is
available too: `SERVICE_KRILL_SPILL_RATIO=0.2` sends ~20% of requests to krill even
while sub2api has room.
- **Provider ceilings**: `SERVICE_KRILL_MAX_IN_FLIGHT` (15) / `SERVICE_SUB2API_MAX_IN_FLIGHT` (8)
cap how many jobs each provider holds at once — that capacity is what triggers the
spill to the next tier. Occupancy is derived from live job
rows, so a restart reconstructs it instead of oversubscribing the upstream; jobs
over the ceiling wait in `queued` and spend nothing.
## Troubleshooting
- **"OPENAI_API_KEY is not set"** — add it to the `env` block of your MCP config.
- **A tool result is a `job_id`, not an image** — that is the normal path. Poll `get_image_job` every few seconds until `state` is `completed` or `failed`. (Set `GPT_IMAGE_2_ASYNC_AFTER_MS=0` to make calls block instead.)
- **"Unknown or expired image job"** — the server restarted, or the job aged past `GPT_IMAGE_2_JOB_TTL_MS`. A finished result cannot be recovered; call `list_image_jobs` to see what is still held, and re-run only if the image is really gone (images already written to disk stay on disk).
- **`continue_edit_session` says the session is busy** — its previous turn is still running in the background. Poll the `pending_job_id` from `list_edit_sessions`, then continue. Sessions are also lost on server restart.
- **`403 / organization verification`** — gpt-image-2 may require Organization Verification on your OpenAI org. Check the dashboard.
- **`429`** — you hit the IPM (images per minute) cap for your tier. Lower `n`, or wait.
- **Size / quality / format came back different** — the origin normalized them (a ChatGPT-subscription gateway does this: see [Sizes](#sizes)). `applied` and `notes` report what happened; point `OPENAI_BASE_URL` at an API-key upstream to have your settings honored.
- **Image doesn't appear in the client** — check the file path in the text block; the image is saved regardless of inline display.
- **Images land in an unexpected directory** — the default is keyed on the *server's* working directory; set `GPT_IMAGE_2_OUTPUT_DIR` to an absolute path.
- **Client times out anyway** — raise `OPENAI_TIMEOUT_MS` (if the *proxy* is slow) or lower `GPT_IMAGE_2_ASYNC_AFTER_MS` (if your host's tool timeout is short).
- **Protocol disconnects silently** — something printed to stdout. Check `src/**/*.ts` — all logs must use `utils/logger.ts` (stderr). This is the single biggest MCP footgun.
## Getting help
- Bugs and feature requests: [github.com/FullBars/gpt-image-2-mcp/issues](https://github.com/FullBars/gpt-image-2-mcp/issues)
- Source, changelog, and the parameter-matrix write-up: [github.com/FullBars/gpt-image-2-mcp](https://github.com/FullBars/gpt-image-2-mcp)
- `GPT_IMAGE_2_MCP_DEBUG=1` on stderr shows every API call, route, and job transition.
## Smoke test against the real API
`tests/smoke/features/*.feature` describe the capability in Gherkin and run it
against the real endpoint — generation, editing, multi-turn sessions, and the
background-job queue:
```bash
pnpm run smoke # every feature (~4 min, a few cents of images)
pnpm run smoke -- editing # one feature, or scenarios whose name matches
```
Each step reports pass/fail on its own line; artifacts (the generated images)
land in a per-run temp directory printed at the top — set `SMOKE_OUT_DIR` to
put them somewhere stable. Requires `OPENAI_API_KEY` (and `OPENAI_BASE_URL` if
you go through a proxy).
Two deliberate choices: the suite runs queue-first (the default window), so
every call exercises the background-job path — set `GPT_IMAGE_2_ASYNC_AFTER_MS`
to a larger value to run the inline path instead (the queue feature is then
skipped); and it asserts what the upstream actually returned rather than
assuming — a proxy that ignores `output_format` or `size` shows up as a
reported observation, not a false failure. These tests are not part of
`pnpm test` or CI, which run without credentials.
`SMOKE_TRANSPORT=stdio pnpm run smoke` drives the built server
(`pnpm run build` first) over stdio — a real MCP client, the same way your host
calls it — instead of the default in-process transport.
### Which parameters the origin honors
`pnpm run smoke:matrix` runs a one-parameter-at-a-time matrix — a baseline plus
`model`, `size`, `quality`, `background`, `output_format`, and `n`, then the
same for `edit_image` — and compares what the API *reports* (`applied`) with
what the bytes *are* (container, dimensions, alpha). A control only counts as
honored when the response agrees **and** the bytes differ from the baseline row:
a gateway can echo a field back and still ignore it.
It writes `outcomes.json` (re-render with `--report <file>`) and
`parameter-matrix.md` into the artifacts directory. It is a study, not a gate:
it takes minutes and costs a few cents. Filter rows with an argument, e.g.
`pnpm run smoke:matrix -- size`.
## Development
```bash
pnpm run dev # tsx watch
pnpm run typecheck # tsc --noEmit for src, plus tsconfig.test.json for tests
pnpm run test # offline unit + integration tests (node:test, no credentials)
pnpm run build # compile to build/
pnpm run inspect # launch MCP Inspector
```
`pnpm run scale` (optionally `-- --jobs=10000`) seeds a throwaway SQLite file with
thousands of jobs and prints per-query timings plus `EXPLAIN QUERY PLAN` output, so a
missing index shows up as a scan-and-sort before it shows up in production.
`pnpm test` is the fast, offline suite — every test runs against a local HTTP
mock, and CI runs it on Node 24 (the version the service and its `node:sqlite`
driver require) with a Redis service container, so the queue integration tests run
too. `pnpm run smoke` (below) is the *live* suite and spends money.
### Releasing
Bump `version` in `package.json`, add a CHANGELOG section for the new version,
commit, then tag it:
```bash
git tag v0.5.1 && git push origin v0.5.1 # must match package.json
```
The `Release` workflow verifies the tag matches `package.json`, runs typecheck +
tests + build, publishes with `--provenance`, and opens the GitHub release. It
needs an `NPM_TOKEN` repository secret (npm automation token for the
`@speed-fullbar` scope). Running the workflow manually defaults to a `--dry-run`
publish, and tagging a version that is already on npm skips the publish (with a
notice) instead of failing — useful when a release was published by hand.
## Credits
This is a fork of [Borys520/gpt-image-2-mcp](https://github.com/Borys520/gpt-image-2-mcp)
by Borys Kusmirek (MIT) — the original server, tool surface, and research brief
are theirs, and the copyright notice in [LICENSE](LICENSE) is retained. This
fork adds the background job queue (`get_image_job`, `list_image_jobs`), edit
sessions that inherit their settings, the transparent-PNG and parameter-matrix
work, the BDD smoke suite, and the packaging/docs here.
## License
MIT
TDQS
Scored across 8 tools
Each tool targets a distinct concern: one-shot generation, one-shot editing, multi-turn session lifecycle, and job polling/listing. The overlap between edit_image and continue_edit_session is clarified by the former being a single edit and the latter being an iterative turn within a session.
All tool names follow a consistent snake_case verb_noun pattern: generate_image, edit_image, get_image_job, list_image_jobs, start_edit_session, continue_edit_session, end_edit_session, list_edit_sessions. The verb choices clearly reflect the action for each resource.
Eight tools is well-scoped for the server's purpose: image generation and editing plus the supporting async-job and edit-session infrastructure. Each tool earns its place and there is no redundant surface.
The tool set covers the full workflow: generate and edit images, poll and list background jobs, and start/continue/list/end iterative edit sessions. No critical dead ends remain for the stated domain.