chatgpt-desktop-image-mcp
Drives the signed-in OpenAI Codex desktop app over Chrome DevTools Protocol to generate images with ChatGPT. Provides tools to create PNG images from prompts, return the image bytes and dimensions, and optionally reuse a conversation thread to iterate on previous generations.
chatgpt-desktop-image-mcp
Give Claude Code, Claude Desktop, or any MCP client the ability to generate images with ChatGPT — on the subscription you already pay for.
It works by driving the already-signed-in Codex desktop app over the Chrome DevTools Protocol. No API key, and no Codex quota.

No API key.
No Codex quota. The app has a
Chat/Workmode switch.Workruns the Codex agent against your Codex allowance;Chatis plain ChatGPT. This project refuses to run unless the app reportscurrent mode: ChatGPT.Real PNGs land on disk, and the image comes back in the tool result.
Iterate.
thread: "reuse"keeps generations in one conversation, so "make it purple instead" works. The default is a fresh, isolated conversation.
The MCP server itself registers under the shorter name codeximg.
Why not the Images API?
Because you already pay for ChatGPT, and API image credits are metered separately.
Related MCP server: GPTImage
Why not automate chatgpt.com in Chrome?
That works too, but it needs a dedicated Chrome profile plus a one-time login, and
since Chrome 136 --remote-debugging-port is ignored on the default profile
(background). The desktop
app is already logged in, so there is nothing to set up.
Quick start
# 1. start the app with a debug port (closes any running instance first)
powershell -File launch-app.ps1
# 2. check the bridge is usable
node generate.mjs --status
# 3. generate
node generate.mjs "a single solid red circle on a plain white background" -o ./out -n circleRegister the MCP server
Claude Code. The claude CLI is often not on PATH — when Claude Code ships
inside the Claude desktop app it lives under a versioned path that changes on update:
& "$env:APPDATA\Claude\claude-code\<version>\claude.exe" mcp add -s user codeximg -- node C:/path/to/codeximg/mcp-server.mjsThat writes mcpServers into ~/.claude.json.
Claude Desktop. It is an MSIX package, so its config is not at
%APPDATA%\Claude\ — it is virtualised:
%LOCALAPPDATA%\Packages\Claude_<publisherid>\LocalCache\Roaming\Claude\claude_desktop_config.jsonAdd the server there and restart the app:
{
"mcpServers": {
"codeximg": {
"command": "node",
"args": ["C:/path/to/codeximg/mcp-server.mjs"]
}
}
}Use absolute paths, with forward slashes. Then ask the agent for an image: it calls
generate_image and gets the PNG back.
Before wiring anything up, it is worth running node smoke-test.mjs once — it proves
the whole chain without touching your client config.
MCP tools
Tool | Arguments | Notes |
|
| 15–60 s. Calls are serialized, because they share one browser window. Returns the path, the dimensions, and the image itself. |
| — | Read-only. Reports whether the port is open, whether the app is in Chat mode, whether the composer is present, and which conversation is open. |
thread: new vs reuse
|
| |
Conversation | a fresh one on every call | the one this tool last used |
Sees earlier generations | no | yes — "make it purple" works |
Sidebar rows created | one per call | one, total |
reuse is what lets an agent iterate on its own output. It is verified, not guessed:
the id of the open conversation is read from the composer before and after navigating,
and the call refuses to post if it did not land where it asked to.
node generate.mjs --reuse "make it purple instead" # iterate
node generate.mjs "a red bicycle" # fresh conversation (default)
node generate.mjs --focus # just open the saved one
Panels 2 and 3 are the same image with only the hue changed — identical dimensions,
identical pixel counts, different colour. Blue was produced in a fresh conversation
(the fallback path); purple with thread: "reuse". That is what carrying context buys
you, and it is why the two modes exist.
How it works
mcp-server.mjs / generate.mjs
│ ws://127.0.0.1:9222
▼
Codex desktop app (Electron / Chromium 154)
└── target: app://-/index.html ← the ChatGPT UI
├── 1. guard: button[aria-label^="Switch mode"] must say "current mode: ChatGPT"
├── 2. click [aria-label="New chat"] ← Temporary chats cannot generate images
├── 3. focus div.ProseMirror[role="textbox"], Input.insertText(prompt)
├── 4. click button[aria-label="Send"] ← Enter as fallback
├── 5. wait for img[alt^="Generated image"]
└── 6. capture bytes via Network.responseReceived + Network.getResponseBody
▼
out/chatgpt-<timestamp>.pngStep 6 exists because the rendered image src is a blob: URL, which cannot be
fetched from page context. Intercepting the network response is the reliable path.
Requirements
OS | Windows |
App | OpenAI Codex desktop app (MSIX package |
Node | 22+ (uses the global |
Shell | Windows PowerShell 5.1 (built into Windows) — only to launch the app |
Environment variables
Variable | Default | Purpose |
|
| Default output directory |
|
| DevTools port |
|
| ms to wait for an image |
|
| Set to |
|
| Set to |
Auto-launch behaviour
The MCP server starts the app for you, but only when it is not already running.
launch-app.ps1 force-closes any instance, so firing it at a running app would throw
away your session. If the app is running without a debug port, generate_image
fails with instructions instead of killing it.
Layout
Path | Purpose |
| MCP stdio server |
| Core: endpoint handling, mode guard, generation, locking |
| CLI wrapper |
| Dependency-free CDP client: |
| Starts the app via AUMID with |
| Speaks MCP to the server; |
| Read-only reconnaissance tools used to reverse the UI |
recon/
Read-only reverse-engineering tools. They exist so the selectors below can be
re-derived after a ChatGPT UI update instead of guessed at. Always start with
targets.mjs.
Tool | Purpose |
| Fingerprints every CDP target — the main window is the one whose URL is exactly |
| Dumps the composer, mode switch, model picker, and the buttons around the composer |
| Census of the page's |
| Sidebar sections, collapsed state, rendered rows, scroll geometry |
| Find elements by text or attribute, or print a matched element's ancestor chain |
| Click a sidebar conversation by title |
Gotchas discovered the hard way
Symptom | Cause / fix |
| Process is at Low integrity. Packaged-app COM activation needs Medium. |
| PowerShell will not cast the returned RCW to a |
App starts but is not logged in | You launched |
CDP port never opens | An instance was already running. Close it first. |
Message sends, but no image ever appears | You were in a Temporary chat — those do not support image generation. |
In-page |
|
| There are several. The main window is the one whose URL is exactly |
libuv assertion ( | Calling |
A saved conversation cannot be found again | The sidebar holds two unrelated row families — |
Posting into the wrong conversation | Never trust a title match on its own. |
How this was figured out
None of it is documented anywhere. docs/reverse-engineering.md
records the findings: that the Codex app is an Electron shell around the full ChatGPT
UI, that its Chat/Work switch is a quota and safety boundary, how to launch a
packaged Electron app with arguments, why the accessibility tree is unreliable, and
why the sidebar's conversation ids cannot be joined to the composer's.
Security
While the debug port is open, any local process can fully control that app. It binds
to 127.0.0.1 only, and the app must be relaunched with a special flag to open it at
all. Close the app when you are done.
Automating the app may conflict with OpenAI's terms depending on your use. This is built for your own account, at human-ish rates. For production or commercial volume, use the official Images API.
Available Tools
2 toolsgenerate_imageADestructive
Generate an image with ChatGPT and save it as a PNG on disk. Returns the file path, the dimensions, and the image itself.
Use for any visual asset the user asks for: a picture, illustration, icon, logo, mockup or photo. Do not call it speculatively, and do not use it for diagrams or charts that text already conveys.
Not idempotent: the same prompt twice yields two different images and two files. Takes 15-60 seconds, and calls are serialized because they share one application window. Writes a new PNG every time, and overwrites an existing file if "filename" collides with one. With thread "new" it also adds a conversation to the user's ChatGPT sidebar.
It drives the already-signed-in Codex desktop app over the DevTools protocol, so it needs no API key and consumes no Codex agent quota. It refuses to run if the app is in Work mode, because that would spend Codex usage.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Subject, style, palette, mood, composition. Be descriptive and specific. For transparency or text-in-image, say so explicitly in the prompt. | |
| thread | No | "new" (default) starts a fresh conversation: fully isolated, but each call adds a row to the user's sidebar and the model cannot see earlier generations. "reuse" posts into the conversation this tool last used, so you can iterate ("same image but blue"). Use "reuse" when refining, "new" for a fresh subject. | |
| filename | No | Output file name without the .png extension. Defaults to chatgpt-<timestamp>. | |
| output_dir | No | Directory to write into. Prefer an absolute path, or a path relative to this server's working directory. Defaults to <codeximg>/out. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the annotations: non-idempotency with a concrete example, 15-60s latency, serialization due to a shared window, file overwrite behavior, sidebar side-effect for thread 'new', auth model (drives signed-in Codex app, no API key, no quota), and a refusal condition in Work mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool does and returns, then usage, then behavioral caveats. Three tight paragraphs where every sentence carries actionable information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description states the return values (file path, dimensions, the image). Combined with latency, serialization, overwrite and refusal details, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a genuine cross-parameter side effect not in the schema: 'With thread "new" it also adds a conversation to the user's ChatGPT sidebar.' It also adds filename collision semantics beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Generate an image with ChatGPT and save it as a PNG on disk') with scope of use spelled out. Nothing overlaps with the only sibling, image_status, which is a status check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('any visual asset: a picture, illustration, icon, logo, mockup or photo') and when-not ('do not call it speculatively', 'not for diagrams or charts that text already conveys'). Routing is fully determined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_statusARead-onlyIdempotent
Report whether image generation is usable right now: whether the debug port is open, whether the app is in Chat mode, whether the composer is present, and which conversation is currently open.
Call this to diagnose a generate_image failure, or before the first generation of a session, rather than guessing at the cause. Read-only, and safe to call at any time.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is largely covered; the description's "Read-only, and safe to call at any time" partly restates that. It does add genuine behavioral value by enumerating the specific conditions probed, telling the agent what a pass/fail actually depends on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is checked and followed by the usage trigger. The enumerated checks are load-bearing (they define the diagnosis) rather than padding, so every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing what is returned, and it does so by naming the four status conditions. It stops short of describing the return shape (booleans vs. a status string), a small remaining gap for a notification-free diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter semantics to compensate for. The description correctly describes a no-argument probe rather than implying any configurable inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Report") plus the exact resource and scope: whether image generation is usable, enumerated as debug port, Chat mode, composer presence, and open conversation. This clearly separates it from the sibling generate_image, which performs the generation rather than checking readiness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call it: to diagnose a generate_image failure, or before the first generation of a session, and explicitly frames it as preferable to guessing at the cause. The alternative (calling generate_image blindly) is named, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.2.0- First observed
generate_image - First observed
image_status
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: generate_image produces an asset, image_status is a read-only diagnostic. The status tool is explicitly framed as the preflight/troubleshooting companion, so there is no risk of misselection.
Both names are snake_case and share the 'image' subject, but the conventions differ slightly: generate_image is verb_noun while image_status is noun_noun. Still predictable and readable, a minor deviation rather than an inconsistency.
Two tools is on the thin side, but the server's scope (generate one image, verify readiness) is genuinely narrow and both tools earn their place. Slightly under the typical 3-15 range but reasonable for the stated purpose.
Generation plus a status check covers the core workflow, and the tool notes edge cases like overwrites and serialization. Gaps are minor: no way to list or delete previously written PNGs, and no parameters for size/style beyond filename.
Maintenance
Related MCP Connectors
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Generate AI images and videos from any compatible MCP client.
Generate AI music via the Lacuna Music API from MCP clients like Claude Desktop & Code.
Related MCP Servers
- AlicenseAqualityBmaintenanceMulti-provider image generation MCP server that enables image generation from Claude Desktop, Claude Code, or any MCP client using OpenAI, Google Gemini, Stable Diffusion, or a placeholder provider.1071 PyPI1MIT
- AlicenseNot gradedqualityFmaintenanceImage generation for Claude Code via ChatGPT subscription token. No API key needed.3MIT
- AlicenseBqualityAmaintenanceLocal MCP server that uses Playwright browser automation to enable Claude Code to generate images, create variations, expand, and remove backgrounds via Adobe Firefly, requiring manual sign-in once.116 npmMIT
- AlicenseNot gradedqualityAmaintenanceGenerates and edits images through the Codex CLI using an existing ChatGPT subscription, with queued jobs, progress reporting, and artifact URLs for MCP clients.538 npmMIT