Skip to main content
Glama
kuchris

chatgpt-desktop-image-mcp

chatgpt-desktop-image-mcp

Give Claude Code, Claude Desktop, or any MCP client the ability to generate images with ChatGPT — on the subscription you already pay for.

It works by driving the already-signed-in Codex desktop app over the Chrome DevTools Protocol. No API key, and no Codex quota.

codeximg generating an image

  • No API key.

  • No Codex quota. The app has a Chat / Work mode switch. Work runs the Codex agent against your Codex allowance; Chat is plain ChatGPT. This project refuses to run unless the app reports current mode: ChatGPT.

  • Real PNGs land on disk, and the image comes back in the tool result.

  • Iterate. thread: "reuse" keeps generations in one conversation, so "make it purple instead" works. The default is a fresh, isolated conversation.

The MCP server itself registers under the shorter name codeximg.

Why not the Images API?

Because you already pay for ChatGPT, and API image credits are metered separately.

Related MCP server: GPTImage

Why not automate chatgpt.com in Chrome?

That works too, but it needs a dedicated Chrome profile plus a one-time login, and since Chrome 136 --remote-debugging-port is ignored on the default profile (background). The desktop app is already logged in, so there is nothing to set up.

Quick start

# 1. start the app with a debug port (closes any running instance first)
powershell -File launch-app.ps1

# 2. check the bridge is usable
node generate.mjs --status

# 3. generate
node generate.mjs "a single solid red circle on a plain white background" -o ./out -n circle

Register the MCP server

Claude Code. The claude CLI is often not on PATH — when Claude Code ships inside the Claude desktop app it lives under a versioned path that changes on update:

& "$env:APPDATA\Claude\claude-code\<version>\claude.exe" mcp add -s user codeximg -- node C:/path/to/codeximg/mcp-server.mjs

That writes mcpServers into ~/.claude.json.

Claude Desktop. It is an MSIX package, so its config is not at %APPDATA%\Claude\ — it is virtualised:

%LOCALAPPDATA%\Packages\Claude_<publisherid>\LocalCache\Roaming\Claude\claude_desktop_config.json

Add the server there and restart the app:

{
  "mcpServers": {
    "codeximg": {
      "command": "node",
      "args": ["C:/path/to/codeximg/mcp-server.mjs"]
    }
  }
}

Use absolute paths, with forward slashes. Then ask the agent for an image: it calls generate_image and gets the PNG back.

Before wiring anything up, it is worth running node smoke-test.mjs once — it proves the whole chain without touching your client config.

MCP tools

Tool

Arguments

Notes

generate_image

prompt (required), filename, output_dir, thread

15–60 s. Calls are serialized, because they share one browser window. Returns the path, the dimensions, and the image itself.

image_status

—

Read-only. Reports whether the port is open, whether the app is in Chat mode, whether the composer is present, and which conversation is open.

thread: new vs reuse

new (default)

reuse

Conversation

a fresh one on every call

the one this tool last used

Sees earlier generations

no

yes — "make it purple" works

Sidebar rows created

one per call

one, total

reuse is what lets an agent iterate on its own output. It is verified, not guessed: the id of the open conversation is read from the composer before and after navigating, and the call refuses to post if it did not land where it asked to.

node generate.mjs --reuse "make it purple instead"   # iterate
node generate.mjs "a red bicycle"                    # fresh conversation (default)
node generate.mjs --focus                            # just open the saved one

three generations: green, blue, purple

Panels 2 and 3 are the same image with only the hue changed — identical dimensions, identical pixel counts, different colour. Blue was produced in a fresh conversation (the fallback path); purple with thread: "reuse". That is what carrying context buys you, and it is why the two modes exist.

How it works

mcp-server.mjs / generate.mjs
   │  ws://127.0.0.1:9222
   ▼
Codex desktop app (Electron / Chromium 154)
   └── target: app://-/index.html        ← the ChatGPT UI
        ├── 1. guard: button[aria-label^="Switch mode"] must say "current mode: ChatGPT"
        ├── 2. click [aria-label="New chat"]      ← Temporary chats cannot generate images
        ├── 3. focus div.ProseMirror[role="textbox"], Input.insertText(prompt)
        ├── 4. click button[aria-label="Send"]    ← Enter as fallback
        ├── 5. wait for img[alt^="Generated image"]
        └── 6. capture bytes via Network.responseReceived + Network.getResponseBody
   ▼
out/chatgpt-<timestamp>.png

Step 6 exists because the rendered image src is a blob: URL, which cannot be fetched from page context. Intercepting the network response is the reliable path.

Requirements

OS

Windows

App

OpenAI Codex desktop app (MSIX package OpenAI.Codex)

Node

22+ (uses the global WebSocket)

Shell

Windows PowerShell 5.1 (built into Windows) — only to launch the app

Environment variables

Variable

Default

Purpose

OUT_DIR

./out

Default output directory

CDP_PORT

9222

DevTools port

TIMEOUT

180000

ms to wait for an image

CODEXIMG_AUTOLAUNCH

1

Set to 0 to never auto-start the app

CODEXIMG_RETURN_IMAGE

1

Set to 0 to return only the file path, not the image bytes

Auto-launch behaviour

The MCP server starts the app for you, but only when it is not already running. launch-app.ps1 force-closes any instance, so firing it at a running app would throw away your session. If the app is running without a debug port, generate_image fails with instructions instead of killing it.

Layout

Path

Purpose

mcp-server.mjs

MCP stdio server

chatgpt.mjs

Core: endpoint handling, mode guard, generation, locking

generate.mjs

CLI wrapper

cdp.mjs

Dependency-free CDP client: attach, connect, evaluate, waitFor

launch-app.ps1

Starts the app via AUMID with --remote-debugging-port

smoke-test.mjs

Speaks MCP to the server; --generate also does a real run

recon/*.mjs

Read-only reconnaissance tools used to reverse the UI

recon/

Read-only reverse-engineering tools. They exist so the selectors below can be re-derived after a ChatGPT UI update instead of guessed at. Always start with targets.mjs.

Tool

Purpose

targets.mjs

Fingerprints every CDP target — the main window is the one whose URL is exactly app://-/index.html

inspect.mjs

Dumps the composer, mode switch, model picker, and the buttons around the composer

attrs.mjs [regex]

Census of the page's data-* attributes and sample values

sidebar-map.mjs

Sidebar sections, collapsed state, rendered rows, scroll geometry

find.mjs text|id|shape <v>

Find elements by text or attribute, or print a matched element's ancestor chain

open-thread.mjs "<title>"

Click a sidebar conversation by title

Gotchas discovered the hard way

Symptom

Cause / fix

IApplicationActivationManager returns E_ACCESSDENIED

Process is at Low integrity. Packaged-app COM activation needs Medium.

Cannot convert __ComObject to IApplicationActivationManager

PowerShell will not cast the returned RCW to a [ComImport] interface inline. Make the call inside a C# helper (launch-app.ps1 does).

App starts but is not logged in

You launched ChatGPT.exe by path, losing package identity and the virtualised APPDATA. Always go through the AUMID.

CDP port never opens

An instance was already running. Close it first.

Message sends, but no image ever appears

You were in a Temporary chat — those do not support image generation.

In-page fetch(img.src) fails

src is a blob: URL. Use the CDP Network domain instead.

Runtime.evaluate on the wrong target

There are several. The main window is the one whose URL is exactly app://-/index.html, not ?initialRoute=/avatar-overlay.

libuv assertion (async.c) on shutdown

Calling process.exit() while stdio handles are still closing. Let the event loop drain instead.

A saved conversation cannot be found again

The sidebar holds two unrelated row families — [data-sidebar-chatgpt-conversation-key] (a ChatGPT conversation) and [data-app-action-sidebar-thread-row] (an app-local thread) — and the composer id lives in a different namespace than either (local-chatgpt:<uuid> vs chatgpt:conversation:<uuid>). Only the title joins them, and rows render lazily, so you have to scroll the sidebar while looking.

Posting into the wrong conversation

Never trust a title match on its own. reuse re-reads the composer's conversation id after navigating, and refuses if it does not match what was requested.

How this was figured out

None of it is documented anywhere. docs/reverse-engineering.md records the findings: that the Codex app is an Electron shell around the full ChatGPT UI, that its Chat/Work switch is a quota and safety boundary, how to launch a packaged Electron app with arguments, why the accessibility tree is unreliable, and why the sidebar's conversation ids cannot be joined to the composer's.

Security

While the debug port is open, any local process can fully control that app. It binds to 127.0.0.1 only, and the app must be relaunched with a special flag to open it at all. Close the app when you are done.

Automating the app may conflict with OpenAI's terms depending on your use. This is built for your own account, at human-ish rates. For production or commercial volume, use the official Images API.

Available Tools

2 tools
generate_imageA
Destructive

Generate an image with ChatGPT and save it as a PNG on disk. Returns the file path, the dimensions, and the image itself.

Use for any visual asset the user asks for: a picture, illustration, icon, logo, mockup or photo. Do not call it speculatively, and do not use it for diagrams or charts that text already conveys.

Not idempotent: the same prompt twice yields two different images and two files. Takes 15-60 seconds, and calls are serialized because they share one application window. Writes a new PNG every time, and overwrites an existing file if "filename" collides with one. With thread "new" it also adds a conversation to the user's ChatGPT sidebar.

It drives the already-signed-in Codex desktop app over the DevTools protocol, so it needs no API key and consumes no Codex agent quota. It refuses to run if the app is in Work mode, because that would spend Codex usage.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesSubject, style, palette, mood, composition. Be descriptive and specific. For transparency or text-in-image, say so explicitly in the prompt.
threadNo"new" (default) starts a fresh conversation: fully isolated, but each call adds a row to the user's sidebar and the model cannot see earlier generations. "reuse" posts into the conversation this tool last used, so you can iterate ("same image but blue"). Use "reuse" when refining, "new" for a fresh subject.
filenameNoOutput file name without the .png extension. Defaults to chatgpt-<timestamp>.
output_dirNoDirectory to write into. Prefer an absolute path, or a path relative to this server's working directory. Defaults to <codeximg>/out.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: non-idempotency with a concrete example, 15-60s latency, serialization due to a shared window, file overwrite behavior, sidebar side-effect for thread 'new', auth model (drives signed-in Codex app, no API key, no quota), and a refusal condition in Work mode.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what the tool does and returns, then usage, then behavioral caveats. Three tight paragraphs where every sentence carries actionable information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description states the return values (file path, dimensions, the image). Combined with latency, serialization, overwrite and refusal details, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds a genuine cross-parameter side effect not in the schema: 'With thread "new" it also adds a conversation to the user's ChatGPT sidebar.' It also adds filename collision semantics beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Generate an image with ChatGPT and save it as a PNG on disk') with scope of use spelled out. Nothing overlaps with the only sibling, image_status, which is a status check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('any visual asset: a picture, illustration, icon, logo, mockup or photo') and when-not ('do not call it speculatively', 'not for diagrams or charts that text already conveys'). Routing is fully determined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image_statusA
Read-onlyIdempotent

Report whether image generation is usable right now: whether the debug port is open, whether the app is in Chat mode, whether the composer is present, and which conversation is currently open.

Call this to diagnose a generate_image failure, or before the first generation of a session, rather than guessing at the cause. Read-only, and safe to call at any time.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is largely covered; the description's "Read-only, and safe to call at any time" partly restates that. It does add genuine behavioral value by enumerating the specific conditions probed, telling the agent what a pass/fail actually depends on.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is checked and followed by the usage trigger. The enumerated checks are load-bearing (they define the diagnosis) rather than padding, so every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing what is returned, and it does so by naming the four status conditions. It stops short of describing the return shape (booleans vs. a status string), a small remaining gap for a notification-free diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and there is no parameter semantics to compensate for. The description correctly describes a no-argument probe rather than implying any configurable inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ("Report") plus the exact resource and scope: whether image generation is usable, enumerated as debug port, Chat mode, composer presence, and open conversation. This clearly separates it from the sibling generate_image, which performs the generation rather than checking readiness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call it: to diagnose a generate_image failure, or before the first generation of a session, and explicitly frames it as preferable to guessing at the cause. The alternative (calling generate_image blindly) is named, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.2.0
    • First observedgenerate_image
    • First observedimage_status

TDQS

A4.6/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: generate_image produces an asset, image_status is a read-only diagnostic. The status tool is explicitly framed as the preflight/troubleshooting companion, so there is no risk of misselection.

Naming Consistency4/5

Both names are snake_case and share the 'image' subject, but the conventions differ slightly: generate_image is verb_noun while image_status is noun_noun. Still predictable and readable, a minor deviation rather than an inconsistency.

Tool Count4/5

Two tools is on the thin side, but the server's scope (generate one image, verify readiness) is genuinely narrow and both tools earn their place. Slightly under the typical 3-15 range but reasonable for the stated purpose.

Completeness4/5

Generation plus a status check covers the core workflow, and the tool notes edge cases like overwrites and serialization. Gaps are minor: no way to list or delete previously written PNGs, and no parameters for size/style beyond filename.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers