Skip to main content
Glama

openai-image-mcp

A local, stdio-transport MCP server that lets Claude Code and Claude Desktop generate and edit images with OpenAI's Images API (gpt-image-2 by default). Finished images are written into the current project (default ./assets/generated) so they can be used directly as site assets, and every tool call returns a small JPEG preview so the calling model can inspect the result and iterate. Images only; no video.

Tools

Tool

Purpose

generate_image

New image(s) from a text prompt.

edit_image

New image guided by 1–8 reference images (consistent style, modifications, optional mask).

list_generated_images

Files in the output directory with dimensions, size and the prompt from each .json sidecar.

Each generated image gets a <name>.json sidecar next to it holding the prompt, model, requested and final size, quality, background, output format, reference paths (for edits) and an ISO timestamp. Files are never overwritten: a taken name gets -2, -3, … appended.

Related MCP server: RunPod Image MCP Server

Prerequisites

  • Python ≥ 3.11 and, ideally, uv.

  • An OpenAI API key with access to the GPT Image models.

  • Organization verification. OpenAI gates GPT Image models behind API Organization Verification. If your organization has not completed it, calls fail with a 403 and this server reports "Your OpenAI organization is not verified for gpt-image-2 …". Verification can take a few minutes to propagate after it is approved.

Install

With uv (recommended):

git clone https://github.com/mostmark/openai-image-mcp.git
cd openai-image-mcp
uv sync
uv run openai-image-mcp        # starts the server on stdio; Ctrl-C to stop

With plain pip:

python -m venv .venv && source .venv/bin/activate
pip install -e .
python -m openai_image_mcp

Environment variables

Variable

Required

Default

Purpose

OPENAI_API_KEY

yes

API key. The server exits at startup with a clear message if it is missing.

OPENAI_BASE_URL

no

OpenAI default

Route calls through an OpenAI-compatible gateway.

OPENAI_IMAGE_MODEL

no

gpt-image-2

Model for both generation and edits.

IMAGE_OUTPUT_DIR

no

./assets/generated

Where images are written, relative to the server's working directory. Created on first write.

IMAGE_TIMEOUT_SECONDS

no

180

Per-request timeout. High-quality generations can take over a minute.

See .env.example. All logging goes to stderr; stdout is reserved for the MCP protocol.

Register with Claude Code

The server writes images relative to its working directory, and Claude Code starts MCP servers in the project directory, so a user-scoped registration puts assets into whichever project you are working in. Use uv run --project (not --directory) so the working directory stays your project:

claude mcp add openai-image --scope user \
  -e OPENAI_API_KEY=sk-... \
  -- uv run --project /absolute/path/to/openai-image-mcp openai-image-mcp

Equivalent project-scoped setup in the project's .mcp.json (commit it, keep the key out of it by using an environment reference):

{
  "mcpServers": {
    "openai-image": {
      "command": "uv",
      "args": ["run", "--project", "/absolute/path/to/openai-image-mcp", "openai-image-mcp"],
      "env": {
        "OPENAI_API_KEY": "${OPENAI_API_KEY}",
        "IMAGE_OUTPUT_DIR": "./public/images/generated"
      }
    }
  }
}

Verify with claude mcp list, then ask Claude Code to "generate a hero image for the landing page".

Register with Claude Desktop

Add to claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json, Windows: %APPDATA%\Claude\claude_desktop_config.json). Claude Desktop does not start servers inside a project directory, so set IMAGE_OUTPUT_DIR to an absolute path:

{
  "mcpServers": {
    "openai-image": {
      "command": "uv",
      "args": ["run", "--project", "/absolute/path/to/openai-image-mcp", "openai-image-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "IMAGE_OUTPUT_DIR": "/absolute/path/to/your/project/assets/generated"
      }
    }
  }
}

If Desktop cannot find uv, use its absolute path (which uv) as command.

Getting started: example prompts

Once the server is registered, you talk to Claude in plain language and it picks the tool and parameters. Everything below is something you can type into Claude Code (inside your project) or Claude Desktop. Generated files land in IMAGE_OUTPUT_DIR, and Claude sees a preview of each result, so you can follow up with "make it warmer" or "try a version without text".

First image

Generate a hero image for the landing page: a lighthouse on a rocky coast at dawn, soft pastel sky, flat vector illustration, no text. Save it as hero-lighthouse.

Claude calls generate_image with size: "hero", filename: "hero-lighthouse", and reports the path, for example assets/generated/hero-lighthouse.png, plus a preview.

Quick drafts, then a final render

Give me three low-quality draft ideas for a blog card about database migrations, landscape format.

generate_image with n: 3, quality: "low", size: "landscape". Pick one, then:

I like the second one. Regenerate it at high quality with the same prompt and call it blog-migrations.

Icons and logos with transparent backgrounds

Create a transparent PNG app icon of a paper airplane, single flat blue shape, centered, 512x512.

generate_image with background: "transparent", output_format: "png", size: "512x512". Transparent output needs png or webp; Claude gets a clear error if it asks for jpeg.

Make a webp banner for the newsletter header, 1400x500, our mascot waving, warm colours.

generate_image with output_format: "webp" and size: "1400x500". Dimensions are rounded to multiples of 16, and the result reports the size actually produced (here 1408x496). Anything wider than 3:1 (or taller than 1:3) is rejected with a message explaining the limit.

A consistent set of assets

First create a brand board: a style sheet with a navy, amber and off-white palette, flat vector illustration style, soft shadows, rounded geometry, no text. Landscape, high quality, filename brand-board.

Then, for every further asset:

Using assets/generated/brand-board.png as the reference, create a feature illustration of a developer at a desk with two monitors, landscape, filename feature-workspace.

Same reference: a transparent PNG icon of a cloud with an upload arrow, square, filename icon-upload.

Each of these is an edit_image call with reference_paths: ["assets/generated/brand-board.png"]. The section below shows how to make this the default via CLAUDE.md.

Editing an existing image

Take public/images/team-photo.jpg and turn it into a flat illustration in the same composition, keep the number of people and their poses.

Using assets/generated/hero-lighthouse.png as reference, produce a night-time version with a lit lamp and a starry sky, same framing.

Combine assets/logo.png and assets/generated/hero-lighthouse.png: place the logo bottom-right on the hero as a small watermark.

All three are edit_image calls; the last one passes both files in reference_paths. With a mask:

Using assets/generated/hero-lighthouse.png and the mask assets/masks/sky.png, replace the sky with a dramatic thunderstorm and leave the rest untouched.

Masking is prompt-guided rather than pixel-exact, so keep the description of the change in the prompt.

Picking up where you left off

What images have already been generated in this project?

list_generated_images returns each file's path, dimensions, size in KB, and the prompt from its sidecar, so a new session can reuse a brand board or earlier asset without regenerating it.

Wiring the result into a site

Generate an Open Graph image for the pricing page, 1200x630, and add it to the page's <meta> tags.

Claude generates the file, then edits your HTML or framework config to reference the returned path. Because the file is already in your project tree, no copying step is needed.

Size presets

Preset

Pixels

Typical use

square

1024x1024

Icons, avatars, social tiles (default)

landscape

1536x1024

Blog and card images

portrait

1024x1536

Posters, mobile screens

hero

1920x1088

Full-width hero sections

banner

1536x512

Headers, email banners

auto

model decides

When the prompt implies a shape

Explicit WIDTHxHEIGHT is also accepted. Each dimension is rounded to the nearest multiple of 16 (minimum 256); the aspect ratio must be within 1:3 to 3:1 and neither side may exceed 3840. The tool result reports the size actually produced. background="transparent" requires png or webp output.

Consistent assets

To keep one look across a whole set of assets, generate a brand board first and then pass it as a reference for every subsequent image. A suggested snippet for a project's CLAUDE.md:

## Image assets

- Images are generated with the `openai-image` MCP server into `assets/generated/`.
- Run `list_generated_images` first; if `assets/generated/brand-board.png` exists, reuse it.
- If it does not exist, create it once with `generate_image`
  (filename `brand-board`, size `landscape`, quality `high`): a style sheet showing our palette
  (#0B3D91 navy, #F2A900 amber, off-white), flat vector illustration style, soft shadows,
  rounded geometry, no text.
- Create every other asset with `edit_image`, passing
  `reference_paths: ["assets/generated/brand-board.png"]` and describing the new subject
  in the prompt ("In the style of the reference: …"). Use `background: transparent`
  with `png` for icons and logos.
- Inspect the returned preview; if it is off-brand, adjust the prompt and regenerate rather
  than editing the file by hand.

Errors

Common OpenAI failures are translated into short, actionable messages: missing organization verification, safety-system (content policy) rejections, rate limits and quota, and timeouts. The OpenAI request ID is included when available. Transient errors are retried by the OpenAI SDK's default policy; the server adds no retry loop of its own.

Development

uv sync                          # installs dev dependencies (pytest)
uv run pytest                    # unit tests + a stdio round-trip; no network access
uv run python scripts/smoke.py   # one real low-quality generation; needs OPENAI_API_KEY

Layout: src/openai_image_mcp/server.py (tool definitions), client.py (OpenAI calls and error translation), images.py (size handling, file naming, sidecars, previews).

Available Tools

3 tools
edit_imageA

Create a new image that is guided by one or more existing images (1-8 reference files) plus a text prompt, save it to the project's generated-assets directory, and return a preview. The model sees every reference image and the prompt together, so use this tool whenever the result must stay consistent with images you already have: to keep one visual style across a whole set of assets, pass the same brand-board or style-reference image in reference_paths for every asset and describe the new subject in the prompt (for example 'In the style of the reference: a hero illustration of a lighthouse at dawn'). Also use it to modify an image (change colours, add or remove elements, restyle it) or to combine several images into one. For a brand-new image with no reference, use generate_image. MASKING: mask_path is an optional PNG the same size as the first reference whose transparent pixels mark the region to change. Masking is prompt-guided rather than pixel-exact: the model treats the mask as strong guidance and may adjust areas near the edge, so always describe the desired change in the prompt too. Paths may be absolute or relative to the server's working directory. SIZE: pass a preset name or explicit WIDTHxHEIGHT. Presets: 'square' (1024x1024), 'landscape' (1536x1024), 'portrait' (1024x1536), 'hero' (1920x1088), 'banner' (1536x512), 'auto' (model picks). Explicit sizes are rounded to the nearest multiple of 16 (minimum 256); the aspect ratio must be between 1:3 and 3:1 and neither side may exceed 3840. The result text reports the size actually produced, so check it if the exact pixel size matters. BACKGROUND: 'transparent' produces an alpha channel and requires output_format 'png' or 'webp' (never 'jpeg'); use it for logos, icons and cut-out illustrations. QUALITY: 'low' is fast and cheap (good for drafts and iteration), 'high' is slow (can take over a minute) and detailed; 'auto' lets the model choose. RETURNS: for each image, a text block with the absolute file path, final width x height, model, quality and the path of a .json sidecar holding the prompt and settings, followed by a small JPEG preview (longest side 512px) so you can inspect the result. The full-resolution file is already on disk; reference it by the returned path (for example as a site asset).

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoPreset name (square, landscape, portrait, hero, banner, auto) or WIDTHxHEIGHT.square
promptYesWhat the new image should show; describe how the references should influence it.
qualityNoauto, low (fast draft), medium, or high (slow, detailed).auto
filenameNoOptional file name without extension. Defaults to a slug of the prompt plus a timestamp. Never overwrites.
mask_pathNoOptional PNG mask for the first reference; transparent pixels mark the area to change.
backgroundNoauto, opaque, or transparent (transparent requires png or webp).auto
output_formatNoFile format: png, jpeg, or webp.png
reference_pathsYes1-8 image files (png, jpeg, or webp) that guide the result.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and succeeds impressively. It discloses masking is 'prompt-guided rather than pixel-exact' with edge adjustment, size rounding to multiples of 16 with aspect-ratio and 3840 limits, background/output-format coupling ('transparent... requires output_format png or webp'), quality latency tradeoffs ('high... can take over a minute'), and that the full-resolution file is already on disk to be referenced by the returned path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool's complexity (8 params, masking nuance, size constraints, format coupling) justifies the length. It is well-sectioned with labeled blocks (MASKING, SIZE, BACKGROUND, QUALITY, RETURNS) and front-loads the core purpose and use cases before parameter details. Minor redundancy exists with schema fields like the non-overwrite filename behavior, so it is not zero-waste, but it is efficiently organized for its scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, this description covers everything an agent needs to invoke correctly: return format (absolute path, final dimensions, model, quality, .json sidecar, JPEG preview), size behavior verification ('check it if the exact pixel size matters'), mask semantics, format limits, and path resolution. For an 8-parameter generative tool, the completeness is exceptional.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds real value beyond the schema: exact preset dimensions ('square' 1024x1024, 'hero' 1920x1088), rounding/constraint rules for explicit sizes, the mask being 'the same size as the first reference,' and the quality speed implications. This meaningfully deepens the agent's understanding without repeating schema boilerplate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Create a new image that is guided by one or more existing images (1-8 reference files) plus a text prompt, save it... and return a preview.' It clearly distinguishes itself from the sibling generate_image by centering on reference-guided generation and explicitly framing style-consistency and image-modification use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'use this tool whenever the result must stay consistent with images you already have,' including style batching, modification, and combining images. It also names the exclusion condition and alternative: 'For a brand-new image with no reference, use generate_image.' This is model-level routing with nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate one or more brand-new images from a text prompt using OpenAI's image model, save them as files in the project's generated-assets directory, and return a small preview of each. Use this when no existing image needs to guide the result. If you need the new image to match the look of images you already have (a brand board, an earlier asset, a photo to modify), use edit_image instead and pass those files as references. Write prompts as a full description: subject, style, composition, lighting, colour palette, and any text that must appear. SIZE: pass a preset name or explicit WIDTHxHEIGHT. Presets: 'square' (1024x1024), 'landscape' (1536x1024), 'portrait' (1024x1536), 'hero' (1920x1088), 'banner' (1536x512), 'auto' (model picks). Explicit sizes are rounded to the nearest multiple of 16 (minimum 256); the aspect ratio must be between 1:3 and 3:1 and neither side may exceed 3840. The result text reports the size actually produced, so check it if the exact pixel size matters. BACKGROUND: 'transparent' produces an alpha channel and requires output_format 'png' or 'webp' (never 'jpeg'); use it for logos, icons and cut-out illustrations. QUALITY: 'low' is fast and cheap (good for drafts and iteration), 'high' is slow (can take over a minute) and detailed; 'auto' lets the model choose. RETURNS: for each image, a text block with the absolute file path, final width x height, model, quality and the path of a .json sidecar holding the prompt and settings, followed by a small JPEG preview (longest side 512px) so you can inspect the result. The full-resolution file is already on disk; reference it by the returned path (for example as a site asset).

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoHow many variations to generate (1-4).
sizeNoPreset name (square, landscape, portrait, hero, banner, auto) or WIDTHxHEIGHT.square
promptYesFull description of the image to create.
qualityNoauto, low (fast draft), medium, or high (slow, detailed).auto
filenameNoOptional file name without extension. Defaults to a slug of the prompt plus a timestamp. Never overwrites: a taken name gets -2, -3, ...
backgroundNoauto, opaque, or transparent (transparent requires png or webp).auto
output_formatNoFile format: png, jpeg, or webp.png

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and meets it impressively. It discloses concrete side effects: files land in the generated-assets directory, a .json sidecar is written, and 'a taken name gets -2, -3, ...' (never overwrites). It reveals timing behavior ('high' can take over a minute), size-correctness caveats (rounded to multiples of 16, aspect-ratio bounds, result text reports actual size), and transparency/format coupling. It also states the return contract in detail, compensating for the missing output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place; there is no filler. It is front-loaded with purpose and usage routing, then uses labeled SHOUTED sections (SIZE, BACKGROUND, QUALITY, RETURNS) that make dense parameter logic scannable. The closing note about referencing the saved path tells the agent what to do with the result, closing an otherwise easy-to-miss loop.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no annotations and no output schema, the description is complete. It covers the operation, side effects (file writes, sidecar), parameter semantics with constraints, timing expectations, and a detailed return contract (absolute path, dimensions, model, quality, preview). Nothing an agent needs to invoke it correctly and interpret results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial meaning beyond the schema: exact pixel dimensions for every size preset, rounding-to-16 rules, the 1:3–3:1 aspect-ratio bound, and the 3840-pixel maximum are absent from the schema. It adds the transparent-requires-png/webp constraint, explains quality trade-offs (low = fast/cheap for drafts, high = slow over a minute), and the filename never-overwrite behavior. This goes well beyond the baseline 3 justified by full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate one or more brand-new images from a text prompt... save them as files... return a small preview.' It explicitly contrasts with edit_image ('if you need the new image to match the look of images you already have'), so an agent can distinguish create-from-prompt from edit-from-reference without opening schemas. The phrase 'brand-new' signals no input image is required, which is the core differentiator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when/when-not rule: 'Use this when no existing image needs to guide the result. If you need the new image to match the look of images you already have... use edit_image instead and pass those files as references.' This names the sibling alternative and the exact condition that selects it — nothing is left to inference. The prompt-writing guidance (subject, style, composition, lighting, colour palette, text) is also actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_generated_imagesA

List the images this server has written to the project's generated-assets directory (IMAGE_OUTPUT_DIR, default ./assets/generated): absolute path, pixel dimensions, file size in KB, and the prompt that produced each one when its .json sidecar exists. Use it at the start of a session to discover assets (such as a brand board) created earlier, so you can reuse them or pass them to edit_image as references.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the responsibility for behavioral disclosure. It transparently explains that only images written by the server are listed, that the prompt is included only when a .json sidecar exists, and that the default directory is ./assets/generated. This gives the agent a good understanding of what to expect, though it does not describe edge cases like an empty directory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads the core listing behavior and return fields, while the second adds practical session-start guidance. Every clause contributes necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity, the presence of an output schema, and its sibling tools, the description is complete. It covers what is listed, where from, the conditional behavior, and how to use the results, so an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema coverage, so there is no parameter ambiguity. The description adds context by confirming there are no options and that calling it simply lists all generated assets, which is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: listing images written to the generated-assets directory. It also specifies exactly what data is returned (absolute path, dimensions, size, prompt) and frames the purpose as discovering reusable assets, which separates it from the sibling generation and editing tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: use it at the start of a session to discover previously created assets and to pass them as references to edit_image. It does not explicitly state when not to use it versus generate_image, but the intended workflow is clear enough to select this tool appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.1.0
    • First observededit_image
    • First observedgenerate_image
    • First observedlist_generated_images

TDQS

A4.8/5.0
Disambiguation5/5

The three tools have clearly distinct purposes: generate_image creates new images from text only, edit_image modifies or creates images guided by reference images, and list_generated_images discovers previously created assets. No overlap or ambiguity exists between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern: generate_image, edit_image, list_generated_images. The verbs clearly indicate the action and the nouns consistently refer to images.

Tool Count5/5

Three tools is well-scoped for an image generation server: create, edit, and list. Each tool earns its place and there is no bloat or missing core functionality.

Completeness5/5

The tool set covers the full lifecycle for this domain: generating new images, editing existing ones (including masking and style references), and listing previously generated assets for reuse. No critical dead end or missing operation is apparent.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/mostmark/openai-image-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server