GPT Image Playground MCP
GPT Image Playground MCP
让你的 Agent 通过浏览器使用 GPT Image Playground。
无需修改 Playground 源码。安装浏览器扩展并完成一次连接设置后,Agent 就可以提交图片生成任务、查看任务进度,并将生成的原图保存到本机。
本项目由 CookSleep/gpt_image_playground 提供 Playground 页面基础能力。本项目围绕页面上可见的操作提供 MCP 桥接,不包含也不修改 Playground 源码。
[!NOTE] Playground 仍然负责调用图像 API 并管理自己的登录状态。这个工具只负责传递任务和操作页面。
你可以这样使用
让 Agent 根据一句描述生成图片;
查询任务是否已经完成;
下载指定任务的原图;
让多个 Agent 共享同一个浏览器任务队列;
使用本地参考图参与生成。
对应的 MCP 工具为:generate_image、get_task_status 和 download_image。
Related MCP server: openai-gpt-image-1-mcp
工作方式
Agent
-> MCP stdio
-> 本机桥接服务
-> Chromium 扩展
-> Playground 页面
-> 生成并返回图片任务会由本机桥接服务按提交顺序逐个处理。图片生成没有固定的 20 秒完成期限;扩展领取任务后会持续等待页面完成,并保持连接状态。不同图片的生成时间可能不同,这不会导致任务重复提交或自动重排。
浏览器关闭后,页面任务无法继续执行。重新打开浏览器后,请先查看原任务状态,再决定是否提交新的任务。
开始之前
请准备好:
Node.js 18 或更高版本;
Chrome、Chromium 或其他支持 Manifest V3 的 Chromium 浏览器;
一个可以正常打开 GPT Image Playground 的浏览器页面;
一个支持 MCP 的 Agent 客户端。
快速开始
1. 获取并构建项目
从 GitHub 获取项目:
git clone https://github.com/PEKI7483/image-playground-mcp.git
cd image-playground-mcp
npm install
npm run build如果你将项目放在其他目录,后文中的 <项目根目录> 就是克隆出的 image-playground-mcp 目录。
构建完成后,dist/ 目录中应包含 server.js、bridge.js 和 bridgeMain.js。
2. 生成连接码
连接码用于保护本机桥接服务,不是 Playground API Key。请生成一个随机连接码,并妥善保管:
node -e "console.log(require('node:crypto').randomBytes(32).toString('hex'))"扩展和每个 Agent 客户端都要使用同一个连接码。请不要将连接码提交到代码仓库或公开日志中。
默认桥接地址为 http://127.0.0.1:8787。如果端口已被占用,可以改用其他端口,例如 8790,但扩展和所有 Agent 必须保持一致。
3. 安装浏览器扩展
打开 Chrome 或 Chromium,访问
chrome://extensions。开启右上角的“开发者模式”。
点击“加载已解压的扩展程序”。
选择:
<项目根目录>/extension点击浏览器工具栏中的扩展图标,打开扩展小窗口。
点击顶部的“连接设置”。
将“图片工具地址”设置为
http://127.0.0.1:8787,将“连接码”设置为刚才生成的连接码。点击“保存并检查”。连接成功后,窗口会回到“运行概览”。
设置过程只在扩展小窗口内完成,不会打开新的标签页。扩展保存的是桥接地址和连接码,不保存 Playground API Key。
4. 选择你的 Agent
先完成扩展设置,再选择下面与你使用的 Agent 对应的方式。所有 Agent 可以共享一个桥接服务,但必须使用相同的连接码和端口。
Agent | 添加方式 | 适合的配置位置 |
Codex CLI / Codex 应用 |
| 共享 MCP 配置 |
Claude Code |
| 用户或项目配置 |
Gemini CLI | MCP 设置文件 |
|
Cursor | MCP 设置或 JSON |
|
Cline | MCP 设置 | MCP Servers 页面 |
Roo Code | MCP 设置 | MCP Servers 页面 |
Windsurf | MCP 设置或 JSON | MCP 配置文件 |
Claude Desktop | JSON 配置 |
|
Codex CLI 和 Codex 应用
Codex CLI、Codex 应用和 IDE 扩展共享 MCP 配置。将 <项目根目录> 和连接码替换为实际值后执行:
codex mcp add gpt-image-playground \
--env MCP_BRIDGE_TOKEN=替换为连接码 \
--env MCP_BRIDGE_PORT=8787 \
-- node "<项目根目录>/dist/server.js"检查配置:
codex mcp list更多选项请参阅 Codex MCP 官方文档。
Claude Code
claude mcp add --transport stdio gpt-image-playground \
--env MCP_BRIDGE_TOKEN=替换为连接码 \
--env MCP_BRIDGE_PORT=8787 \
-- node "<项目根目录>/dist/server.js"Gemini CLI
Gemini CLI 通常通过 MCP 设置文件配置本地服务。请在其 settings.json 的 mcpServers 中加入以下内容:
{
"gpt-image-playground": {
"command": "node",
"args": ["<项目根目录>/dist/server.js"],
"env": {
"MCP_BRIDGE_TOKEN": "替换为连接码",
"MCP_BRIDGE_PORT": "8787"
}
}
}Cursor、Cline、Roo Code、Windsurf 和 Claude Desktop
这些客户端可以在 MCP 设置页面中添加本地 STDIO 服务,也可以编辑对应的 JSON 配置。将下面的 gpt-image-playground 对象合并到已有的 mcpServers 中,请保留其他服务器配置:
{
"gpt-image-playground": {
"command": "node",
"args": ["<项目根目录>/dist/server.js"],
"env": {
"MCP_BRIDGE_TOKEN": "替换为连接码",
"MCP_BRIDGE_PORT": "8787"
}
}
}常见位置如下:
Cursor:项目目录中的
.cursor/mcp.json,或Cursor Settings > MCP;Cline:扩展中的
MCP Servers页面;Roo Code:扩展中的
MCP Servers页面;Windsurf:MCP 设置页面或其 MCP 配置文件;
Claude Desktop:
claude_desktop_config.json中的mcpServers。
保存后,请重启客户端或重新加载 MCP 配置。
MCP 客户端启动 dist/server.js 时,会通过 stdio 提供 MCP 协议,同时启动本机桥接服务。通常不需要手工运行 npm start。
启动并确认连接
打开 GPT Image Playground,并保持普通画廊模式页面处于打开状态。
确认提示词输入框和图片上传控件在页面中可见。
启动或重启 Agent 客户端,使它加载刚才的 MCP 配置。
点击浏览器工具栏中的扩展图标。
在“运行概览”中确认桥接服务、扩展连接和 Playground 页面均处于正常状态。
也可以使用健康检查接口确认连接:
curl -H 'X-MCP-Bridge-Token: 替换为连接码' \
http://127.0.0.1:8787/v1/health健康状态示例:
{
"bridge": "running",
"extensionConnected": true,
"playgroundTabCount": 1,
"queueLength": 0
}调用示例
生成图片
调用 generate_image,至少传入一个提示词:
{
"prompt": "一只戴红色围巾的橘猫,工作室摄影风格"
}工具会等待页面完成,不以 20 秒作为完成期限。多个 Agent 可以同时连接同一个桥接地址,任务会进入同一个 FIFO 队列并逐个执行。
使用参考图
参考图可以直接通过 MCP 参数传入,不需要手工点击 Playground 的上传控件:
{
"prompt": "保留参考图的构图,改成水彩插画",
"reference_image_paths": [
"/absolute/path/reference.png",
"/absolute/path/style.jpg"
]
}支持 PNG、JPEG/JPG、WebP、GIF 和 AVIF;最多 16 张,单张不超过 8 MiB,总大小不超过 24 MiB。MCP 服务读取这些本地文件后,扩展会注入 Playground 已有的多文件上传控件,图片处理仍由 Playground 页面完成。
下载生成结果
generate_image 成功后,使用返回的 task_id 调用 download_image:
{
"task_id": "上一步返回的 task_id",
"output_path": "/absolute/path/generated.png",
"image_index": 0
}output_path 必须是绝对路径。已有文件只有在内容完全相同时才会幂等成功;如果内容不同,工具会提示冲突,不会覆盖原文件。
详情弹窗中的原图可能比任务卡片缩略图更晚加载。扩展会兼容普通 img 图片,最多等待 60 秒,并在下载完成或失败后关闭详情弹窗。
安全与隐私
[!IMPORTANT] 你的 Playground 登录状态和 API Key 仍由 Playground 自己管理。这个工具只操作页面,不会读取或保存 Cookie、
localStorage、IndexedDB、页面 JavaScript 变量或 API 响应中的密钥。
连接码只用于保护本机桥接接口,不是 Playground API Key;
桥接服务默认只监听本机回环地址;
图片生成通过 Playground 的可见页面完成,不直接调用图片 API;
参考图由 MCP 服务读取后交给 Playground 页面处理;
多个 Agent 可以共用服务,但应使用同一个连接码和端口。
常见问题
扩展显示“服务不可达”
请按下面的顺序检查:
Agent 客户端是否已经启动并加载了
dist/server.js?扩展中的桥接地址是否为
http://127.0.0.1:8787?扩展中的连接码是否与 Agent 配置完全一致?
修改扩展代码后,是否在
chrome://extensions点击了“重新加载”?是否已经打开 GPT Image Playground 的普通画廊页面?
如果仍然没有连接,请使用上面的健康检查命令,并保留返回结果,便于进一步定位。
Playground 页面数量为 0
请打开 GPT Image Playground 的普通画廊模式页面,并确认提示词输入框可见。扩展不会通过直接调用图像 API 工作,也不会处理没有对应页面控件的页面。
生成时间较长,Agent 似乎没有响应
图片生成时间会因提示词、参考图和页面状态而不同。请先查看 Playground 任务卡片和扩展运行概览,不要重复点击生成。已领取的任务会继续等待页面完成,必要时可以使用 get_task_status 查询。
参考图没有出现
请确认:
文件路径是绝对路径;
文件格式受支持;
单张图片不超过 8 MB,总大小不超过 24 MB;
图片数量不超过 16 张;
Playground 页面中存在多文件上传控件。
下载原图失败
请先确认任务已经完成、任务卡片仍然可见,并使用新的绝对输出路径。下载读取的是详情弹窗中的原图,不会通过浏览器存储读取图片。
配置项
环境变量 | 默认值 | 说明 |
| 启动时随机生成 | MCP 与扩展之间的本机连接码;多个 Agent 需要固定为同一个值 |
|
| 本机桥接端口 |
|
| 默认关闭;开启后只会释放心跳失联任务,不会自动重试生成 |
如需单独启动桥接服务进行调试,可以使用:
MCP_BRIDGE_TOKEN='替换为连接码' MCP_BRIDGE_PORT=8787 npm run bridge通常不需要手工运行这条命令。MCP 客户端启动 dist/server.js 时,会同时启动本机桥接服务。
反馈与帮助
如果你遇到问题,欢迎提交 Issue。为了帮助我们更快了解情况,请一并提供:
使用的 Agent 客户端及版本;
浏览器及版本;
扩展运行概览中的状态;
健康检查接口返回结果;
相关任务的错误信息。
请在分享日志前移除连接码、个人路径和其他敏感信息。
设计边界
本项目不修改 GPT Image Playground 源码;
API Key 留在 Playground 自己的浏览器会话中,MCP 和扩展不读取它;
任务只通过本机回环地址提供服务,并要求连接码;
浏览器关闭后,页面任务无法继续执行;
任务会保持串行处理,避免多个 Agent 同时操作同一个 Playground 页面。
感谢你花时间尝试这个工具。希望它能让图片生成工作更顺手,也让 Agent 与 Playground 之间的协作更自然。
Available Tools
3 toolsdownload_imageB
把已完成任务中的图片从 Playground 页面保存到 MCP 主机的本地绝对路径。
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ||
| image_index | No | ||
| output_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states the action (saving a file) without mentioning side effects, overwriting behavior, permission requirements, or error handling. The agent has no information about what happens if the task is incomplete or if the output path already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no unnecessary words. It front-loads the key action and scope, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (3 parameters, no output schema, no annotations), the description is too minimal. It lacks essential context for parameter usage, expected behavior on failure, and any safety considerations, making it incomplete for an agent to confidently invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explain any of the three parameters (task_id, image_index, output_path). The agent must infer meaning solely from parameter names, which is insufficient, especially for optional image_index which could be ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: downloading an image from a completed task to a local absolute path. It uses a specific verb (save/download) and resource (image from completed task), distinguishing it from sibling tools like generate_image and get_task_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly mentions the precondition that the task must be completed, giving clear context for when to use this tool. However, it does not mention alternatives or when not to use it, though the distinct purpose makes the use case clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
通过已打开的 GPT Image Playground 页面生成图片,可先注入本地参考图。请求会被浏览器桥接扩展串行处理,直到页面报告完成或失败;没有 20 秒固定上限。
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| reference_image_paths | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
没有注解,描述承担了行为披露的全部责任。描述说明了请求会被串行处理、等待页面完成报告、没有20秒超时限制,这些是有用的行为细节,让代理了解执行模型。但未涉及权限、副作用或错误处理等,考虑到无注解,整体披露较充分。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述简洁,两个句子传达了核心功能和行为特征,没有冗余信息。第二句补充了重要的执行细节(串行处理、无固定超时),结构合理,但可再稍微结构化一点。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
工具接受2个参数,无输出schema,描述提供了行为的额外细节(串行处理、异步等待)。但对于兄弟工具(get_task_status、download_image)的关系没有提及,缺少集成上下文的说明。考虑到相对简单,描述基本够用,但可以更全面。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
schema覆盖率为0%,描述仅提到了'注入本地参考图'对应reference_images参数,但核心参数prompt没有额外解释,仅从字段名推断其含义。描述未能完全补偿schema参数的缺乏,特别是prompt的格式、要求或约束未提及。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确说明了工具的动作是生成图片,资源是GPT Image Playground页面,并提到可注入本地参考图。与兄弟工具get_task_status和download_image在功能上明显区分,目的清晰具体。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述暗示了使用场景(需要通过已打开的页面生成图片),但没有明确说明何时应使用此工具而非其他兄弟工具,也未提及限制条件(如页面必须已打开、注入参考图的步骤等)。缺少明确的'何时使用/何时不使用'指导。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_statusB
查询 Playground 页面中某个生成任务的状态。
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It only states that the tool queries a status, implying a read-only operation, but it does not disclose return format, error behavior, or any Playground-specific constraints or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with no filler, directly stating the tool's purpose. It is front-loaded and appropriately sized for a simple status-query tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one parameter and no output schema, yet the description does not explain what the status response looks like or how it relates to the sibling generate/download workflow. This leaves the agent with gaps in knowing how to interpret the result or integrate the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the task_id parameter. While '某个生成任务' hints at the parameter's purpose, the description adds no concrete meaning beyond the schema's bare field name and type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb '查询' (query) and a clear resource '生成任务的状态' (status of a generation task), scoped to the Playground page. This distinguishes it from sibling tools that generate or download images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by the nature of the tool—checking status after generation—but there is no explicit guidance on when to use it vs. alternatives, no exclusions, and no mention of prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
download_image - First observed
generate_image - First observed
get_task_status
TDQS
Scored across 3 tools
Each tool handles a distinct phase of the image generation workflow: generating, checking status, and downloading. There is no overlap in their purposes, making it clear which tool to use at each step.
All three tool names follow a consistent verb_noun pattern (generate_image, get_task_status, download_image), using snake_case and clear action-first naming. The pattern is uniform across the set.
With only 3 tools, the server is tightly scoped to a single workflow (generate, track, download). This is appropriate for a focused purpose and each tool is necessary for the complete flow.
The workflow covers the essential lifecycle: generation initiation, status polling, and downloading results. A minor gap is the lack of a cancellation or listing tool, but agents can work around this by waiting or using status checks.
Maintenance
Related MCP Connectors
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Generate images, video, and audio with Glif's media-generation agent
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables image generation and editing using OpenAI's GPT Image API (gpt-image-1, 1.5, 2) with support for multi-image generation, history management, and batch processing.12125 npm1MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI agents and coding assistants with image generation and editing capabilities using OpenAI's GPT-image-1 model, with support for local or Supabase storage.3MIT
- AlicenseBqualityBmaintenanceEnables generating images from text or transforming existing images using GPT-Image-compatible APIs, with support for OpenAI and Agnes AI backends.2MIT
- AlicenseNot gradedqualityCmaintenanceEnables image generation and editing via ChatGPT web without API keys, saving images locally with conversation-aware editing.11 npm2MIT