Micu Image MCP
米醋画图 MCP
从第一张图开始记忆,接着聊、接着改。

快速安装 · 上下文记忆 · 调试与验证 · v0.4.1 Release
把 米醋 的图像接口接入 Claude Code、Claude Desktop、Codex、Cursor 等 MCP 客户端。 生成、连续编辑、批处理、多图参考共用五个工具。v0.4.1 新增持久图片会话:保存成功版本、指令和有效约束, 让同一张图可以逐轮修改,也让一个聊天里的不同创作主题各自保留记忆。
**新版重点:**首次生成就建立记忆;沿用 ID 继续编辑;新主题新建会话;重启 MCP 后仍可用原 ID 继续。
上下文记忆
先生成一只小猫,再换飞行装备;中途生成独立海景,最后回到小猫改日落。 客户端选对会话 ID,MCP 就会携带对应图片与编辑上下文。

你说的话 | 图片会话 | MCP 的处理 |
“生成一只小猫跳伞。” | A · r1 | 建立新会话,保存首张图和描述 |
“给它换一种飞行装备。” | A · r2 | 上传 A 的上一版图片,沿用主角和有效约束 |
“另外生成一张独立海景人像。” | B · r1 | 新建 B,使用这个主题自己的图片与约束 |
“回到刚才的小猫,把背景改成日落。” | A · r3 | 返回 A,从 A 的成功历史继续编辑 |
这四张图来自 Osaka 上的 OpenCode + DeepSeek Flash → micu-image-mcp → 米醋主站真实调用, 对应四次成功操作。海景有一次上游失败,调整提示后重试成功;失败没有推进会话版本。 第二张实际画出了硬质滑翔翼,随后人工补充软翼形态约束,得到海报中的 A · r4。 海报与四步图分别标明版本,完整来源见配图说明。
记忆如何接上下一轮
首次生成就开始记忆:
image_generate或image_edit传session_id="new",记下返回的session.id。**继续改同一张图:**沿用这个 ID;省略
model、size、quality时继承所选版本,并携带成功编辑历史和当前约束。**换主题或回到旧图:**新主题用
"new";回到旧图用原 ID。A、B是上图简称,实际 ID 形如edit_<32位十六进制>。**重启也能继续:**会话保存在用户电脑;保留原 ID、输出根目录和对应图片,即可接着编辑。指定
parent_revision可从旧版本开分支。
选择哪条会话由客户端或模型负责,MCP 根据显式 ID 读取记忆;一次聊天可以拥有多条图片会话。
省略 session_id 仍按原有单次调用处理。每次成功才提交新版本,同一会话并发修改会返回忙。
图片记忆是本地版本记录;支持范围和存储管理见连续编辑指南、会话存储说明。
Related MCP server: gpt-image-2-mcp
功能
Tool | 说明 |
| 文生图;默认 Flare,当前主站 quality 使用 |
| 单图参考/编辑;默认 Sunburst,走 |
| 多张图逐张同指令处理;默认 Sunburst 串行 |
| 2-10 张参考图融合成 1 张新图;默认 Sunburst |
| 查看 base URL、模型、size 规则、重试策略、安全约束 |
第一次使用前,让 LLM 调一次 server_info,可以看到当前运行时配置和可用能力。
使用教程
面向 Cursor / Claude Code / Codex 用户的完整 MCP 使用指南见 docs/MCP使用教程.md, 涵盖工具选型、尺寸规则、环境变量与故障排查(含 Clash/Surge fake-ip 落盘问题)。
Claude Desktop 兼容性
v0.4.1 已修复 issue #11 的 Connected · 0 tools。
升级实际启动的 binary、重新连接 MCP,再调用 server_info 确认版本。
Desktop 的 Code 模式读取 ~/.claude.json;普通 Chat 模式使用独立的 claude_desktop_config.json,
当前安装器不会自动写入后者。兼容性指南包含根因、两种模式的配置和测试范围。
当前模型范围
当前支持 gpt-image-2.5-flare、gpt-image-2.5-sunburst、gpt-image-2、gpt-image-2-openai。
文生图默认 Flare,编辑、批处理和多图参考默认 Sunburst;四个图像工具都接受这四个模型。
实际可用模型以当前凭据的 /v1/models 为准。
2026-10-09 主站的 1K 生成和单图编辑实测:Flare / Sunburst 接受 low、auto 或省略
quality;medium / high / xhigh / max 均返回 HTTP 400(上游明确只支持 low)。
gpt-image-2 接受 low / medium / high / auto 或省略。本次凭据未列出 gpt-image-2-openai,
因此旧 gpt-image-2 的 2K/4K 自动切换模型前需要确认该凭据支持目标模型;Flare / Sunburst
的 2K/4K 保持所选模型并进入高分辨率串行队列。参数完整选项仍保留,用于兼容其他线路或后续支持。
Grok 相关实现继续保持休眠。
Windows 中文提示词:MCP 会以原生 UTF-8 JSON 发送中文。自行编写 PowerShell 测试脚本时,先设置
$OutputEncoding = [Console]::OutputEncoding = [System.Text.UTF8Encoding]::new(),避免旧 PowerShell 在管道中把中文变成?。
安装
从 v0.3.0 起,main 与推荐安装入口是 Rust 原生单文件 MCP server:
默认 STDIO serve,无参数即可运行;
运行时不需要 Python、pip、httpx 或 Pillow;
提供
install/reset/doctor/version;Python v0.2.0 reference 永久保留在
python-reference分支, main 的运行时、测试和工程编排均为 Rust;历史实现不参与当前构建或发布。
一键安装(推荐)
macOS / Linux:
curl -fsSL https://raw.githubusercontent.com/Subaru486desuwa/micu-image-mcp/main/scripts/install.sh | shWindows PowerShell:
irm https://raw.githubusercontent.com/Subaru486desuwa/micu-image-mcp/main/scripts/install.ps1 | iex一键脚本会识别当前平台,从最新 GitHub Release 下载 Rust binary,使用 Release 中的
SHA256SUMS 校验后调用 binary 自带的 install --yes。默认自动写入 Codex 与 Claude Code 的
micu-image MCP 配置(Claude Code 使用 ~/.claude.json);API key 不会写入配置文件。installer 会先检查 MICU_API_KEY 与系统
安全凭据存储;首次安装且两者都没有时,会在交互式终端中隐藏输入地询问一次 API key。
只有通过 sk- 前缀、20–512 字符长度和 ASCII 字符集校验后才会写入 macOS Keychain、Windows
Credential Manager 或 Linux Secret Service。之后 MCP 启动时自动读取,不需要重复输入。默认
endpoint 仍是 https://www.micuapi.ai;需要自定义时使用 --baseurl 或 MICU_BASEURL。
macOS / Linux 需要只配置某个客户端时,可把参数传给内置 installer:
curl -fsSL https://raw.githubusercontent.com/Subaru486desuwa/micu-image-mcp/main/scripts/install.sh \
| sh -s -- --no-claude脚本源码在 scripts/install.sh 和
scripts/install.ps1,可先审阅再执行。
手动下载 Rust binary
从 最新 Release 下载平台对应文件并
核对 SHA256SUMS:
chmod +x /absolute/path/micu-image-mcp # macOS/Linux
MICU_SAVE_DIR="$HOME/Pictures/micu-out" \
/absolute/path/micu-image-mcp install --yesinstall 会把当前 binary 原子复制到稳定的 per-user data-local 目录,再让 Codex/Claude 指向该
副本;配置不会指向仓库的 target/release,移动仓库或 cargo clean 不会使 MCP 失效:
macOS:
~/Library/Application Support/micu-image-mcp/bin/micu-image-mcpLinux:
~/.local/share/micu-image-mcp/bin/micu-image-mcpWindows:
%LOCALAPPDATA%\micu-image-mcp\bin\micu-image-mcp.exe
Rust CLI:
micu-image-mcp # 等同 serve,STDIO MCP
micu-image-mcp serve
micu-image-mcp install --yes --no-claude
micu-image-mcp install --yes --binary-path /path/to/downloaded/binary
micu-image-mcp install --yes --dev --binary-path "$PWD/target/release/micu-image-mcp"
micu-image-mcp reset --yes
micu-image-mcp doctor
micu-image-mcp version源码开发时可显式配置已编译 binary:
cargo build --release
MICU_SAVE_DIR="$HOME/Pictures/micu-out" \
target/release/micu-image-mcp install --yes --dev \
--binary-path "$PWD/target/release/micu-image-mcp"macOS Keychain
Rust binary 在启动时按 service/account 从 macOS Keychain 取 key,稳定 binary 可作为纯
command。保留的 Keychain launcher 也只启动 Rust:
security add-generic-password \
-U -a "$USER" -s ai.micuapi.mcp \
-l "Micu Image MCP API Key" \
-T /usr/bin/security -w[mcp_servers.micu-image]
command = "/Users/you/Library/Application Support/micu-image-mcp/bin/micu-image-mcp"
args = []
[mcp_servers.micu-image.env]
MICU_KEYCHAIN_SERVICE = "ai.micuapi.mcp"
MICU_KEYCHAIN_ACCOUNT = "your-macos-account"
MICU_SAVE_DIR = "/Users/you/Pictures/micu-out"
MICU_SAVE_DIR_ROOT = "/Users/you/Pictures/micu-out"Codex 桌面、CLI 和 IDE 扩展共享 ~/.codex/config.toml;修改后重启客户端。
验证与回滚
安装后让客户端调用 server_info,确认 available_models、base URL、save root 和
api_key_configured。Rust 还可先运行:
micu-image-mcp doctor
cargo xtask smoke --binary /absolute/path/micu-image-mcp回滚时恢复可信 binary 与 installer 的配置备份,保留图片和会话。历史 Python reference 在独立分支,当前 main 不再提供 Python 入口。完整说明见 迁移与回滚,开发、测试和工程工具见 Rust-only 开发。
Size 规则
旧 GPT Image 2 路径:
W/H 必须是 16 的倍数
最长边不超过 3840;长宽比不超过 3:1
总像素必须在 655,360 到 8,294,400 之间
2K/4K 自动切
gpt-image-2-openai2K/4K 强制
n=1并加跨进程锁,避免多个 MCP 同时打爆高质量队列
GPT Image 2.5 的 Flare 与 Sunburst 已实测支持 1024x1024、2048x1152 和
3840x2160,且返回像素与请求一致。2K/4K 保持所选 2.5 模型,强制 n=1 并使用跨进程锁。
推荐 size:
档位 | 推荐值 |
1K |
|
2K |
|
4K |
|
尺寸与路由行为
/v1/images/edits负责单图编辑、多图参考与批量编辑。gpt-image-2的 2K/4K 请求会自动切换到gpt-image-2-openai。GPT Image 2.5 的 2K/4K 请求保持所选 Flare / Sunburst 模型。
≥2K 请求强制
n=1,并通过进程内与跨进程锁串行访问高质量队列。返回的真实像素以响应中的
saved.actual_size为准。
环境变量
变量 | 默认值 | 说明 |
| 空 | 米醋 image2 token |
|
| 米醋 base URL |
|
| 系统安全凭据 service;兼容旧自定义 Keychain 项 |
|
| 系统安全凭据 account |
| 空 | 可选全局覆盖;未设置时生成默认 Flare,编辑类默认 Sunburst |
|
| 默认输出目录 |
| 同输出目录 | 输出安全根目录 |
| 空(不限制) | 可选输入图片白名单根;启用后阻止路径/符号链接逃逸 |
|
| 设为 |
|
|
|
|
| 可信 CDN host,逗号分隔 |
|
| 仅 trusted host 可放行 198.18.0.0/15 fake-ip |
路径在 server 启动时只解析一次:相对 MICU_SAVE_DIR 和 tool save_dir 都以 save root 为基准;
设置 MICU_INPUT_ROOT 时,相对输入路径以 input root 为基准,否则以启动时捕获的 cwd 为基准。
只展开精确的 ~、~/...、Windows ~\...,~someone 会被拒。Python/Rust 兼容期共用
~/.cache/micu-image/bigsize.lock。
手动配置
Claude Code:
{
"mcpServers": {
"micu-image": {
"command": "/absolute/path/micu-image-mcp",
"args": [],
"env": {
"MICU_SAVE_DIR": "/Users/you/Pictures/micu-out",
"MICU_SAVE_DIR_ROOT": "/Users/you/Pictures/micu-out"
}
}
}
}Codex:
[mcp_servers.micu-image]
command = "/absolute/path/micu-image-mcp"
args = []
[mcp_servers.micu-image.env]
MICU_SAVE_DIR = "/Users/you/Pictures/micu-out"
MICU_SAVE_DIR_ROOT = "/Users/you/Pictures/micu-out"不要手工把 Windows 路径拼进 TOML 字符串。Rust installer 使用 toml_edit AST,临时写入后会
再用 TOML parser 校验 command/args/env 的 PathBuf round-trip;单引号 literal string 和正确
转义的双引号 basic string 都合法,关键是 parser 回读值完全一致。API key 不持久化到上述
JSON/TOML;由客户端进程环境、macOS Keychain 或 tool 的既有 api_key 参数提供。
迁移期若要手动使用 Python reference,把 command 改为 Python、args 改为绝对
server.py 路径即可;五工具 schema 保持相同。
调试与验证
这次更新从“小猫改图后还能否回到同一条记忆”出发,逐步测试会话、真实 harness、协议兼容和发布包。
调试问题 | 定位与修正 | 验证结果 |
连续编辑容易丢失主角、约束或版本来源 | 显式 ID、持久成功历史、分支约束;失败不提交,同会话加锁 | 生成、跨工具编辑、重启恢复、失败恢复与会话隔离均有离线覆盖 |
DeepSeek 有时省略 | 把“新主题新建、继续编辑复用 ID”放到工具说明最前面 | 三种原生 harness 组合在两组固定四轮场景中通过 24/24 个路由检查 |
Claude Desktop 连接成功却显示零工具 | 复刻缺少 | 真实 Claude 客户端首次注册五工具并调用 |
干净 CI 与本地结果不同 | 仅归一化 Pydantic 帮助链接版本;测试夹具使用标准化路径 | Python 3.10/3.13、Rust 最低版本和四个平台原生 CI 全部通过 |
视觉对象的细分形态不准确 | 检查真实输出,追加软翼与无金属支架的明确约束 | 同一会话继续修正,保留小猫与日落;人工精修单独标注 |
工具说明从 23,127 → 3,256 字节;DeepSeek 首轮工具上下文从 9,924 → 3,607 tokens。 固定 200 ms 本地接口替身下,四个独立会话串行约 953 ms,并发约 248 ms;增加工作线程没有可测收益。 这些是指定场景的实测结果,真实图像渲染、网络与模型等待另计。
发布前在 Osaka 验证 Rust 90 项、Python 402 passed / 5 skipped;
完整 CI与
发布流程均通过。
MacBook 已安装并校验正式包,本地五工具、新旧协议与 server_info 调用通过;Desktop 窗口生图验收仍由用户手动完成。
完整调试过程见v0.4.1 调试笔记和原生 harness 验收。
工程与验证文档
README 展示主要功能、真实案例、安装方法与调试结论;完整实现说明和验证记录见 docs/ 与 CI:
Star History
Available Tools
5 toolsimage_batch_editA
批量图像编辑:N 张输入图 → N 张输出图,每张独立应用同一指令。
[WHAT] 对 image_paths 里的每一张图分别调用 image_edit,统一 prompt 与 size,结果合并返回。
[WHEN TO USE]
用户提供多张图且每张要做"同样的修改"(如批量加水印 / 统一换底 / 统一调色)→ 用此 tool。
如果是"用多张图作风格参考画 1 张新图" → 这不是此 tool,暂未实现。
如果只有 1 张图 → 用 image_edit。
[并发策略]
gpt-image-2:5 并发(HTML 网页同款)。
gpt-image-2-openai:串行 + 1.5s gap(高质量线路并发更容易被限流)。
任意一张失败不影响其他张;返回 results 里逐张标 ok/error。
[LIMITS]
与 image_edit 一致支持 1K/2K/4K;2K/4K 自动切高质量线路并逐张串行。
image_paths 长度建议 2-20 张;高分辨率批次成本与耗时按图片数量线性增加。
Args: prompt: 应用到每张图的修改指令。例:"add a subtle watermark in bottom-right". image_paths: 输入图路径列表(绝对或相对)。 size: 输出 size,支持 1K/2K/4K;≥2K 自动使用高质量线路。默认 "1024x1024"。 model: "gpt-image-2" / "gpt-image-2-openai"。留空按 size 自动选。 save_dir: 输出目录(必须在安全根目录之下)。文件名 batch__.png。 api_key: 覆盖 MICU_API_KEY;base_url 已锁在启动期 env,运行期不接受。
Returns: dict 含: ok (bool): True 表示至少 1 张成功。 total (int): 输入图总数。 succeeded (int): 成功张数。 failed (int): 失败张数。 concurrency (int): 实际用的并发度(5 或 1)。 results (list[dict]): 每张图的详细结果(含 input 路径、saved.path、可能的 error)。
Examples: image_batch_edit( prompt="convert to pencil sketch style", image_paths=["/p/a.jpg", "/p/b.jpg", "/p/c.jpg"], size="1024x1024", )
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | 1024x1024 | |
| model | No | ||
| prompt | Yes | ||
| api_key | No | ||
| save_dir | No | ||
| image_paths | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers rich behavioral detail: concurrency strategy (5 for gpt-image-2, serial+1.5s gap for the openai line), failure isolation ('任意一张失败不影响其他张'), high-resolution auto-switching to serial processing, rate-limit risk disclosure, and note that cost/time scale linearly with batch size. This exceeds what any annotation set would typically provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with clear section headers ([WHAT], [WHEN TO USE], [并发策略], [LIMITS], Args, Returns, Examples) and front-loaded summary. Though long, every block earns its place — this is a complex 6-parameter batch tool with output-format documentation and routing logic; the length is proportionate to the complexity. No filler or redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, 0% schema coverage, and an output schema, the description covers everything an agent needs: input parameter semantics, return format (ok/total/succeeded/failed/concurrency/results), concurrency and failure behavior, limits, and a working example. The presence of an output schema relaxes the burden on return-value explanation, and the description still documents it. Complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does so thoroughly. The Args section gives each parameter real meaning beyond the schema titles: prompt gets an example, size gets the 1K/2K/4K values and default, model gets the two accepted values, save_dir gets the filename pattern batch_<ts>_<idx>.png and the security-root constraint, api_key gets its MICU_API_KEY override behavior with base_url locking noted. Full compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a precise verb+resource+scope: 'N 张输入图 → N 张输出图,每张独立应用同一指令' (batch edit N images, each independently applying the same instruction). The [WHAT] section plainly states it calls image_edit per image and merges results. It differentiates itself from siblings by explicitly declaring '用多张图作风格参考画 1 张新图' is NOT this tool, and routes single images to image_edit. Unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The [WHEN TO USE] section gives explicit selection criteria: use this tool when the user has multiple images all needing the same modification (watermark, background, color grading), use image_edit for a single image, and explicitly states the multi-reference case is not implemented. It also adds concurrency strategy per model line. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_editA
图像编辑(image-to-image,单张输入)。当前线路支持 1K/2K/4K。
[WHAT] 接受 1 张本地图片 + 修改指令,输出修改后的图。
[WHEN TO USE]
用户提供 1 张图(路径或刚刚生成的图)且要"改 / 替换 / 加 / 去掉某部分" → 用此 tool。
如果用户没提供图想从零生成 → 改用 image_generate。
如果用户提供了多张图想"批量改"(每张做同样操作)→ 改用 image_batch_edit。
如果用户用多张图作风格参考想画一张新的 → 用 image_multi_reference。
[尺寸能力](2026-08-14 当前线路实测)
1K:gpt-image-2 与 gpt-image-2-openai 的 1024×1024 edits 均成功并精确返回。
2K:自动切 gpt-image-2-openai;2048×1152 edits 成功并精确返回。
4K:自动切 gpt-image-2-openai;3840×2160 edits 成功并精确返回。
[当前线路]
参考图 4K 的旧线路硬阻断已移除;1K/2K/4K 均统一走 /v1/images/edits。
2K/4K 自动使用 gpt-image-2-openai,并通过跨进程锁串行请求高质量队列。
始终通过 saved.actual_size 核对后端实际返回像素。
[路由实现](实测确定)
所有尺寸统一走 /v1/images/edits multipart(米醋唯一真正消费输入图的端点)。 Images API 返回错误时直接报错,不把图像模型转发到不兼容的 /v1/chat/completions。
mask 现已在所有尺寸支持(不再区分 1K/2K)。
[MASK 工作原理]
mask_path 指向一张 PNG,尺寸应与 image_path 一致。
mask 中 alpha=0(透明) 的像素 = 要修改的区域。
alpha=255(不透明)的像素 = 要保持原样。
不传 mask 则模型自由决定改哪里。
Args: prompt: 修改指令,越具体越好。例:"change the background to deep navy with stars, keep the subject pixel-identical". image_path: 输入图的绝对或相对路径。PNG / JPG / WebP 都支持。 mask_path: 可选 alpha mask PNG 路径,透明区即编辑区。所有尺寸均生效。 size: 输出 size。W/H 必须是 16 的倍数;总像素和长宽比规则见 server_info。 "1024x1024" "1280x720" "1024x1536" "1536x1024" "720x1280" ← 1K 档 "2048x2048" "2048x1152" "1152x2048" ← 2K 档(自动高质量线路) "3840x2160" / "2160x3840" ← 4K 档(自动高质量线路) 默认 "1024x1024"。 model: "gpt-image-2"(默认)/ "gpt-image-2-openai"(高质量线路,≥2K 自动切)。 save_dir: 输出目录(必须在安全根目录之下)。默认 ~/Pictures/micu-out 或 MICU_SAVE_DIR。 basename: 文件名前缀(仅 [A-Za-z0-9_-.])。默认 "edit_"。 api_key: 覆盖 MICU_API_KEY;base_url 已锁在启动期 env,运行期不接受。
Returns: dict 含: ok (bool): 是否成功。 model (str): 实际用的模型。 size (str): 请求 size。 used_fallback (bool): 为兼容既有返回结构保留;当前 Image2 模型固定为 False。 saved (dict): { path, size_bytes, actual_size, actual_megapixels }。 notes (list[str]): 决策与提示。
Examples: # 换背景 image_edit(prompt="replace background with a sunset beach", image_path="/p/portrait.jpg")
# 局部修改(mask 生效)
image_edit(prompt="change hair color to silver", image_path="/p/x.png", mask_path="/p/x_mask.png")
# 升细节(2K 自动使用高质量线路)
image_edit(prompt="enhance to cinematic detail, preserve composition", image_path="/p/draft.png", size="2048x2048")
# 4K 参考图编辑(自动使用高质量线路)
image_edit(prompt="preserve composition and refine every detail", image_path="/p/draft.png", size="3840x2160")Common errors: "image_path 不存在" → 检查路径,建议用绝对路径。 "HTTP 524" → 当前高质量队列繁忙;自动策略仍失败时请稍后再试。
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | 1024x1024 | |
| model | No | ||
| prompt | Yes | ||
| api_key | No | ||
| basename | No | ||
| save_dir | No | ||
| mask_path | No | ||
| image_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it delivers: discloses the underlying /v1/images/edits endpoint, the automatic 2K/4K switch to gpt-image-2-openai with cross-process lock serialization, the removal of the old 4K hard-block, mask alpha semantics (alpha=0 = edit region, alpha=255 = preserve), actual_size verification, and explicit error conditions (HTTP 524 queue busy, missing image_path). No contradiction with annotations since none are provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but every section earns its place and the header-based structure ([WHAT], [WHEN TO USE], [尺寸能力], [MASK 工作原理], Args, Returns, Examples, Common errors) makes it highly scannable with the purpose front-loaded. Minor deduction for the dated 尺寸能力 and 路由实现 sections, which are somewhat redundant with the args and could be trimmed; overall this is efficient organization for a high-complexity tool, not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, 0% schema coverage, no annotations, and moderate routing complexity, the description is complete: all parameters defined with valid values, usage context with sibling routing, behavioral specifics, a full Returns dict specification, concrete examples for each use case (background swap, mask edit, 2K upscale, 4K refine), and common errors with remediation. Despite the output schema existing, the description also documents the return structure — a bonus that exceeds the baseline requirement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate — and it does. Every one of the 8 parameters is documented beyond the schema: prompt includes a concrete example and advice (the more specific the better); image_path lists supported formats (PNG/JPG/WebP); size enumerates exact valid values per tier with the 16-multiple constraint; model documents auto-switch behavior; save_dir notes the safety-root restriction; basename specifies the character whitelist; api_key explains the base_url lock. This is exemplary compensation for a zero-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb (edit), resource (single input image), and mutation type (modify/replace/add/remove parts). Clearly names and excludes siblings: image_generate (from scratch), image_batch_edit (batch), image_multi_reference (multi-style-reference). An agent can unambiguously route to this tool based on the WHAT and WHEN TO USE sections alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The [WHEN TO USE] section provides crisp selection criteria with explicit alternatives and exclusion conditions: use this when 1 image + edit intent; switch to image_generate if no image; to image_batch_edit for batch ops; to image_multi_reference for style-reference synthesis. Every branch names the sibling and the condition that selects it — nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_generateA
文本生成图像(text-to-image)。米醋代理 + gpt-image-2 系列。
[WHAT] 把一段文字 prompt 渲染成 1 张或 N 张图像,落盘到本地。
[WHEN TO USE]
用户要"画 / 生成 / 创建一张图"且没有提供任何参考图 → 用此 tool。
如果用户提供了 1 张参考图要"修改 / 编辑 / 替换某部分" → 改用 image_edit。
如果用户提供了多张参考图要"按它们的风格画一张新的" → 用 image_multi_reference。
如果不知道怎么选 size:先调 server_info() 看 recommended_sizes。
[SIZE 选取建议]
默认 None:MCP 自动从 prompt 关键字推断(4K/UHD → 3840x2160;1080p/2K → 2048x1152; 正方形/logo/头像 → 1024x1024;竖屏/9:16 → 1024x1536;横屏/16:9 → 1536x1024 等)。 推断不出来 fallback 1024x1024。
强烈推荐:如果你(LLM)已经从用户消息读出确定的 size 偏好,直接显式传 size,比关键字推断准。
用户提到"高清/4K/海报/壁纸" → "3840x2160"(横)或 "2160x3840"(竖),自动用 gpt-image-2-openai。
用户提到"FullHD/1080p/横屏视频封面" → "2048x1152"(横)或 "1152x2048"(竖); 这两个尺寸均满足当前 16 像素对齐和总像素约束。
W 与 H 必须都是 16 的倍数;最长边 ≤3840;长宽比 ≤3:1;总像素 655,360-8,294,400。
2K/4K 自动走高质量线路:≥2K 自动切 gpt-image-2-openai。2026-08-14 实测其 1536×1024、 2048×1152、3840×2160 均按请求像素返回;gpt-image-2 的自定义宽高可能被后端重映射。
[PROMPT 写法建议]
中英文混合可。gpt-image-2 文本渲染近完美,可大段嵌字(中英标点都行)。
越具体越好:风格 / 视角 / 光线 / 主体 / 细节程度。
Args: prompt: 图像描述。1-2000 字符。例:"A minimalist sushi mascot logo, soft pastel palette". size: "WxH" 字符串或 None。留 None 让 MCP 从 prompt 推(弱 LLM 兜底用); 强 LLM 已知偏好时直接显式传更准。W 和 H 都必须是 16 的倍数。常用: "1024x1024" "1280x720" "1024x1536" "1536x1024" "720x1280" ← 1K 档 "2048x2048" "2048x1152" "1152x2048" ← 2K 档(自动 gpt-image-2-openai) "3840x2160" "2160x3840" ← 4K 档(自动 gpt-image-2-openai) 默认 None(推断后兜底 1024x1024)。 n: 张数 1-10。1K 时 N>1 自动 5 并发;≥2K 强制 N=1(代理限流)。默认 1。 model: 显式指定模型。留空时按 size 自动选(max edge ≥1600 用 gpt-image-2-openai,否则 gpt-image-2)。 可选值:"gpt-image-2"(标准线路)/ "gpt-image-2-openai"(高质量线路)。 quality: 可选质量参数:"auto" / "low" / "medium" / "high";留空则使用后端默认值。 save_dir: 输出目录。必须在安全根目录 MICU_SAVE_DIR_ROOT 之下(默认 ~/Pictures/micu-out); 传 root 之外路径会被拒。留空使用默认。 basename: 文件名前缀(不带扩展名),仅允许 [A-Za-z0-9_-.]。 含 / .. 或路径分量会被拒。默认 "gen_"。 api_key: 覆盖 MICU_API_KEY 环境变量。一般留空。 注意:base_url 已锁在启动时 env,运行期不接受 tool 参数(防 key 外泄到攻击者 host)。
Returns: dict 含以下字段: ok (bool): 至少有 1 张成功才为 True。 model (str): 实际用的模型 id。 size (str): 请求的 size。 requested_n (int): 实际生成的张数。 saved (list[dict]): 每张成功的图。每项含 path(绝对路径)/ size_bytes / actual_size(PNG header 读出的真实像素)/ actual_megapixels。 errors (list[str]): 失败请求的错误描述。 notes (list[str]): 路由 / 自动决策 / 实测尺寸偏差的说明。
Examples: # 最简:默认 1024x1024 单张 image_generate(prompt="a red apple on white")
# 4K 壁纸
image_generate(prompt="cyberpunk Tokyo at night", size="3840x2160")
# 一次出 4 张候选(1K 自动并发)
image_generate(prompt="cute sticker of a cat", size="1024x1024", n=4)Common errors and what to do: "size W/H 必须是 16 的倍数" → 客户端入口拒;例如 1920×1080 应改为 1920×1088 或推荐的 2048×1152。 "HTTP 524: timeout" → 已自动重试 3 次仍失败,建议改小 size 或稍后再试。 "未配置 API key" → 设置 MICU_API_KEY 环境变量或传 api_key 参数。
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| size | No | ||
| model | No | ||
| prompt | Yes | ||
| api_key | No | ||
| quality | No | ||
| basename | No | ||
| save_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden – and it excels. It discloses automatic size inference from prompt keywords, fallback to 1024x1024, auto model selection based on max edge ≥1600, forced N=1 for ≥2K, 5-way concurrency for 1K with N>1, security constraints (save_dir must be under MICU_SAVE_DIR_ROOT, basename whitelist, base_url locked at startup to prevent key leakage), and measured deviations from requested sizes on certain models. It even documents retry behavior for timeouts. This is far beyond typical transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but extremely well-structured with clear section headers ([WHAT], [WHEN TO USE], [SIZE], [PROMPT]) and bullet points. The most critical info (what it does, when to use) is front-loaded. Examples and common errors are placed at the end where they belong. While it's verbose, every section adds distinct value – no redundant filler. The length is justified by the tool's complexity (8 params, routing logic, safety checks). It earns a 4, not 5, only because it is genuinely long and might overwhelm a quick scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema (returns dict), the description still explains each return field (ok, model, size, requested_n, saved, errors, notes) and their semantics, which is helpful. It covers the full decision tree (size selection, model routing, concurrency), all constraints (16-pixel multiples, aspect ratio limits, pixel bounds), and common error scenarios with remediation. It also includes practical examples and edge-case notes (like 1920×1080 → 2048×1152). For a tool with this many parameters and automatic behaviors, the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (all 8 parameters lack schema descriptions), so the description must provide full parameter semantics – and it does. Each parameter gets its own line with default, constraints, and examples: prompt with character limit and example, size with valid formats and auto-inference logic, n with range and concurrency implications, model with optional values and selection rule, quality with allowed values, save_dir with security path restriction, basename with regex whitelist, and api_key with env override note. The description fully compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: '文本生成图像(text-to-image)' – it generates images from text prompts and saves them to disk. It directly distinguishes itself from siblings by stating that image_edit is for editing with a reference image and image_multi_reference is for style transfer from multiple references. An agent can unambiguously route to this tool versus alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The [WHEN TO USE] section explicitly states when to use this tool ('画/生成/创建一张图' without reference), names two sibling alternatives with exact conditions (image_edit for single ref, image_multi_reference for multiple refs), and points to server_info for size selection when uncertain. It also gives negative guidance (use other tools when refs are provided), making the routing decision completely deterministic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_multi_referenceA
多图融合参考 → 输出 1 张新图;当前线路支持 1K/2K/4K。
[WHAT] 输入 2-10 张参考图 + prompt,模型综合所有图的视觉信息后画 1 张全新的图。 与 image_batch_edit 的本质区别:batch 是 N 进 N 出(每张独立改),此 tool 是 N 进 1 出(综合参考)。
[WHEN TO USE]
用户:"这几张是同一产品的不同角度,按这个风格画一个新角度" → 用此 tool。
用户:"这些是我喜欢的风格,画一张类似风格的 X" → 用此 tool。
用户:"这是 logo 主图,这是辅助图,做成海报" → 用此 tool。
如果用户只想"逐张修改" → 改用 image_batch_edit。
如果用户只有 1 张图 → 改用 image_edit。
如果用户没提供任何参考图 → 改用 image_generate。
[当前线路]
参考图 4K 的旧线路硬阻断已移除;所有尺寸统一走 /v1/images/edits + image[]。
2K/4K 自动使用 gpt-image-2-openai,并通过跨进程锁串行请求高质量队列。
[路由实现]
固定走 /v1/images/edits + 多个 image[] 字段。米醋唯一真正消费输入图的端点 (实测 image_tokens 线性 = 560×N)。旧的 generations + image_urls 被米醋静默忽略 (image_tokens=0,等于纯文生图,参考图不起作用),已弃用。
自动切高质量线路:max edge ≥1600 → gpt-image-2-openai
Images API 返回错误时直接报错,不把图像模型转发到不兼容的 /v1/chat/completions。
[LIMITS](当前真实状态,会变化)
image_paths 长度 2-10 张。
1K 档:多图 N=2..10 历史实测成功,参考图真消费;实际像素以 saved.actual_size 为准。
2K/4K:自动切 gpt-image-2-openai + edits/image[];不再有本地尺寸硬阻断。 高分辨率多图融合的耗时会随参考图数量增加,成功后以 saved.actual_size / size_honored 核对真实像素。
米醋多图间歇拒绝时会按重试策略处理,仍失败则直接返回 Images API 错误。
单张参考图建议 ≤2MB;总输入 ≤8MB(米醋代理上限实测约 10MB)。
Args: prompt: 综合指令。例:"combine the colors from img1 and the composition from img2 into a sunset cityscape". image_paths: 2-10 张参考图路径(绝对或相对)。 size: 输出 size。支持 1K/2K/4K;≥2K 自动切高质量线路。 成功时以 saved.actual_size 和 size_honored 核对真实像素。默认 "1024x1024"。 model: "gpt-image-2"(默认)/ "gpt-image-2-openai"(高质量线路,≥2K 自动切换)。 save_dir: 输出目录(必须在安全根目录之下)。 basename: 文件名前缀(仅 [A-Za-z0-9_-.],含 / .. 会被拒)。默认 "multiref_"。 api_key: 覆盖 MICU_API_KEY;base_url 已锁在启动期 env,运行期不接受。
Returns: dict 含: ok (bool): 是否成功。 model (str): 实际用的模型。 n_references (int): 实际嵌入的参考图张数。 saved (dict): { path, size_bytes, actual_size, actual_megapixels }。 notes (list[str]): 决策与提示。
Examples: # 1K 综合参考 image_multi_reference( prompt="combine these into a single cinematic poster", image_paths=["/p/sketch.png", "/p/character.png", "/p/background.png"], )
# 2K 综合参考(高质量线路)
image_multi_reference(
prompt="merge the architecture style from img1 with the lighting from img2",
image_paths=["/p/img1.jpg", "/p/img2.jpg"],
size="2048x2048",
)
# 4K 综合参考(自动使用高质量线路)
image_multi_reference(
prompt="combine the product references into one 4K campaign visual",
image_paths=["/p/front.jpg", "/p/side.jpg"],
size="3840x2160",
)Common errors: "至少需要 2 张参考图" → 1 张请用 image_edit。 "请求体超 X MB" → 减少图片数量或先压缩。 "HTTP 524" → 当前高质量队列繁忙;自动策略仍失败时请稍后再试。
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | 1024x1024 | |
| model | No | ||
| prompt | Yes | ||
| api_key | No | ||
| basename | No | ||
| save_dir | No | ||
| image_paths | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the exact endpoint (/v1/images/edits), the deprecated endpoint that is silently ignored, automatic high-quality line switching for ≥2K, retry behavior for intermittent rejections, error propagation, and concrete size limits (2-10 images, ≤8MB total). It even explains the token linearity (560×N). This is exceptionally transparent about runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though long, every section earns its place: [WHAT] states the core purpose, [WHEN TO USE] gives routing, [当前线路] covers current routing, [路由实现] explains implementation details, [LIMITS] lists constraints, and parameter descriptions are structured. The most critical info (purpose and routing) is front-loaded, and the rest is organized with clear headers, making it scannable despite length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, high-resolution handling, routing logic, limits), the description is complete. It covers input constraints, output verification, error handling, and alternatives. It even provides three examples demonstrating 1K, 2K, and 4K use cases. Nothing an agent needs to correctly select and invoke this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does comprehensively. It explains every parameter: prompt with an example, image_paths with count constraints, size with 1K/2K/4K semantics and default, model with the high-quality line mapping, save_dir safety restriction, basename pattern rule and default, and api_key override behavior. It also documents the return dict fields. This is far beyond minimal compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with '多图融合参考 → 输出 1 张新图' (multi-image fusion reference → output one new image), explicitly stating the verb (fuse/combine), resource (multiple reference images + prompt), and output (one new image). It then contrasts with image_batch_edit (N-in-N-out vs N-in-1-out), clearly differentiating it from its sibling. This is a precise, distinguishing definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A dedicated [WHEN TO USE] section lists concrete user requests and maps each to this tool or its alternatives (image_batch_edit for batch edits, image_edit for single image, image_generate when no references). It also includes explicit 'if...use...' conditions and a 'Common errors' section with troubleshooting guidance. This provides unambiguous routing for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_infoB
返回当前 Image2 模型、参数约束、路由和安全边界。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It indicates the tool returns information but does not disclose whether it is a read-only operation, any authentication requirements, or side effects. While it is likely a safe read, the lack of explicit disclosure and absence of annotations leaves this gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that directly states the core information returned. Zero waste, perfectly front-loaded, and appropriately sized for a parameterless tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the key content areas (model, parameter constraints, routing, security boundaries), which is sufficient for a simple info tool. An output schema exists to document the return structure, so the description does not need to explain return values. The only gap is lack of explicit usage context, but given the simplicity, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter documentation. The description does not need to add parameter meaning since there are none. Baseline 4 applies as the description is not compromised by missing parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (返回/return) and resource (当前 Image2 模型、参数约束、路由和安全边界), making the purpose clear. It is distinguishable from sibling tools that perform editing or generation actions, though it does not explicitly name alternatives. The resource is distinct enough that an agent can infer the purpose without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus the sibling action tools. The description only states what information is returned, but does not give context on when an agent should call it (e.g., before editing/generating to check constraints). Usage is implicitly obvious but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.3.0- Changed
image_generate1 field changed- added
Input schema / properties / qualityAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Quality" +}
5 tool updates
v0.1.0- First observed
image_batch_edit - First observed
image_edit - First observed
image_generate - First observed
image_multi_reference - First observed
server_info
TDQS
Scored across 5 tools
Each tool has a clearly distinct purpose: generate (text-to-image), edit (single image with optional mask), batch edit (N-to-N same operation), multi-reference (N-to-1 style fusion), and server_info (metadata). The WHEN TO USE sections explicitly disambiguate edge cases, making misselection unlikely.
The image tools follow a consistent image_ prefix pattern (image_generate, image_edit, image_batch_edit, image_multi_reference). However, server_info breaks the verb_noun convention, and image_multi_reference uses a noun rather than a verb, creating slight inconsistency. Still, the pattern is predictable and readable.
With 5 tools, the set is well-scoped for an image generation/editing server. Each tool covers a distinct workflow (single generation, single edit, batch edit, multi-reference fusion, and information), and none feel redundant or unnecessary. This is an ideal size for the domain.
The tool surface fully covers the core image workflows: generation from text, editing with masks and prompts, batch processing, multi-image reference fusion, and server configuration/limits. There are no apparent dead ends—any user request for image creation or modification can be routed to an appropriate tool. Missing operations like upscaling or dedicated background removal are achievable through existing tools (e.g., image_edit with mask or size parameters).
Maintenance
Related MCP Connectors
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
MCP server for Qwen Image 3 AI image generation
MCP server for OpenAI API (chat completions, image generation, embeddings) via AceDataCloud
MCP server for Pixapi: check live credit pricing and balance, then generate images and video.
Related MCP Servers
- AlicenseAqualityFmaintenanceProduction-grade MCP server for image and video understanding and generation across Gemini, OpenAI, and Grok.54Apache 2.0
- AlicenseAqualityBmaintenanceExposes OpenAI's gpt-image-2 (image generation and editing) as an MCP server for tools like generate_image, edit_image, and iterative edit sessions.638 npm12MIT
- FlicenseAqualityDmaintenanceWraps Google Gemini's image generation API as an MCP server, enabling text-to-image, image editing, and grounded search workflows from any MCP client.2-
- AlicenseNot gradedqualityDmaintenanceWraps Flow2API / OpenAI-compatible image generation upstream into an MCP service, providing image generation, history, and caching tools.12MIT