zimage-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@zimage-mcpgenerate a hero banner for the homepage, save to ./public/hero.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
zimage-mcp
A local MCP server for generating web-dev images with the
Z-Image Turbo model running on the ComfyUI box. It hard-codes the verified Z-Image graph
(UNETLoader + CLIPLoader lumina2 + ModelSamplingAuraFlow shift 3 + KSampler 8/cfg1/res_multistep/simple),
downloads the result, and saves it to a local path at exact pixel dimensions.
It runs alongside the existing @peleke.s/comfyui-mcp (which stays for SDXL/FLUX); tools are
namespaced mcp__zimage__*.
Tools
generate_image(prompt, output_path, size="square", width?, height?, seed?, steps?=8, negative?="", n?=1)
Generates and saves image(s).
output_path— local path, e.g../public/img/hero.png. Parent dirs are created. Format follows the extension (.png/.jpg/.jpeg/.webp).size— a preset (below), or pass exactwidth+height. Because Z-Image needs dimensions that are multiples of 16, any exact size is generated at the nearest valid size (≥ requested, matching aspect) then resize-cover + center-crop to your exact pixels.seed— omitted = random; echoed back in the result for reproducibility.n— 1–4 variants. Files get_1,_2, … suffixes; seeds arebase, base+1, ….Returns
{ images, seed, gen_size, output_size, seconds }.
list_presets()
Returns the preset table (name → WxH).
health()
ComfyUI pre-flight: reachable, comfyui_version, gpu, vram_free_mb.
Related MCP server: image-studio-mcp
Size presets
preset | output | use |
| 1024×1024 | cards, avatars, icons |
| 1216×832 | general content images |
| 832×1216 | tall cards |
| 1280×720 | 16:9 hero/banner |
| 1536×640 | wide banner (12:5) |
| 1200×630 | OpenGraph/social share |
| 768×1344 | mobile-first portrait |
Z-Image is tuned around ~1 MP. Very large custom dimensions (>~1.3 MP) may reduce quality or slow down; for big hero art, generate at a preset and upscale separately (a 2K upscaler template exists on the box). Note: text-in-image is weak — use FLUX/Qwen for legible baked-in text.
Configuration (env vars)
var | default |
|
|
|
|
|
|
|
|
|
|
|
|
Run / register
Requires uv. Add to ~/.claude.json under mcpServers:
"zimage": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "/Users/jaimie/projects/zimage-mcp", "server.py"],
"env": { "COMFYUI_URL": "http://10.0.0.109:8188" }
}Restart Claude Code; the tools appear as mcp__zimage__generate_image, etc.
Develop / test
uv sync
uv run pytest -q # unit tests (no network)
ZIMAGE_TEST_LIVE=1 uv run pytest -q # + live tests against the boxAvailable Tools
3 toolsgenerate_imageA
Generate web-dev image(s) with Z-Image Turbo and save to output_path.
size: one of square|landscape|portrait|hero|wide|og|mobile, or pass width+height for exact pixels (image is generated at the nearest valid size then resized/cropped). n: 1-4 variants (paths get _1,_2,... ); seeds are base_seed, base_seed+1, ... Returns {images, seed, gen_size, output_size, seconds}.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| size | No | square | |
| steps | No | ||
| width | No | ||
| height | No | ||
| prompt | Yes | ||
| negative | No | ||
| output_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It explains generation behavior (resizing/cropping to nearest valid size, multiple variants using seeds, saving to output_path, return format). It does not mention overwrite behavior, auth requirements, or computational cost, but covers essential behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a lead sentence defining the action, followed by concise bullet-like explanations for key parameters (size, n, seeds, returns). Every sentence adds value, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description includes return format. It covers size flexibility, multi-image generation, and output saving. However, it lacks details on step parameter semantics, negative prompt effects, and potential edge cases (e.g., path conflicts). Still, it is largely complete for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so description compensates by explaining size (named presets or width+height), n and seed inference, and the return object. It does not detail steps or negative prompt semantics, but most parameters are covered sufficiently for usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Generate web-dev image(s) with Z-Image Turbo and save to output_path', clearly identifying the verb (generate), resource (web-dev images), and output. Sibling tools (health, list_presets) are completely unrelated, so no confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit when/when-not guidance, but the purpose is distinct from siblings, and the parameter explanations implicitly define usage context. It would benefit from stating 'Use this to generate images, not for health checks or listing presets'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
healthA
Quick ComfyUI pre-flight: version, GPU, free VRAM.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, but the description discloses that the tool checks version, GPU, and free VRAM. For a simple read-only health check, this is sufficient behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no wasted words. Front-loaded with purpose and key details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple health check with no parameters and no output schema, the description covers the essential outputs. Could add units or format, but not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters with 100% schema coverage, so baseline is 4. The description adds meaning by specifying what data the tool returns, even without output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks version, GPU, and free VRAM, which are specific health indicators. It distinguishes from siblings 'generate_image' and 'list_presets' as a diagnostic pre-flight tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is before operations like image generation, as a pre-flight check. No explicit when-not-to-use or alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_presetsA
Return available size presets as name -> 'WxH'.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description does not disclose behavioral traits beyond the output format ('name -> 'WxH''). There is no mention of side effects, permissions, or state changes, which is insufficient for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded with the verb and resource. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (zero parameters, no nested objects, no output schema), the description is nearly complete. It could mention that the tool is safe to call, but it is sufficient for a straightforward info-retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so schema coverage is 100%. The description adds meaning by specifying the output format, which is above baseline. Baseline for 0 params is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return available size presets') and the resource, and distinguishes from sibling tools like 'generate_image' and 'health'. It is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, but the purpose is clear enough that usage is implied. A more explicit note would improve this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
generate_image - First observed
health - First observed
list_presets
TDQS
Scored across 3 tools
Each tool has a distinct purpose: generate_image for creation, health for system status, list_presets for configuration. No overlap or ambiguity.
Most tools follow verb_noun pattern (generate_image, list_presets), but 'health' is a noun alone. Minor deviation but still clear and readable.
Three tools is well-scoped for an image generation server: core generation, health check, and preset listing. No unnecessary bloat.
The tool set covers the essential workflow: generate images, check system readiness, and retrieve presets. No obvious missing functionality for the stated purpose.
Maintenance
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
MCP server for Luma Dream Machine AI video generation
MCP server for Flux AI image generation
MCP server for NanoBanana AI image generation and editing
Related MCP Servers
- FlicenseCqualityDmaintenanceAn MCP server that generates images based on text prompts using Black Forest Lab's FLUX model, allowing for customized image dimensions, prompt upsampling, safety settings, and batch generation.31-
- AlicenseBqualityDmaintenanceA local MCP server for generating and editing images using OpenAI-compatible APIs. It provides text-to-image generation and image editing capabilities with configurable endpoints and saves output directly to local files.213 npmMIT
- AlicenseAqualityDmaintenanceA local MCP server for AI image generation using ComfyUI, Claude Code, and OpenClaw, with a built-in library of over 1,300 prompts for private, fast image creation.8600 npm8MIT
- AlicenseNot gradedqualityAmaintenanceMCP server for generating and editing images using Amazon Nova Canvas, Stable Diffusion 3.5 Large, and Stability AI services through Amazon Bedrock.200 PyPIApache 2.0