Skip to main content
Glama

zimage-mcp

A local MCP server for generating web-dev images with the Z-Image Turbo model running on the ComfyUI box. It hard-codes the verified Z-Image graph (UNETLoader + CLIPLoader lumina2 + ModelSamplingAuraFlow shift 3 + KSampler 8/cfg1/res_multistep/simple), downloads the result, and saves it to a local path at exact pixel dimensions.

It runs alongside the existing @peleke.s/comfyui-mcp (which stays for SDXL/FLUX); tools are namespaced mcp__zimage__*.

Tools

generate_image(prompt, output_path, size="square", width?, height?, seed?, steps?=8, negative?="", n?=1)

Generates and saves image(s).

  • output_path — local path, e.g. ./public/img/hero.png. Parent dirs are created. Format follows the extension (.png / .jpg / .jpeg / .webp).

  • size — a preset (below), or pass exact width + height. Because Z-Image needs dimensions that are multiples of 16, any exact size is generated at the nearest valid size (≥ requested, matching aspect) then resize-cover + center-crop to your exact pixels.

  • seed — omitted = random; echoed back in the result for reproducibility.

  • n — 1–4 variants. Files get _1, _2, … suffixes; seeds are base, base+1, ….

  • Returns { images, seed, gen_size, output_size, seconds }.

list_presets()

Returns the preset table (name → WxH).

health()

ComfyUI pre-flight: reachable, comfyui_version, gpu, vram_free_mb.

Related MCP server: image-studio-mcp

Size presets

preset

output

use

square

1024×1024

cards, avatars, icons

landscape

1216×832

general content images

portrait

832×1216

tall cards

hero

1280×720

16:9 hero/banner

wide

1536×640

wide banner (12:5)

og

1200×630

OpenGraph/social share

mobile

768×1344

mobile-first portrait

Z-Image is tuned around ~1 MP. Very large custom dimensions (>~1.3 MP) may reduce quality or slow down; for big hero art, generate at a preset and upscale separately (a 2K upscaler template exists on the box). Note: text-in-image is weak — use FLUX/Qwen for legible baked-in text.

Configuration (env vars)

var

default

COMFYUI_URL

http://10.0.0.109:8188

ZIMAGE_UNET

z_image_turbo_bf16.safetensors

ZIMAGE_CLIP

qwen_3_4b.safetensors

ZIMAGE_VAE

ae.safetensors

ZIMAGE_DEFAULT_STEPS

8

ZIMAGE_TIMEOUT_S

180

Run / register

Requires uv. Add to ~/.claude.json under mcpServers:

"zimage": {
  "type": "stdio",
  "command": "uv",
  "args": ["run", "--directory", "/Users/jaimie/projects/zimage-mcp", "server.py"],
  "env": { "COMFYUI_URL": "http://10.0.0.109:8188" }
}

Restart Claude Code; the tools appear as mcp__zimage__generate_image, etc.

Develop / test

uv sync
uv run pytest -q                       # unit tests (no network)
ZIMAGE_TEST_LIVE=1 uv run pytest -q    # + live tests against the box

Available Tools

3 tools
generate_imageA

Generate web-dev image(s) with Z-Image Turbo and save to output_path.

size: one of square|landscape|portrait|hero|wide|og|mobile, or pass width+height for exact pixels (image is generated at the nearest valid size then resized/cropped). n: 1-4 variants (paths get _1,_2,... ); seeds are base_seed, base_seed+1, ... Returns {images, seed, gen_size, output_size, seconds}.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNo
seedNo
sizeNosquare
stepsNo
widthNo
heightNo
promptYes
negativeNo
output_pathYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It explains generation behavior (resizing/cropping to nearest valid size, multiple variants using seeds, saving to output_path, return format). It does not mention overwrite behavior, auth requirements, or computational cost, but covers essential behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a lead sentence defining the action, followed by concise bullet-like explanations for key parameters (size, n, seeds, returns). Every sentence adds value, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description includes return format. It covers size flexibility, multi-image generation, and output saving. However, it lacks details on step parameter semantics, negative prompt effects, and potential edge cases (e.g., path conflicts). Still, it is largely complete for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description compensates by explaining size (named presets or width+height), n and seed inference, and the return object. It does not detail steps or negative prompt semantics, but most parameters are covered sufficiently for usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Generate web-dev image(s) with Z-Image Turbo and save to output_path', clearly identifying the verb (generate), resource (web-dev images), and output. Sibling tools (health, list_presets) are completely unrelated, so no confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit when/when-not guidance, but the purpose is distinct from siblings, and the parameter explanations implicitly define usage context. It would benefit from stating 'Use this to generate images, not for health checks or listing presets'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthA

Quick ComfyUI pre-flight: version, GPU, free VRAM.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, but the description discloses that the tool checks version, GPU, and free VRAM. For a simple read-only health check, this is sufficient behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no wasted words. Front-loaded with purpose and key details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple health check with no parameters and no output schema, the description covers the essential outputs. Could add units or format, but not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters with 100% schema coverage, so baseline is 4. The description adds meaning by specifying what data the tool returns, even without output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks version, GPU, and free VRAM, which are specific health indicators. It distinguishes from siblings 'generate_image' and 'list_presets' as a diagnostic pre-flight tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implied usage is before operations like image generation, as a pre-flight check. No explicit when-not-to-use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_presetsA

Return available size presets as name -> 'WxH'.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description does not disclose behavioral traits beyond the output format ('name -> 'WxH''). There is no mention of side effects, permissions, or state changes, which is insufficient for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded with the verb and resource. Every word earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (zero parameters, no nested objects, no output schema), the description is nearly complete. It could mention that the tool is safe to call, but it is sufficient for a straightforward info-retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is 100%. The description adds meaning by specifying the output format, which is above baseline. Baseline for 0 params is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return available size presets') and the resource, and distinguishes from sibling tools like 'generate_image' and 'health'. It is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives, but the purpose is clear enough that usage is implied. A more explicit note would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedgenerate_image
    • First observedhealth
    • First observedlist_presets

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: generate_image for creation, health for system status, list_presets for configuration. No overlap or ambiguity.

Naming Consistency4/5

Most tools follow verb_noun pattern (generate_image, list_presets), but 'health' is a noun alone. Minor deviation but still clear and readable.

Tool Count5/5

Three tools is well-scoped for an image generation server: core generation, health check, and preset listing. No unnecessary bloat.

Completeness5/5

The tool set covers the essential workflow: generate images, check system readiness, and retrieve presets. No obvious missing functionality for the stated purpose.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers