MeiGen AI Image Generation MCP
MeiGen MCP server lets AI assistants generate images, videos, and ecommerce assets using MeiGen Cloud, OpenAI-compatible APIs, or local ComfyUI—plus inspiration search, prompt enhancement, and workflow management.
Image generation: create images with models such as GPT Image 2, Nanobanana, Seedream, Midjourney V8.1, Flux 2 Klein, and Grok Imagine; supports reference images, aspect ratio, resolution, and quality settings.
Video generation: create videos with Seedance 2.0, Happyhorse, Veo 3.1, and Grok Video; supports first/last frames, reference video continuation, audio, tiers, durations, and resolutions.
Five ecommerce Skills: remove backgrounds, generate product detail images, create marketing posters, generate AI backgrounds, and upscale original images—all billed via MeiGen credits.
Inspiration & prompts: search the gallery semantically, get full prompts/images, and enhance simple ideas into professional prompts (offline/free).
Lookups & recovery: list models, list Skills and prices, check generation status, recover interrupted jobs, and upload reference images.
Local/ComfyUI control: manage ComfyUI workflows, and read/update local preferences (default style, model, ratio, style notes, favorites).
CLI mode: run one-shot generation from shell scripts or CI, with JSON output, local file upload, and no-wait submission options.
Includes an automation hook that automatically opens generated images in the macOS Preview application for immediate visual feedback and a more efficient design workflow.
Supports image generation by connecting to any OpenAI-compatible API endpoint, allowing users to leverage various external models and providers using a standardized interface.
New in 2.0.1:
generate_videoaccepts multiple reference videos and reference audio (referenceVideos/referenceAudios; local.mp4/.mov/.wav/.mp3files are uploaded automatically). Per-model limits come fromlist_models.New in 2.0.0: five guided ecommerce/image Skills, original-image enhancement, upload validation, request recovery and Codex setup. Use a Skill · Upgrade guide · HTTP API
What Is This?
An MCP server that helps your AI assistant create images, videos and ecommerce assets. Connect to the remote server (14 tools) or install the local npm server (17 tools). Works in Claude Code, Cursor, Codex, Windsurf, Roo Code, OpenClaw, Hermes Agent, and other compatible MCP hosts.
Version 2.0.0 adds five guided Skills: background removal, Product Detail Images, Marketing Poster, AI Backgrounds and Upscale. These run on MeiGen Cloud and require a MeiGen API key and purchased credits. The local server also supports OpenAI-compatible APIs and ComfyUI for general image generation; those providers do not run the five Skills.
Local general image generation supports three backends: MeiGen Cloud, OpenAI-compatible APIs, or local ComfyUI. Use
list_modelsfor current capabilities.Built-in 1,446 curated prompt templates from nanobanana-trending-prompts plus style-aware prompt enhancement
Callable image/video steps for upstream workflows, optional creative helpers, and a standalone CLI for shell scripts and CI
Related MCP server: MCP Doubao Seedream 4.0
See It in Action
Product Photo — 4 Directions in Parallel
"Create 4 product display images for this perfume, one of which should feature a model."
Process — AI uploads the reference image, crafts 4 distinct prompts, then generates all 4 in parallel:
Result — 4 creative directions delivered in under 2 minutes:
Generated images:
Quick Start
Ask your AI assistant to install
Copy this entire block into Codex, Claude Code, Cursor, or another AI assistant that can configure MCP servers. It can set up the connection and guide you through credentials. ChatGPT web uses the separate setup below.
Install MeiGen MCP for this AI client using this guide:
https://github.com/jau123/MeiGen-AI-Design-MCP#quick-start
Detect the current client and inspect its MCP configuration without printing
credentials. Preserve other servers and reuse an existing MeiGen entry.
Prefer Streamable HTTP at https://www.meigen.ai/api/mcp. For Codex, use
codex mcp add meigen --url https://www.meigen.ai/api/mcp
Add bearer_token_env_var = "MEIGEN_API_TOKEN" only after that local
environment variable is configured; otherwise leave authentication unset.
If I need automatic local-file preparation, ComfyUI, or local-only tools, use
the stdio command npx -y meigen@2.0.1 instead. Check that this exact npm
version exists before configuring it; report an unavailable version.
Guide me to enter my MeiGen key in local credentials settings or the launch
environment, never in this chat. Public lookups can be tested without a key.
Reload/reconnect, inspect the actual tool list, and call list_skills to verify.
If you cannot configure or reconnect this client, give me the exact manual
steps and say what remains unverified. If list_skills is missing, report it.
Do not upload images, generate anything, or spend credits during installation.For manual setup, jump to Codex / ChatGPT desktop, ChatGPT web, or the client-specific instructions below.
1. Prepare your account
You can browse inspiration and inspect models or Skill prices without a key. To generate with MeiGen:
Open API Keys in a desktop browser, sign in, and create a key. The key starts with
meigen_sk_; the mobile site currently redirects this page.Open your profile and select Top Up to add purchased credits to the same account; on mobile use Premium. API generation does not use daily free credits; the five Skills have no free attempts, including the first cutout. Use
list_skillsfor current Skill prices.Add the key to your MCP host's connection settings. Keep it out of chat messages and shared config files.
2. Choose one connection
Remote MCP — recommended | Local npm MCP 2.0.1 | |
Connection | Streamable HTTP at | Node.js process over stdio |
Tools | 14: MeiGen generation, gallery and the five Skills | The same 14, plus prompt enhancement, preferences and ComfyUI management |
Reference images | Public image link, an existing MeiGen URL, or actual attachment bytes readable by the host | Local files and public image links are prepared automatically |
Local extras | Result URLs; the host handles preview/download | General generation saves files; also CLI, offline prompt library and local ComfyUI |
Plugin extras | A bare MCP connection does not install commands, agents, output styles or hooks | The Claude Code plugin adds these; a bare npm connection also does not include them |
Updates | Backend changes arrive after server deployment; refresh/reconnect if the host caches tools | Local tool changes require an npm release and a client update |
Choose one server entry for this host to avoid duplicate tools. The two entries use the same MeiGen account and purchased credits.
Install for your client: Remote settings · Codex · Claude Code · Cursor / VS Code / Windsurf / Roo · OpenClaw · ChatGPT web.
3. Try your first Skill
After connecting, ask: “List the available Skills and their current prices.” This does not generate an image or spend generation credits. Then provide the required material and describe the result you want:
Skill / tool | Example request | Required | Optional | Output and cost |
Background removal — | “Remove this background and give me a transparent PNG.” | One source image | No creative brief needed | One cutout; charged from the first request |
Product Detail Images — | “Use this product photo to make a main shot, a detail close-up and a lifestyle image.” | One product image; resolve the desired count/modules | Product name, selling points, copy, logo, model photo, up to two extra product photos, language and marketplace | 1–6 images; one paid image per module |
Marketing Poster — | “Make one poster for a weekend coffee tasting.” | A brand, event, campaign or topic | Display copy, logo, up to three product images, one style reference, language and style | One poster; images are optional |
AI Backgrounds — | “Put this product on a sunlit stone counter.” | One product image; desired setting for custom mode | White/smart/custom mode, ratio and quality | One product image with a new background; not a transparent cutout |
Upscale — | “Make this original product photo clearer while keeping its appearance.” | One original still PNG/JPEG/WebP image | crisp preserves structure (default); creative reconstructs details and requires acceptance of changes | One enhanced image; video enhancement is not exposed |
The assistant chooses the tool, prepares images, checks status and presents preview/download links. You do not need to write prompts, UUIDs or API parameters. It asks only for missing essential information. An explicit request for a specified count, modules or quality already confirms that scope.
Product-detail batches: choose 1–6 modules in total. The presets are main shot (hero), close-up (detail), lifestyle (scene), texture/craft (material), how-to-use (usage) and brand story (brand); custom modules are also supported. MCP requires the assistant to pass modules explicitly, preventing extra images when an argument is omitted. You can specify only the count and let the assistant choose modules. Direct HTTP API calls still default to three images (hero, detail and scene) when modules are omitted. The assistant should calculate the batch cost from the live per-image price and requested count before submitting an unresolved batch.
Copy and quality: ask to preserve your wording when exact copy matters; otherwise the service can draft copy from your brief. State the desired text language. Product details and posters default to Fast; Pro costs more. AI Backgrounds defaults to smart/Fast; white mode uses a fixed output specification and ignores ratio/quality options. Current options and prices come from list_skills.
Poster fields: with autoCopy: true, content is a brief; with autoCopy: false, it is the visible wording to preserve (the selected language may translate it). Put style/layout/design directions in extraNotes or customStyle, not in verbatim content. extraNotes may also contain verified facts or explicitly requested display copy; design instructions in it are directions, not text to print verbatim. Use a catalog preset ID for styleId, not its display label. Omit both written style fields for Auto; nonempty customStyle overrides styleId. styleImage is the primary visual reference, with written style only as a compatible supplement; do not copy its products, wording or layout. Logo and product references preserve identity.
Output and timing: current Product Detail and Poster Fast/Pro tiers both use the 2K preset; quality is not resolution. Actual pixels depend on ratio and provider output. Use current list_skills specifications/prices and do not send an unsupported resolution argument. Queueing, planning, provider execution and image count affect completion time; no fixed number of seconds is guaranteed. Polling intervals and HTTP timeouts are not ETAs.
Valid MCP call example: this illustrative request already supplies the event copy and time. Only content below is the exact visible wording; extraNotes describes layout and preservation requirements, not additional text to print. Pass it to client.callTool(...). Use list_skills({skill: "brand-poster"}) for live options and price when choosing. It creates one paid poster without image material. The caller generates and saves a UUID for each real attempt, reuses it for recovery and does not blindly rerun this example.
{
"name": "generate_marketing_poster",
"arguments": {
"requestId": "8f729f7e-934e-4e2c-bae3-bf23a782f964",
"brand": "Coffee tasting",
"content": "Coffee tasting\nSaturday, 10:00–12:00",
"autoCopy": false,
"extraNotes": "Keep the supplied time. Use a clear headline and a small schedule block.",
"styleId": "minimalist",
"language": "en",
"ratio": "4:5",
"quality": "low"
}
}Images: for the local npm server, provide an actual file path or a public direct HTTPS image link. For remote MCP, the assistant uses upload_skill_image for external links or readable attachment bytes, then passes its imageUrl to the Skill. Existing images.meigen.ai, images.meigen.art or pbs.twimg.com HTTPS URLs can be used directly. upload_skill_image and /api/skills/upload require a positive purchased-credit balance for both local and remote MCP connections. Uploading does not start generation or spend generation credits. If your host cannot read an attachment, provide a direct image URL or use the local npm server.
Local preprocessing accepts source files up to 32 MiB and prepares references up to 4096px / 8 MiB, preserving PNG/WebP transparency. Remote uploads accept a public direct HTTPS image up to 8 MiB, or real base64 bytes up to 3 MiB decoded. Private-network URLs, redirects, authenticated links and IPv6-only sources are unsupported. These are reference image limits, not generated output specifications.
Upscale uses the original: pass an original local file or public direct HTTPS URL to upscale_image; do not resize it with a generic reference uploader. For readable attachment bytes, use upload_skill_image with purpose: "upscale". Originals may be up to 64 MiB / 64 million pixels; base64 remains limited to 3 MiB decoded. Only still JPEG/PNG/WebP is supported. Local uploads fully decode, auto-orient, remove metadata and preserve alpha and dimensions; encoding must fit below 9,500,000 bytes. If it cannot, provide a public original URL.
A source larger than 4096px on either edge or 16 million pixels returns upscale_resize_required before generation or charging. Explain that resizing can produce a result smaller than the original with limited clarity gain; after acceptance, use allowDownscale: true and a new requestId. Both MCP transports require confirmedCredits on the first call too: use the live list_skills quote within the accepted user or upstream workflow budget. This is a pre-dispatch check, not an atomic spending cap. For price_changed, obtain acceptance of the updated quote before submitting a new ID with the accepted confirmedCredits. Explain possible detail changes before choosing creative mode.
Direct HTTP API: developers without an MCP host can use the complete five-Skill API guide, including upload, run and recovery examples. GET /api/skills supplies capabilities and current prices.
Remote MCP Endpoint (zero-install, recommended)
In your host's MCP settings, select Streamable HTTP, enter https://www.meigen.ai/api/mcp, and add the HTTP header Authorization: Bearer YOUR_MEIGEN_API_KEY. The exact settings screen varies by host. Claude Code also supports:
# MEIGEN_API_TOKEN must already be set in this terminal's environment.
claude mcp add --transport http meigen https://www.meigen.ai/api/mcp \
--header "Authorization: Bearer $MEIGEN_API_TOKEN"The remote endpoint supports stateless Streamable HTTP, including the 2026-07-28 protocol and compatible 2025 clients. Stateless means no persistent MCP session is required. Accepted generation jobs and their billing records still live on the server, so an interrupted conversation can recover them.
There is no npm install or local server process. Changes to remote tools require a backend deployment, not an npm release; the client may need to reload its tool list. Remote generation still requires an internet connection. Local npm updates remain necessary when local tools or file-handling behavior change.
Codex / ChatGPT desktop (local Codex host)
Use this setup for Codex CLI, the Codex IDE extension, and the desktop app when using a local Codex host. These clients share MCP configuration on the same host. ChatGPT web has a different connection flow. See the official Codex MCP guide.
Remote — recommended for MeiGen Cloud and Skills:
codex mcp add meigen --url https://www.meigen.ai/api/mcp \
--bearer-token-env-var MEIGEN_API_TOKENThe equivalent configuration below also sets timeouts suitable for Skills. Merge it into ~/.codex/config.toml; if you used the command above, edit its existing meigen table instead of adding another one. Preserve your other servers.
[mcp_servers.meigen]
url = "https://www.meigen.ai/api/mcp"
bearer_token_env_var = "MEIGEN_API_TOKEN"
startup_timeout_sec = 30
tool_timeout_sec = 240To try public lookups before configuring a key, omit --bearer-token-env-var from the command and bearer_token_env_var from the TOML. Add them after setting the variable.
Set MEIGEN_API_TOKEN locally in the environment that launches Codex before using authenticated tools. This is your MeiGen key, not an OpenAI API key. Codex does not automatically load a project's .env.local. If a desktop launch does not inherit your terminal variables, enter the Authorization: Bearer … header through its private MCP connection settings when available, or configure http_headers.Authorization yourself in your private user-level config. When using a direct Authorization header, remove bearer_token_env_var so the connection does not depend on that missing environment variable. Keep the key out of chat and shared project files.
Local — for automatic local-file preparation, ComfyUI and local-only tools: use this entry instead of the remote entry. Node.js 22 or newer is recommended.
[mcp_servers.meigen]
command = "npx"
args = ["-y", "meigen@2.0.1"]
env_vars = ["MEIGEN_API_TOKEN"]
startup_timeout_sec = 90
tool_timeout_sec = 240This forwards your locally configured MEIGEN_API_TOKEN to the npm process. To register just the local command with the CLI, use codex mcp add meigen -- npx -y meigen@2.0.1, then add the environment forwarding and timeout settings shown above. meigen init codex is not supported; use Codex's own MCP configuration.
Restart/reconnect after setup. In Codex CLI, codex mcp list checks registration and /mcp shows connection status. Then ask “List MeiGen's available Skills and current prices” to verify an actual tool call without generating or spending credits. A saved configuration alone does not prove that the server connected. Longer video jobs may need a longer tool timeout.
ChatGPT web (public lookups only)
For accounts and workspaces with custom MCP access, enable Developer mode, create a custom remote app/plugin, enter https://www.meigen.ai/api/mcp, and choose No Authentication. After connecting, select it in a conversation and request a public model, Skill-price or gallery lookup. Follow OpenAI's Developer mode setup for the current settings and availability. A chat message alone cannot install a local npm server into ChatGPT web.
Paid MeiGen tools are not supported through this ChatGPT web connection yet. OpenAI's hosted MCP client cannot send custom API keys, while MeiGen currently requires a Bearer API key for generation, image upload and check_skill recovery (check_generation by known generationId remains public; requestId recovery requires a key). Full support needs a MeiGen OAuth integration. Use Codex or another client that supports Bearer headers for those tools; do not put the key in chat or in the server URL.
Local npm MCP (Node.js)
Node.js 22 or newer is recommended. The examples below pin meigen@2.0.1. After changing an installed version or connection settings, restart or reconnect the host. The five Skills still call MeiGen Cloud and need the account setup above.
Claude Code Plugin (npm, local tools)
# Add the plugin marketplace
/plugin marketplace add jau123/MeiGen-AI-Design-MCP
# Install
/plugin install meigen@meigen-marketplaceRestart Claude Code after installation (close and reopen, or open a new terminal tab).
Alternative marketplace — also available via wshobson/agents (30k+ stars):
/plugin marketplace add wshobson/agents
/plugin install meigen-ai-design@claude-code-workflowsThis marketplace doesn't bundle MCP server config. After installing, add to your project's
.mcp.json:{ "mcpServers": { "meigen": { "command": "npx", "args": ["-y", "meigen@2.0.1"] } } }
First-Time Setup
Free features work immediately after restart — try:
"Search for some creative inspiration"
The Claude Code plugin includes a setup command:
/meigen:setupFor the five Skills, choose MeiGen Cloud and configure your MeiGen key in the connection settings. For general image generation, the wizard also offers ComfyUI and OpenAI-compatible APIs. Restart Claude Code after changing configuration. Do not paste secrets into a conversation.
Cursor / VS Code / Windsurf / Roo Code
One command to set up MeiGen for any supported AI coding tool:
npx -y meigen@2.0.1 init cursor # Cursor
npx -y meigen@2.0.1 init vscode # VS Code / GitHub Copilot
npx -y meigen@2.0.1 init windsurf # Windsurf
npx -y meigen@2.0.1 init roo # Roo Code
npx -y meigen@2.0.1 init claude # Claude Code (project-level)This writes the correct MCP config file with the right format and path for your tool. If a config file already exists, MeiGen is merged in without overwriting your other servers.
init writes a configuration that follows the default npm release tag; the manual examples above pin 2.0.1. To pin an initialized connection too, change its args to ["-y", "meigen@2.0.1"] and restart the host.
OpenClaw
Install the full plugin from ClawHub (includes Skills and an explicit MCP connection; other features depend on the loader):
openclaw plugins install clawhub:meigen-ai-designOr install only the skill (no commands/agents):
npx clawhub@latest install creative-toolkitUse as CLI (no MCP host required)
For shell scripts, CI pipelines, or anyone who wants AI image generation without an MCP host, MeiGen ships a one-shot gen command in the same npm package.
# Set your token locally (create it at https://www.meigen.ai/profile/api-keys on desktop)
export MEIGEN_API_TOKEN=meigen_sk_...
# Generate
npx -y meigen@2.0.1 gen --prompt "a calico cat in a sunlit kitchen"
# With a specific model + aspect ratio
npx -y meigen@2.0.1 gen -p "tech logo" -m midjourney-v8.1 -r 1:1
# With a reference image (local file auto-uploaded)
npx -y meigen@2.0.1 gen -p "product hero shot" --ref ~/Desktop/bottle.jpg
# Submit only — print generationId without polling (good for CI)
npx -y meigen@2.0.1 gen -p "..." --no-wait
# Machine-readable output (good for jq pipes)
npx -y meigen@2.0.1 gen -p "..." --json | jq -r '.imageUrls[0]'CLI image output is saved to ~/Pictures/meigen/ (override with MEIGEN_OUTPUT_DIR). The five Skills return result links; your host can preview or download them.
meigen gen --help lists all flags.
Other MCP-Compatible Hosts
Add to your MCP config (e.g. .mcp.json, claude_desktop_config.json):
{
"mcpServers": {
"meigen": {
"command": "npx",
"args": ["-y", "meigen@2.0.1"],
"env": {
"MEIGEN_API_TOKEN": "meigen_sk_..."
}
}
}
}Free features (inspiration search, prompt enhancement, model listing) work without any API key.
Hermes Agent (NousResearch)
Hermes Agent is a first-class MCP client — add MeiGen to ~/.hermes/config.yaml:
mcp_servers:
meigen:
command: "npx"
args: ["-y", "meigen@2.0.1"]
env:
MEIGEN_API_TOKEN: "meigen_sk_..."
timeout: 2700 # generate_video polls until the server reports a terminal state (long videos can run 15+ min) — default 120s is not enough
connect_timeout: 120 # first npx download can take a minuteThe
timeout: 2700andconnect_timeout: 120overrides are important — Hermes defaults (120s / 60s) are tuned for short-running tools and will time out on video generation or first-run npx downloads.
Compose with an existing workflow
An upstream Skill can write N scripts, call MeiGen for each first frame, then pass completed frame URLs to generate_video(firstFrame=...). Reference clips (referenceVideos, referenceAudios) can be mixed in the same call and addressed from the prompt as "Video 1" / "Audio 1". It owns prompts, models/providers, ratios, approved count/budget and presentation. Creative planning and plugin agents are optional; resolved requests do not need repeated approval at every step.
For MeiGen jobs, persist one UUID requestId and exact inputs per logical step. Use wait: false for an immediate task handle; local npm also accepts download: false. Existing local defaults remain wait: true, download: true; asynchronous calls skip download. Remote MCP returns URLs and has no download setting. Recover with check_generation using the original requestId or generationId. Request lookup requires a MeiGen key belonging to the same account; known generation-ID status remains public remotely. Read structuredContent for status, handles, URLs, errors and polling advice.
Local npm bounds concurrent API submissions at four, with polling/downloads outside those slots; ComfyUI executes one job at a time. The caller also bounds outstanding work and reserves in-flight costs within its approved budget. Actual backend rate limits and Retry-After remain authoritative. Parallel videos are allowed within authorized scope; no ten-image total or atomic batch spending guarantee applies.
Local waiting retries temporary status-query failures within its observation budget; cancellation stops waiting while IDs remain recoverable. A completed media mismatch keeps the actual result and returns requestedMediaType plus review_media_type. A missing or unrecognized recovery route returns endpoint_unavailable / check_backend, never permission to submit again. Deploy the matching backend first and preserve recovery APIs during rollback.
See persistent step IDs, frame/video calls and recovery rules.
MCP Tools
Both entries expose the following 14 cloud tools. The local npm entry adds three local tools, for 17 total. Read-only lookups do not spend generation credits; check_skill still requires the key that owns the request.
Tool | Entry | Billing / purpose |
| Both | No generation charge; search inspiration with image previews, at most 3 per call. With a MeiGen key configured the call is authenticated and counts against that account's daily search quota instead of the shared per-IP budget. Local npm also bundles 1,446 prompts. |
| Both | No generation charge; full prompt, images and metadata for a gallery entry. |
| Both | No generation charge; current supported models and options. |
| Both | Generate an image. Remote uses MeiGen purchased credits; local also supports configured BYOK/ComfyUI providers. |
| Both | MeiGen key and purchased credits; use the current model options from |
| Both | No generation charge; recover by generationId or authenticated requestId. |
| Both | No key or generation charge; current Skill inputs, defaults and prices. |
| Both | MeiGen key; prepares a reference without spending generation credits. |
| Both | MeiGen purchased credits; one transparent cutout. |
| Both | MeiGen purchased credits; 1–6 images, billed per module. |
| Both | MeiGen purchased credits; one poster, optional image references. |
| Both | MeiGen purchased credits; one product image with a white, smart or custom background. |
| Both | MeiGen purchased credits; faithful or creative enhancement from the original image. |
| Both | Same MeiGen key; no additional generation charge. Returns completed images, failed modules and refund states. |
| Local only | Local prompt enhancement; no generation charge. |
| Local only | Read/write local preferences; no generation charge. |
| Local only | Manage local ComfyUI workflows; no MeiGen generation charge. |
Local synchronous image/video generation saves files by default; use download=false to skip, or wait=false to return a handle without downloading. Skills return preview/download links in both entries. Tools with the same name can have transport-specific input schemas; clients should read the connected server's tool list.
Slash Commands
These commands require the Claude Code plugin; connecting a bare MCP server does not install them.
Command | Description |
| Quick generate — skip conversation, go straight to image |
| Search 1,446 curated prompts for inspiration |
| Browse and switch AI models for this session |
| Interactive provider configuration wizard |
Standalone CLI Mode
For shell scripts, CI pipelines, and terminal users who don't run an MCP host:
export MEIGEN_API_TOKEN=meigen_sk_...
npx -y meigen@2.0.1 gen --prompt "a calico cat in a sunlit kitchen"
npx -y meigen@2.0.1 gen -p "logo design" -m midjourney-v8.1 -r 1:1 --jsonSee Use as CLI (no MCP host required) for the full flag list.
Smart Agents
The Claude Code plugin includes optional helpers for general image generation; direct tool calls remain supported. The five Skills handle their own planning and do not require these agents:
Agent | Purpose |
| Optional executor; preserves caller parameters and returns task handles/results |
| Writes multiple distinct prompts for batch generation (runs on Haiku for cost efficiency) |
| Deep gallery exploration without cluttering the main conversation (runs on Haiku) |
Output Styles
Switch creative modes with /output-style:
Creative Director — Art direction mode with visual storytelling, mood boards, and design thinking
Minimal — Just images and file paths, no commentary. Ideal for batch workflows
Automation Hooks
Automatic Preview is off by default. Set MEIGEN_AUTO_OPEN=1 in the plugin host environment to open saved images on macOS; workflow callers otherwise control presentation. Async/no-download calls have no saved image to open.
Config Check — Validates provider configuration on session start, guides setup if missing
Optional preview — Set
MEIGEN_AUTO_OPEN=1to open saved images in Preview (macOS)
The local npm server supports three backends for generate_image. Configure one or multiple. The remote server and all five Skills use MeiGen Cloud; Skills do not accept a BYOK provider or run on ComfyUI.
ComfyUI — Local & Free
Run generation on your own GPU with full control over models, samplers, and workflow parameters. Import any ComfyUI API-format workflow — MeiGen auto-detects KSampler, CLIPTextEncode, EmptyLatentImage, and LoadImage nodes.
{
"comfyuiUrl": "http://localhost:8188",
"comfyuiDefaultWorkflow": "txt2img"
}Useful for models you run locally. Generation can stay on your machine when you use local files and a local workflow. The MeiGen Skills still use cloud services.
MeiGen Cloud
Cloud API with multiple models: GPT Image 2.0, Nanobanana 2, Seedream 5.0, and more. No GPU required.
Get your API key and credits:
Sign in and open API Keys in a desktop browser.
Create a key starting with
meigen_sk_and save it in your MCP connection settings.On the same account, open Profile → Top Up or mobile Premium to buy credits.
{ "meigenApiToken": "meigen_sk_..." }General image resolution & quality — generate_image accepts model-dependent options. Use list_models for the current default and supported values:
resolution: e.g."1K"/"2K"/"4K"— upgrade for posters, prints, wallpapersquality: e.g."low"/"medium"/"high"— use"low"for quick drafts and thumbnails
Seedance 2.0 video now renders native 4K — but only on the pro tier (mini/fast cap at 480p/720p); pass tier: "pro" for 1080p/4K output.
Each model exposes its own supported resolutions and quality tiers — run list_models to see what's available. For up-to-date pricing across all models, see meigen.ai/model-comparison.
Bring Your Own API (OpenAI-Compatible)
Connect any image generation API that follows the OpenAI format — Together AI, Fireworks AI, DeepInfra, SiliconFlow, or your own endpoint. Just provide your key, base URL, and model name:
{
"openaiApiKey": "sk-...",
"openaiBaseUrl": "https://api.together.xyz/v1",
"openaiModel": "black-forest-labs/FLUX.1-schnell"
}All three providers support reference images. MeiGen and OpenAI-compatible APIs accept URLs directly; ComfyUI accepts both URLs and local file paths, injecting them into LoadImage nodes in your workflow.
Configuration
Claude Code Plugin Setup
/meigen:setupThis command belongs to the Claude Code plugin. Other MCP hosts should use their own connection settings. For the five Skills, configure a MeiGen key; OpenAI-compatible credentials and ComfyUI only apply to general image generation. Store credentials in connection settings or the local config file, not a chat message.
Config File
Configuration is stored at ~/.config/meigen/config.json. ComfyUI workflows are stored at ~/.config/meigen/workflows/.
Environment Variables
Environment variables take priority over the config file.
Variable | Description |
| MeiGen API key; required for all five Skills and MeiGen generation |
| Local server API origin; default |
| Private local receipt directory (default |
| Local reference upload gateway; default |
| Your API key (any OpenAI-compatible provider) |
| API base URL — change this to use Together AI, Fireworks AI, etc. |
| Model ID supported by your endpoint |
| ComfyUI server URL (default: |
| Override the local save directory for generated images (default: |
| Override the local save directory for generated videos (default: |
| Linux only — when |
| Linux only — same logic as |
Privacy
MeiGen MCP respects your privacy. Here's what happens with your data:
ComfyUI (local) — A local workflow with local files can run without cloud generation. Gallery queries and MeiGen Skills still use external services.
MeiGen Cloud and Skills — Prompts and reference images are processed by MeiGen and its generation providers; result images are stored on Cloudflare R2. See MeiGen Privacy Policy.
OpenAI-compatible — Prompts and reference images are sent to the configured API endpoint. See your provider's privacy policy.
Reference image upload — Local files use the configured upload gateway (default
gen.meigen.ai) and Cloudflare R2. Local MCP ordinary generation and standard Skill references target 4096px / 8 MiB; the standalonemeigen genCLI retains its 2 MiB target. Upscale keeps original dimensions and uses the limits above. Skill preparation removes metadata and preserves transparency; GIF references use the first frame. Remote Skill uploads require a MeiGen key. Reference URLs are accessible to anyone with the link. Preserve accepted URLs for retries, keep your originals and download results you need; URLs are not promised as permanent archival storage. ComfyUI can use local paths without uploading.Gallery search — With a search query, the MeiGen API is queried (your query text is sent to
www.meigen.ai); category browsing and offline fallback use bundled local data. Prompt enhancement runs locally with no external calls.
The local npm server adds no telemetry. Requests sent to MeiGen are subject to its service and privacy policies.
Custom Storage Backend
If you prefer to use your own S3/R2 bucket for reference image uploads, set the UPLOAD_GATEWAY_URL environment variable or uploadGatewayUrl in ~/.config/meigen/config.json to point to your own presign endpoint. The endpoint must implement:
POST /upload/presign
Content-Type: application/json
Request: { "filename": "photo.jpg", "contentType": "image/jpeg", "size": 123456 }
Response: { "success": true, "presignedUrl": "https://...", "publicUrl": "https://..." }The presignedUrl is used for a PUT upload, and publicUrl is the publicly accessible URL returned to the user. This option applies to local uploads. For standard Skills, the local server automatically prepares a custom CDN URL through authenticated /api/skills/upload before submission. Upscale passes a public original URL to the backend for source validation and preparation. Custom gateway images must remain publicly reachable without authentication or redirects.
Troubleshooting
Problem | Next step |
New Skills do not appear | Remote: refresh/reconnect the tool list after the backend is deployed. Local: confirm the connection runs |
Invalid or missing key | Create/check the key in desktop API Keys and update the MCP connection's header or |
Insufficient credits, but the website shows a balance | API calls use purchased credits only, never daily free credits. Use Profile → Top Up on the same account; then ask the assistant to continue. A rejected call should not be polled. |
The image cannot be uploaded | Use a real local file (local npm) or a public direct HTTPS image link. Check format/size and remove login or redirect requirements. If the host cannot read attachments, use a direct link. Upload failure has not started a generation. |
A tool timed out or the host restarted | Recover Skills with |
Some product-detail modules failed | Keep and display completed images; show the failed modules and their reported refund state. Do not regenerate the whole batch or add paid replacements automatically. |
The request is rate-limited | Follow the returned waiting instruction. A daily request limit is separate from the credit balance; repeated polling or topping up does not reset it. |
I configured OpenAI or ComfyUI, but a Skill still asks for a MeiGen key | Those backends support local general image generation. The five Skills use MeiGen Cloud and purchased credits. |
For Skill client implementers: generate requestId internally, keep it with the original inputs, and query check_skill after interruptions. Follow nextAction; processing responses normally suggest a 10-second polling interval. Reuse returned retryParameters exactly, including uploaded URLs. A new ID represents a new paid attempt; changing inputs under an existing ID is rejected. These IDs belong in the client, not in questions to the user.
Upgrading from 1.4.0
Remote MCP: keep the endpoint and your valid MeiGen key. After backend deployment, reconnect and call
list_skills; expect five Skills and 14 tools. Updating npm alone does not deploy the APIs.Local npm: change pinned configurations to
meigen@2.0.1and restart; global installations can runnpm install -g meigen@2.0.1. Expect 17 tools. Check that the version is available on npm first.Plugin users: update the Claude marketplace plugin, OpenClaw native plugin or standalone ClawHub Skill separately. Updating npm alone does not replace installed instruction files. Do not add a second MCP entry when the plugin already supplies one.
Composable calls: local
wait: true/download: trueremain defaults. New workflows should persist UUIDrequestId, usewait: falseand recover by that ID. Remote legacyattemptIdremains accepted; older receipts cannot retroactively prove every historical parameter mismatch. Update plugin instructions as well as the server to get the optional creative flow.Existing configuration and jobs: general generation keeps its MeiGen/OpenAI/ComfyUI configuration; the five Skills need a MeiGen key and purchased credits. Preserve IDs and inputs for interrupted jobs and recover them; an upgrade is not a reason to resubmit a paid request.
Input changes: local Skill paths must be absolute,
~/orfile://; relative paths are rejected instead of being resolved against a hidden process directory. Images lose metadata, GIF references use the first frame, and Upscale needs a still original. Keep original assets and download results you need; result links are not a permanent-storage guarantee.
Releasing
The npm package version 2.0.1 is separate from the MCP protocol date and SDK version.
Maintainers: follow RELEASING.md for the 2.0.1 build, package checks and publishing process. Store NPM_TOKEN only in this repository's ignored .env.local as described there. It authorizes npm publishing and is separate from the MEIGEN_API_TOKEN used by customers. Never include either credential in commits or the published package.
License
MIT — free for personal and commercial use.
Available Tools
17 toolscheck_generationARead-only
Read an existing MeiGen job without new charges. Provide generationId OR the caller requestId UUID. requestId lookup works after a lost submit response, process restart or another host. Follow structured status and nextAction; do not turn an uncertain status into a new paid UUID.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | No | Original workflow step UUID supplied to generate_image/video; authenticated lookup across hosts. Provide exactly one identifier. | |
| generationId | No | Accepted generation ID. Provide exactly one of generationId and requestId. | |
| requestedMediaType | No | Original workflow step intent, when known. Preserve it from nextAction.arguments to detect a completed result of the other media type; this does not change or resubmit the job. |
Output Schema
| Name | Required | Description |
|---|---|---|
| urls | Yes | |
| error | No | |
| status | Yes | |
| deduped | No | |
| modelId | No | |
| success | Yes | |
| imageUrl | No | |
| provider | No | |
| videoUrl | No | |
| mediaType | No | |
| requestId | No | |
| savedPath | No | |
| nextAction | No | |
| creditsUsed | No | |
| generationId | No | |
| creditsStatus | No | |
| receiptWarning | No | |
| downloadWarning | No | |
| observationEnded | No | |
| pollAfterSeconds | No | |
| requestedMediaType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already include readOnlyHint=true, and the description adds value beyond that by guaranteeing no new charges and explaining that requestId lookup survives restarts and host changes. It also instructs the agent to follow structured status and nextAction, giving practical behavioral guidance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four concise sentences, each earning its place: purpose, identifier rule, recovery scenario, and action guidance. The most important information is front-loaded, with no redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema, full parameter documentation, and readOnly annotation, the description covers the essential operational details: what the tool reads, how to identify the job, recovery scenarios, and a caution against unnecessary paid actions. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds useful semantic guidance by clarifying the OR relationship between generationId and requestId and explaining when requestId lookup is valuable, which is not fully captured in the schema. It does not add much for requestedMediaType, but the schema already documents that parameter well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific action and resource: 'Read an existing MeiGen job without new charges.' This clearly distinguishes a read/status operation from the generation siblings like generate_image and generate_video, and clarifies the non-billing aspect upfront.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete when-to-use context: after a lost submit response, process restart, or another host. It also explicitly warns against converting an uncertain status into a new paid UUID, effectively telling the agent when not to fall back to generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_skillARead-only
Read all skill images, partial failures and refund states. No new charges. Follow nextAction: wait afterSeconds before polling (normally 10s); recover only when it says retry_request, with its exact parameters. Payment/auth/input failures and daily limits require the indicated action instead of polling. Show completed resource links.
| Name | Required | Description | Default |
|---|---|---|---|
| skill | Yes | ||
| requestId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=true, and the description meaningfully extends this with 'No new charges,' concrete retry timing, and a conditional recovery rule. This gives the agent operational expectations beyond the structured annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core read behavior and side-effect note, followed by precise polling and recovery rules. No filler; the closing instruction about resource links adds actionable output guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description covers what the response should surface ('Show completed resource links'), how to poll, when to stop polling, and how to handle retryable versus non-retryable failures. The remaining gap is explicit parameter guidance, already penalized separately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never names or explains the two parameters, skill and requestId. It implies a request context via 'retry_request, with its exact parameters,' but does not clarify that requestId identifies the operation or how skill selects the resource, leaving the agent to rely solely on field names and enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Read') and resource ('skill images, partial failures and refund states'), and clarifies the operation has no side effects ('No new charges'). It does not explicitly contrast with sibling check_generation, but the skill-specific wording makes the tool's role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit polling instructions: 'wait afterSeconds before polling (normally 10s)' and 'recover only when it says retry_request, with its exact parameters.' It also directly handles when-not-to-poll by stating that payment/auth/input failures and daily limits require a different action instead of polling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comfyui_workflowADestructive
Manage ComfyUI workflow templates: list, view parameters, import from file, modify settings, or delete.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Workflow name. Required for view/modify/delete. For import, used as the save name (defaults to filename). | |
| input | No | Input field name to modify (for modify action). E.g. "steps", "cfg", "sampler_name", "ckpt_name". | |
| value | No | New value as JSON (for modify action). Examples: "30", "\"euler\"", "7.5", "true". | |
| action | Yes | Action to perform on ComfyUI workflows | |
| nodeId | No | Node ID to modify (for modify action). Use "view" action first to see available node IDs. | |
| filePath | No | Path to a ComfyUI API-format workflow JSON file (for import action). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive action (destructiveHint=true, readOnlyHint=false), and the description's action list is consistent with that. However, the description adds no extra behavioral details such as permanence of deletion, side effects of modification, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the primary verb and lists all supported actions without fluff. It is appropriately sized for the tool's scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a complex multi-action interface with no output schema, and the description does not mention what actions return (e.g., whether view returns the workflow JSON or list returns names). While the param descriptions are thorough, the lack of output/return info leaves a gap for agents.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage with detailed descriptions for all six parameters, including the enum for action and examples for modify values. The description itself adds no parameter semantics beyond the schema, so baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool manages ComfyUI workflow templates and enumerates the five supported actions (list, view, import, modify, delete). This distinguishes it from sibling tools focused on generation or search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for workflow template management but does not explicitly state when to use this tool over alternatives, nor does it mention any prerequisites or exclusions. However, the action list provides clear context for when each action is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enhance_promptARead-only
Transform a simple idea into a professional image generation prompt. Use when the user provides a brief description (e.g., "a cat in a garden") and needs a detailed, high-quality prompt. Combine with gallery inspiration for best results. Free, no API key needed.
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | Target visual style: realistic (photorealistic), anime (2D/Japanese), illustration (concept art). Use "realistic" for general/photorealistic generation (GPT Image, Nanobanana, Seedream, Midjourney V8.1 in default mode, etc.). Use "anime" when the user wants anime/illustration output — V8.1 and most general-purpose models follow the prompt and benefit from explicit anime trigger words; the default "realistic" produces prompts poorly suited for stylized output. | realistic |
| prompt | Yes | The simple prompt to enhance (e.g., "a cat in a garden") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description's mention of 'free, no API key needed' adds minor behavioral context. It does not detail rate limits or failure modes, but the tool's simplicity and annotation coverage make this acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each serving a purpose: stating the transformation, providing a use case, and offering a tip. No wasted words, front-loaded with the main function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with two parameters and no output schema, the description covers the core functionality, usage context, and accessibility. It is sufficiently complete for an AI agent to understand and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions. The description's mention of 'simple idea' aligns with the prompt parameter, but adds no new information beyond the schema. Baseline 3 is appropriate given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: transforming a simple idea into a professional image generation prompt. It provides a concrete example ('a cat in a garden') and implies its role as a prompt enhancer distinct from image generation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool ('when the user provides a brief description and needs a detailed, high-quality prompt') and suggests combining with gallery inspiration. It does not explicitly list when not to use or compare to siblings, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ai_backgroundA
Create one product photo with a new background. Required: one actual product photo. White mode produces a fixed white background; smart chooses a scene; custom needs a background description, inferred from the request when present (e.g. a beach). Background references are not required. Use remove_background for transparent PNG cutouts. Requires a MeiGen API key configured for this local npm server (MEIGEN_API_TOKEN or saved local configuration), and purchased credits only. Local files and external public HTTPS URLs are uploaded automatically. Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation. Preserve caller-supplied inputs. Infer omitted optional settings from the request and defaults; when only a requested count is known, choose suitable modules unless the caller selected them. Ask only for missing required material or unresolved scope. An explicit user request or authorized upstream workflow establishes its count, quality and budget: do not reconfirm that scope or add paid images. The caller persists a requestId for each logical step; never ask an end user for technical IDs. Use live list_skills prices and account for in-flight charges when planning within a budget; a batch is not atomic and an estimate is not a server-enforced spending cap. Recovery actions take precedence over generic retry advice: check an interrupted submission, preserve exact retryParameters and do not create a new ID or replace failed modules automatically. Retry a temporary upload at most once after the suggested wait. Return structured status, handles, errors and completed URLs to the caller; it owns progress, previews, downloads and final presentation. If interacting directly with the user, explain the problem and a concrete next step in their language. Describe image details only after actual inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | white=fixed white-background output; smart=AI chooses a suitable scene (default); custom=customPrompt required. | |
| ratio | No | Supported output ratio from list_skills for this Skill; omit for its default. Smart/custom only; auto matches source proportions. Ignored in white mode. | |
| quality | No | Smart/custom only: fast=1K default, hd=2K; ignored in white mode. | |
| requestId | Yes | Generate a new UUID for a new paid request; reuse the SAME requestId and inputs on retry. Use check_skill after interruptions. | |
| customPrompt | No | Required in custom mode: describe the desired background and lighting. | |
| productImage | Yes | Required product photo: preserve the product while replacing its surroundings. Do not pre-remove its background. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden and meets it thoroughly: it discloses authentication requirements (MEIGEN_API_TOKEN or saved local config), purchased-credit gating, automatic upload of local files and HTTPS URLs, non-atomic batch behavior, lack of a server-enforced spending cap, exact retry semantics (same requestId, at most one retry after suggested wait), and recovery-action precedence over generic retry advice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the opening sentences are tight, but the description becomes a dense wall of operational policy — budget planning, requestId persistence, user-interaction language, and progress ownership — with some redundancy (e.g., asking only for missing material vs. explaining problems to the user). Most sentences carry information, but the length strains conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 6-parameter tool with no annotations and no output schema, the description is unusually complete: it covers auth, cost accounting, upload behavior, retry/recovery semantics, and return expectations ('structured status, handles, errors and completed URLs'). The only minor gap is that it doesn't describe the concrete shape of the success/failure return structure, but it says the caller owns presentation, which covers the essentials.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, and the description adds real value on top: mode/customPrompt interdependence ('custom=customPrompt required'), ratio and quality being smart/custom-only and ignored in white mode, and nuanced productImage semantics (do not pre-remove background, no credentials/fragments/custom ports, GIF first frame, never invent paths). This exceeds the baseline meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource — 'Create one product photo with a new background' — and the three modes (white/smart/custom) are clearly enumerated. It distinguishes itself from siblings by explicitly naming remove_background for transparent PNG cutouts, so an agent can differentiate it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: 'Use remove_background for transparent PNG cutouts', 'Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation', and 'Use live list_skills prices'. It also directs recovery flows to check_skill and names the sibling for pricing lookups, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageADestructive
Generate an image using AI. Supports MeiGen platform, local ComfyUI, or OpenAI-compatible APIs. Tip: get prompts from get_inspiration() or enhance_prompt(), and use gallery image URLs as referenceImages for style guidance. For Midjourney V8.1, an optional style reference can be passed by appending --sref <code> at the end of the prompt — only when the user provides a Midjourney style code (numeric or text). Do NOT pass URLs or local paths via --sref; for any image-based reference, use the referenceImages parameter instead.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Image size for OpenAI-compatible providers: "1024x1024", "1536x1024", "auto". MeiGen/ComfyUI: use aspectRatio instead. | |
| wait | No | MeiGen only: false returns the accepted generation ID immediately; poll check_generation separately. Default true waits for completion. Submit-only requires requestId. | |
| model | No | Model name. For OpenAI-compatible providers: any model ID your endpoint supports. For MeiGen: use model IDs from list_models (e.g. "gpt-image-2", "grok-image" = xAI Grok Imagine Quality, 1K/2K, supports image-to-image, "nanobanana-2", "seedream-4.5", "flux2-klein"). | |
| prompt | Yes | The image generation prompt | |
| modelId | No | Alias of model for portable workflow calls. If both are present they must match. | |
| quality | No | Image quality. MeiGen gpt-image-2: "low" / "medium" / "high". OpenAI-compatible providers also accept "high". | |
| download | No | Save the completed result locally (default true). Set false for URL-only workflows. Ignored when wait=false. Other providers may return inline image content when local saving is disabled. | |
| provider | No | Which provider to use. Auto-detected from configuration if not specified. | |
| workflow | No | ComfyUI workflow name to use (from comfyui_workflow list). Uses default workflow if not specified. | |
| requestId | No | Persistent UUID for this workflow step. Required when wait=false. Reuse with identical inputs after interruption, including after MCP restart; use a new UUID for a new generation. Omit only for a new interactive generation. | |
| resolution | No | Resolution tier. MeiGen: "1K" / "2K" / "3K" / "4K" — each model supports a subset (list_models reports resolutions when applicable). OpenAI: not used (use size instead). | |
| aspectRatio | No | Aspect ratio for MeiGen provider. Use "auto" (recommended, default when omitted) to let MeiGen infer the best ratio from the prompt content. Explicit values: "1:1", "3:4", "4:3", "16:9", "9:16", "21:9", "2:3", "3:2", "4:5", "5:4", etc. (model-dependent). ComfyUI: use comfyui_workflow modify to adjust dimensions before generating. | |
| modelVariant | No | Optional model variant from live list_models, such as a supported GPT Image 2.5 variant. Forwarded unchanged to MeiGen. | |
| negativePrompt | No | Negative prompt for OpenAI-compatible providers. ComfyUI: use comfyui_workflow modify to set negative prompt in the workflow before generating. | |
| referenceImages | No | Image references for style/content guidance. Accepts direct public HTTPS URLs without credentials/fragments or accessible absolute local paths. Relative paths are rejected. Local PNG/JPEG/WebP/GIF references up to 32 MiB and 64 million pixels are fully decoded, stripped of metadata and prepared up to 4096px, preserving transparency. For ComfyUI: local files are passed directly to the workflow (requires LoadImage node). Sources: gallery URLs from search_gallery/get_inspiration, URLs from previous generate_image results, or local file paths. |
Output Schema
| Name | Required | Description |
|---|---|---|
| urls | Yes | |
| error | No | |
| status | Yes | |
| deduped | No | |
| modelId | No | |
| success | Yes | |
| imageUrl | No | |
| provider | No | |
| videoUrl | No | |
| mediaType | No | |
| requestId | No | |
| savedPath | No | |
| nextAction | No | |
| creditsUsed | No | |
| generationId | No | |
| creditsStatus | No | |
| receiptWarning | No | |
| downloadWarning | No | |
| observationEnded | No | |
| pollAfterSeconds | No | |
| requestedMediaType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnlyHint=false and destructiveHint=true; the description adds a valuable behavioral rule: only append --sref for Midjourney when a style code is provided, and never pass URLs or local paths through it. This prevents a real misuse and clarifies referenceImages as the image-reference pathway. It does not explicitly discuss costs or file-writing side effects in the main description, but the annotations and schema cover part of that burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences are front-loaded with the core operation and platform support. The tip and sref warning earn their place; there is no filler and no redundant restatement of what the schema already documents.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 15 parameters and multiple provider backends, the description plus the rich schema covers the critical call decisions: provider, async wait behavior, reference images, and Midjourney style handling. It does not enumerate every provider quirk, but an output schema exists and the parameter descriptions are detailed, so the definition is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents 100% of the 15 parameters, so the baseline is 3. The description adds meaning beyond the schema by explaining the Midjourney --sref modifier (only when a style code exists) and by explicitly routing image-based references to referenceImages instead of the prompt.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation ('Generate an image using AI') and names the supported backends (MeiGen, local ComfyUI, OpenAI-compatible APIs). It is generic enough to be distinguished from specialized siblings like generate_marketing_poster or generate_ai_background, but it does not explicitly name those alternatives for direct differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tip points the agent to get_inspiration/enhance_prompt for prompt quality and to gallery URLs for style reference, giving useful context about how to compose a call. However, it does not state when to choose this general generator over the specialized generation siblings, nor does it give exclusions or alternative routing conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_marketing_posterA
Design one poster for a brand, shop, event, promotion or topic. Required: only the subject. An image is not required. Exact wording, dates, offers, logo, product photos and style references are optional. Keep supplied copy with autoCopy=false; never invent event details or offers. The service plans the layout and writes the prompt. Requires a MeiGen API key configured for this local npm server (MEIGEN_API_TOKEN or saved local configuration), and purchased credits only. Local files and external public HTTPS URLs are uploaded automatically. Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation. Preserve caller-supplied inputs. Infer omitted optional settings from the request and defaults; when only a requested count is known, choose suitable modules unless the caller selected them. Ask only for missing required material or unresolved scope. An explicit user request or authorized upstream workflow establishes its count, quality and budget: do not reconfirm that scope or add paid images. The caller persists a requestId for each logical step; never ask an end user for technical IDs. Use live list_skills prices and account for in-flight charges when planning within a budget; a batch is not atomic and an estimate is not a server-enforced spending cap. Recovery actions take precedence over generic retry advice: check an interrupted submission, preserve exact retryParameters and do not create a new ID or replace failed modules automatically. Retry a temporary upload at most once after the suggested wait. Return structured status, handles, errors and completed URLs to the caller; it owns progress, previews, downloads and final presentation. If interacting directly with the user, explain the problem and a concrete next step in their language. Describe image details only after actual inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| logo | No | Optional exact brand logo to reproduce accurately, not a style reference. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. | |
| brand | Yes | Poster subject: brand, event, shop, campaign or topic. | |
| ratio | No | Supported output ratio from list_skills for this Skill; omit for its default. | |
| content | No | Poster brief when autoCopy=true; exact visible wording when autoCopy=false (the selected language may translate it). Put style, layout and design directions in extraNotes or customStyle, not in verbatim content. | |
| quality | No | low=Fast (default); medium=Pro. These select rendering quality, not output resolution. Read list_skills for current output specifications and purchased-credit prices; no fixed completion time is guaranteed. | |
| styleId | No | Preset ID, not its display label; choose from list_skills style options. Omit styleId and customStyle for Auto. Nonempty customStyle overrides this preset. | |
| autoCopy | No | Default true: compose copy from the brief. False: use supplied wording faithfully; selected language may translate it. | |
| language | No | Copy language, e.g. auto, en, zh; see list_skills for supported values | |
| uiLocale | No | Optional UI locale fallback; language controls the text inside images. | |
| requestId | Yes | Generate a new UUID for a new paid request; reuse the SAME requestId and inputs on retry. Use check_skill after interruptions. | |
| extraNotes | No | Additional verified facts, explicitly requested display copy, or layout/design constraints, up to 500 characters. Treat design instructions as directions, not text to print verbatim. Preserve supplied details; do not invent dates, prices, offers or claims. | |
| styleImage | No | Optional PRIMARY visual-style reference: palette, lighting, typography and mood. Do not copy its content, products, text or layout; written style is only a compatible supplement. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. | |
| customStyle | No | Optional written visual direction, up to 200 characters. Nonempty text overrides styleId; omit both for Auto. If styleImage is supplied, supplement its visual style without conflicting with it. | |
| productImages | No | Optional product/subject photos, up to three; these identify what the poster depicts, not its visual style. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses authentication/credit requirements, automatic file uploads, non-atomic batch behavior, retry semantics, budget estimation caveats, and the caller's ownership of progress and presentation. It also warns against inventing event details or describing images before actual inspection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and essential constraints, and the guidance is topically organized from scope to prerequisites to execution to recovery. However, it is a long single paragraph with some operational details that could be condensed or split for easier scanning, though virtually every sentence carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 14 parameters and no output schema, the description covers the full call lifecycle: prerequisites, file upload, budget planning, retry handling, caller responsibilities, and response shape. It also correctly points to list_skills for dynamic details like prices, ratios, and style options, leaving no critical gap for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds meaningful semantic guidance beyond the schema by explaining copy preservation rules, omitted-setting inference, confirmation boundaries, and requestId reuse/retry behavior. It does not redefine every parameter, but it clarifies how parameter choices should be made in context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Design one poster for a brand, shop, event, promotion or topic.' It clearly states the tool's unique scope (a single poster) and differentiates it from generic image generation by noting an image is not required and the service plans the layout and writes the prompt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use the dedicated skill directly, skip prompt enhancement, preference loading and delegation, and ask only for missing required material. It does not explicitly name sibling alternatives like generate_image for non-poster requests, but it does define the tool's boundaries and prerequisites, including API key and purchased credits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_product_detail_imagesA
Create 1–6 coordinated ecommerce listing images. Required: one actual product photo and a resolved image count. Product name, verified selling points, logo, model photo and extra angles are optional. Each module creates one paid image. If only the count is specified, choose suitable modules; MCP calls must explicitly select modules to match that count. Requires a MeiGen API key configured for this local npm server (MEIGEN_API_TOKEN or saved local configuration), and purchased credits only. Local files and external public HTTPS URLs are uploaded automatically. Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation. Preserve caller-supplied inputs. Infer omitted optional settings from the request and defaults; when only a requested count is known, choose suitable modules unless the caller selected them. Ask only for missing required material or unresolved scope. An explicit user request or authorized upstream workflow establishes its count, quality and budget: do not reconfirm that scope or add paid images. The caller persists a requestId for each logical step; never ask an end user for technical IDs. Use live list_skills prices and account for in-flight charges when planning within a budget; a batch is not atomic and an estimate is not a server-enforced spending cap. Recovery actions take precedence over generic retry advice: check an interrupted submission, preserve exact retryParameters and do not create a new ID or replace failed modules automatically. Retry a temporary upload at most once after the suggested wait. Return structured status, handles, errors and completed URLs to the caller; it owns progress, previews, downloads and final presentation. If interacting directly with the user, explain the problem and a concrete next step in their language. Describe image details only after actual inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| logo | No | Optional exact brand logo to reproduce accurately. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. | |
| modules | Yes | Required selection; one paid image each: hero=main shot, detail=close-up, scene=lifestyle, material=texture/craft, usage=how to use, brand=brand story. Set [] for custom modules only; otherwise select modules explicitly. Match the requested count. | |
| quality | No | low=Fast (default); medium=Pro. These select rendering quality, not output resolution. Read list_skills for current output specifications and purchased-credit prices; no fixed completion time is guaranteed. | |
| autoCopy | No | Default true: draft copy from the supplied product brief. False: preserve supplied wording, subject to the selected-language translation. | |
| language | No | Copy language, e.g. auto, en, zh; see list_skills for supported values | |
| platform | No | Marketplace preset from list_skills; default amazon. | |
| uiLocale | No | Optional UI locale fallback; language controls the text inside images. | |
| requestId | Yes | Generate a new UUID for a new paid request; reuse the SAME requestId and inputs on retry. Use check_skill after interruptions. | |
| modelImage | No | Optional person/model reference for hero and scene modules: show that person wearing or using the product. Not an image-generation model identifier. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. | |
| aspectRatio | No | Supported output ratio from list_skills; default 4:5. | |
| productName | No | Optional supplied product name; do not invent a brand or model. | |
| productImage | Yes | Required main product photo: preserve the actual shape, color, packaging and readable labels. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. | |
| customModules | No | ||
| sellingPoints | No | Optional verified benefits or specifications from the caller; do not invent product claims. | |
| extraRequirements | No | Optional additional copy, layout or product presentation requirements from the caller; preserve factual constraints. | |
| extraProductImages | No | Optional extra product angles or detail photos, up to two; these supplement the main productImage. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so extensively. It discloses that each module creates one paid image, that batches are not atomic, that estimates are not spending caps, that uploads are automatic, that retries should preserve retryParameters, and that recovery actions take precedence over generic retry advice. It also explains the requestId reuse semantics and that image details should only be described after actual inspection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core purpose, but it is quite long and covers many operational details that could arguably live in a separate workflow document. Every sentence does add some value, but the density makes it harder to parse quickly. It is structured as a block of prose rather than scannable sections, which hurts usability for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex paid-image-generation tool with 16 parameters, no annotations, and no output schema, the description is remarkably complete. It covers required inputs, optional inputs, pricing/credit behavior, retry semantics, error recovery, caller responsibilities, and user interaction guidelines. The only minor gap is that it doesn't describe the exact return structure, but it explicitly says it returns 'structured status, handles, errors and completed URLs,' which is sufficient given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 94%, so the schema already documents most parameters well. The description adds meaningful context beyond the schema: it explains the relationship between modules and count ('MCP calls must explicitly select modules to match that count'), clarifies that quality selects rendering quality not resolution, and emphasizes that productName and sellingPoints must not be invented. It doesn't add per-parameter syntax details, but the schema already covers those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create 1–6 coordinated ecommerce listing images.' It clearly distinguishes this tool from siblings like generate_marketing_poster or generate_ai_background by focusing on coordinated listing images with modules. It also names the dedicated skill and explicitly says to skip delegation, which further differentiates it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation.' It also states when to ask the user ('Ask only for missing required material or unresolved scope') and when not to reconfirm scope. It references list_skills for prices and check_skill for interruptions, providing clear alternatives and recovery paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoADestructive
Generate a MeiGen video using a required live model ID from list_models. Reference videos and reference audio are passed as referenceVideos / referenceAudios arrays (images.meigen.ai URLs, or local files which are uploaded for you — other hosts are rejected); per-model counts and second budgets come from list_models, and reference audio is never billed. Preserve the caller’s resolved prompt, parameters and authorized scope. Set requestId, wait=false and download=false for workflow submission; then query check_generation. At most four submissions run concurrently per MCP process; the backend quota and Retry-After remain authoritative. Video generation consumes purchased credits.
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | Optional model tier. Use list_models for the selected model's live tier values. | |
| wait | No | MeiGen only: false returns the accepted generation ID immediately; poll check_generation separately. Default true waits for completion. Submit-only requires requestId. | |
| model | No | Video model ID. REQUIRED; call list_models for the live lineup and capabilities. | |
| prompt | Yes | The video generation prompt. Describe motion, scene, and style — not just the still image. | |
| modelId | No | Alias of model. Provide at least one; both must match when supplied. | |
| download | No | Save the completed result locally (default true). Set false for URL-only workflows. Ignored when wait=false. Other providers may return inline image content when local saving is disabled. | |
| duration | No | Video duration in seconds. Use list_models for the model's enum/range; when omitted the server uses the omitted-request default shown there. | |
| lastFrame | No | Optional last-frame image for a model that supports it. Accepts a public URL or local file path; requires firstFrame. Use list_models for the live model contract. | |
| requestId | No | Persistent UUID for this workflow step. Required when wait=false. Reuse with identical inputs after interruption, including after MCP restart; use a new UUID for a new generation. Omit only for a new interactive generation. | |
| firstFrame | No | First-frame image when required or supported by the selected model. Accepts a public URL or local file path (auto-uploaded); use list_models and let the server enforce the live model contract. | |
| resolution | No | Output resolution. Use list_models for the selected model/tier's live values. | |
| aspectRatio | No | Aspect ratio: "16:9", "9:16", "1:1", "4:3", "3:4", "21:9", "auto", "adaptive" (model-dependent). Defaults to "auto" when omitted. | |
| referenceVideo | No | Deprecated single-clip alias of referenceVideos. Still accepted forever; when referenceVideos is also supplied it must equal its first entry. Prefer referenceVideos. | |
| referenceAudios | No | Reference audio clips for a model whose list_models entry shows a "Reference audio" line. Each entry is either an https://images.meigen.ai/... URL or a local .wav/.mp3 path, which is uploaded for you (local paths require MEIGEN_API_TOKEN). Other hosts are rejected: pass the local file and the server uploads it. Per-model limits (clip count, per-clip seconds, total seconds, accepted formats and per-file size) come from list_models. On a model whose line says it requires a visual reference (Seedance 2.0), the request must also carry at least one reference image or reference video — audio alone is rejected. Refer to a clip in the prompt as "Audio 1", "Audio 2" … numbered in the order given here. Reference audio seconds are never billed. | |
| referenceVideos | No | Reference video clips for a model that advertises reference-video support in list_models. Each entry is either an https://images.meigen.ai/... URL — typically a clip MeiGen generated earlier, passed through unchanged — or a local .mp4/.mov path, which is uploaded for you (local paths require MEIGEN_API_TOKEN). Other hosts are rejected: the server only probes clips it can fetch from that CDN, so pass the local file instead. Per-model limits — maximum number of clips, per-clip seconds and the maximum SUM of clip seconds — come from list_models; the server enforces them and rejects an over-limit request before charging. IMPORTANT — prompt requirement: to make the new clip semantically continue a reference, the `prompt` MUST explicitly say "extend" / "continue" (e.g. "Extend this video with the following plot:"). Without that, the model treats the clips as visual reference only. Refer to a specific clip in the prompt as "Video 1", "Video 2" … numbered in the order given here. Output behavior: the output is only the configured `duration` of new content — reference clips are never concatenated into it. Billing counts the SUM of the server-probed input video seconds plus the output; do not estimate it from client-side metadata. | |
| referenceVideoDuration | No | Deprecated compatibility hint. Ignored because the server probes the authoritative duration of every clip. |
Output Schema
| Name | Required | Description |
|---|---|---|
| urls | Yes | |
| error | No | |
| status | Yes | |
| deduped | No | |
| modelId | No | |
| success | Yes | |
| imageUrl | No | |
| provider | No | |
| videoUrl | No | |
| mediaType | No | |
| requestId | No | |
| savedPath | No | |
| nextAction | No | |
| creditsUsed | No | |
| generationId | No | |
| creditsStatus | No | |
| receiptWarning | No | |
| downloadWarning | No | |
| observationEnded | No | |
| pollAfterSeconds | No | |
| requestedMediaType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only signal readOnlyHint=false and destructiveHint=true. The description goes further by disclosing credit consumption, the fact that reference audio is never billed, the concurrency limit of four submissions, and that backend quota/Retry-After are authoritative. It also clarifies output behavior — only configured duration, no concatenation — and that billing uses server-probed seconds, which is valuable beyond the annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five dense sentences front-load the core action and then immediately cover dependencies, workflow submission, concurrency, and billing. The parentheticals are information-dense rather than padding, and each sentence contributes a distinct, necessary fact for a complex generation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with concurrency, billing, file-upload, and workflow-state constraints, the description covers all key decisions: model selection, reference handling, polling pattern, quota authority, and credit consumption. With an output schema present, not restating return values is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the model, prompt, reference arrays, wait, download, and requestId semantics. The description restates key workflow conditions such as "Set requestId, wait=false and download=false" and "reference audio is never billed," but these are also present in the input schema. This matches the baseline where the schema carries the parameter-meaning burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with "Generate a MeiGen video" — a specific verb and resource that clearly distinguishes it from sibling tools like generate_image and generate_marketing_poster. It also names the required dependency on list_models for a live model ID, making the scope and prerequisite explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use the tool: after selecting a model via list_models, and for workflow submission with requestId, wait=false, download=false, followed by polling with check_generation. It does not explicitly compare against sibling tools or state when not to use it, so it stops short of a full alternatives statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_inspirationARead-only
Get the full prompt and all image URLs for a gallery entry. Show the images to the user as visual examples. The prompt can be used directly with generate_image(), and image URLs can be passed as referenceImages for style transfer.
| Name | Required | Description | Default |
|---|---|---|---|
| imageId | Yes | Image/prompt ID from search_gallery results |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the tool is safe. Description adds that it returns full prompt and image URLs and suggests showing images to users, which is helpful context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences that front-load the main action and provide immediate value with downstream usage tips. No redundant or extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieval tool with one parameter and no output schema, the description fully explains what is returned and how to use it in conjunction with sibling tools. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for the single parameter imageId. The tool description does not add further detail about the parameter, so baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves the full prompt and all image URLs for a gallery entry, distinguishing it from search_gallery (which lists entries) and generate_image (which creates images).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides direct guidance on using the prompt with generate_image() and image URLs as referenceImages for style transfer, and implies it should be used after selecting an entry from search_gallery. Lacks explicit when-not-to-use, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsARead-only
List available AI image generation models and their capabilities. For up-to-date pricing, see https://www.meigen.ai/model-comparison.
| Name | Required | Description | Default |
|---|---|---|---|
| activeOnly | No | Only show active models (default: true) |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| models | Yes | |
| success | Yes | |
| configuredProviders | Yes | |
| executionCapabilities | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds the context that it lists models and their capabilities, and points to an external URL for up-to-date pricing, which is useful. It doesn't describe pagination or response format, but the output schema exists to cover that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The core purpose is front-loaded, and the pricing link is a useful addition that doesn't clutter the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and an output schema, the description is nearly complete. The only minor gap is not explicitly stating that it returns a list of models with capabilities, but that's implied by 'List available AI image generation models and their capabilities.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the activeOnly parameter. The description doesn't add much beyond the schema, but it does mention 'capabilities' which hints at what the output contains. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists available AI image generation models and their capabilities, which is a specific verb+resource. It distinguishes itself from siblings like list_skills by explicitly scoping to AI image generation models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for discovering available models and capabilities, and the activeOnly parameter provides context for filtering. However, it doesn't explicitly state when to use this tool versus alternatives like list_skills or check_skill, nor does it mention exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_skillsCRead-only
Choose a Skill by use case and required materials; read account setup, top-up instructions, live prices and upload options. Set skill to inspect one workflow. Returned inputSchema is the HTTP API schema; this local MCP also accepts actual file paths and external HTTPS image URLs using its tool schemas.
| Name | Required | Description | Default |
|---|---|---|---|
| skill | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint is true and the description does not contradict it; it mentions reading account setup, prices, etc. The description adds useful context by noting that the returned schema is the HTTP API schema and that local MCP accepts file paths/URLs, which goes beyond the annotation. However, it does not elaborate on other behavioral aspects like authentication or rate limits, but with readOnlyHint present the bar is lower.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences but not front-loaded with a clear verb+resource. The first sentence mixes selection guidance with a list of informational items; the second is technical but tangential for a listing tool. It is not concise because the core purpose is obscured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter and no output schema, the description is incomplete. It does not clarify whether the tool returns a list of available skills or details for a selected skill, and the mention of 'inputSchema' is confusing. An agent would struggle to anticipate the exact return shape or how to use the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It partially explains the 'skill' parameter by saying 'Set skill to inspect one workflow' and hints at selection by use case, adding meaning beyond the enum values. However, it does not explain the specific enum options or how the parameter affects the output, leaving gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description does not clearly state that the tool lists skills. It says 'Choose a Skill by use case and required materials' and 'Set skill to inspect one workflow,' which suggests selection/inspection rather than a simple list. The verb 'list' in the name is not echoed clearly, and it is ambiguous whether this returns a list of skills or details for one skill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus siblings like check_skill or list_models. The note about 'Returned inputSchema' and accepting file paths/URLs provides some contextual usage, but it does not explain when to choose this tool over alternatives or any exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_preferencesA
Read or update user preferences: default style, aspect ratio, model, style notes, and favorite prompts. Call with action "get" at conversation start to load preferences.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | remove_favorite: 0-based index of the favorite to remove | |
| model | No | set: preferred default model name | |
| style | No | set: preferred default style (e.g. "realistic", "anime", "illustration") | |
| action | Yes | Action to perform: "get" reads all preferences, "set" updates defaults/styleNotes, "add_favorite" saves a prompt, "remove_favorite" removes by index | |
| prompt | No | add_favorite: the prompt text to save | |
| provider | No | set: preferred default provider | |
| styleNotes | No | set: free-text style notes (e.g. "cinematic lighting, shallow DOF, brand colors #1A1A2E") | |
| aspectRatio | No | set: preferred default aspect ratio. Use "auto" (recommended) to let MeiGen infer per-prompt, or pin a value like "16:9", "1:1", "9:16". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations set readOnlyHint=false, implying mutation, and the description confirms 'Read or update'. No additional behavioral traits are disclosed (e.g., side effects, permissions, rate limits). The description adds minimal value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, no wasted words. Efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks details about return values for each action, especially since no output schema is provided. It mentions one usage scenario (get at conversation start) but not the semantics of other actions. Overall adequate but with noticeable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The tool description lists the preferences categories but does not add significant new meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads or updates user preferences and lists specific preferences (style, aspect ratio, model, style notes, favorite prompts). It is a specific verb+resource combination that distinguishes it from sibling tools like generate_image or list_models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises calling with action 'get' at conversation start to load preferences. However, it does not provide guidance on when to use 'set', 'add_favorite', or 'remove_favorite', nor does it mention when not to use this tool or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_backgroundA
Create one transparent PNG cutout for compositing, catalog assets or logos. Required: one actual source photo containing the subject. No prompt or product facts needed. This removes the background; it does not create a new scene. Requires a MeiGen API key configured for this local npm server (MEIGEN_API_TOKEN or saved local configuration), and purchased credits only. Local files and external public HTTPS URLs are uploaded automatically. Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation. Preserve caller-supplied inputs. Infer omitted optional settings from the request and defaults; when only a requested count is known, choose suitable modules unless the caller selected them. Ask only for missing required material or unresolved scope. An explicit user request or authorized upstream workflow establishes its count, quality and budget: do not reconfirm that scope or add paid images. The caller persists a requestId for each logical step; never ask an end user for technical IDs. Use live list_skills prices and account for in-flight charges when planning within a budget; a batch is not atomic and an estimate is not a server-enforced spending cap. Recovery actions take precedence over generic retry advice: check an interrupted submission, preserve exact retryParameters and do not create a new ID or replace failed modules automatically. Retry a temporary upload at most once after the suggested wait. Return structured status, handles, errors and completed URLs to the caller; it owns progress, previews, downloads and final presentation. If interacting directly with the user, explain the problem and a concrete next step in their language. Describe image details only after actual inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | Generate a new UUID for a new paid request; reuse the SAME requestId and inputs on retry. Use check_skill after interruptions. | |
| productImage | Yes | Required source containing the subject to cut out; the subject may be a product, person or logo. The output has a transparent background. An accessible absolute local path (Windows drive/UNC, POSIX, ~/, or file://) or public direct HTTPS image URL without credentials, fragments or custom ports. Relative paths are ambiguous and rejected. Files are fully decoded, metadata removed, transparency preserved; GIF uses the first frame. Never invent attachment paths or URLs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so extensively: it discloses upload behavior (local files and public HTTPS URLs uploaded automatically), auth requirements (MeiGen API key, MEIGEN_API_TOKEN or saved config), credit requirements (purchased credits only), retry semantics (at most once after suggested wait, preserve retryParameters), batch non-atomicity, and the fact that estimates are not spending caps. It also discloses that files are decoded, metadata removed, and GIF uses the first frame. The only minor gap is that it doesn't explicitly state the output format beyond 'transparent PNG cutout' and 'completed URLs', but the description is unusually transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core purpose, but it is quite long and includes a large amount of operational policy (budget planning, recovery actions, caller ownership of progress) that, while valuable, could be trimmed or moved to a more structured location. Every sentence does earn its place in terms of content, but the sheer length makes it harder to parse quickly. It's not bloated with fluff, but it's at the edge of what an agent can absorb efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 2 parameters, no output schema, and no annotations, the description is remarkably complete. It covers input requirements, auth, upload behavior, retry semantics, error handling, caller responsibilities, and interaction guidance. An agent has everything it needs to invoke this tool correctly and handle failures appropriately. The only thing not specified is the exact response shape, but the description says it returns 'structured status, handles, errors and completed URLs', which is sufficient given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both parameters well. The description adds meaningful context beyond the schema: it clarifies that productImage must be an actual source photo containing the subject, that relative paths are rejected, that files are decoded and metadata removed, and that GIF uses the first frame. It also reinforces the requestId reuse semantics. This goes beyond the baseline 3 by adding operational meaning to the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create one transparent PNG cutout' for compositing, catalog assets, or logos. It explicitly distinguishes itself from scene creation ('it does not create a new scene'), which separates it from siblings like generate_ai_background and generate_marketing_poster. The purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: requires one actual source photo, no prompt or product facts needed, and it explicitly says to use the dedicated skill directly while skipping prompt enhancement, preference loading, and delegation. It also names the alternative path (list_skills for prices, check_skill for interruptions) and clarifies when to ask the user. This is comprehensive routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_galleryARead-only
Search AI image prompts with semantic understanding — finds visually and conceptually similar results, not just keyword matches. Returns at most 3 entries per call; larger limits are clamped. With a MeiGen API key configured, searches are authenticated and counted against that account's daily search quota instead of the shared per-IP budget. Results include image URLs — render them as markdown images () so users can visually browse and pick styles. Use when users need inspiration, want to explore styles, or say "generate an image" without a specific idea.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Requested number of results. The server returns at most 3; larger values from existing automations are accepted and clamped rather than rejected. | |
| query | No | Search keywords (e.g., "cyberpunk", "product photo", "portrait"). Supports semantic search — natural language descriptions work well. Leave empty to browse by category or get random picks. | |
| offset | No | Pagination offset | |
| sortBy | No | Sort order when browsing without search query (default: rank) | rank |
| category | No | Filter by category. Available: Photography, Illustration & 3D, Product & Brand, Food & Drink, Poster Design, UI & Graphic |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral traits: returns at most 3 entries per call, larger limits are clamped, and authentication affects quota counting. It also mentions that results include image URLs and instructs to render them as markdown images. The readOnlyHint annotation is consistent with the description's read-only nature, and the description adds context about quotas and clamping beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose, then adds key behavioral details and usage guidance. Every sentence adds value, and it is well-structured for an agent to quickly grasp the tool's function and constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 parameters, no output schema), the description covers the essential behavioral aspects: result limits, authentication, rendering instructions, and usage scenarios. It doesn't detail the exact response structure, but since there is no output schema, the description's mention of image URLs and markdown rendering is sufficient for an agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds value by explaining the semantic search capability, the clamping behavior for limit, and the browsing behavior when query is empty. It also clarifies that natural language descriptions work well, which is not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: searching AI image prompts with semantic understanding, distinguishing it from keyword-only search. It also explicitly mentions the return of image URLs and how to render them, which is a specific and actionable purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: when users need inspiration, want to explore styles, or say 'generate an image' without a specific idea. It also contrasts with the shared per-IP budget and mentions authentication with MeiGen API key, giving clear context for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_skill_imageB
Prepare a reference image and return imageUrl. External URLs and local file paths can also be passed directly to the generation skill. Use actual accessible bytes/URLs only; never fabricate base64 or attachment paths. No generation is started.
| Name | Required | Description | Default |
|---|---|---|---|
| purpose | No | Use upscale for enhancement attachments: preserves source dimensions until the backend asks for resizing acceptance. | reference |
| sourceUrl | No | Actual public direct HTTPS image URL to prepare; choose either sourceUrl or imageBase64. Local paths go directly to a generation Skill image field, not this URL field. | |
| imageBase64 | No | Actual raw base64 image bytes, max 3 MiB decoded. Exactly one input is required. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It adds a key safety trait ('No generation is started') and the instruction to use 'actual accessible bytes/URLs only,' which helps agents avoid invalid inputs. It does not disclose storage, lifetime, permissions, or other side effects of preparing the image, so transparency is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with the core purpose and return value front-loaded, followed by concise caveats about real inputs and no generation. The direct-passing note is slightly ambiguous but not padding, and the overall length is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The definition covers the essential contract: it returns imageUrl, does not start generation, and requires real accessible bytes/URLs. However, without an output schema or annotations, it leaves the purpose/upscale distinction and the exact relationship to direct generation partly ambiguous, so an agent may still need to infer important call context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline 3 applies. The description reinforces the no-fabrication rule and notes that external URLs/local paths can bypass the tool, but it adds little new meaning about purpose, sourceUrl, or imageBase64 beyond what the property descriptions already provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete outcome ('Prepare a reference image and return imageUrl') and explicitly distinguishes the tool from generation by adding 'No generation is started.' The verb 'prepare' is slightly vague, but the tool name, the return-value mention, and the sibling context make the resource and result reasonably clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives partial routing guidance by noting that external URLs and local file paths 'can also be passed directly to the generation skill' and by warning against fabricated inputs. However, it never clearly states when upload_skill_image is required versus when to skip it, and the sourceUrl parameter makes the 'external URLs can be passed directly' advice ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_imageA
Enhance one still image. Crisp preserves structure; creative regenerates detail for blurry images. Pass the original public URL directly, including external URLs; avoid generic reference-image resizing. The shared backend prepares the image. Inputs above 4096px or 16 MP require acceptance of resizing; the output may be smaller than the original. Source safety cap: 64 MiB / 64 MP. Video enhancement is not exposed through this API. Requires a MeiGen API key configured for this local npm server (MEIGEN_API_TOKEN or saved local configuration), and purchased credits only. Local files and external public HTTPS URLs are uploaded automatically. Use the dedicated skill directly; skip unrequested prompt enhancement, preference loading and delegation. Preserve caller-supplied inputs. Infer omitted optional settings from the request and defaults; when only a requested count is known, choose suitable modules unless the caller selected them. Ask only for missing required material or unresolved scope. An explicit user request or authorized upstream workflow establishes its count, quality and budget: do not reconfirm that scope or add paid images. The caller persists a requestId for each logical step; never ask an end user for technical IDs. Use live list_skills prices and account for in-flight charges when planning within a budget; a batch is not atomic and an estimate is not a server-enforced spending cap. Recovery actions take precedence over generic retry advice: check an interrupted submission, preserve exact retryParameters and do not create a new ID or replace failed modules automatically. Retry a temporary upload at most once after the suggested wait. Return structured status, handles, errors and completed URLs to the caller; it owns progress, previews, downloads and final presentation. If interacting directly with the user, explain the problem and a concrete next step in their language. Describe image details only after actual inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | crisp=faithful enhancement (default); creative=AI reconstructs details and may change them. Use creative only when the user accepts those changes. | crisp |
| imageUrl | Yes | Original still PNG/JPEG/WebP: absolute local path, ~/, file://, or public direct HTTPS URL. Maximum 64 MiB/64 million pixels. Local uploads remove metadata, preserve alpha and dimensions, and compress below the gateway limit; if that fails, provide a public original URL. The backend asks before any resizing. | |
| requestId | Yes | Generate a new UUID for a new paid request; reuse the SAME requestId and inputs on retry. Use check_skill after interruptions. | |
| allowDownscale | No | True only after the user accepts resizing a large source and potentially receiving a smaller result. On upscale_resize_required, no charge occurred; after acceptance use a new requestId. | |
| confirmedCredits | Yes | Required on every MCP submission, including the first: current list_skills quote within the user or upstream workflow accepted budget. Reuse explicit acceptance. On price_changed, accept the updated quote before a new requestId. This pre-dispatch check is not an atomic spending cap. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden and does so thoroughly: it discloses upload behavior, metadata removal, 64 MiB/64 MP caps, possible downscaling, non-atomic batches, quote-vs-spending-cap semantics, retry rules, requestId reuse, and the structured return contract. This goes well beyond simple 'upscale this image' and gives an agent realistic expectations of side effects and failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but it is front-loaded with the core purpose and then parcels planning, retry, and cost guidance into useful clauses. It earns most of its length, though a few operational policies could be tightened without losing meaning; this keeps it just shy of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paid, upload-capable, retry-sensitive tool with no output schema, the description is remarkably complete: it covers prerequisites, cost handling, failure recovery, interaction boundaries, and what the caller receives. An agent has enough context to invoke it correctly and to coordinate with check_skill and list_skills.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description enriches every parameter meaningfully: imageUrl gets path formats, size limits, and metadata behavior; requestId gets retry semantics; confirmedCredits gets quote and budget context; allowDownscale gets a clear charge-related caveat; mode gets user-acceptance considerations. This is far beyond the baseline 3 for complete schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource — 'Enhance one still image' — and immediately distinguishes the two operating modes, crisp vs creative. It also explicitly separates this from video enhancement and general image generation, making it easy to differentiate from siblings like generate_image and generate_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete when-to-use guidance: pass original public URLs, avoid generic reference-image resizing, do not use this for video, and use the dedicated skill directly rather than routing through prompt enhancement or delegation. It references check_skill and list_skills as supporting tools, though it does not explicitly name sibling tools as alternatives for similar still-image operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v2.0.1- Added
check_generation - Added
check_skill - Added
generate_ai_background - Changed
generate_image8 fields changed- added
Input schema / properties / downloadAdded value: +{ + "default": true, + "description": "Save the completed result locally (default true). Set false for URL-only workflows. Ignored when wait=false. Other providers may return inline image content when local saving is disabled.", + "type": "boolean" +} - added
Input schema / properties / model / minLengthAdded value: +1 - added
Input schema / properties / modelIdAdded value: +{ + "description": "Alias of model for portable workflow calls. If both are present they must match.", + "minLength": 1, + "type": "string" +} - added
Input schema / properties / modelVariantAdded value: +{ + "description": "Optional model variant from live list_models, such as a supported GPT Image 2.5 variant. Forwarded unchanged to MeiGen.", + "type": "string" +} - changed
Input schema / properties / referenceImages / descriptionPrevious value: -"Image references for style/content guidance. Accepts both public URLs (http/https) and local file paths. Local files are automatically compressed and uploaded when needed. For ComfyUI: local files are passed directly to the workflow (requires LoadImage node). Sources: gallery URLs from search_gallery/get_inspiration, URLs from previous generate_image results, or local file paths."New value: +"Image references for style/content guidance. Accepts direct public HTTPS URLs without credentials/fragments or accessible absolute local paths. Relative paths are rejected. Local PNG/JPEG/WebP/GIF references up to 32 MiB and 64 million pixels are fully decoded, stripped of metadata and prepared up to 4096px, preserving transparency. For ComfyUI: local files are passed directly to the workflow (requires LoadImage node). Sources: gallery URLs from search_gallery/get_inspiration, URLs from previous generate_image results, or local file paths." - added
Input schema / properties / requestIdAdded value: +{ + "description": "Persistent UUID for this workflow step. Required when wait=false. Reuse with identical inputs after interruption, including after MCP restart; use a new UUID for a new generation. Omit only for a new interactive generation.", + "format": "uuid", + "type": "string" +} - added
Input schema / properties / waitAdded value: +{ + "default": true, + "description": "MeiGen only: false returns the accepted generation ID immediately; poll check_generation separately. Default true waits for completion. Submit-only requires requestId.", + "type": "boolean" +} - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "creditsStatus": { + "type": "string" + }, + "creditsUsed": { + "type": "number" + }, + "deduped": { + "type": "boolean" + }, + "downloadWarning": { + "type": "string" + }, + "error": { + "additionalProperties": false, + "properties": { + "available": { + "type": "number" + }, + "code": { + "type": "string" + }, + "httpStatus": { + "type": "number" + }, + "message": { + "type": "string" + }, + "required": { + "type": "number" + }, + "retryAfterSeconds": { + "type": "number" + }, + "retryable": { + "type": "boolean" + } + }, + "required": [ + "code", + "message", + "retryable" + ], + "type": "object" + }, + "generationId": { + "type": "string" + }, + "imageUrl": { + "type": "string" + }, + "mediaType": { + "enum": [ + "image", + "video" + ], + "type": "string" + }, + "modelId": { + "type": "string" + }, + "nextAction": { + "additionalProperties": false, + "properties": { + "afterSeconds": { + "type": "number" + }, + "arguments": { + "additionalProperties": {}, + "type": "object" + }, + "message": { + "type": "string" + }, + "tool": { + "type": "string" + }, + "type": { + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + "observationEnded": { + "type": "boolean" + }, + "pollAfterSeconds": { + "type": "number" + }, + "provider": { + "enum": [ + "meigen", + "openai", + "comfyui" + ], + "type": "string" + }, + "receiptWarning": { + "type": "string" + }, + "requestId": { + "type": "string" + }, + "requestedMediaType": { + "enum": [ + "image", + "video" + ], + "type": "string" + }, + "savedPath": { + "type": "string" + }, + "status": { + "enum": [ + "processing", + "completed", + "failed", + "error", + "unknown" + ], + "type": "string" + }, + "success": { + "type": "boolean" + }, + "urls": { + "items": { + "type": "string" + }, + "type": "array" + }, + "videoUrl": { + "type": "string" + } + }, + "required": [ + "success", + "status", + "urls" + ], + "type": "object" +}
- Added
generate_marketing_poster - Added
generate_product_detail_images - Changed
generate_video16 fields changed- added
Input schema / properties / downloadAdded value: +{ + "default": true, + "description": "Save the completed result locally (default true). Set false for URL-only workflows. Ignored when wait=false. Other providers may return inline image content when local saving is disabled.", + "type": "boolean" +} - changed
Input schema / properties / duration / descriptionPrevious value: -"Video duration in seconds. seedance-2-0 / happyhorse-1.0 currently accept ~3–15s (any integer in range). veo-3.1 accepts exactly 4, 6, or 8 (default 4) — other values will be rejected. Defaults to the model's default duration. Call list_models for the current allowed values per model."New value: +"Video duration in seconds. Use list_models for the model's enum/range; when omitted the server uses the omitted-request default shown there." - changed
Input schema / properties / firstFrame / descriptionPrevious value: -"First-frame image to control where the video starts. Accepts public URL or local file path (auto-uploaded). REQUIRED for grok-video (image-to-video only — backend rejects it without a firstFrame). For seedance/happyhorse/veo it is optional: with no first frame they do pure text-to-video."New value: +"First-frame image when required or supported by the selected model. Accepts a public URL or local file path (auto-uploaded); use list_models and let the server enforce the live model contract." - changed
Input schema / properties / lastFrame / descriptionPrevious value: -"Optional last-frame image to also control where the video ends. Used by seedance-2-0 and veo-3.1; happyhorse-1.0 ignores this field. Accepts public URL or local file path. Requires firstFrame to also be provided — passing lastFrame alone is rejected."New value: +"Optional last-frame image for a model that supports it. Accepts a public URL or local file path; requires firstFrame. Use list_models for the live model contract." - changed
Input schema / properties / model / descriptionPrevious value: -"Video model ID. Use list_models to see available video models. Common (as of writing): \"seedance-2-0\" (multi-tier general purpose), \"happyhorse-1.0\" (cost-effective i2v/t2v), \"veo-3.1\" (Google Veo with two tiers, 4/6/8s, native audio), \"grok-video\" (xAI Grok Imagine 1.5 — IMAGE-TO-VIDEO ONLY: firstFrame REQUIRED, pure text-to-video is rejected; native audio; 4-15s; 480p/720p)."New value: +"Video model ID. REQUIRED; call list_models for the live lineup and capabilities." - added
Input schema / properties / modelIdAdded value: +{ + "description": "Alias of model. Provide at least one; both must match when supplied.", + "minLength": 1, + "type": "string" +} - added
Input schema / properties / referenceAudiosAdded value: +{ + "description": "Reference audio clips for a model whose list_models entry shows a \"Reference audio\" line. Each entry is either an https://images.meigen.ai/... URL or a local .wav/.mp3 path, which is uploaded for you (local paths require MEIGEN_API_TOKEN). Other hosts are rejected: pass the local file and the server uploads it. Per-model limits (clip count, per-clip seconds, total seconds, accepted formats and per-file size) come from list_models. On a model whose line says it requires a visual reference (Seedance 2.0), the request must also carry at least one reference image or reference video — audio alone is rejected. Refer to a clip in the prompt as \"Audio 1\", \"Audio 2\" … numbered in the order given here. Reference audio seconds are never billed.", + "items": { + "type": "string" + }, + "maxItems": 10, + "type": "array" +} - changed
Input schema / properties / referenceVideo / descriptionPrevious value: -"Optional reference video URL for Seedance 2.0 \"video continuation\". Must be a publicly accessible HTTPS URL (typically a previous generation result `videoUrl`); local paths are not supported. Only seedance-2-0 accepts this — passing it with other models will fail. IMPORTANT — prompt requirement: to make the new clip semantically continue the reference, the `prompt` MUST explicitly say \"extend\" / \"continue\" (e.g. prefix with \"Extend this video with the following plot:\"). Without that, the model treats the video as visual reference only and the new clip may drift from a true continuation. Output behavior: the output is ONLY your `duration` seconds (4-15s) of new content — the reference video is NOT concatenated into the output. To get a single \"original + new\" clip the user must stitch them locally. Billing: credits are charged at the With-reference-video rate, with `billable_seconds = max(reference_duration + duration, min_billable[duration])`. Total cost is often higher than direct generation of the same output length. Always pass `referenceVideoDuration` alongside this field — omitting it causes underbilling and broken continuation behavior."New value: +"Deprecated single-clip alias of referenceVideos. Still accepted forever; when referenceVideos is also supplied it must equal its first entry. Prefer referenceVideos." - changed
Input schema / properties / referenceVideoDuration / descriptionPrevious value: -"Duration of the reference video in seconds (typically 2–15 — backend validates the current allowed range). REQUIRED whenever `referenceVideo` is set; if omitted the backend treats it as 0, leading to undercharged credits and misconfigured generation. Pass the actual duration of the clip at `referenceVideo`."New value: +"Deprecated compatibility hint. Ignored because the server probes the authoritative duration of every clip." - added
Input schema / properties / referenceVideosAdded value: +{ + "description": "Reference video clips for a model that advertises reference-video support in list_models. Each entry is either an https://images.meigen.ai/... URL — typically a clip MeiGen generated earlier, passed through unchanged — or a local .mp4/.mov path, which is uploaded for you (local paths require MEIGEN_API_TOKEN). Other hosts are rejected: the server only probes clips it can fetch from that CDN, so pass the local file instead. Per-model limits — maximum number of clips, per-clip seconds and the maximum SUM of clip seconds — come from list_models; the server enforces them and rejects an over-limit request before charging. IMPORTANT — prompt requirement: to make the new clip semantically continue a reference, the `prompt` MUST explicitly say \"extend\" / \"continue\" (e.g. \"Extend this video with the following plot:\"). Without that, the model treats the clips as visual reference only. Refer to a specific clip in the prompt as \"Video 1\", \"Video 2\" … numbered in the order given here. Output behavior: the output is only the configured `duration` of new content — reference clips are never concatenated into it. Billing counts the SUM of the server-probed input video seconds plus the output; do not estimate it from client-side metadata.", + "items": { + "type": "string" + }, + "maxItems": 10, + "type": "array" +} - added
Input schema / properties / requestIdAdded value: +{ + "description": "Persistent UUID for this workflow step. Required when wait=false. Reuse with identical inputs after interruption, including after MCP restart; use a new UUID for a new generation. Omit only for a new interactive generation.", + "format": "uuid", + "type": "string" +} - changed
Input schema / properties / resolution / descriptionPrevious value: -"Output resolution. Common: \"480p\" / \"720p\" / \"1080p\" / \"4k\" (model-dependent; e.g. Seedance Pro adds 1080p and 4k, while Fast/Mini are 480p/720p only). Use list_models to see what each model supports. Higher resolutions cost more credits per second."New value: +"Output resolution. Use list_models for the selected model/tier's live values." - changed
Input schema / properties / tier / descriptionPrevious value: -"Quality tier — only for models that support tiers. seedance-2-0 accepts \"mini\" (default, cheapest; 480p/720p, no reference video), \"fast\" (480p/720p), or \"pro\" (highest fidelity; native 1080p and 4K); veo-3.1 accepts \"fast\" (default) or \"pro\". Tiers may be added by the platform — call list_models to see what each model exposes."New value: +"Optional model tier. Use list_models for the selected model's live tier values." - added
Input schema / properties / waitAdded value: +{ + "default": true, + "description": "MeiGen only: false returns the accepted generation ID immediately; poll check_generation separately. Default true waits for completion. Submit-only requires requestId.", + "type": "boolean" +} - changed
Input schema / requiredPrevious value: -[ - "prompt", - "model" -]New value: +[ + "prompt" +] - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "creditsStatus": { + "type": "string" + }, + "creditsUsed": { + "type": "number" + }, + "deduped": { + "type": "boolean" + }, + "downloadWarning": { + "type": "string" + }, + "error": { + "additionalProperties": false, + "properties": { + "available": { + "type": "number" + }, + "code": { + "type": "string" + }, + "httpStatus": { + "type": "number" + }, + "message": { + "type": "string" + }, + "required": { + "type": "number" + }, + "retryAfterSeconds": { + "type": "number" + }, + "retryable": { + "type": "boolean" + } + }, + "required": [ + "code", + "message", + "retryable" + ], + "type": "object" + }, + "generationId": { + "type": "string" + }, + "imageUrl": { + "type": "string" + }, + "mediaType": { + "enum": [ + "image", + "video" + ], + "type": "string" + }, + "modelId": { + "type": "string" + }, + "nextAction": { + "additionalProperties": false, + "properties": { + "afterSeconds": { + "type": "number" + }, + "arguments": { + "additionalProperties": {}, + "type": "object" + }, + "message": { + "type": "string" + }, + "tool": { + "type": "string" + }, + "type": { + "type": "string" + } + }, + "required": [ + "type" + ], + "type": "object" + }, + "observationEnded": { + "type": "boolean" + }, + "pollAfterSeconds": { + "type": "number" + }, + "provider": { + "enum": [ + "meigen", + "openai", + "comfyui" + ], + "type": "string" + }, + "receiptWarning": { + "type": "string" + }, + "requestId": { + "type": "string" + }, + "requestedMediaType": { + "enum": [ + "image", + "video" + ], + "type": "string" + }, + "savedPath": { + "type": "string" + }, + "status": { + "enum": [ + "processing", + "completed", + "failed", + "error", + "unknown" + ], + "type": "string" + }, + "success": { + "type": "boolean" + }, + "urls": { + "items": { + "type": "string" + }, + "type": "array" + }, + "videoUrl": { + "type": "string" + } + }, + "required": [ + "success", + "status", + "urls" + ], + "type": "object" +}
- Changed
list_models1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "configuredProviders": { + "items": { + "enum": [ + "meigen", + "openai", + "comfyui" + ], + "type": "string" + }, + "type": "array" + }, + "error": { + "additionalProperties": false, + "properties": { + "code": { + "type": "string" + }, + "message": { + "type": "string" + }, + "retryable": { + "type": "boolean" + } + }, + "required": [ + "code", + "message", + "retryable" + ], + "type": "object" + }, + "executionCapabilities": { + "additionalProperties": false, + "properties": { + "concurrency": { + "type": "string" + }, + "download": { + "type": "boolean" + }, + "maxConcurrentSubmissions": { + "type": "number" + }, + "submitOnly": { + "type": "boolean" + } + }, + "required": [ + "submitOnly", + "download", + "maxConcurrentSubmissions", + "concurrency" + ], + "type": "object" + }, + "models": { + "items": { + "additionalProperties": true, + "properties": { + "id": { + "type": "string" + }, + "name": { + "type": "string" + } + }, + "required": [ + "id", + "name" + ], + "type": "object" + }, + "type": "array" + }, + "success": { + "type": "boolean" + } + }, + "required": [ + "success", + "models", + "configuredProviders", + "executionCapabilities" + ], + "type": "object" +}
- Added
list_skills - Added
remove_background - Changed
search_gallery2 fields changed- changed
Input schema / properties / limit / defaultPrevious value: -5New value: +3 - changed
Input schema / properties / limit / descriptionPrevious value: -"Number of results (1-20, default 5)"New value: +"Requested number of results. The server returns at most 3; larger values from existing automations are accepted and clamped rather than rejected."
- Added
upload_skill_image - Added
upscale_image
2 tool updates
v1.3.3- Changed
generate_image1 field changed- changed
Input schema / properties / model / descriptionPrevious value: -"Model name. For OpenAI-compatible providers: any model ID your endpoint supports. For MeiGen: use model IDs from list_models."New value: +"Model name. For OpenAI-compatible providers: any model ID your endpoint supports. For MeiGen: use model IDs from list_models (e.g. \"gpt-image-2\", \"grok-image\" = xAI Grok Imagine Quality, 1K/2K, supports image-to-image, \"nanobanana-2\", \"seedream-4.5\", \"flux2-klein\")."
- Changed
generate_video4 fields changed- changed
Input schema / properties / firstFrame / descriptionPrevious value: -"Optional first-frame image to control where the video starts. Accepts public URL or local file path (auto-uploaded). Highly recommended for image-to-video; with no first frame the model does pure text-to-video."New value: +"First-frame image to control where the video starts. Accepts public URL or local file path (auto-uploaded). REQUIRED for grok-video (image-to-video only — backend rejects it without a firstFrame). For seedance/happyhorse/veo it is optional: with no first frame they do pure text-to-video." - changed
Input schema / properties / model / descriptionPrevious value: -"Video model ID. Use list_models to see available video models. Common (as of writing): \"seedance-2-0\" (multi-tier general purpose), \"happyhorse-1.0\" (cost-effective i2v/t2v), \"veo-3.1\" (Google Veo with two tiers, 4/6/8s, native audio)."New value: +"Video model ID. Use list_models to see available video models. Common (as of writing): \"seedance-2-0\" (multi-tier general purpose), \"happyhorse-1.0\" (cost-effective i2v/t2v), \"veo-3.1\" (Google Veo with two tiers, 4/6/8s, native audio), \"grok-video\" (xAI Grok Imagine 1.5 — IMAGE-TO-VIDEO ONLY: firstFrame REQUIRED, pure text-to-video is rejected; native audio; 4-15s; 480p/720p)." - changed
Input schema / properties / resolution / descriptionPrevious value: -"Output resolution. Common: \"480p\" / \"720p\" / \"1080p\" (model-dependent). Use list_models to see what each model supports. Higher resolutions cost more credits per second."New value: +"Output resolution. Common: \"480p\" / \"720p\" / \"1080p\" / \"4k\" (model-dependent; e.g. Seedance Pro adds 1080p and 4k, while Fast/Mini are 480p/720p only). Use list_models to see what each model supports. Higher resolutions cost more credits per second." - changed
Input schema / properties / tier / descriptionPrevious value: -"Quality tier — only for models that support tiers. seedance-2-0 and veo-3.1 currently accept \"fast\" (default, cheaper) or \"pro\" (higher fidelity). Tiers may be added by the platform — call list_models to see what each model exposes."New value: +"Quality tier — only for models that support tiers. seedance-2-0 accepts \"mini\" (default, cheapest; 480p/720p, no reference video), \"fast\" (480p/720p), or \"pro\" (highest fidelity; native 1080p and 4K); veo-3.1 accepts \"fast\" (default) or \"pro\". Tiers may be added by the platform — call list_models to see what each model exposes."
8 tool updates
v1.3.1- Added
comfyui_workflow - Added
enhance_prompt - Added
generate_image - Added
generate_video - Added
get_inspiration - Added
list_models - Added
manage_preferences - Added
search_gallery
TDQS
Scored across 17 tools
Most tools are clearly separated by purpose, but generate_image overlaps with the specialized generators (generate_marketing_poster, generate_product_detail_images, generate_ai_background), creating possible selection ambiguity. Also, check_skill vs check_generation and list_skills vs list_models require careful reading, though the descriptions help clarify boundaries.
The vast majority of tools follow a verb_noun pattern (generate_*, list_*, check_*, manage_*, enhance_*), making the naming predictable. The main outlier is comfyui_workflow, which lacks a verb and breaks the otherwise consistent convention.
17 tools is on the heavier end for an MCP server, though the breadth of image generation, video generation, skills, gallery search, ComfyUI workflows, and preferences mostly justifies the count. The set feels slightly large but not bloated to the point of being unwieldy.
The server covers the core workflow well: skill discovery, image upload, specialized generation, generic generation, video generation, status polling, prompt enhancement, inspiration, gallery search, and preferences. Minor gaps exist such as no explicit cancel operation, no image editing/remix tool, and no direct account/billing management, but these do not create dead ends.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Images, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Pay as you go.
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
Self-hosted AI prompt library: prompts, collections, tags, teams, chains. 29 MCP tools for agents.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI assistants to generate and edit images using Google's Gemini 2.5 Flash Image API with intelligent prompt enhancement. Supports text-to-image generation, image editing with natural language instructions, and advanced features like character consistency and multi-image blending.12,925 npm166MIT
- FlicenseNot gradedqualityFmaintenanceEnables AI image generation using Doubao's Seedream 4.0 model through natural language prompts. Automatically downloads generated images to local directories with configurable parameters like resolution, watermarks, and batch generation.8-
- AlicenseAqualityFmaintenanceEnables AI assistants to generate images from text prompts and transform existing images using Google Gemini's nano banana model through the Nanana AI service. Supports both text-to-image generation and image-to-image transformation capabilities.2117 npm10MIT
- AlicenseAqualityAmaintenanceEnables AI image generation through Volcano Engine's Seedream 4.0 API, supporting text-to-image, image-to-image, multi-image fusion, and sequential generation with automatic local saving and Markdown support.5183 PyPI22MIT