opencode-image-mcp
Provides image generation using Google Gemini's image models via its OpenAI-compatible endpoint, with model auto-discovery and file output.
Provides image generation using OpenAI's image models (e.g., gpt-image, dall-e) via the Images API, with model auto-discovery and file output.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@opencode-image-mcpMake me a corgi astronaut in 16:9."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
opencode-image-mcp
A multi-provider image generation MCP server. Point it at any OpenAI-compatible API, let it auto-detect the image models each provider offers, pick the model you ask for, then generate and wait for the result — saving the image to disk and returning its path (optionally inline).
It is a generalized, hardened take on the popular single-provider
openrouter-image-mcp: instead of one hardcoded vendor, it speaks both common
OpenAI-compatible dialects and works with OpenAI, OpenRouter, Google Gemini,
xAI, Together, DeepInfra, SiliconFlow, Fireworks, or any custom endpoint you
configure.
Bring your own endpoint
No preset required. Hand the agent a base URL and it registers the endpoint, discovers its models, and remembers it:
Use
https://api.example.com/v1with keysk-abc123and call itexample.
The agent calls add_provider, which:
probes
https://api.example.com/v1/models,keeps only the models that can produce images,
saves the endpoint (and key) to the gitignored config file,
sets that model as the default.
From then on, in this session and every future run:
Make me a corgi astronaut in 16:9.
...resolves straight to your endpoint. If the endpoint exposes several image
models, the agent inspects the list and calls set_default_model to pick one.
Endpoints without a /models route still work — pass default_model and
discover=false.
Related MCP server: Image Gen Pro MCP
Features
Any OpenAI-compatible provider — built-in presets plus arbitrary custom base URLs.
Two API dialects, auto-selected — the classic Images API (
/images/generations,/images/edits) and the chat-completionsmodalities: ["image","text"]style used by aggregators. Includes automatic chat fallback when a provider rejects the Images endpoint.Model auto-discovery — queries each provider's model endpoint, filters to image-capable models, and caches the result on disk. OpenRouter's dedicated
/images/modelscapability metadata is used when available.Flexible model selection —
provider:model, bare model ids searched across providers, or user-defined aliases.Waits for slow jobs — long timeouts, retry with backoff on
429/5xx(honouringRetry-After), and generic polling for providers that return a job id instead of an image.Robust output handling — base64, data URLs and hosted URLs; MIME/extension detection from magic bytes; multi-image
_2/_3suffixes; never clobbers existing files.Friendly errors — 401/402/404/413/429/5xx and timeouts all become clear, actionable messages naming the provider and model.
Inline previews on demand — return file paths by default, or embed the generated image in the tool result so the model can see it.
Small footprint — Python + the MCP SDK +
httpx. Nothing else.
Supported providers (optional presets)
These presets are just convenient shortcuts — none of them are required, and the custom-endpoint flow above works with any OpenAI-compatible API.
Provider | Dialect | Model list endpoint | Notes |
| images |
|
|
| images + chat |
| capability-aware; Gemini/GPT/Flux/Seedream via one key |
| images |
| Google's OpenAI-compatible endpoint |
| images |
| Grok Imagine; returns hosted URLs |
| images |
| FLUX and friends ( |
| images |
| FLUX ( |
| images |
|
|
| images |
|
|
| any | configurable | LM Studio, vLLM, ComfyUI gateway, proxies… |
A provider becomes active the moment its API key env var is set. Unconfigured
presets are listed as configured: false and skipped.
Install
Requires Python 3.10+ and (recommended) uv.
From a git URL
uv tool install --from git+https://github.com/you/opencode-image-mcp opencode-image-mcpFrom a local checkout
git clone https://github.com/you/opencode-image-mcp
cd opencode-image-mcp
uv tool install --reinstall --from . opencode-image-mcpTwo ways to launch it:
the console script
opencode-image-mcpthe module entry point
python -m opencode_image_mcp(more reliable on Windows, where the generated.exelauncher can miss the venv's site-packages)
Find the tool's venv Python with uv tool dir, then use it in your client
config as shown below.
Run from source without installing
uv venv
uv pip install -e ".[dev]"
python -m opencode_image_mcpConfigure
1. API keys (environment variables)
Set any subset; each one enables that provider:
OPENAI_API_KEY=sk-...
OPENROUTER_API_KEY=sk-or-v1-...
GEMINI_API_KEY=... # GOOGLE_API_KEY also accepted
XAI_API_KEY=xai-...
TOGETHER_API_KEY=...
DEEPINFRA_API_KEY=...
SILICONFLOW_API_KEY=...
FIREWORKS_API_KEY=...2. Server settings (optional)
Variable | Default | Purpose |
| the workspace (current directory) | Default save directory |
|
| Per-request timeout in seconds |
|
| Retry attempts for |
|
| Model-list cache TTL in seconds |
|
| Where the model cache lives |
|
| Extra config file |
| — | Configure the |
| — | Point the |
|
| Logging level (logs go to stderr) |
3. Custom providers, aliases and defaults (providers.json)
Copy providers.example.json and edit. Only the keys you set are changed.
{
"default_model": "openrouter:google/gemini-3.1-flash-image-preview",
"aliases": {
"nb2": "openrouter:google/gemini-3.1-flash-image-preview",
"fast": "gemini:gemini-2.5-flash-image"
},
"settings": { "output_dir": "./generated", "timeout_s": 300 },
"providers": {
"local": {
"base_url": "http://localhost:8000/v1",
"dialect": "images",
"api_key_env": null,
"default_model": "flux-schnell",
"include_models": ["flux", "sdxl"],
"aspect_ratio_mode": "size"
},
"comfy": {
"base_url": "http://localhost:8188/openai/v1",
"dialect": "chat",
"default_model": "comfy-image",
"poll": {
"url_template": "{base_url}/jobs/{id}",
"status_field": "status",
"done_values": ["completed", "done", "succeeded"],
"failed_values": ["failed", "error", "cancelled"],
"result_url_field": "result_url",
"interval_s": 2,
"max_interval_s": 10
}
},
"together": null
}
}Set a provider to null to disable a preset. Setting api_key_env to null
marks a provider as anonymous (no key required) — handy for local servers.
Where runtime additions are stored: add_provider writes to
IMAGE_MCP_CONFIG when set, otherwise ~/.config/opencode-image-mcp/providers.json.
A key passed to add_provider is stored inline in that file. It is gitignored
and blocked by the pre-commit hook, so it never reaches the cloud — but keep the
file out of shared/backup folders if you would rather the key not sit on disk.
Provider keys: base_url, dialect, default_model, label, api_key,
api_key_env, alt_api_key_envs, auth_header, auth_scheme, models_path,
image_models_path, generation_path, edit_path, chat_path, edit_style,
aspect_ratio_mode, size_style, size_param, send_response_format,
chat_modalities, include_models, exclude_models, extra_headers,
extra_query, allow_chat_fallback, poll. Unknown keys are rejected so typos
surface immediately.
Field | Meaning |
|
|
|
|
|
|
|
|
| regex allow/deny applied to model ids |
| e.g. Azure-style |
| e.g. |
| enables waiting for async job APIs |
Wire it into your MCP client
Replace the path with the tool venv's Python (uv tool dir → .../Scripts/python.exe
on Windows, .../bin/python elsewhere).
opencode / Claude Code (~/.claude.json)
{
"mcpServers": {
"opencode-image": {
"command": "C:/Users/YOU/AppData/Roaming/uv/tools/opencode-image-mcp/Scripts/python.exe",
"args": ["-m", "opencode_image_mcp"],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-...",
"OUTPUT_DIR": "C:/Users/YOU/Pictures/generated"
}
}
}
}Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"opencode-image": {
"command": "/Users/YOU/.local/share/uv/tools/opencode-image-mcp/bin/python",
"args": ["-m", "opencode_image_mcp"],
"env": { "OPENAI_API_KEY": "sk-..." }
}
}
}Cursor (.cursor/mcp.json)
{
"mcpServers": {
"opencode-image": {
"command": "/path/to/uv/tools/opencode-image-mcp/bin/python",
"args": ["-m", "opencode_image_mcp"],
"env": { "GEMINI_API_KEY": "..." }
}
}
}Generic stdio client
Set the env vars in the parent shell and run python -m opencode_image_mcp.
Tools
list_providers
No arguments. Returns every known provider with configured, key_env,
dialect, default_model, supports_edit and polls_async_jobs, plus the
effective default model and aliases.
list_models(provider=None, refresh=False)
Discovers image models from configured providers, cached on disk. Each entry
has provider, id, full_id, name, source and, when the provider exposes
them, capabilities and pricing. Errors per provider are returned under
errors. Use refresh=true to bypass the cache.
add_provider(...)
Register any OpenAI-compatible image endpoint at runtime. It probes
{base_url}/models, filters to image models, persists the provider to the
config file, and makes it usable immediately and in future sessions.
Parameter | Type | Default | Description |
| string | required | Short handle for |
| string | required | API root, e.g. |
| string |
| Bearer token, stored in the gitignored config file |
| string |
| Read the key from an env var instead (ignored if |
| string |
| Model to use by default; auto-picked if exactly one image model is found |
| string |
|
|
| bool |
| Also set this as the global default model |
| object |
| Extra request headers |
| string[] |
| Regex allow/deny applied to discovered model ids |
| string |
|
|
| string |
|
|
| bool |
| Set |
Returns the discovered models, the chosen default_model, key_source,
warnings and the saved_to path. The key is never echoed back.
set_default_model(model)
Persist the model used when generate_image is called without one, e.g.
example:flux-schnell. Saved to the config file, so it applies to future runs.
remove_provider(id)
Remove a provider (including a built-in preset) from the config file. Clears the default model if it pointed at that provider.
generate_image(...)
Parameter | Type | Default | Description |
| string | required | Image description, or an edit instruction with input images |
| string | config default |
|
| string |
| e.g. |
| string |
| Explicit size, e.g. |
| string |
|
|
| int |
| Number of images |
| string |
| Exact file, a directory, or omitted for the workspace directory |
| string |
| Stem for auto-generated names |
| string |
| One reference image → edit mode |
| string[] |
| Multiple references (where supported) |
| string |
|
|
| string |
|
|
| int |
| Deterministic generation (where supported) |
| bool |
| Also embed the image in the result |
| float |
| Override the configured timeout |
| object |
| Escape hatch: merged into the request body |
Return shape (also available as structured_content):
{
"file_paths": ["C:/.../corgi_20260926_120000.png"],
"provider": "openrouter",
"model_id": "google/gemini-3.1-flash-image-preview",
"dialect": "images",
"mode": "generate",
"prompt": "a corgi astronaut",
"aspect_ratio": "16:9",
"image_count": 1,
"elapsed_ms": 8421,
"usage": { "prompt_tokens": 22, "completion_tokens": 1500, "cost": 0.06834 },
"cost_usd": 0.06834
}Example prompts
Generate a cinematic corgi astronaut in 16:9 using
openrouter:google/gemini-3.1-flash-image-preview.
List the image models my providers offer, then use the cheapest one.
Take
./photo.jpgand remove the background, save to./cutout.png.
How model selection works
provider:model— explicit and unambiguous. No discovery needed.Bare model id — the server searches every configured provider's cached catalogue (exact, then case-insensitive, then suffix match). If several providers match, it lists the candidates and asks you to pick one.
Alias — any key from
aliasesinproviders.json.Omitted —
default_modelfrom config, else the first configured provider's default.
Discovery uses OpenRouter's authoritative /images/models capability metadata
when available; otherwise it filters model ids by image-family heuristics
(flux, dall-e, gpt-image, seedream, imagen, grok-imagine,
qwen-image, …) refined by your include_models/exclude_models regexes.
Waiting for the response
Requests use a generous timeout (default 300s) — image generation is slow.
429/5xxand network errors are retried with exponential backoff and jitter, honouringRetry-After.If a provider answers with a job id instead of an image, configure
polland the server will poll until the job completes, then fetch the result (inline or viaresult_url).
Troubleshooting
Symptom | Fix |
| No API keys visible. Set them in the client's |
| Wrong/expired key for that provider; check the named env var. |
| Unsupported parameter for that model — drop |
| Model uses the other dialect; try another model/provider (chat fallback is automatic where possible). |
| The model answered with text (refusal or chatter). Rephrase or switch model. |
Wrong extension | Never happens by design — extensions come from magic bytes, not the URL. |
| Use |
Development
The test suite is local only — tests/ is gitignored and never committed.
uv venv
uv pip install -e ".[dev]"
python -m pytestLayout:
src/opencode_image_mcp/
config.py # presets, providers.json read/write, env, Settings
registry.py # model discovery, cache, resolution
dialects.py # images/chat request builders + response parsers
transport.py # httpx client, retries/backoff, downloads, polling
images.py # MIME sniffing, decode, save
errors.py # friendly error mapping
server.py # MCP toolsThe tests mock HTTP with httpx.MockTransport, so they run offline and cost
nothing.
Repository safety (keeping keys out of the cloud)
.env and providers.json are gitignored, so real keys and custom endpoints
never get committed. That is backed by a pre-commit hook in .githooks/ which
also catches git add -f and any staged diff that looks like a live API key.
The tests/ directory is gitignored too: it stays on your machine for local
runs and is never committed.
The hooks path is local git config, so enable it once per clone:
git config core.hooksPath .githooksThen verify:
git check-ignore -v .env providers.json # both should be listedBypass only if you are certain: git commit --no-verify.
Companion skill
skills/image-generation/SKILL.md is an optional agent skill with a model
selection decision tree, prompt-engineering tips, aspect-ratio guidance and
common pitfalls. Install it by copying the folder into your client's skills
directory.
License
MIT — see LICENSE.
Available Tools
6 toolsadd_providerAdd a custom providerAIdempotent
Register any OpenAI-compatible image endpoint and remember it for later.
Use this when the user names an endpoint like https://api.example.com/v1.
The server probes {base_url}/models, figures out which models can produce
images, saves the provider to the config file, and makes it available to
generate_image immediately and on every future run.
Parameters
id:
Short handle used in provider:model specs (e.g. example).
Lowercase letters, digits, - and _.
base_url:
The OpenAI-compatible API root, e.g. https://api.example.com/v1.
api_key:
Optional bearer token, stored in the gitignored config file so it
survives restarts. Prefer api_key_env if you would rather keep the
secret in an environment variable.
api_key_env:
Name of an env var holding the key (ignored when api_key is given).
Leave both unset for endpoints that need no auth.
default_model:
Model id to use by default. If omitted and exactly one image model is
found, it is selected automatically.
dialect:
auto (default), images (/images/generations) or chat
(/chat/completions with image modalities). auto uses the Images
API and transparently falls back to chat-completions on 404/405.
make_default:
Also set this provider's model as the global default model.
discover:
Set false to skip the /models probe (required for endpoints
that do not implement it).
Returns the discovered models plus any warnings; inspect models and, if
there are several, call set_default_model to choose one.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| api_key | No | ||
| dialect | No | auto | |
| base_url | Yes | ||
| discover | No | ||
| size_style | No | ||
| api_key_env | No | ||
| make_default | No | ||
| default_model | No | ||
| extra_headers | No | ||
| exclude_models | No | ||
| include_models | No | ||
| aspect_ratio_mode | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate non-read-only, open-world, and idempotent behavior. The description adds substantial useful details: it probes `/models`, falls back from Images API to chat-completions on 404/405, saves the provider to the config file, stores API keys in a gitignored file, and survives restarts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured, front-loads the main use case, and uses a parameter list for clear scanning. Each section adds necessary behavior or context, though a few bullet items could be tightened to reduce overall length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with zero schema descriptions, the description covers the core workflow, required parameters, authentication, discovery, fallback behavior, persistence, and next steps via `set_default_model`. It is not fully complete because several optional parameters remain undocumented, and the behavior for duplicate provider IDs is not stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description carries the parameter-documentation burden and does so well for 8 of 13 parameters, adding constraints, examples, defaults, and interaction rules such as `api_key_env` being ignored when `api_key` is given. However, `size_style`, `extra_headers`, `include_models`, `exclude_models`, and `aspect_ratio_mode` receive no explanation, leaving some gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource: 'Register any OpenAI-compatible image endpoint and remember it for later.' It also clarifies the tool's role relative to siblings by stating it makes the provider available to `generate_image` immediately and on future runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: 'Use this when the user names an endpoint like `https://api.example.com/v1`.' It also gives concrete guidance on auth alternatives, skipping discovery with `discover=false`, and tells the agent to call `set_default_model` if multiple models are returned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageGenerate or edit an imageA
Generate an image, or edit existing images, via an OpenAI-compatible provider.
Model selection
model accepts three forms:
provider:model— explicit, e.g.openrouter:google/gemini-3.1-flash-image-preview,openai:gpt-image-1,xai:grok-imagine-image-2.0. Fastest and unambiguous.a bare model id — searched across every configured provider's catalogue (run
list_modelsfirst). Ambiguous ids raise an error listing candidates.an alias — defined in
providers.json(e.g.nb2).
If omitted, the configured default model is used.
Editing
Pass input_image_path (and optionally input_image_paths) and the
prompt becomes an edit instruction. Requirements vary by provider.
Output
Images are always written to disk and their paths returned. Set
return_image_content=true to additionally embed the images in the tool
result so the model can look at them.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| size | No | ||
| model | No | ||
| prompt | Yes | ||
| quality | No | ||
| timeout_s | No | ||
| background | No | ||
| extra_body | No | ||
| output_path | No | ||
| aspect_ratio | No | ||
| output_format | No | ||
| filename_prefix | No | ||
| input_image_path | No | ||
| input_image_paths | No | ||
| return_image_content | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful behavioral detail beyond the annotations: images are always written to disk, paths are returned, return_image_content can embed results for the model, ambiguous models raise an error, and provider requirements vary. This gives an agent a clear picture of side effects and output behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with clear sections for model selection, editing, and output. Every section earns its place, and the most action-critical information is front-loaded. It is longer than average but appropriately so for a tool with this many parameters and modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core usage modes, output behavior, and model-selection nuances well. With no output schema and only sparse annotations, it leaves some parameter semantics to inference, but the most important behavioral and decision-relevant context is present. A small gap remains for the unmentioned parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter-semantics burden. It explains model, prompt, input_image_path(s), output_path, and return_image_content well, but leaves most of the other 16 parameters (n, seed, size, quality, background, extra_body, etc.) undocumented. The value added is solid but incomplete given the parameter count.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: generating or editing images through an OpenAI-compatible provider. It clearly distinguishes this from the sibling tools, which are all provider/model management operations rather than image operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete usage guidance: how to specify models, when aliases are valid, and to run list_models first for bare model ids. It also explains the edit mode via input_image_path. It does not explicitly discuss when not to use the tool, but the sibling tools are different enough that no exclusion is strictly needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList image modelsARead-onlyIdempotent
Discover image-capable models offered by configured providers.
Queries each provider's model endpoint, filters to image models, and caches
the result on disk. Set refresh=true to bypass the cache. Pass
provider to query just one provider.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| provider | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavior beyond annotations: it queries each provider's endpoint, filters to image models, and caches on disk. It also explains the refresh parameter as cache bypass. Annotations already cover read-only and idempotent safety, so the added detail is valuable and non-contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no fluff. The purpose is front-loaded, and behavioral/parameter details follow efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (which covers return values) and read-only annotations, the description fully explains what the tool does, how it behaves (caching, filtering), and how to control it via parameters. Nothing essential is missing for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully explain parameters. It does: 'refresh' is defined as bypassing the cache, and 'provider' as limiting to one provider. This adds meaning beyond the bare type/default schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Discover image-capable models') with a clear resource ('offered by configured providers') and distinguishes it from siblings like list_providers and generate_image. It leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context on when to use (discover available image models) and how to adjust behavior via 'refresh' and 'provider'. It does not explicitly name alternatives or exclusion cases, but the purpose is self-evident in the sibling context, so this meets the 'clear context, no exclusions' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersList image providersARead-onlyIdempotent
List configured image providers and whether their API key is present.
A provider only appears as configured: true once its key env var is set.
Use the returned id values as the prefix in provider:model specs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds a meaningful behavioral rule: a provider only appears with configured:true once its key env var is set. This explains a conditional in the output that wouldn't be obvious from the schema alone, going beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exceptionally lean: two sentences. The first states the purpose, the second adds a behavioral condition and a usage tip. Every phrase earns its place, and the most important information (what it lists) is front-loaded. No redundancy, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with an output schema present, the description covers everything an agent needs: it states the purpose, explains the meaning of configured:true, and tells how to use the returned ids. Return structure and field details are presumably handled by the output schema, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the input schema is empty with 100% coverage. Per the rubric, a zero-parameter tool gets a baseline of 4. The description doesn't add parameter-specific semantics because there are none to explain, and it correctly stays silent rather than inventing unnecessary detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List configured image providers' with the additional detail about API key presence. It differentiates cleanly from siblings like list_models (models, not providers) and mutations like add_provider/remove_provider, so an agent can identify this as the pure read-only listing tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a practical downstream use for the result ('Use the returned id values as the prefix in provider:model specs'), which tells the agent how to leverage the output. It doesn't explicitly contrast with sibling tools (e.g., 'when you need models, use list_models'), but the distinctive title and purpose provide clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_providerRemove a providerAIdempotent
Remove a provider (including built-in presets) from the config file.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnlyHint=false and idempotentHint=true, and the description adds meaningful context by specifying that removal affects the config file and that even built-in presets can be removed. There is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-structured sentence that front-loads the action and object, then adds the built-in presets nuance parenthetically. Every word contributes useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with an output schema and annotations, the description is minimally adequate. However, it leaves the semantics of the sole 'id' parameter and any prerequisites unstated, so it is not fully complete on its own.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for the single 'id' parameter and schema coverage is 0%. The description does not explain what id should reference or how to obtain valid provider ids, so the agent must infer the parameter's meaning from the parameter name alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove') and a clear resource ('a provider'), and adds the important nuance that built-in presets are included. This makes the tool's function unambiguous and distinguishes it from sibling operations like add_provider and list_providers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: use this tool when you need to delete a provider from the config file. However, it does not explicitly state when not to use it or how it compares to add_provider/set_default_model, leaving the usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_default_modelSet the default modelAIdempotent
Persist the model used when generate_image is called without one.
model is a provider:model spec or an alias, e.g.
example:flux-schnell. The choice is saved to the config file so it
applies to future runs.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and idempotentHint=true. The description adds meaningful behavioral context by disclosing that the choice is persisted to the config file and applies to future runs, making the side effect concrete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus an example, with no filler. The core behavior is front-loaded and the parameter format is explained immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter configuration tool, the description covers the purpose, the effect, the persistence behavior, and the parameter format. There is an output schema, so return-value documentation is not needed here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only says 'model' is a string with 0% description coverage. The description compensates fully by defining the expected format as a 'provider:model' spec or alias and giving a concrete example ('example:flux-schnell').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Persist') and resource (the default model), and explicitly ties it to generate_image calls made without a model argument. This clearly differentiates it from sibling tools like list_models or add_provider.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool to set the model that generate_image will fall back to when no model is provided. It does not explicitly list when not to use it or name alternatives, but the scoped usage is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
add_provider - First observed
generate_image - First observed
list_models - First observed
list_providers - First observed
remove_provider - First observed
set_default_model
TDQS
Scored across 6 tools
Each tool serves a distinct function: discovery (list_models, list_providers), provider lifecycle management (add_provider, remove_provider), configuration (set_default_model), and the core action (generate_image). There is no overlap or ambiguity in purpose.
All tool names follow a consistent snake_case verb_noun pattern: list_models, list_providers, add_provider, remove_provider, set_default_model, generate_image. The naming style is uniform and predictable.
Six tools is well-scoped for an image generation MCP server. Each tool earns its place and covers provider management, model discovery, default selection, and generation without redundancy.
The tool surface covers the full lifecycle: configure providers (add/remove/list), discover models (list_models), set defaults, and generate/edit images. No obvious gaps exist for the server's stated purpose.
Maintenance
Related MCP Connectors
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Generate AI images and videos from any compatible MCP client.
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables image generation via OpenRouter API, supporting models like Gemini 2.5 Flash Image Preview with options to save files locally.21Do What The F*ck You Want To Public
- AlicenseAqualityBmaintenanceEnables AI assistants to generate real images via multiple models (OpenAI, Gemini, Recraft, Seedream, Grok, Arrow) and returns usable file paths instead of base64 data.157 npmMIT
- AlicenseAqualityBmaintenanceEnables image analysis, OCR, and text-to-image generation through OpenAI-compatible APIs. Supports local paths, URLs, or base64 images with configurable models and backup endpoints.3275 npmMIT
- FlicenseAqualityCmaintenanceEnables generating and editing images through OpenAI image models, with configurable output as saved files or returned URLs.2-