openrouter-image-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@openrouter-image-mcpgenerate a watercolor image of a fox in a snowy forest and save it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
openrouter-image-mcp
A local MCP server that lets Claude generate and edit images with any model on OpenRouter. Models are discovered live, nothing is hard-coded, and results are saved with a metadata sidecar and the cost shown: edits next to the input image, new images in ~/Pictures/OpenRouter Images by default.
It runs on your machine over stdio and is installed with uvx straight from this repo. You sign in once in the browser; you never see or paste an API key.
Tools
Tool | What it does |
| Shows whether you are signed in, the key's label, its usage (today, this week, this month, total) and its spending limit. It does not show the workspace. |
| Starts browser sign-in. Returns at once; you finish in the browser. |
| Deletes the stored key from this machine. |
| Lists image models live from OpenRouter, with filters for search, image input, author and sort order. |
| Shows one model's real aspect ratios, resolutions, quality values, limits and pricing. |
| Text to image. Saves to |
| One or more input images plus a prompt to a new image. Supports masks and exact output sizes. |
The skills/ folder holds two companion skills: openrouter-image (model choice, cost habits, masks, moderation) and architectural-render-polish.
Related MCP server: Image MCP
Prerequisites
Install uv and Git (uv uses Git to fetch the server from this repo):
winget install astral-sh.uv # Windows
winget install Git.Git # Windows
brew install uv git # macOS (or run `xcode-select --install` for Apple's Git)Open a new terminal afterwards so uvx and git are on your PATH, and restart Claude (quit it fully) after installing uv or Git so it sees them too.
Install
Claude Code
/plugin marketplace add skelly-77/openrouter-image-mcp
/plugin install openrouter-image@skelly-77-toolsThis installs the server config and both skills.
Claude Desktop (Windows)
Print a config snippet filled in with this machine's paths:
uvx --from git+https://github.com/skelly-77/openrouter-image-mcp@v0.1.0 openrouter-image-mcp print-config --client desktopOpen the config file through Settings → Developer → Edit Config and merge the
openrouter-imageentry into itsmcpServersobject (back the file up first). Use this route on the Microsoft Store (MSIX) build in particular: edits made by hand under%APPDATA%\Claudecan be redirected or vanish there. The command only prints; it never edits files.Restart Claude Desktop fully (quit from the tray icon).
Why the UV_* variables? The Microsoft Store (MSIX) build of Claude Desktop virtualizes writes under AppData, so uv's default cache, tools and Python folders can be hidden or discarded. The snippet sets UV_PYTHON_INSTALL_DIR, UV_CACHE_DIR and UV_TOOL_DIR to real folders under %USERPROFILE%\.uv, written out literally because the config file does not expand variables.
Skills for Desktop: download the repo for the release tag (v0.1.0 zip), unzip it, then zip each folder under skills/ on its own (so each zip contains the skill folder with its SKILL.md) and upload them in Settings → Capabilities.
GitHub Copilot Code (to be confirmed)
Not yet verified: the config file location, its exact format and the skills folder are still open. The standard stdio entry is below. print-config --client copilot prints the same JSON to stdout (so you can pipe or paste it) and an "unverified" note to stderr.
{
"mcpServers": {
"openrouter-image": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/skelly-77/openrouter-image-mcp@v0.1.0",
"openrouter-image-mcp"
]
}
}
}If its sandbox cannot open a browser or reach the credential store, sign in from a terminal (below).
First sign-in
Either run this in a terminal:
uvx --from git+https://github.com/skelly-77/openrouter-image-mcp@v0.1.0 openrouter-image-mcp loginor just ask Claude to run auth_login, then approve in the browser. The key is stored in your operating system's credential store (Windows Credential Manager, macOS Keychain) and nowhere else. Other commands: status, logout, login --switch (different account).
Organization billing and privacy
Keys are locked to your firm's OpenRouter organization's image workspace, so every call spends org credits, not personal ones.
Credits must be added by an org admin (account switcher → the organization → Credits) before anyone can generate. Only org admins buy credits or see billing.
Organizations are limited to 10 members by default (contact OpenRouter support for more).
Privacy: members see request metadata (model, cost, tokens, creator) for everyone's requests in the workspace, but prompt and response content only for their own requests. Org admins can view everyone's prompts and outputs.
Client renders and other uploaded images pass through OpenRouter to the model provider (for example OpenAI or Google). Get the client's or project lead's OK before uploading client imagery. The skills make Claude ask first.
Admins can set guardrails on the workspace: allowed models, budgets and data policy.
Members who leave: a member who leaves or is removed from the organization loses access to its credits and API keys. A member who still has active API keys in a workspace can't be removed from that workspace until those keys are deleted — delete their keys first (Workspace → API Keys).
Configuration
All settings are optional environment variables. Set them in the env block of the server's MCP config entry.
Variable | Default | Meaning |
| the organization's image workspace | Workspace that sign-in creates keys in. An empty string means your personal account. |
|
| Where |
|
| Input images are downscaled to this longest edge (pixels) before upload. |
|
| Seconds to wait for one generation. |
|
| How close (3%) a model's aspect ratio must be to the input's for |
Cost visibility
Each generation reports its cost and the sidecar records it. Failed or cancelled generations are not charged. Ask Claude for account_status for today's and this month's usage. Larger n and max-quality settings cost more, so iterate cheaply first.
Troubleshooting
Symptom | Fix |
401 or "not signed in" | The key was deleted or revoked. Ask Claude to run |
Server won't start in Desktop, or uv errors about permissions or missing folders | Add the |
Sign-in doesn't stick | Check Windows Credential Manager for an entry named |
"Git executable not found" or "Failed to clone" when the server starts | Install Git ( |
"Couldn't reach the OpenRouter catalog" | A network or proxy problem fetching the model list. Retry; check you can open https://openrouter.ai/api/v1/images/models in a browser. |
Moderation refusal | The model's provider declined the prompt or image. Reword the prompt, drop the flagged element, or try another model. |
Future work
A hosted remote MCP (streamable HTTP plus OAuth) would remove the need for a local uv install, but it means running a service that holds user keys, so it is out of scope for v1.
Development
uv sync
uv run pytest # unit tests, no network
uv run pytest -m live # hits real OpenRouter; sign in first
uv run ruff check
uv run python scripts/sync_plugin_skills.py # after editing skills/, refresh plugin/skills/CI runs ruff and the unit tests on Windows, macOS and Ubuntu.
License
MIT, see LICENSE.
Available Tools
7 toolsaccount_statusA
Show whether this computer is signed in to OpenRouter, with the key's spending.
Reports the key label, usage today / this week / this month / in total, and the
key's spending limit and remaining credit. If a sign-in is in progress, shows its
status and link. Use it after auth_login to confirm sign-in finished, or when the
user asks about cost, usage or credits. Takes no parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it largely does: it enumerates what is reported (key label, usage today/week/month/total, spending limit, remaining credit) and handles the in-progress sign-in case with a status and link. It stops short of stating prerequisites (e.g. behavior when not signed in) or read-only guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose followed by returned data, usage triggers, and the no-parameter note. Every sentence adds something, though the return-value inventory is somewhat list-heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must explain returns, which it does in detail, and it supplies usage triggers. It is close to complete for a zero-parameter status tool, missing only edge-case behavior such as the not-signed-in response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which sets a baseline of 4, and the description usefully confirms this with 'Takes no parameters.' There is nothing further to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show whether this computer is signed in to OpenRouter, with the key's spending') and is clearly distinguishable from siblings like auth_login/auth_logout, which are actions on the same auth state. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers: 'after auth_login to confirm sign-in finished, or when the user asks about cost, usage or credits.' It names the sibling it follows but does not state when *not* to use it or contrast it against any alternative status source, so it falls short of the 5 bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_loginA
Sign in to OpenRouter in the user's browser.
Opens the OpenRouter sign-in page and returns at once with a link. Show the link to
the user in case the browser didn't open (it contains no secret). The user finishes
in the browser; then call account_status to confirm. The new key is stored in the
OS credential store and is never shown. Does nothing if already signed in, unless
switch_account is true.
Args: switch_account: true to sign in as a different account or workspace. The current key is kept until the new one is stored.
| Name | Required | Description | Default |
|---|---|---|---|
| switch_account | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so: it returns at once with a link, the link contains no secret, the new key goes to the OS credential store and is never shown, and it is idempotent unless switch_account is true. It even discloses that the current key is retained during a switch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and the link-return contract, then the follow-up step, then the parameter note. Slightly verbose in places (the secret/link caveat), but nearly every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description explains what the tool returns (an immediate link) plus the follow-up confirmation flow. An agent has everything needed to invoke it correctly and handle the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description fully compensates for the single parameter: switch_account signs in as a different account or workspace, and the current key is kept until the new one is stored. This is meaningful semantics beyond the schema's bare boolean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (sign in to OpenRouter in the user's browser) and describes the actual mechanism (opens the sign-in page, returns a link). It also distinguishes itself from siblings by naming account_status as the confirmation step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance: finish in the browser, then call account_status to confirm; show the link if the browser didn't open; no-op if already signed in unless switch_account is true. The alternative action (auth_logout) is implicitly separated by the sign-in framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_logoutA
Sign out: delete the stored OpenRouter key from this computer.
The key stays valid on OpenRouter until it is deleted on the keys dashboard; the result links that page. Takes no parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers a non-obvious behavioral fact: the local key deletion does NOT invalidate the key on OpenRouter, which must be deleted on the dashboard. It also notes the result links that page. It stops short of describing any failure modes or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action verb and core effect, followed by the important caveat about remote key validity. Every sentence adds value; the 'Takes no parameters' line is mildly redundant given the schema but harmless.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter logout tool, the description covers what happens locally, what does not happen remotely, and what the result contains. No output schema exists, yet it adequately signals the return. Complete enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description reinforces this with 'Takes no parameters,' which is consistent with the empty schema and leaves no ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Sign out') and the concrete action ('delete the stored OpenRouter key from this computer'), making the resource and effect unambiguous. It is clearly distinguishable from its counterpart auth_login and the other image-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intent (sign out) is implied clearly enough that an agent will know when to call it, but there is no explicit when-to-use vs when-not guidance or reference to alternatives such as auth_login. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageA
Edit an image (or create one from reference images) with an OpenRouter image model.
Call list_image_models first to choose a model, picking one whose inputs column
says "text+image". Options are checked against the model before anything is sent.
The input images are uploaded to a third-party model provider through OpenRouter: get the user's OK before sending client or project images to third-party providers. Each call costs money on the user's OpenRouter account; the result reports the cost.
All paths must be absolute (or start with "~"); relative paths are rejected. Results are saved next to the first input image (never overwriting it) with a JSON sidecar, unless output_dir is given. The result lists the saved paths, notes, failures, model, provider, time and cost, plus JPEG previews.
Args:
prompt: The change to make, e.g. "make it dusk with warm window light".
model: Model id from list_image_models.
images: Absolute paths of the input images. The first is the image being
edited; any others are references.
mask_path: Optional absolute path of a mask for the first image: white = change,
black = keep. The edit is composited locally, so only the white area
changes, but there may be a visible lighting seam at the mask edge where
new and old pixels meet; a soft mask edge or mask_feather_px helps.
Requires fit="preserve".
mask_feather_px: Blur radius in pixels for the mask edge; by default it is
chosen from the image size.
fit: "preserve" (default) returns each result at exactly the first input's
pixel size: the server picks the model's closest aspect ratio, then scales
and crops back (padding the input first if no ratio is close). With
fit="preserve" don't pass aspect_ratio or size. "model" keeps the size the
model returns and lets you choose aspect_ratio.
n: Number of images, 1 to 10. Above the model's max n, the server splits them
into several calls (each billed) and says so in the notes.
aspect_ratio: One of the model's aspect ratios (only with fit="model").
resolution: One of the model's resolution tiers.
size: Exact pixel size such as "1024x1024" (only with fit="model").
quality: One of the model's quality values.
seed: Integer for repeatable results, for models that support it.
background: One of the model's background values, e.g. "transparent".
output_format: One of the model's output formats, e.g. "png".
output_dir: Folder to save into instead of next to the first input. Must be
an absolute path or start with "~".
filename_prefix: File name stem; defaults to the first input's name.
provider_options: Provider passthrough options as a flat dict, e.g.
{"moderation": "low"}. Keys must be in the model's passthrough parameters
(see get_image_model); the server sends each key to every provider that
allows it. Any other key is rejected before anything is spent.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| fit | No | preserve | |
| seed | No | ||
| size | No | ||
| model | Yes | ||
| images | Yes | ||
| prompt | Yes | ||
| quality | No | ||
| mask_path | No | ||
| background | No | ||
| output_dir | No | ||
| resolution | No | ||
| aspect_ratio | No | ||
| output_format | No | ||
| filename_prefix | No | ||
| mask_feather_px | No | ||
| provider_options | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden and does so: third-party upload path, monetary cost per call on the user's account, absolute-path enforcement, non-overwriting save behavior with JSON sidecar, n splitting into multiple billed calls, provider_options validation before spending, and the mask seam artifact. This is well beyond what the schema declares.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then the critical prerequisites, then args. It is long, but with 17 undocumented parameters the length is largely earned; a few arg entries (e.g. resolution/quality) could be compressed further.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description describes the return payload (saved paths, notes, failures, model, provider, time, cost, JPEG previews) and covers the full parameter surface, consent, cost, and side effects. Nothing material an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 17 parameters, and the description compensates almost completely: white=change/black=keep mask semantics, feather default derivation, fit=preserve crop/pad behavior, aspect_ratio/size only under fit='model', seed repeatability caveat, and array ordering ('first is the image being edited; any others are references').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('edit an image') and broadens scope to 'create one from reference images' with a named backend ('OpenRouter image model'). An agent can tell this apart from generate_image in practice, but the sibling is never named and the create-from-references overlap is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit prerequisites and ordering: 'Call `list_image_models` first to choose a model, picking one whose inputs column says "text+image"', plus a consent gate ('get the user's OK before sending client or project images') and a conditional constraint ('With fit="preserve" don't pass aspect_ratio or size'). It states when-not for several parameters, not just when.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate new images from a text prompt with an OpenRouter image model.
Call list_image_models first to choose a model, and get_image_model for its exact
allowed values. Options are checked against the model before anything is sent: an
invalid value fails with the allowed values and costs nothing.
Each call costs money on the user's OpenRouter account; the result reports the cost. Prompts go to a third-party model provider through OpenRouter: get the user's OK before sending client or project images or details to third-party providers.
Images are saved (never overwriting) to the configured output folder with a JSON sidecar of the settings used. The result lists the saved paths, any notes and failures, the model, provider, time and cost, plus JPEG previews.
Args:
prompt: What to create.
model: Model id from list_image_models.
n: Number of images, 1 to 10. Above the model's max n, the server splits them
into several calls (each billed) and says so in the notes.
aspect_ratio: One of the model's aspect ratios, e.g. "16:9".
resolution: One of the model's resolution tiers.
size: Exact pixel size such as "1024x1024", for models that support it.
quality: One of the model's quality values.
seed: Integer for repeatable results, for models that support it.
background: One of the model's background values, e.g. "transparent".
output_format: One of the model's output formats, e.g. "png" or "svg".
output_dir: Folder to save into instead of the default. Must be an absolute
path or start with "~".
filename_prefix: File name stem; defaults to a slug of the prompt.
provider_options: Provider passthrough options as a flat dict, e.g.
{"moderation": "low"}. Keys must be in the model's passthrough parameters
(see get_image_model); the server sends each key to every provider that
allows it. Any other key is rejected before anything is spent.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| size | No | ||
| model | Yes | ||
| prompt | Yes | ||
| quality | No | ||
| background | No | ||
| output_dir | No | ||
| resolution | No | ||
| aspect_ratio | No | ||
| output_format | No | ||
| filename_prefix | No | ||
| provider_options | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that validation happens before any spend ('an invalid value fails with the allowed values and costs nothing'), that each call is billed to the user's account, that n above the model max splits into multiple billed calls, and that files are saved never overwriting plus a JSON sidecar. It also describes the return payload in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose, the pre-call workflow and the cost/consent warnings are front-loaded, which is the right ordering. The Args block is long, but each line carries non-redundant semantics for a 13-parameter tool with zero schema coverage; only slight tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description enumerates the result contents (saved paths, notes and failures, model, provider, time, cost, JPEG previews), covers the cost/consent/privacy dimensions, and routes to the two prerequisite lookup tools. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate entirely, and it documents all 13 parameters with meaning beyond the schema: n's 1-10 range and server-side splitting behavior, output_dir's absolute-or-tilde path rule, filename_prefix defaulting to a prompt slug, and provider_options' passthrough validation rules. This is far more than the bare schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate new images from a text prompt') plus the backend ('OpenRouter image model'), which cleanly distinguishes it from siblings edit_image, list_image_models and get_image_model. An agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the workflow ('Call list_image_models first... and get_image_model for its exact allowed values') and the human-in-the-loop condition, plus a concrete when-not ('get the user's OK before sending client or project images or details to third-party providers'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_image_modelA
Show full details for one image model.
Returns the description, the capabilities, and one section per provider endpoint
with its supported parameters and allowed values, the passthrough parameters usable
in provider_options, and its pricing line items.
Args:
model_id: Exact model id from list_image_models.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the return content in detail, but says nothing about permissions, authentication requirements, rate limits, or whether missing ids raise an error versus returning empty. For a no-annotation tool that is a real, if modest, gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a scannable enumeration of the payload and a short Args block. No filler, though the multi-line return breakdown is slightly longer than needed for a one-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing the returned sections is exactly the right compensation and is done well. The main omission is any behavioral/prerequisite context (auth, error behavior), which matters more given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single param is undocumented in the schema, so the description must compensate. It does: it defines model_id as the exact id and tells the agent where to obtain it (list_image_models), which is meaning beyond the bare 'Model Id' title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show full details for one image model') and enumerates what comes back (description, capabilities, provider endpoint sections, passthrough params, pricing). This clearly separates it from the sibling list_image_models, which returns a collection rather than a single model's full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the model_id arg is described as 'Exact model id from `list_image_models`', which hints you should list first. There is no explicit statement of when to prefer this over list_image_models or any exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_image_modelsA
List the image models available on OpenRouter right now, as a markdown table.
Use this before generate_image or edit_image to choose a model and see which
options it takes. The catalog is live (cached for 10 minutes), so don't rely on
remembered model ids. Columns: id, name, inputs ("text", or "text+image (max N)"
input images), aspect ratios, resolutions, quality values, max n per call, seed
support, whether the provider moderates prompts (yes/no/?), and a price summary
($/img-unit per image-output unit, $/image, or $/tok per completion token).
Call get_image_model for one model's per-provider details.
Args:
query: Case-insensitive text to find in the model id or name.
accepts_images: true for models that take input images (needed by
edit_image); false for text-only models.
author: Only this author's models: the part of the id before "/".
sort: "newest" (default, newest first) or "price" (cheapest first).
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | newest | |
| query | No | ||
| author | No | ||
| accepts_images | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and mostly does: it discloses that the catalog is live, cached for 10 minutes, and that moderation status may be unknown ('?'). It does not mention auth requirements or rate limits, but for a read-only listing that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then usage, then a column inventory, then args. The column enumeration is lengthy but earns its place because there is no output schema; slightly over-detailed for a listing tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description supplies the return shape by naming every column and its semantics (inputs, aspect ratios, resolutions, quality, max n, seed, moderation, price units). Combined with the cache/liveness note and full param docs, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it does, defining all four args: case-insensitive query scope, accepts_images tied to edit_image needs, author as the id prefix before '/', and both sort values with their defaults and ordering.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (image models on OpenRouter), plus the output shape (markdown table) and liveness of the catalog. It is clearly separable from sibling get_image_model, which is explicitly scoped to 'one model's per-provider details.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule: 'Use this before generate_image or edit_image to choose a model and see which options it takes.' It also routes the agent elsewhere when only one model is needed ('Call get_image_model for one model's per-provider details'), and warns against stale remembered model ids.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
account_status - First observed
auth_login - First observed
auth_logout - First observed
edit_image - First observed
generate_image - First observed
get_image_model - First observed
list_image_models
TDQS
Scored across 7 tools
Each tool maps to a clearly distinct action: auth_login/auth_logout are opposite operations, account_status is a read-only check, list_image_models vs get_image_model cleanly separate browsing from detail lookup, and generate_image vs edit_image separate creation from modification. No meaningful overlap exists.
Most tools follow a readable verb_noun pattern (list_image_models, get_image_model, generate_image, edit_image, auth_login, auth_logout). Minor deviation: the auth group mixes prefixes ('auth_' for login/logout vs 'account_' for status), but it remains predictable.
Seven tools is well-scoped for an image-generation service: three for auth/session, two for model discovery, two for generation/editing. Every tool earns its place with no redundancy or bloat.
The surface covers the full lifecycle: sign in, verify session/spend, sign out, discover models, inspect model details, and both generate and edit images with rich parameter control. No obvious gaps for the stated domain.
Maintenance
Related MCP Connectors
- lightgenOAuthapp.lightgen
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
AI image and video generation, talking avatars, consistent characters and photo packs from Claude.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI image generation, editing, composition, and style transfer in Claude conversations using Google's Gemini 2.5 Flash model. Automatically saves generated images to a local directory.469 npm11MIT
- AlicenseAqualityBmaintenanceConnects Claude Code to image generation models (OpenAI GPT Image 2, Google Nano Banana) for generating, editing, and converting images via natural language.638 npmMIT
- AlicenseAqualityCmaintenanceEnables Claude to generate and edit images using OpenAI's GPT Image models, with automatic model selection and local file saving.3MIT
- AlicenseNot gradedqualityBmaintenanceExposes OpenRouter's image generation, video generation, image-to-video, and model discovery to local AI clients such as Claude Code, OpenCode, OMP, and Codex over stdio. It lets clients list available models, generate images and videos, track and download video jobs, and resume in-flight jobs without blocking past client tool timeouts.MIT