Northwestern AI image MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Northwestern AI image MCPcheck my account status and sign me in to the Class of 2027 workspace"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Northwestern AI image MCP
Local image generation and editing for Northwestern University AI classes. Each student signs in individually to the Northwestern OpenRouter organization. Claude supplies image analysis; this server retains its eight image and account tools.
Instructor testing only. Student deployment is on hold until OpenRouter increases organization capacity.
Repository: jrenaldi79/northwestern-image-mcp (private). Adapted from skelly-77/openrouter-image-mcp, with the original MIT license and commit history retained.
Cohorts
Cohort | Students | Required OpenRouter workspace |
2027 | Capstone, full-time and second-year | Class of 2027 ( |
2028 | First-year | Class of 2028 ( |
Set OPENROUTER_IMAGE_COHORT=2027 or 2028. Missing, empty, unknown or conflicting selection fails before sign-in. OPENROUTER_IMAGE_WORKSPACE_ID can alternatively select either ID above. There is no personal-account or original-firm fallback.
Related MCP server: openrouter-imgen-mcp
Local Claude Desktop setup
Use this checkout and Python 3.12 or newer. Prepare an environment with uv:
uv venv --python 3.13 .venv
uv pip install --python .venv\Scripts\python.exe --group dev -e .
.venv\Scripts\python.exe -m openrouter_image_mcp.cli print-config --cohort 2027On macOS/Linux use .venv/bin/python instead. print-config produces one northwestern-images entry using this interpreter and checkout. Use --cohort 2028 for a first-year student. Use Settings → This computer → Developer → Edit config to locate the active file; the usual Windows location is %APPDATA%\Claude\claude_desktop_config.json.
Fully exit Claude Desktop through File → Exit before editing. Closing its window can leave it running. Back up the configuration and merge these mcpServers entries, preserving other settings and servers. Reopen Claude Desktop and confirm the single Northwestern entry shows Running in Developer settings. Manually configured local servers are managed there, separately from remote connector setup in Customize. Keep the checkout and environment in place.
Maintain one server and one installation entry. Each student receives the same server configured with their confirmed cohort; the two class workspaces remain in OpenRouter. Student release instructions will follow instructor testing and the capacity increase. The original upstream release does not contain these changes.
Instructor test
Select the 2027 tools in Claude Desktop and ask for
account_status. Confirm the configured workspace above.Ask for
auth_loginon that server and finish browser consent with your Northwestern organization account. Organization/workspace membership must be granted separately; this server does not enroll users.Ask for
account_statusagain and inspect the created key in the corresponding OpenRouter workspace. The target is configuration; a key label alone does not independently prove billing ownership.To test 2028 later, fully exit Claude, rerun the installer with --cohort 2028, and reopen it. This replaces the same entry. Saved credentials remain isolated by workspace.
List image models and inspect pricing. Before a paid test, agree on one image (
n=1), the model and likely cost. Startup and tool discovery do not require paid generation.
Terminal testing selects a cohort first:
$env:OPENROUTER_IMAGE_COHORT = "2027"
.venv\Scripts\python.exe -m openrouter_image_mcp.cli status
.venv\Scripts\python.exe -m openrouter_image_mcp.cli login
.venv\Scripts\python.exe -m openrouter_image_mcp.cli logoutChange the cohort to 2028 for that profile. If OPENROUTER_IMAGE_WORKSPACE_ID is also set, it must match. login --switch replaces only this cohort's credential. Keys use the OS credential store under northwestern-openrouter-image-mcp, indexed by workspace ID. Legacy credentials are neither copied nor reused. Logout removes the local key; revoke it separately in OpenRouter if needed.
Tools
Tool | What it does |
| Shows configured Northwestern cohort/workspace, sign-in state, key label, spending and available limits. |
| Starts browser sign-in constrained to this workspace. |
| Removes this cohort's saved credential. |
| Discovers available image models and pricing. |
| Shows model parameters, providers and prices. |
| Generates images with metadata and reported costs. |
| Edits local images using references, masks and output sizing. |
| Re-blends a saved masked edit locally, with no upload or model charge. |
New images default to ~/Pictures/Northwestern AI/Class of 2027 or Class of 2028. Edits save beside the primary input. Existing explicit output overrides remain available. Files are not overwritten; JSON sidecars record prompts, settings and reported costs. Keep those files in mind when sharing coursework.
Prompts and input images go through OpenRouter to third-party providers. Get explicit permission before uploading other people's images, research materials or confidential project information. Start with one low-cost draft and relay reported spending. Use the live catalog for model choice and prices. Organization/workspace budgets are administered in OpenRouter; this server does not create or increase them.
Inline image previews
The existing server includes an MCP Apps gallery for generate_image, edit_image,
and remask_image. Hosts advertising MCP Apps HTML support receive a bundled viewer
with JPEG previews, filenames, saved paths, reported usage/cost, notes and failures.
Original files still save locally. Clients without that capability keep receiving
ordinary text and image content. SVGs and files beyond the preview limit show their
saved filenames without an inline preview. The gallery needs no network connection.
After updating the checkout, fully exit and reopen Claude Desktop, then start a new chat to refresh tool metadata. Actual inline rendering depends on the installed host's MCP Apps support for local servers. Claude's web app cannot reach this local stdio server directly.
To check the gallery without paying for a model call, prepare a synthetic fixture:
.venv\Scripts\python.exe scripts\prepare_gallery_test.pyPaste the printed remask_image request into the new Claude Desktop chat. It creates
a local preview with a $0.00 cost and uploads nothing. An existing masked edit can
also be re-masked for free; a plain generated image lacks the required mask sidecar.
Companion skills and local plugin
skills/openrouter-image carries the general workflow; skills/architectural-render-polish remains optional. Run python scripts/sync_plugin_skills.py after source skill changes. The bundled Claude Code plugin contains one instructor entry configured for 2027 and uses uv run --no-sync against the prepared checkout; it requires an installed environment and uv on the client's PATH. It is not a published Northwestern distribution.
Development
.venv\Scripts\python.exe -m pytest -q
.venv\Scripts\ruff.exe check .Live tests are excluded by default because they use OpenRouter and may cost money. Original project provenance remains in the Git remote and MIT license. Keep student roster data in the roster spreadsheet.
The gallery HTML is committed in the Python package. Node is only needed to change or rebuild it, not to run the server:
cd ui
npm ci
npm run build
npm testBrowser tests use Playwright Chromium, or set PLAYWRIGHT_CHROMIUM_EXECUTABLE_PATH to an
existing Chrome executable. They exercise the bundled viewer with a synthetic
image and a mock MCP Apps host, without OpenRouter calls.
Available Tools
8 toolsaccount_statusA
Show whether this computer is signed in to OpenRouter, with the key's spending.
Reports the key label, usage today / this week / this month / in total, and the
key's spending limit and remaining credit. If a sign-in is in progress, shows its
status and link. Use it after auth_login to confirm sign-in finished, or when the
user asks about cost, usage or credits. Takes no parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it largely does: it enumerates what is reported (key label, usage today/week/month/total, spending limit, remaining credit) and handles the in-progress sign-in case with a status and link. It stops short of stating prerequisites (e.g. behavior when not signed in) or read-only guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose followed by returned data, usage triggers, and the no-parameter note. Every sentence adds something, though the return-value inventory is somewhat list-heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must explain returns, which it does in detail, and it supplies usage triggers. It is close to complete for a zero-parameter status tool, missing only edge-case behavior such as the not-signed-in response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which sets a baseline of 4, and the description usefully confirms this with 'Takes no parameters.' There is nothing further to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show whether this computer is signed in to OpenRouter, with the key's spending') and is clearly distinguishable from siblings like auth_login/auth_logout, which are actions on the same auth state. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers: 'after auth_login to confirm sign-in finished, or when the user asks about cost, usage or credits.' It names the sibling it follows but does not state when *not* to use it or contrast it against any alternative status source, so it falls short of the 5 bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_loginA
Sign in to OpenRouter in the user's browser.
Opens the OpenRouter sign-in page and returns at once with a link. Show the link to
the user in case the browser didn't open (it contains no secret). The user finishes
in the browser; then call account_status to confirm. The new key is stored in the
OS credential store and is never shown. Does nothing if already signed in, unless
switch_account is true.
Args: switch_account: true to sign in as a different account or workspace. The current key is kept until the new one is stored.
| Name | Required | Description | Default |
|---|---|---|---|
| switch_account | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so: it returns at once with a link, the link contains no secret, the new key goes to the OS credential store and is never shown, and it is idempotent unless switch_account is true. It even discloses that the current key is retained during a switch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and the link-return contract, then the follow-up step, then the parameter note. Slightly verbose in places (the secret/link caveat), but nearly every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description explains what the tool returns (an immediate link) plus the follow-up confirmation flow. An agent has everything needed to invoke it correctly and handle the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description fully compensates for the single parameter: switch_account signs in as a different account or workspace, and the current key is kept until the new one is stored. This is meaningful semantics beyond the schema's bare boolean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (sign in to OpenRouter in the user's browser) and describes the actual mechanism (opens the sign-in page, returns a link). It also distinguishes itself from siblings by naming account_status as the confirmation step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance: finish in the browser, then call account_status to confirm; show the link if the browser didn't open; no-op if already signed in unless switch_account is true. The alternative action (auth_logout) is implicitly separated by the sign-in framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_logoutA
Sign out: delete the stored OpenRouter key from this computer.
The key stays valid on OpenRouter until it is deleted on the keys dashboard; the result links that page. Takes no parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers a non-obvious behavioral fact: the local key deletion does NOT invalidate the key on OpenRouter, which must be deleted on the dashboard. It also notes the result links that page. It stops short of describing any failure modes or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action verb and core effect, followed by the important caveat about remote key validity. Every sentence adds value; the 'Takes no parameters' line is mildly redundant given the schema but harmless.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter logout tool, the description covers what happens locally, what does not happen remotely, and what the result contains. No output schema exists, yet it adequately signals the return. Complete enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description reinforces this with 'Takes no parameters,' which is consistent with the empty schema and leaves no ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Sign out') and the concrete action ('delete the stored OpenRouter key from this computer'), making the resource and effect unambiguous. It is clearly distinguishable from its counterpart auth_login and the other image-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intent (sign out) is implied clearly enough that an agent will know when to call it, but there is no explicit when-to-use vs when-not guidance or reference to alternatives such as auth_login. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageA
Edit an image (or create one from reference images) with an OpenRouter image model.
Call list_image_models first to choose a model, picking one whose inputs column
says "text+image". Options are checked against the model before anything is sent.
The input images are uploaded to a third-party model provider through OpenRouter: get the user's OK before sending client or project images to third-party providers. Each call costs money on the user's OpenRouter account; the result reports the cost.
All paths must be absolute (or start with "~"); relative paths are rejected. Results are saved next to the first input image (never overwriting it) with a JSON sidecar, unless output_dir is given. The result lists the saved paths, notes, failures, model, provider, time and cost, plus JPEG previews.
Args:
prompt: The change to make, e.g. "make it dusk with warm window light".
model: Model id from list_image_models.
images: Absolute paths of the input images. The first is the image being
edited; any others are references.
mask_path: Optional absolute path of a mask for the first image: white = change,
black = keep. The edit is composited locally, so only the white area
changes, but there may be a visible lighting seam at the mask edge where
new and old pixels meet; a soft mask edge or mask_feather_px helps.
Requires fit="preserve". A masked edit also saves the model's image
before the blend as <result>.unmasked.png, so remask_image can redo
the blend with a different mask for free.
mask_feather_px: Blur radius in pixels for the mask edge; by default it is
chosen from the image size.
fit: "preserve" (default) returns each result at exactly the first input's
pixel size: the server picks the model's closest aspect ratio, then scales
and crops back (padding the input first if no ratio is close). With
fit="preserve" don't pass aspect_ratio or size. "model" keeps the size the
model returns and lets you choose aspect_ratio.
n: Number of images, 1 to 10. Above the model's max n, the server splits them
into several calls (each billed) and says so in the notes.
aspect_ratio: One of the model's aspect ratios (only with fit="model").
resolution: One of the model's resolution tiers.
size: Exact pixel size such as "1024x1024" (only with fit="model").
quality: One of the model's quality values.
seed: Integer for repeatable results, for models that support it.
background: One of the model's background values, e.g. "transparent".
output_format: One of the model's output formats, e.g. "png".
output_dir: Folder to save into instead of next to the first input. Must be
an absolute path or start with "~".
filename_prefix: File name stem; defaults to the first input's name.
provider_options: Provider passthrough options as a flat dict, e.g.
{"moderation": "low"}. Keys must be in the model's passthrough parameters
(see get_image_model); the server sends each key to every provider that
allows it. Any other key is rejected before anything is spent.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| fit | No | preserve | |
| seed | No | ||
| size | No | ||
| model | Yes | ||
| images | Yes | ||
| prompt | Yes | ||
| quality | No | ||
| mask_path | No | ||
| background | No | ||
| output_dir | No | ||
| resolution | No | ||
| aspect_ratio | No | ||
| output_format | No | ||
| filename_prefix | No | ||
| mask_feather_px | No | ||
| provider_options | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: third-party upload and consent requirement, per-call monetary cost billed to the user's OpenRouter account, absolute-path enforcement with relative paths rejected, non-overwriting save behavior plus JSON sidecar, and the n-above-max splitting into multiple billed calls. It even discloses an edge-case artifact (visible lighting seam at the mask edge) and how to mitigate it, plus the <result>.unmasked.png artifact saved for remask_image.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, then a prerequisites/behavior paragraph, then a clean per-argument Args block that mirrors the parameter order. It is long, but the length is largely justified by 17 undocumented parameters; a few behavioral details (mask seam, previews) could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating the returned payload (saved paths, notes, failures, model, provider, time, cost, JPEG previews) and where files land. Given 17 params, cost implications, and a third-party data flow, nothing an agent needs to call this correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 17 parameters, so the description must compensate and it does so for every argument: prompt example, images ordering semantics (first = edited, rest = references), mask white/black convention, mask_feather_px default derivation, fit modes and their interactions, n range and billing split, provider_options key validation. This adds substantial meaning beyond the bare type/default schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (edit an image, or create one from reference images) with the engine named (OpenRouter image model). It differentiates itself from siblings by pointing to list_image_models for model choice, get_image_model for passthrough params, and remask_image for re-blending, so an agent can separate it from generate_image without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit prerequisites (call list_image_models first, pick a model whose inputs column says text+image), explicit consent rule (get the user's OK before sending client/project images to third-party providers), and explicit when-nots (don't pass aspect_ratio or size with fit="preserve"; mask_path requires fit="preserve"). It also routes to remask_image for free re-blends, covering alternatives rather than leaving them to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate new images from a text prompt with an OpenRouter image model.
Call list_image_models first to choose a model, and get_image_model for its exact
allowed values. Options are checked against the model before anything is sent: an
invalid value fails with the allowed values and costs nothing.
Each call costs money on the user's OpenRouter account; the result reports the cost. Prompts go to a third-party model provider through OpenRouter: get the user's OK before sending client or project images or details to third-party providers.
Images are saved (never overwriting) to the configured output folder with a JSON sidecar of the settings used. The result lists the saved paths, any notes and failures, the model, provider, time and cost, plus JPEG previews.
Args:
prompt: What to create.
model: Model id from list_image_models.
n: Number of images, 1 to 10. Above the model's max n, the server splits them
into several calls (each billed) and says so in the notes.
aspect_ratio: One of the model's aspect ratios, e.g. "16:9".
resolution: One of the model's resolution tiers.
size: Exact pixel size such as "1024x1024", for models that support it.
quality: One of the model's quality values.
seed: Integer for repeatable results, for models that support it.
background: One of the model's background values, e.g. "transparent".
output_format: One of the model's output formats, e.g. "png" or "svg".
output_dir: Folder to save into instead of the default. Must be an absolute
path or start with "~".
filename_prefix: File name stem; defaults to a slug of the prompt.
provider_options: Provider passthrough options as a flat dict, e.g.
{"moderation": "low"}. Keys must be in the model's passthrough parameters
(see get_image_model); the server sends each key to every provider that
allows it. Any other key is rejected before anything is spent.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| seed | No | ||
| size | No | ||
| model | Yes | ||
| prompt | Yes | ||
| quality | No | ||
| background | No | ||
| output_dir | No | ||
| resolution | No | ||
| aspect_ratio | No | ||
| output_format | No | ||
| filename_prefix | No | ||
| provider_options | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that validation happens before any spend ('an invalid value fails with the allowed values and costs nothing'), that each call is billed to the user's account, that n above the model max splits into multiple billed calls, and that files are saved never overwriting plus a JSON sidecar. It also describes the return payload in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose, the pre-call workflow and the cost/consent warnings are front-loaded, which is the right ordering. The Args block is long, but each line carries non-redundant semantics for a 13-parameter tool with zero schema coverage; only slight tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description enumerates the result contents (saved paths, notes and failures, model, provider, time, cost, JPEG previews), covers the cost/consent/privacy dimensions, and routes to the two prerequisite lookup tools. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate entirely, and it documents all 13 parameters with meaning beyond the schema: n's 1-10 range and server-side splitting behavior, output_dir's absolute-or-tilde path rule, filename_prefix defaulting to a prompt slug, and provider_options' passthrough validation rules. This is far more than the bare schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generate new images from a text prompt') plus the backend ('OpenRouter image model'), which cleanly distinguishes it from siblings edit_image, list_image_models and get_image_model. An agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the workflow ('Call list_image_models first... and get_image_model for its exact allowed values') and the human-in-the-loop condition, plus a concrete when-not ('get the user's OK before sending client or project images or details to third-party providers'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_image_modelA
Show full details for one image model.
Returns the description, the capabilities, and one section per provider endpoint
with its supported parameters and allowed values, the passthrough parameters usable
in provider_options, and its pricing line items.
Args:
model_id: Exact model id from list_image_models.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the return content in detail, but says nothing about permissions, authentication requirements, rate limits, or whether missing ids raise an error versus returning empty. For a no-annotation tool that is a real, if modest, gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a scannable enumeration of the payload and a short Args block. No filler, though the multi-line return breakdown is slightly longer than needed for a one-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing the returned sections is exactly the right compensation and is done well. The main omission is any behavioral/prerequisite context (auth, error behavior), which matters more given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single param is undocumented in the schema, so the description must compensate. It does: it defines model_id as the exact id and tells the agent where to obtain it (list_image_models), which is meaning beyond the bare 'Model Id' title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show full details for one image model') and enumerates what comes back (description, capabilities, provider endpoint sections, passthrough params, pricing). This clearly separates it from the sibling list_image_models, which returns a collection rather than a single model's full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the model_id arg is described as 'Exact model id from `list_image_models`', which hints you should list first. There is no explicit statement of when to prefer this over list_image_models or any exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_image_modelsA
List the image models available on OpenRouter right now, as a markdown table.
Use this before generate_image or edit_image to choose a model and see which
options it takes. The catalog is live (cached for 10 minutes), so don't rely on
remembered model ids. Columns: id, name, inputs ("text", or "text+image (max N)"
input images), aspect ratios, resolutions, quality values, max n per call, seed
support, whether the provider moderates prompts (yes/no/?), and a price summary
($/img-unit per image-output unit, $/image, or $/tok per completion token).
Call get_image_model for one model's per-provider details.
Args:
query: Case-insensitive text to find in the model id or name.
accepts_images: true for models that take input images (needed by
edit_image); false for text-only models.
author: Only this author's models: the part of the id before "/".
sort: "newest" (default, newest first) or "price" (cheapest first).
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | newest | |
| query | No | ||
| author | No | ||
| accepts_images | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and mostly does: it discloses that the catalog is live, cached for 10 minutes, and that moderation status may be unknown ('?'). It does not mention auth requirements or rate limits, but for a read-only listing that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then usage, then a column inventory, then args. The column enumeration is lengthy but earns its place because there is no output schema; slightly over-detailed for a listing tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description supplies the return shape by naming every column and its semantics (inputs, aspect ratios, resolutions, quality, max n, seed, moderation, price units). Combined with the cache/liveness note and full param docs, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it does, defining all four args: case-insensitive query scope, accepts_images tied to edit_image needs, author as the id prefix before '/', and both sort values with their defaults and ordering.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (image models on OpenRouter), plus the output shape (markdown table) and liveness of the catalog. It is clearly separable from sibling get_image_model, which is explicitly scoped to 'one model's per-provider details.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule: 'Use this before generate_image or edit_image to choose a model and see which options it takes.' It also routes the agent elsewhere when only one model is needed ('Call get_image_model for one model's per-provider details'), and warns against stale remembered model ids.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remask_imageA
Redo the mask blend of a masked edit_image result with a new mask: free and offline.
Use it when a masked edit came back with a good change that blends badly, such as
a seam or a ghosted edge where the model drew past the mask (a limb or shadow cut
off at the mask edge), or when the mask should have been bigger or smaller. Draw a
new mask and call this instead of paying for another edit_image: it re-blends the
model's image that the edit already returned, on this computer. Nothing is
uploaded, no model is called and it costs nothing.
It needs a result made by edit_image with mask_path by openrouter-image-mcp
0.2.0 or later: those keep the model's image before the blend as
<result>.unmasked.png, next to the result and its .json sidecar. The original
input image must still be in place and unchanged. A re-masked result can itself
be re-masked.
All paths must be absolute (or start with "~"); relative paths are rejected. The
new image is saved as a PNG next to image (never overwriting) with a JSON
sidecar, unless output_dir is given. The result lists the saved path, plus a JPEG
preview.
Args:
image: Absolute path of a masked edit_image result (or of an earlier
remask_image result).
mask_path: Absolute path of the new mask for the original input: white = take
the model's image, black = keep the original. Any size; it is stretched
to the input's size.
mask_feather_px: Blur radius in pixels for the mask edge; by default it is
chosen from the image size.
output_dir: Folder to save into instead of next to image. Must be an
absolute path or start with "~".
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | ||
| mask_path | Yes | ||
| output_dir | No | ||
| mask_feather_px | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so: it states that nothing is uploaded, no model is called, no cost is incurred, and the work happens on this computer. It also discloses the hard prerequisites (requires an `edit_image` result made with `mask_path` on openrouter-image-mcp 0.2.0+, the sibling `<result>.unmasked.png`, and the original input still in place/unchanged), and the save behavior (PNG next to `image`, never overwriting, JSON sidecar, plus a JPEG preview in the result).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the routing decision, then prerequisites, then args; the structure is easy to scan. It is somewhat longer than necessary, repeating the free/offline/no-cost point in three separate places ("free and offline", "it costs nothing", "no model is called"), which is mild redundancy rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no annotations, no output schema, and 0% schema coverage, everything an agent needs is present: when to call it, prerequisites that will make it fail, path-format rules, side effects, and the shape of the returned result (saved path + JPEG preview).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all four parameters — and it does. It defines `image` (masked `edit_image` or earlier `remask_image` result), the white/black semantics of `mask_path` plus that any size is stretched to the input, the default derivation of `mask_feather_px` from image size, and the `output_dir` override — all beyond the bare titles in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: re-doing the mask blend of a masked `edit_image` result with a new mask, locally and for free. It explicitly distinguishes itself from the sibling `edit_image` ("call this instead of paying for another `edit_image`"), so an agent can route between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions (a masked edit that blends badly: a seam, a ghosted edge, a limb/shadow cut off at the mask edge, or a mask that should have been bigger/smaller) and names the alternative plus the reason to prefer this one (free, offline vs. paying for another `edit_image`). It also states prerequisites and that a re-masked result can itself be re-masked.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
account_status - First observed
auth_login - First observed
auth_logout - First observed
edit_image - First observed
generate_image - First observed
get_image_model - First observed
list_image_models - First observed
remask_image
TDQS
Scored across 8 tools
Each tool targets a distinct action or resource: authentication (auth_login, auth_logout, account_status), model discovery (list_image_models, get_image_model), and image operations (generate_image, edit_image, remask_image). The only pair with potential overlap is list_image_models vs get_image_model, but they are clearly differentiated as list vs detail. No tool boundaries are confusing.
All tool names follow a consistent verb_noun pattern: auth_login, list_image_models, get_image_model, generate_image, auth_logout, account_status, edit_image, remask_image. The prefix 'auth_' and 'image' nouns are used predictably.
Eight tools are well-scoped for an image generation and editing server with authentication. Each tool has a clear purpose and no redundancy; the set covers the full workflow from sign-in to model discovery to image generation, editing, and local re-masking.
The surface covers the complete lifecycle: authentication (login, logout, status), model discovery (list, get), image generation, editing with masking, and free local re-masking. It even anticipates advanced use cases like mask adjustments without extra cost. No obvious gaps for the stated domain.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
Image and video AI tools and your own pipelines, run from any AI assistant.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables image generation via OpenRouter API, supporting models like Gemini 2.5 Flash Image Preview with options to save files locally.21Do What The F*ck You Want To Public
- AlicenseAqualityCmaintenanceGives AI assistants image generation and editing capabilities through OpenRouter, supporting multiple models, style presets, variations, and batch operations.553 npm2MIT
- AlicenseNot gradedqualityDmaintenanceProvides image generation tools powered by OpenRouter, running on Cloudflare Workers.26 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables image and video generation from text, image editing, and text-to-video workflows using AI models via OpenRouter and fal.ai.MIT