OpenRouter Image Generation MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@OpenRouter Image Generation MCP Servergenerate an image of a sunset over the mountains"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
OpenRouter Image Generation MCP Server
An MCP (Model Context Protocol) server that provides image generation capabilities through the OpenRouter API, supporting models like Gemini 2.5 Flash Image Preview.
Features
Image Generation: Generate images using Google Gemini 2.5 Flash Image Preview
Flexible Options:
Save generated images to local files
Related MCP server: OpenRouter MCP Multimodal Server
Installation
Clone the repository:
git clone https://github.com/yourusername/openrouter-image-gen-mcp.git
cd openrouter-image-gen-mcpInstall dependencies:
npm installBuild the TypeScript code:
npm run buildSet up your OpenRouter API key:
export OPENROUTER_API_KEY="your-api-key-here"You can get an API key from OpenRouter.
Configuration for Claude Desktop
Add the following to your Claude Desktop configuration file:
macOS/Linux
Location: ~/.config/claude/claude_desktop_config.json
Windows
Location: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"openrouter-image-gen": {
"command": "node",
"args": ["/path/to/openrouter-image-gen-mcp/dist/index.js"],
"env": {
"OPENROUTER_API_KEY": "your-api-key-here"
}
}
}
}Replace /path/to/openrouter-image-gen-mcp with the actual path to your installation directory.
Available Tools
1. generate_image
Generate images using AI models.
Parameters:
prompt(required): Text description of the image to generatemodel: Model to use (default:google/gemini-2.5-flash-image-preview:free)n: Number of images to generate (1-4, default: 1)size: Image dimensions (default:1024x1024)save_to_file: Save images locally (default: false)filename: Base filename for saved imagesshow_full_response: Include full base64 data in response (default: false, returns concise info only)
Example:
{
"prompt": "A serene Japanese garden with cherry blossoms",
"model": "google/gemini-2.5-flash-image-preview:free",
"save_to_file": true,
"filename": "japanese_garden"
}Note: Gemini image generation works through the chat completions API. The model will generate an image based on your prompt and return it as a URL or base64 data in the response. The size parameter is not used for Gemini models.
2. list_models
List all available image generation models.
Development
Build
npm run buildRun in development mode
npm run devStart the server
npm startAPI Documentation
Troubleshooting
401 Authentication Error
If you get a 401 error, check:
Your API key is correctly set in the environment or Claude Desktop config
The API key starts with
sk-or-(OpenRouter format)The API key is valid and has not expired
You have credits available in your OpenRouter account
Test your API key loading:
node test-api-key.jsCommon Issues
API Key not loading: Make sure the
OPENROUTER_API_KEYis set in your Claude Desktop config'senvsectionModel access denied: Some models require specific permissions or higher tier accounts
Image not generating for Gemini: Gemini uses the chat completions endpoint, not the images endpoint
License
WTFPL - Do What The Fuck You Want To Public License
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Available Tools
2 toolsgenerate_imageA
Generate images using Google Gemini API. Control image style, aspect ratio, and composition through descriptive text in your prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate. Include style details (e.g., "photorealistic", "oil painting"), aspect ratio (e.g., "square image", "landscape"), and composition details directly in the prompt. | |
| save_to_file | No | Save generated image to local file | |
| filename | No | Base filename for saved image (without extension) | |
| show_full_response | No | Show full response including base64 data (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It states that images are generated via Gemini API but does not disclose what the tool returns by default, whether it saves files, whether authentication or costs are involved, or any side effects. The schema hints at save_to_file and show_full_response, but the description itself adds little behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no wasted words. The primary action is front-loaded, and the second sentence gives practical prompt-construction guidance without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters fully documented in the schema, the description covers the core invocation guidance. However, without annotations or an output schema, it leaves gaps around default return behavior, file-saving semantics, and operational prerequisites such as API access.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description reinforces the prompt guidance about style, aspect ratio, and composition, but it does not add meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate images using Google Gemini API.' It clearly differentiates from the sibling tool list_models, which serves a different purpose, and adds useful guidance that style, aspect ratio, and composition can be controlled through the prompt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is for generating images, so an agent can infer when to call it. It does not explicitly discuss when not to use it, but no competing image-generation sibling exists among the listed tools, so the omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsB
Show information about the Gemini image generation model
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Show information' implies a read-only operation, but it does not disclose what data is returned, whether a remote API is called, or whether any rate limits or prerequisites apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler. It loses one point because 'Show information about' is somewhat vague and 'model' is singular, while the tool name list_models suggests listing multiple models.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite being a zero-parameter tool, there is no output schema, no annotations, and no description of what information will be shown or how this tool relates to generate_image. An agent can guess the basic behavior, but important selection and return-value context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero properties and schema coverage is 100%, so there are no parameter semantics to document. The description cannot add meaning to an already complete empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb-resource pair: 'Show information' about 'the Gemini image generation model.' It is distinguishable from the sibling generate_image because one provides model information and the other generates images, though it does not explicitly name or contrast the sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to call this tool instead of generate_image. The only implied context is that listing model information might precede generation, but the description does not state this or any other selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- First observed
generate_image - First observed
list_models
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one generates images and the other lists available models. There is no overlap or ambiguity.
Both tool names follow a consistent verb_noun pattern (generate_image, list_models), making them predictable and easy to understand.
With only two tools, the server feels thin for an image generation service. While the core generation tool is present, the set may lack additional functionality like retrieving generation history or managing images, making it borderline.
The server covers the basic generate operation and model listing, but is missing potential features such as model selection parameter in generation, image deletion, or error handling details. This feels incomplete for a full image generation workflow.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Related MCP Servers
- -licenseAqualityNot gradedmaintenanceEnables generating and editing images using OpenRouter's API with Gemini 2.5 Flash Image model. Supports custom aspect ratios, iterative editing, and reference images for style transfer.623 npm-
- AlicenseNot gradedqualityDmaintenanceEnables chat and image analysis through OpenRouter.ai models. Supports text chat, image generation, and analysis with multiple images and custom questions.787 npmApache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to generate and edit images using multiple models through OpenRouter, with features like style presets, batch operations, and variations.22 npm1MIT
- AlicenseAqualityCmaintenanceGives AI assistants image generation and editing capabilities through OpenRouter, supporting multiple models, style presets, variations, and batch operations.545 npm2MIT