imagengen
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| XAI_API_KEY | No | API key for Grok Image provider (optional) | |
| GEMINI_API_KEY | No | API key for Gemini provider (optional, used if you want to use Gemini image generation) | |
| OPENAI_API_KEY | No | API key for GPT-image provider (optional) | |
| IMAGE_OUTPUT_DIR | No | Directory where generated/edited images are saved (default: ./output) | ./output |
| IMAGE_PROVIDER_DEFAULT | No | Default provider to use when multiple API keys are set. One of: gemini, grok, gpt-image |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_image_providersA | Lists which image providers (gemini, grok, gpt-image, impossibl) are configured via API key, their available models (discovered live from each provider), and which provider/model would be used by default. |
| text-to-imageA | Generates an image from a text prompt using Gemini, Grok Image, GPT-image, or impossibl.com, and saves it to disk. |
| image-to-imageA | Edits or transforms one or more input images according to a text prompt, using Gemini, Grok Image, or GPT-image, and saves the result to disk. (impossibl.com does not support this tool.) |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The three tools are clearly distinct: one lists available providers, one does text-to-image generation, and one does image-to-image editing/transformation. The descriptions make it obvious which tool to use for each task, and the coverage of providers in each is well documented.
The naming pattern is inconsistent: 'list_image_providers' uses snake_case with a verb_noun pattern, while 'text-to-image' and 'image-to-image' use a hyphenated adjective-noun naming convention that describes the task rather than an action. The latter two don't follow a verb-leading pattern.
Three tools is a compact, well-scoped set that covers the core capabilities of an image generation server: discovery/configuration, generation, and editing. It's on the lean side but appropriate for the narrow domain; each tool serves a distinct and necessary purpose.
The server covers the primary workflows (discover providers, generate, edit). However, there are notable gaps: no tool to view/retrieve generated images or their metadata, no batch generation, no image style/variation features, and no way to delete or manage saved images. The core generate/edit flows work but retrieval and management are missing.