Generate images
images_generateGenerate an image from a text prompt and return its URL
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | What the image should depict |
images_generateGenerate an image from a text prompt and return its URL
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | What the image should depict |
Changes observed during successful MCP inspections.
Input schema / properties / prompt / descriptionPrevious value: -"Describe the image to generate"New value: +"What the image should depict"Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide no hints (only a title), so the description carries the full burden. It discloses the primary action and return, but omits any behavioral context such as side effects (image storage), rate limits, content policies, or whether the URL is ephemeral. This is a significant gap for a tool that generates content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no wasted words. It efficiently captures the input, action, and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool without an output schema, the description is adequate, covering input and return type. It could add more detail on image format or URL stability, but this is not critical for basic use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole parameter `prompt` is described as 'What the image should depict'. The description reinforces 'text prompt' but adds no new meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('generate'), resource ('an image'), and output ('return its URL'), clearly distinguishing it from siblings like images_search. It is unambiguous and directly tied to the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives (e.g., images_search). The generation-vs-search nuance is implied by the verb and sibling context, but not stated. A direct mention of when to choose generation over search would improve it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.