veida-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@veida-mcpGenerate a 16:9 image of a lighthouse on a rocky cliff during a storm."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Veida MCP
An MCP server that actually generates images, with no API key and no account.
Most image MCP servers are a wrapper around a key you have to go and buy first. This one talks to the anonymous tier of Veida — a client id in a header instead of a login — so it works the moment it is installed.
> generate a 16:9 cover image of a paper boat on a calm lake at sunrise
→ https://cdn.veida.ai/uploads/kie/image/….webpTools
Tool | What it does |
| Prompt → hosted image URL. Submits, polls, returns. 40–120 seconds. |
| The model shelf, which ones run signed-out, and the page for each. |
| Pages worth sending a person to — editor, upscaler, background remover. |
| The exact ceilings, so an agent does not build something that always fails. |
Related MCP server: NovelAI MCP Server
Install
Claude Desktop / Claude Code
{
"mcpServers": {
"veida": { "command": "npx", "args": ["-y", "veida-mcp"] }
}
}Cursor / Cline / Windsurf
Same block, in that client's MCP config file. There is nothing else to set — no
env, no key.
What the free tier is, measured
The anonymous grant is 4 credits and one image costs 4 — so the server mints a fresh client id per image rather than reusing one.
A ceiling of 30 credits per IP per day sits on top: roughly 7 images a day from one machine, however many ids it invents.
Free output is 1K and watermarked.
🔴 Editing an existing image cannot run anonymously. The editing model costs 30 credits against a 4-credit grant, so an anonymous edit can never be paid for. The tool set says so rather than offering a call that always fails; for edits, sign in and use the image editor or the ChatGPT image editor.
There is no video on this site, so no video tool.
Signing in removes the watermark, lifts the cap and opens the rest of the shelf. The current lineup is on the models the generator offers: GPT Image 2 when the picture has to contain readable text, Nano Banana 2 for instruction following, Seedream 5.0 Lite for colour and composition, and Qwen Image 3 for typography. Pricing lists what each tier costs.
Writing prompts that work
Name the light, the material and the composition. "A product photo of a mug" gives the model nothing; "matte black ceramic mug on pale oak, soft window light from the left, shallow depth of field" gives it a picture. Worked examples live in the prompt library, and the showcase is there if you want to judge output quality before wiring anything up.
The same engine is what the ChatGPT image generator page serves in a browser — this server is the headless door to it.
Develop
npm install
npm run build
node dist/index.js # speaks MCP over stdioLinks
MIT.
Available Tools
4 toolsbrowse_toolsA
List the Veida pages worth sending a person to — generator, editor, upscaler, background remover, prompt library, pricing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. 'List' implicitly suggests a read-only, non-destructive operation, but the description does not explicitly state side-effect-free behavior, whether results are live or static, or if any authentication is needed. For a simple listing tool, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with the core action front-loaded and the specific page list following. Every element adds useful information, and there is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description gives enough context: it names the exact pages the caller should expect. It could be more complete by noting whether the output is page names, URLs, or both, but the enumeration makes the tool usable without further detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema already fully defines the invocation surface. The description correctly focuses on what the tool returns rather than parameter details, matching the baseline for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a clear resource ('Veida pages worth sending a person to'), then enumerates the exact pages included. It distinguishes itself from the sibling tools, which perform actions like image generation or model listing, by framing this as a navigation/discovery tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'worth sending a person to' implies the tool is for finding pages to direct users to, which is a helpful context clue. However, it never explicitly states when to use this tool over siblings like generate_image, list_models, or free_tier_limits, nor does it give exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
free_tier_limitsA
Explain exactly what the anonymous free tier allows, so the caller does not build something that always fails.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It clearly signals a read-only informational tool via 'Explain', with no side effects implied. It does not detail the response format or whether authentication is needed, but for a zero-parameter informational tool the core behavior is transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one efficient sentence that front-loads the action and object, then adds the practical rationale. No words are wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity—zero parameters, no output schema, and no complex behavior—the description covers purpose and usage context well. It could mention what the returned explanation looks like, but the core need is addressed sufficiently for a no-argument explainer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. There are no parameter semantics for the description to clarify, and no schema details to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific action, 'Explain', with a clear resource, 'what the anonymous free tier allows'. It is clearly distinct from siblings like generate_image and list_models, so there is no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: call this before building against the anonymous free tier to avoid constructing something that will always fail. It does not explicitly name alternatives or when-not-to-use conditions, but no sibling fills this role, so the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image from a text prompt and return a hosted image URL. No API key or account is needed — this runs on the anonymous free tier. Output is 1K and watermarked. Takes 40-120 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | What to draw. Naming the light, the material and the composition moves the result far more than adding adjectives. | |
| aspectRatio | No | 1:1 | |
| timeoutSecs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It does well by revealing the anonymous free tier, the 1K watermarked output, and the 40-120 second latency. It could additionally mention rate limits or retention behavior, but the disclosed constraints materially affect invocation expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, each carrying distinct information: function and output, access constraints, and output/latency expectations. The most important functional statement is front-loaded, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a relatively simple three-parameter tool, the description covers the essential operational context: what it returns, authentication requirements, output properties, and latency. It does not mention error scenarios or whether the hosted URL is temporary, but the presence of a timeoutSecs parameter and the sibling free_tier_limits tool suggest those gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only prompt is described). The description adds little beyond 'text prompt' for parameter meaning, though the prompt schema itself gives useful prompt-craft guidance. aspectRatio is self-explanatory via its enum, and timeoutSecs has default/min/max in the schema, so the description need not repeat them but also does not compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Generate an image from a text prompt and return a hosted image URL.' It clearly states the tool's function and output format, and its scope is unambiguous even against the sibling tool names (list_models, browse_tools, free_tier_limits).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when an image needs to be generated from a prompt, with no API key or account needed because it runs on the anonymous free tier. It does not explicitly name alternatives or exclusion criteria, but the context is sufficient for a simple generation task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List the image models Veida serves, which of them a signed-out caller can run, and the page a human can run each one on.
| Name | Required | Description | Default |
|---|---|---|---|
| scene | No | any |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden and does a solid job: it discloses the tool's return content (models, signed-out accessibility, human-facing page). It does not mention output format, ordering, or error behavior, but for a read-only listing tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the main action and lists the key distinctions in parallel structure. Every part adds value, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description captures the core purpose and return value well, but it omits the 'scene' parameter's effect and offers no guidance on how to combine this tool with siblings like free_tier_limits. For a simple read-only tool with one optional parameter, this is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, yet the description never mentions the 'scene' parameter or how it affects the results. The enum values (text-to-image, image-to-image, any) and default are visible, but the description adds no semantic meaning to help the agent decide how to pass this parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'List the image models Veida serves.' It also clarifies the output dimensions (which models are available to signed-out callers and which page each runs on), which distinguishes it from siblings like generate_image and browse_tools without needing to read the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool is appropriate: when a caller needs to know available models, access restrictions for signed-out callers, or the page for each model. It does not explicitly name alternatives or when not to use it, but the context is clear enough for an agent to select this tool over the listed siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
browse_tools - First observed
free_tier_limits - First observed
generate_image - First observed
list_models
TDQS
Scored across 4 tools
generate_image is the only action tool, while list_models, browse_tools, and free_tier_limits are distinct informational resources. There is minor overlap between list_models and browse_tools in that both mention pages a human can visit, but their primary intents differ enough.
Three tools use a clear verb_noun pattern (generate_image, list_models, browse_tools), but free_tier_limits is a noun phrase rather than get_free_tier_limits or explain_free_tier_limits, creating a small inconsistency.
Four tools is a tight, appropriate scope for an anonymous free-tier image generation service. Each tool has a clear purpose and none feel redundant.
The set covers the main programmatic action (generate_image) plus the essential supporting information a caller needs: available models, human pages, and free-tier limits. A model-selection parameter on generate_image or a separate generate_with_model tool would close the most obvious minor gap.
Maintenance
Related MCP Connectors
Image and video AI tools and your own pipelines, run from any AI assistant.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI image generation from text prompts using Pollinations AI. Supports multiple models, customizable dimensions, and direct image saving with URL generation capabilities.88 npm4MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to generate images using NovelAI, supporting text-to-image, image-to-image, and tag suggestions.3MIT
- AlicenseDqualityDmaintenanceEnables image generation, editing, and description using Ideogram AI's models through natural language.453 npm4MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI image generation and editing using Replicate's official models like Flux, SDXL, and Seedream, with tools to search models and generate images.2,788 npm6Apache 2.0