puterMCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@puterMCPGenerate a fantasy landscape with dragons using Flux.1"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
puterMCP ⚠️ ARCHIVED
Status: This project is no longer actively maintained.
Reason: Puter's API is designed for browser-based usage via Puter.js. Server-side direct API calls hit rate limits and cannot be used reliably. The official Puter.js library requires Node.js 24+, which is not widely available.
Historical: A local MCP server that bridged LLM environments with Puter's free AI & Cloud services
puterMCP is a TypeScript/Node.js Model Context Protocol (MCP) server that runs locally via npx and acts as a bridge between any MCP-compatible LLM environment (Claude Desktop, Kilo Code, Trae, Cursor, Windsurf, etc.) and Puter's free, unlimited AI and Cloud APIs.
The first capability shipped is image generation across 30+ models (GPT Image, DALL-E, Gemini Nano Banana, Flux, Stable Diffusion, and more) — all without API keys or per-request costs.
Features
Zero Friction: Install and run with a single
npxcommand.Free Image Generation: Access 30+ models including DALL-E 3, Flux.1, and Stable Diffusion via Puter's free tier.
Secure Authentication: Uses your personal Puter account token, stored locally and securely.
Universal Compatibility: Works with Claude Desktop, Cursor, Trae, and any other MCP client.
Inline Image Generation: Images are returned directly in the chat interface, ready for preview and download.
Smart Fallback: Automatically tries free models (like Flux) if premium models (like DALL-E 3) fail due to quota limits.
Related MCP server: jgkme/kilo-image-gen-mcp
Prerequisites
Installation & Setup
1. Authenticate with Puter
You need to provide your Puter authentication token to the MCP server. This is a one-time setup.
Log in to puter.com.
Open the browser Developer Tools (F12 or Cmd+Option+I) -> Console.
Type
puter.authTokenand press Enter.Copy the string (without quotes).
Run the following command in your terminal:
npx puter-mcp --token <your-token-here>Your token will be securely stored in ~/.puter-mcp/config.json.
2. Configure Your MCP Client
Claude Desktop
Add the following to your claude_desktop_config.json:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"puter": {
"command": "npx",
"args": ["-y", "puter-mcp"]
}
}
}Trae / Cursor / Kilo Code
Add the configuration to your project's MCP settings (e.g., .kilo/mcp.json or via the IDE settings UI):
{
"mcpServers": {
"puter": {
"command": "npx",
"args": ["-y", "puter-mcp"]
}
}
}Usage
Once configured, restart your LLM environment. You can now ask it to generate images:
"Generate a cyberpunk city at night using DALL-E 3"
"Create a logo for a coffee shop using Flux.1 Schnell"
"Show me what models are available"
Available Tools
generate_image: Generate an image from a text prompt.prompt: Description of the image.model: (Optional) Model ID (default:dall-e-3).quality: (Optional) Quality setting (e.g.,hd,standard).
list_models: List all available image generation models.category: (Optional) Filter by category (all,openai,google,flux,stable-diffusion,other).
Development
Clone the repository:
git clone https://github.com/yourusername/puter-mcp.git cd puter-mcpInstall dependencies:
npm installBuild the project:
npm run buildRun locally:
node bin/puter-mcp.mjs
License
MIT
Alternatives
If you need free image generation in your MCP/AI workflows, consider:
OpenRouter - Free tier available with various image models
Together.ai - Free tier with Flux models (10 req/min)
Direct API keys - Use OpenAI, Google, or Anthropic APIs with your own keys
For browser-based applications, the official Puter.js library works well and supports image generation directly from frontend code.
Available Tools
2 toolsgenerate_imageA
Generate an image from a text prompt using Puter's free AI image generation. Supports 30+ models including GPT Image, DALL-E 2/3, Gemini Nano Banana, Flux.1, Stable Diffusion, and more. Falls back across models on quota or availability errors. Returns the image inline as base64 content for direct rendering in MCP clients. No API keys required — uses your Puter account.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Image generation model to use. Options include: "dall-e-3" (default), "gpt-image-1", "gpt-image-1-mini", "gemini-2.5-flash-image-preview" (Nano Banana), "gemini-3-pro-image-preview" (Nano Banana Pro), "gemini-3.1-flash-image-preview" (Nano Banana 2), "black-forest-labs/FLUX.1-schnell", "stabilityai/stable-diffusion-3-medium", and many more. Use list_models tool to see all options. | dall-e-3 |
| prompt | Yes | Text description of the image to generate. Be detailed and specific for best results. | |
| quality | No | Quality setting. For GPT Image: "high", "medium", or "low". For DALL-E 3: "hd" or "standard". Not all models support this. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses fallback on quota/availability errors, return format as base64 for direct rendering, and no API key requirement. It could mention failure behaviour if all models fail, but overall transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each adding value: purpose, features/fallback, return format and auth. No redundancy, front-loaded with core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description explains return format adequately. It covers purpose, models, fallback, auth. Slight gap: no explicit mention of output structure beyond base64, but sufficient for agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% but description adds value by advising 'Be detailed and specific' for prompt, explaining different quality options per model, and recommending list_models for model options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates an image from text prompt, lists supported models, and distinguishes from sibling 'list_models' which lists models instead of generating images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (generation) vs alternative 'list_models' (discovery). It references list_models in the schema for model selection, providing cross-reference. No explicit exclusions but clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List all available image generation models supported by puterMCP via Puter.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Filter models by category. | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It only states what the tool does without disclosing behavior such as pagination, data freshness, rate limits, or any side effects. Minimal behavioral context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded and concise. It contains no wasted words, though it could include slightly more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (list with one optional parameter) and the presence of a sibling tool, the description is fairly complete. It does not explain the output format, but that is often self-evident for a list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with a description for the 'category' parameter. The tool description does not add any additional meaning beyond the schema, meeting the baseline but not exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('list') and the resource ('all available image generation models'), and distinguishes from the sibling tool 'generate_image' which performs a different action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing to know available models, but does not explicitly state when to use this tool vs the sibling tool 'generate_image' or any alternatives. No exclusions or context are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.1- First observed
generate_image - First observed
list_models
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one generates images, the other lists models. There is no overlap or confusion.
Both tools follow a consistent verb_noun pattern in snake_case (generate_image, list_models), which is predictable and clear.
With only two tools for an image generation service that supports 30+ models and includes fallback logic, the tool surface feels too sparse. Typically one would expect additional tools for quota management, image retrieval, or cancellation.
The server lacks obvious operations such as checking user quota, retrieving previously generated images, or managing model preferences. The generate_image tool's fallback behavior suggests quota tracking, but no tool exposes that information, creating a gap.
Maintenance
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
MCP server for Flux AI image generation
MCP server for AI dialogue using various LLM models via AceDataCloud
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
Related MCP Servers
- AlicenseAqualityDmaintenanceAn MCP server that enables AI applications to access 20+ model providers (including OpenAI, Anthropic, Google) through a unified interface for text and image generation.230MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for generating, editing, and processing images via multiple providers including Kilo, OpenRouter, OpenAI, and Gemini, with local tools for background removal, resizing, and cropping.19 npm3MIT
- AlicenseAqualityDmaintenanceMCP server for AI image generation supporting multiple providers (OpenRouter, Together AI, Replicate, fal.ai) and compatible with various MCP agents.233 npm1MIT
- AlicenseAqualityDmaintenanceMCP server for generating images using OpenRouter API, supporting models like Gemini 2.5 Flash. Enables image generation with flexible options like saving to local files.2Do What The F*ck You Want To Public