Draw Things MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Draw Things MCP Servergenerate a serene mountain landscape with a lake at sunrise"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-drawthings
An MCP (Model Context Protocol) server for Draw Things - enabling LLMs to generate images locally on Mac using Stable Diffusion and other AI models.
Features
Text-to-Image Generation - Generate images from text prompts using the currently loaded model in Draw Things
Image-to-Image Transformation - Transform existing images using text prompts
Configuration Access - Query the current Draw Things settings and loaded model
Local Processing - All image generation runs locally on your Mac using Apple Silicon (M1/M2/M3/M4)
Related MCP server: Draw Things MCP Server
Prerequisites
macOS with Apple Silicon (M1/M2/M3/M4)
Draw Things app installed
Node.js 18 or later
Setup
1. Enable Draw Things API Server
Open Draw Things
Click the gear icon (⚙️) to open Settings
Enable API Server / HTTP Server
The server runs on port 7860 by default
Verify the server is running:
curl http://localhost:78602. Configure Your MCP Client
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"drawthings": {
"command": "npx",
"args": ["-y", "mcp-drawthings"]
}
}
}Cursor
Add to .cursor/mcp.json in your project or global config:
{
"mcpServers": {
"drawthings": {
"command": "npx",
"args": ["-y", "mcp-drawthings"]
}
}
}3. Restart Your MCP Client
Restart Claude Desktop or Cursor to load the new MCP server.
Available Tools
check_status
Check if the Draw Things API server is running and accessible.
get_config
Get the current Draw Things configuration including the loaded model and settings.
generate_image
Generate an image from a text prompt.
Parameters:
Parameter | Type | Required | Description |
| string | Yes | Text description of the image to generate |
| string | No | Elements to exclude from the generated image |
| number | No | Image width in pixels (default: 512) |
| number | No | Image height in pixels (default: 512) |
| number | No | Number of inference steps (default: 20) |
| number | No | Guidance scale (default: 7.5) |
| number | No | Random seed for reproducibility (-1 for random) |
| string | No | Custom file path to save the image |
Example:
Generate an image of a futuristic city at sunset with flying carstransform_image
Transform an existing image using a text prompt (img2img).
Parameters:
Parameter | Type | Required | Description |
| string | Yes | Text description of the desired transformation |
| string | No* | Path to the source image file |
| string | No* | Base64-encoded source image |
| string | No | Elements to exclude |
| number | No | Transformation strength 0.0-1.0 (default: 0.75) |
| number | No | Number of inference steps (default: 20) |
| number | No | Guidance scale (default: 7.5) |
| number | No | Random seed (-1 for random) |
| string | No | Custom file path to save the result |
*Either image_path or image_base64 must be provided.
Configuration
Environment Variables
Variable | Default | Description |
|
| Draw Things API host |
|
| Draw Things API port |
|
| Directory for generated images |
Architecture
┌─────────────────┐ stdio ┌──────────────────┐ HTTP ┌─────────────┐
│ MCP Client │◄──────────────►│ mcp-drawthings │◄────────────►│ Draw Things │
│ (Claude/Cursor) │ JSON-RPC │ │ localhost │ App │
└─────────────────┘ └──────────────────┘ :7860 └─────────────┘
│
▼
┌──────────────┐
│ File System │
│ (images) │
└──────────────┘Development
# Clone the repository
git clone https://github.com/james-see/mcp-drawthings
cd mcp-drawthings
# Install dependencies
npm install
# Build
npm run build
# Run in development mode
npm run devTroubleshooting
"Cannot connect to Draw Things API"
Make sure Draw Things is running
Check that the API Server is enabled in Draw Things settings
Verify the server is accessible:
curl http://localhost:7860Check if a different port is configured in Draw Things
Images not generating
Make sure a model is loaded in Draw Things
Check Draw Things for any error messages
Try generating an image directly in Draw Things first
Permission errors saving images
Check that the output directory is writable. You can set a custom directory using the DRAWTHINGS_OUTPUT_DIR environment variable.
License
MIT
Related Projects
Draw Things - The AI image generation app for Mac/iOS
Model Context Protocol - The protocol specification
@modelcontextprotocol/sdk - TypeScript SDK for MCP
Available Tools
4 toolscheck_statusA
Check if the Draw Things API server is running and accessible
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool checks both 'running' and 'accessible' status, which is helpful. However, it doesn't describe the return value format, what indicates success or failure, or any side effects, though for a simple health check this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no filler. It communicates the essential purpose immediately and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, minimal risk), the description is nearly complete. It could ideally state what the response looks like, but for a health-check tool the meaning of 'running and accessible' is largely self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema confirms this with 100% coverage. The description adds no parameter details because none are needed. Baseline 4 is appropriate for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific action ('Check if the Draw Things API server is running and accessible') with a clear resource (the API server). This distinguishes it from siblings like get_config, generate_image, and transform_image, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys a clear diagnostic/health-check purpose, which implies it should be used to verify server availability before other API calls. It doesn't explicitly mention when not to use it or name alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image from a text prompt using the Draw Things app. The image will be saved to disk and the file path returned.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate | |
| negative_prompt | No | Elements to exclude from the generated image | |
| width | No | Width of the generated image in pixels (default: 512) | |
| height | No | Height of the generated image in pixels (default: 512) | |
| steps | No | Number of inference steps (default: 20) | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| seed | No | Random seed for reproducibility (-1 for random) | |
| model | No | Model filename to use for generation (use list_models to see available models) | |
| output_path | No | Custom file path to save the generated image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does state the main side effect and return behavior: an image is generated, saved to disk, and its path is returned. It does not discuss overwrite behavior, blocking behavior, or app availability, but the most important behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the primary action and immediately state the side effect and return value. There is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich, fully documented input schema and no output schema, the description appropriately explains the high-level operation, the external app dependency, and what the caller receives. Minor gaps such as execution time or overwrite semantics are not critical for correct invocation because the schema covers the parameter details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters in adequate detail. The description adds no parameter-specific meaning beyond implicitly referencing the prompt and the resulting file path, which meets the baseline but does not go beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Generate an image from a text prompt using the Draw Things app.' It also declares the key output behavior—saved to disk with the file path returned—which clearly distinguishes it from siblings like check_status, get_config, and transform_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied clearly: use this when you need a new image created from a text prompt. However, the description gives no explicit guidance about when to prefer this tool over transform_image or any other sibling, and it mentions no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_configA
Get the current Draw Things configuration including the loaded model and settings
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Get' and the resource 'configuration' imply a non-mutating read, and 'current' conveys that it reflects live state. However, it does not explicitly state that there are no side effects, nor does it describe output shape or error behavior, which would strengthen transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler. The core action and object are front-loaded, and the specific contents of the configuration are appended compactly. Every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description provides sufficient context: it states what is returned (current configuration) and what it includes (loaded model and settings). It could add a note about the exact output format or that calling it has no side effects, but those are not strictly necessary for an agent to invoke such a simple tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema declares zero parameters with 100% coverage, so there are no parameter semantics for the description to clarify. The baseline for a zero-parameter tool is 4, and the description appropriately avoids inventing parameter-related details that do not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a read operation targeting the current Draw Things configuration and names concrete contents ('loaded model and settings'). It is distinguishable from generate_image and transform_image as those are action-oriented, and from check_status as that targets status rather than configuration. However, it never explicitly contrasts itself with the check_status sibling, so differentiation is implicit rather than direct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus check_status, generate_image, or transform_image. No conditions, exclusions, or alternative routing hints are provided. An agent must infer usage purely from the tool name and resource description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_imageA
Transform an existing image using a text prompt (img2img). Either image_path or image_base64 must be provided.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the desired transformation | |
| image_path | No | Path to the source image file to transform | |
| image_base64 | No | Base64-encoded source image (alternative to image_path) | |
| negative_prompt | No | Elements to exclude from the transformed image | |
| denoising_strength | No | Strength of the transformation (0.0-1.0, default: 0.75). Lower values keep more of the original image. | |
| steps | No | Number of inference steps (default: 20) | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| seed | No | Random seed for reproducibility (-1 for random) | |
| output_path | No | Custom file path to save the transformed image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not mention whether the original image is preserved, whether the tool writes to output_path by default, what side effects occur, or what the return value looks like. For a transformation tool with no annotations, this is a meaningful transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One focused sentence communicates the operation, the input requirement, and the key constraint. Every word earns its place, and the critical either/or input requirement is stated up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema covers all parameter details, so the description does not need to repeat them. However, with no annotations and no output schema, the description should provide more context about output behavior, side effects, or when to choose this tool over generate_image. It is adequate but not fully complete for a 9-parameter tool with no behavioral metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all 9 parameters with 100% coverage, establishing a baseline of 3. The description adds value by explicitly calling out that either image_path or image_base64 must be provided, which is not captured by the schema's required list (only prompt is marked required). This helps agents avoid invalid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Transform'), a specific resource ('an existing image'), and the method ('using a text prompt (img2img)'). This clearly distinguishes it from the sibling generate_image, which is for creating new images rather than modifying existing ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the core use case clear: transform an already-existing image rather than generate a new one. However, it does not explicitly state 'use generate_image for new images' or provide explicit exclusion criteria, so it stops short of fully explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
check_status - First observed
generate_image - First observed
get_config - First observed
transform_image
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: check_status verifies server availability, generate_image creates new images from text, get_config retrieves settings, and transform_image modifies existing images. There is no overlap or ambiguity between these functions, making tool selection straightforward for an agent.
All tool names follow a consistent verb_noun pattern with snake_case formatting: check_status, generate_image, get_config, and transform_image. This uniformity enhances readability and predictability across the toolset.
With 4 tools, this server is well-scoped for its purpose of interacting with the Draw Things app. Each tool serves a specific, essential function—server status, image generation, configuration retrieval, and image transformation—without redundancy or unnecessary complexity.
The toolset covers core workflows for image generation and manipulation, including server checks, configuration, and both text-to-image and image-to-image operations. A minor gap exists in the lack of tools for managing generated images (e.g., deletion or listing), but agents can work around this using file system operations.
Maintenance
Related MCP Connectors
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables local AI image generation on Apple Silicon Macs using MLX and Stable Diffusion. Supports conversational design iteration, asset generation, and wireframe creation with zero API costs through the Model Context Protocol.MIT
- AlicenseAqualityDmaintenanceEnables free local image generation from Claude Desktop/Claude Code using the Draw Things app on Apple Silicon Macs.4110 npmMIT
- AlicenseBqualityBmaintenanceEnables generating images from text or transforming existing images using GPT-Image-compatible APIs, with support for OpenAI and Agnes AI backends.2MIT
- AlicenseNot gradedqualityAmaintenanceProvides local image understanding for text-only LLMs with tools for image analysis, OCR, object detection, and cropping, all processed on-device.2MIT