Draw Things MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Draw Things MCP Servergenerate a serene mountain landscape with a lake at sunrise"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-drawthings
An MCP (Model Context Protocol) server for Draw Things - enabling LLMs to generate images locally on Mac using Stable Diffusion and other AI models.
Features
Text-to-Image Generation - Generate images from text prompts using the currently loaded model in Draw Things
Image-to-Image Transformation - Transform existing images using text prompts
Configuration Access - Query the current Draw Things settings and loaded model
Local Processing - All image generation runs locally on your Mac using Apple Silicon (M1/M2/M3/M4)
Related MCP server: Draw Things MCP Server
Prerequisites
macOS with Apple Silicon (M1/M2/M3/M4)
Draw Things app installed
Node.js 18 or later
Setup
1. Enable Draw Things API Server
Open Draw Things
Click the gear icon (⚙️) to open Settings
Enable API Server / HTTP Server
The server runs on port 7860 by default
Verify the server is running:
curl http://localhost:78602. Configure Your MCP Client
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"drawthings": {
"command": "npx",
"args": ["-y", "mcp-drawthings"]
}
}
}Cursor
Add to .cursor/mcp.json in your project or global config:
{
"mcpServers": {
"drawthings": {
"command": "npx",
"args": ["-y", "mcp-drawthings"]
}
}
}3. Restart Your MCP Client
Restart Claude Desktop or Cursor to load the new MCP server.
Available Tools
check_status
Check if the Draw Things API server is running and accessible.
get_config
Get the current Draw Things configuration including the loaded model and settings.
generate_image
Generate an image from a text prompt.
Parameters:
Parameter | Type | Required | Description |
| string | Yes | Text description of the image to generate |
| string | No | Elements to exclude from the generated image |
| number | No | Image width in pixels (default: 512) |
| number | No | Image height in pixels (default: 512) |
| number | No | Number of inference steps (default: 20) |
| number | No | Guidance scale (default: 7.5) |
| number | No | Random seed for reproducibility (-1 for random) |
| string | No | Custom file path to save the image |
Example:
Generate an image of a futuristic city at sunset with flying carstransform_image
Transform an existing image using a text prompt (img2img).
Parameters:
Parameter | Type | Required | Description |
| string | Yes | Text description of the desired transformation |
| string | No* | Path to the source image file |
| string | No* | Base64-encoded source image |
| string | No | Elements to exclude |
| number | No | Transformation strength 0.0-1.0 (default: 0.75) |
| number | No | Number of inference steps (default: 20) |
| number | No | Guidance scale (default: 7.5) |
| number | No | Random seed (-1 for random) |
| string | No | Custom file path to save the result |
*Either image_path or image_base64 must be provided.
Configuration
Environment Variables
Variable | Default | Description |
|
| Draw Things API host |
|
| Draw Things API port |
|
| Directory for generated images |
Architecture
┌─────────────────┐ stdio ┌──────────────────┐ HTTP ┌─────────────┐
│ MCP Client │◄──────────────►│ mcp-drawthings │◄────────────►│ Draw Things │
│ (Claude/Cursor) │ JSON-RPC │ │ localhost │ App │
└─────────────────┘ └──────────────────┘ :7860 └─────────────┘
│
▼
┌──────────────┐
│ File System │
│ (images) │
└──────────────┘Development
# Clone the repository
git clone https://github.com/james-see/mcp-drawthings
cd mcp-drawthings
# Install dependencies
npm install
# Build
npm run build
# Run in development mode
npm run devTroubleshooting
"Cannot connect to Draw Things API"
Make sure Draw Things is running
Check that the API Server is enabled in Draw Things settings
Verify the server is accessible:
curl http://localhost:7860Check if a different port is configured in Draw Things
Images not generating
Make sure a model is loaded in Draw Things
Check Draw Things for any error messages
Try generating an image directly in Draw Things first
Permission errors saving images
Check that the output directory is writable. You can set a custom directory using the DRAWTHINGS_OUTPUT_DIR environment variable.
License
MIT
Related Projects
Draw Things - The AI image generation app for Mac/iOS
Model Context Protocol - The protocol specification
@modelcontextprotocol/sdk - TypeScript SDK for MCP
Available Tools
4 toolscheck_statusA
Check if the Draw Things API server is running and accessible
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the tool's purpose (connectivity/health check) but doesn't describe behavioral traits like response format, error conditions, timeout behavior, or authentication requirements. It's minimal but accurate for a simple status check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with zero waste - every word contributes essential information. Front-loaded with the core action ('Check'), followed by the target and purpose. No redundant phrases or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no annotations, no output schema), the description is complete enough to understand its basic function. However, for a connectivity check tool, additional context about what 'running and accessible' means (e.g., returns status code, response time) would be helpful despite the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the baseline is 4. The description appropriately doesn't discuss parameters since none exist, and doesn't need to compensate for any schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('check if... is running and accessible') and the target resource ('Draw Things API server'), distinguishing it from sibling tools like generate_image or transform_image. It uses precise language that conveys a diagnostic/health-check function rather than data manipulation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (verifying server availability before attempting operations), but doesn't explicitly state when to use this vs. alternatives or provide exclusions. It suggests a prerequisite check but lacks explicit guidance like 'use before calling generate_image if unsure about connectivity'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageB
Generate an image from a text prompt using the Draw Things app. The image will be saved to disk and the file path returned.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate | |
| negative_prompt | No | Elements to exclude from the generated image | |
| width | No | Width of the generated image in pixels (default: 512) | |
| height | No | Height of the generated image in pixels (default: 512) | |
| steps | No | Number of inference steps (default: 20) | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| seed | No | Random seed for reproducibility (-1 for random) | |
| model | No | Model filename to use for generation (use list_models to see available models) | |
| output_path | No | Custom file path to save the generated image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that the image 'will be saved to disk and the file path returned,' which is useful context about output behavior. However, it doesn't mention potential side effects like disk space usage, performance implications (e.g., time-intensive generation), or error conditions (e.g., invalid prompts).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose. It avoids unnecessary words and gets straight to the point. However, it could be slightly more structured by explicitly separating the action from the output behavior for even clearer readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 9 parameters and no output schema, the description is minimally adequate. It covers the basic purpose and output behavior but lacks details on error handling, performance characteristics, or integration with sibling tools. Without annotations or an output schema, more context would be helpful for an AI agent to use this tool effectively in varied scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, so all parameters are documented in the schema. The description adds no additional parameter semantics beyond what's in the schema. It doesn't explain relationships between parameters (e.g., how 'steps' affects quality vs. speed) or provide usage examples. The baseline of 3 is appropriate given the comprehensive schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an image'), the resource involved ('from a text prompt'), and the tool used ('using the Draw Things app'). It distinguishes itself from sibling tools like 'transform_image' by focusing on generation rather than modification, and from 'check_status' and 'get_config' by being a creative operation rather than informational.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, constraints, or scenarios where other tools might be more appropriate. For example, it doesn't clarify if 'transform_image' should be used for editing existing images instead of generating new ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_configB
Get the current Draw Things configuration including the loaded model and settings
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states what the tool does, not how it behaves. It lacks details on permissions needed, rate limits, error handling, or what 'Get' entails (e.g., is it a read-only operation, does it cache data?). This is a significant gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose without waste. Every word contributes to clarifying the tool's function, making it appropriately sized and well-structured for its simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, no annotations, and no output schema, the description is minimally adequate by stating what configuration is retrieved. However, it lacks details on return values (e.g., format of configuration data) and behavioral context, which could hinder an agent's ability to use it effectively without trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add parameter details, which is appropriate, earning a baseline score of 4 for not introducing unnecessary information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and the resource 'current Draw Things configuration', specifying it includes 'loaded model and settings'. This is specific and unambiguous, though it doesn't explicitly differentiate from sibling tools like 'check_status' which might overlap in checking system state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'check_status' or other siblings. The description implies usage for retrieving configuration details but doesn't specify contexts, prerequisites, or exclusions, leaving the agent to infer based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_imageB
Transform an existing image using a text prompt (img2img). Either image_path or image_base64 must be provided.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the desired transformation | |
| image_path | No | Path to the source image file to transform | |
| image_base64 | No | Base64-encoded source image (alternative to image_path) | |
| negative_prompt | No | Elements to exclude from the transformed image | |
| denoising_strength | No | Strength of the transformation (0.0-1.0, default: 0.75). Lower values keep more of the original image. | |
| steps | No | Number of inference steps (default: 20) | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| seed | No | Random seed for reproducibility (-1 for random) | |
| output_path | No | Custom file path to save the transformed image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the img2img method and the requirement for an image source, but fails to describe critical behaviors such as whether the transformation is destructive to the original image, what permissions or authentication are needed, rate limits, error handling, or the output format (e.g., returns a file path or base64). For a mutation tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, consisting of just two sentences that directly state the purpose and a key requirement. Every word earns its place with no redundancy or fluff, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, mutation operation), lack of annotations, and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., a transformed image file or metadata), error conditions, side effects, or how it differs from siblings. For a tool that modifies images, more behavioral context is needed to ensure correct agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters thoroughly with descriptions, ranges, and defaults. The description adds minimal value beyond the schema by noting the 'Either image_path or image_base64 must be provided' constraint, which isn't explicitly in the schema (though 'prompt' is required). This provides slight additional context, but most parameter semantics are covered by the schema, justifying a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Transform an existing image using a text prompt (img2img).' It specifies the verb ('transform'), resource ('existing image'), and method ('text prompt'), distinguishing it from sibling tools like 'generate_image' (likely text-to-image) and 'check_status'. However, it doesn't explicitly contrast with 'generate_image' beyond implying img2img vs text-to-image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context by stating 'Either image_path or image_base64 must be provided,' which helps the agent understand prerequisites. However, it lacks explicit guidance on when to use this tool versus alternatives like 'generate_image' (e.g., for modifying vs creating from scratch) or 'check_status' (e.g., for monitoring). The usage is implied but not clearly articulated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: check_status verifies server availability, generate_image creates new images from text, get_config retrieves settings, and transform_image modifies existing images. There is no overlap or ambiguity between these functions, making tool selection straightforward for an agent.
All tool names follow a consistent verb_noun pattern with snake_case formatting: check_status, generate_image, get_config, and transform_image. This uniformity enhances readability and predictability across the toolset.
With 4 tools, this server is well-scoped for its purpose of interacting with the Draw Things app. Each tool serves a specific, essential function—server status, image generation, configuration retrieval, and image transformation—without redundancy or unnecessary complexity.
The toolset covers core workflows for image generation and manipulation, including server checks, configuration, and both text-to-image and image-to-image operations. A minor gap exists in the lack of tools for managing generated images (e.g., deletion or listing), but agents can work around this using file system operations.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables local AI image generation on Apple Silicon Macs using MLX and Stable Diffusion. Supports conversational design iteration, asset generation, and wireframe creation with zero API costs through the Model Context Protocol.MIT
- AlicenseNot gradedqualityDmaintenanceEnables free local image generation from Claude Desktop/Claude Code using the Draw Things app on Apple Silicon Macs.136MIT
- AlicenseBqualityBmaintenanceEnables generating images from text or transforming existing images using GPT-Image-compatible APIs, with support for OpenAI and Agnes AI backends.2MIT
- AlicenseNot gradedqualityAmaintenanceProvides local image understanding for text-only LLMs with tools for image analysis, OCR, object detection, and cropping, all processed on-device.2MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/james-see/mcp-drawthings'
If you have feedback or need assistance with the MCP directory API, please join our Discord server