Draw Things MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Draw Things MCP ServerGenerate an image of a cat eating pizza in space"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Draw Things MCP Server
Draw Things アプリを活用して、Claude Desktop / Claude Codeから無料でローカル画像生成ができるMCPサーバーです。
Google Gemini、OpenAIなどの有料APIなしで、MacのApple Silicon(M1〜M4)で直接画像を生成します。
元リポジトリ: james-see/mcp-drawthings このリポジトリは、FLUX.1 Schnellモデル向けにデフォルト設定を最適化したバージョンです。
変更点(元リポジトリとの違い)
デフォルトsteps: 20 → 4 — FLUX.1 Schnellモデルは4 stepsで十分な品質の画像を生成します。20〜30 stepsは不要に遅くなります。
Related MCP server: Draw Things MCP Server
必要要件
macOS(Apple Silicon M1/M2/M3/M4)
Draw Things アプリ(App Store)
Node.js 18以上
ステップ1: Draw Thingsアプリのインストールと設定
1-1. アプリのインストール
App Storeで Draw Things を検索してインストールします。
1-2. 初期設定(モデルのダウンロード)
アプリを初めて起動すると、モデルソースの選択画面が表示されます。
ステップ | 選択内容 |
ステップ 1/3: モデルソース | 「Draw Things経由でモデルをダウンロード」 を選択 |
Cloud Compute | オフ(ローカル実行のみなので不要) |
ステップ 2/3: モデル選択 | 検索欄に |
ステップ 3/3: 保存先フォルダ | デフォルトまたは任意のフォルダを選択 |
補足: FLUX.1 [schnell] 通常版は約11.7GBです。Mac RAMが8GBの場合は5-bit版を選択してください。
ダウンロードが完了するまで待ちます。
1-3. APIサーバーを有効にする
左サイドバーの 「設定」(歯車アイコン)をクリック
上部タブの 「詳細」 をクリック
「APIサーバー」 セクションを探す
以下のように設定:
設定項目 | 値 |
サーバーオンライン | オン(緑色) |
プロトコル | HTTP |
ポート | 7860(デフォルト) |
設定後、ターミナルで確認:
curl http://localhost:7860JSONレスポンスが返ってくれば成功です。
重要: Draw Thingsアプリは常に起動している必要があります。アプリを閉じるとAPIサーバーも停止します。最小化しておけばOKです。
ステップ2: MCPサーバーのインストール
git clone https://github.com/my13each/Drawthings_MCP.git
cd Drawthings_MCP
npm install
npm run buildステップ3: MCPクライアントの設定
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json に追加:
{
"mcpServers": {
"drawthings": {
"command": "node",
"args": ["/絶対パス/Drawthings_MCP/dist/index.js"]
}
}
}設定後、Claude Desktopを再起動します。
Claude Code
claude mcp add --scope user drawthings -- node /絶対パス/Drawthings_MCP/dist/index.js新しい会話を開始すると反映されます。
使い方
Claudeにこう言うだけです:
「宇宙でピザを食べるかわいい猫を描いて」
"Generate an image of a futuristic city at sunset"
英語のプロンプトの方が高品質な画像が生成されます。
MCPツール一覧
ツール | 説明 |
| Draw Things APIの接続状態を確認 |
| 現在ロードされているモデルと設定を取得 |
| テキストから画像を生成 |
| 既存の画像をテキストプロンプトで変換 |
generate_image パラメータ
パラメータ | 型 | 必須 | 説明 |
| string | O | 画像の説明テキスト |
| string | X | 除外する要素 |
| number | X | 画像の幅(デフォルト: 512) |
| number | X | 画像の高さ(デフォルト: 512) |
| number | X | 推論ステップ数(デフォルト: 4) |
| number | X | ガイダンススケール(デフォルト: 7.5) |
| number | X | シード値(-1 = ランダム) |
| string | X | 保存先パス |
transform_image パラメータ
パラメータ | 型 | 必須 | 説明 |
| string | O | 変換の説明テキスト |
| string | * | 元画像のファイルパス |
| string | * | Base64エンコードされた元画像 |
| number | X | 変換強度 0.0〜1.0(デフォルト: 0.75) |
| number | X | 推論ステップ数(デフォルト: 4) |
* image_path または image_base64 のどちらか必須
環境変数
変数 | デフォルト | 説明 |
|
| APIサーバーのホスト |
|
| APIサーバーのポート |
|
| 画像の保存ディレクトリ |
アーキテクチャ
┌─────────────────┐ stdio ┌──────────────────┐ HTTP ┌─────────────┐
│ MCP Client │◄──────────────►│ mcp-drawthings │◄────────────►│ Draw Things │
│ (Claude/Cursor) │ JSON-RPC │ │ localhost │ App │
└─────────────────┘ └──────────────────┘ :7860 └─────────────┘
│
▼
┌──────────────┐
│ File System │
│ (images) │
└──────────────┘トラブルシューティング
「Draw Things APIに接続できません」
Draw Thingsアプリが起動中か確認
設定で「サーバーオンライン」がオンになっているか確認
curl http://localhost:7860でレスポンスを確認アプリを閉じて再起動した場合、APIサーバーを再度オンにする必要がある場合があります
画像が生成されない
Draw Thingsにモデルがロードされているか確認
Draw Thingsアプリで直接画像生成をテスト
アプリのエラーメッセージを確認
ライセンス
MIT
クレジット
元リポジトリ: james-see/mcp-drawthings
Draw Things - Mac/iOS AI画像生成アプリ
Model Context Protocol - MCPプロトコル仕様
Available Tools
4 toolscheck_statusA
Check if the Draw Things API server is running and accessible
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool checks both 'running' and 'accessible' status, which is helpful. However, it doesn't describe the return value format, what indicates success or failure, or any side effects, though for a simple health check this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no filler. It communicates the essential purpose immediately and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, minimal risk), the description is nearly complete. It could ideally state what the response looks like, but for a health-check tool the meaning of 'running and accessible' is largely self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema confirms this with 100% coverage. The description adds no parameter details because none are needed. Baseline 4 is appropriate for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific action ('Check if the Draw Things API server is running and accessible') with a clear resource (the API server). This distinguishes it from siblings like get_config, generate_image, and transform_image, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys a clear diagnostic/health-check purpose, which implies it should be used to verify server availability before other API calls. It doesn't explicitly mention when not to use it or name alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image from a text prompt using the Draw Things app. The image will be saved to disk and the file path returned.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for reproducibility (-1 for random) | |
| model | No | Model filename to use for generation (use list_models to see available models) | |
| steps | No | Number of inference steps (default: 20) | |
| width | No | Width of the generated image in pixels (default: 512) | |
| height | No | Height of the generated image in pixels (default: 512) | |
| prompt | Yes | Text description of the image to generate | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| output_path | No | Custom file path to save the generated image | |
| negative_prompt | No | Elements to exclude from the generated image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does state the main side effect and return behavior: an image is generated, saved to disk, and its path is returned. It does not discuss overwrite behavior, blocking behavior, or app availability, but the most important behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the primary action and immediately state the side effect and return value. There is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich, fully documented input schema and no output schema, the description appropriately explains the high-level operation, the external app dependency, and what the caller receives. Minor gaps such as execution time or overwrite semantics are not critical for correct invocation because the schema covers the parameter details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters in adequate detail. The description adds no parameter-specific meaning beyond implicitly referencing the prompt and the resulting file path, which meets the baseline but does not go beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'Generate an image from a text prompt using the Draw Things app.' It also declares the key output behavior—saved to disk with the file path returned—which clearly distinguishes it from siblings like check_status, get_config, and transform_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied clearly: use this when you need a new image created from a text prompt. However, the description gives no explicit guidance about when to prefer this tool over transform_image or any other sibling, and it mentions no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_configA
Get the current Draw Things configuration including the loaded model and settings
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Get' and the resource 'configuration' imply a non-mutating read, and 'current' conveys that it reflects live state. However, it does not explicitly state that there are no side effects, nor does it describe output shape or error behavior, which would strengthen transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler. The core action and object are front-loaded, and the specific contents of the configuration are appended compactly. Every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description provides sufficient context: it states what is returned (current configuration) and what it includes (loaded model and settings). It could add a note about the exact output format or that calling it has no side effects, but those are not strictly necessary for an agent to invoke such a simple tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema declares zero parameters with 100% coverage, so there are no parameter semantics for the description to clarify. The baseline for a zero-parameter tool is 4, and the description appropriately avoids inventing parameter-related details that do not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a read operation targeting the current Draw Things configuration and names concrete contents ('loaded model and settings'). It is distinguishable from generate_image and transform_image as those are action-oriented, and from check_status as that targets status rather than configuration. However, it never explicitly contrasts itself with the check_status sibling, so differentiation is implicit rather than direct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus check_status, generate_image, or transform_image. No conditions, exclusions, or alternative routing hints are provided. An agent must infer usage purely from the tool name and resource description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_imageA
Transform an existing image using a text prompt (img2img). Either image_path or image_base64 must be provided.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for reproducibility (-1 for random) | |
| steps | No | Number of inference steps (default: 20) | |
| prompt | Yes | Text description of the desired transformation | |
| cfg_scale | No | Classifier-free guidance scale (default: 7.5) | |
| image_path | No | Path to the source image file to transform | |
| output_path | No | Custom file path to save the transformed image | |
| image_base64 | No | Base64-encoded source image (alternative to image_path) | |
| negative_prompt | No | Elements to exclude from the transformed image | |
| denoising_strength | No | Strength of the transformation (0.0-1.0, default: 0.75). Lower values keep more of the original image. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not mention whether the original image is preserved, whether the tool writes to output_path by default, what side effects occur, or what the return value looks like. For a transformation tool with no annotations, this is a meaningful transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One focused sentence communicates the operation, the input requirement, and the key constraint. Every word earns its place, and the critical either/or input requirement is stated up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema covers all parameter details, so the description does not need to repeat them. However, with no annotations and no output schema, the description should provide more context about output behavior, side effects, or when to choose this tool over generate_image. It is adequate but not fully complete for a 9-parameter tool with no behavioral metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all 9 parameters with 100% coverage, establishing a baseline of 3. The description adds value by explicitly calling out that either image_path or image_base64 must be provided, which is not captured by the schema's required list (only prompt is marked required). This helps agents avoid invalid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Transform'), a specific resource ('an existing image'), and the method ('using a text prompt (img2img)'). This clearly distinguishes it from the sibling generate_image, which is for creating new images rather than modifying existing ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the core use case clear: transform an already-existing image rather than generate a new one. However, it does not explicitly state 'use generate_image for new images' or provide explicit exclusion criteria, so it stops short of fully explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.1.2- First observed
check_status - First observed
generate_image - First observed
get_config - First observed
transform_image
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: status checking, configuration retrieval, text-to-image generation, and image-to-image transformation. There is no meaningful overlap between any pair of tools.
All tool names follow a consistent verb_noun pattern: check_status, get_config, generate_image, transform_image. The naming style is uniform and predictable.
Four tools is a well-scoped count for an image generation server. Each tool covers an essential operation without unnecessary redundancy or bloat.
The core workflow of checking server status, viewing configuration, generating images, and transforming images is covered. A possible minor gap is the lack of a tool to update configuration or switch models, but this does not block the primary use case.
Maintenance
Related MCP Connectors
- lightgenOAuthapp.lightgen
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
WHOOP recovery, strain, sleep and workouts in Claude via official WHOOP OAuth. Free, open source.
Let ChatGPT, Claude & Cursor use your Mac: email, calendar, iMessage, Teams, files. Local, free.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables Claude Desktop to interact with local ComfyUI installations for AI-powered image generation, including workflow management, model selection, real-time monitoring, and custom workflow execution through natural language.1443,607 npm18MIT
- AlicenseAqualityFmaintenanceEnables LLMs to generate and transform images locally on Mac using Stable Diffusion through the Draw Things app, supporting text-to-image and image-to-image generation with Apple Silicon acceleration.4106 npm24MIT
- AlicenseAqualityDmaintenanceEnables AI image generation, editing, composition, and style transfer in Claude conversations using Google's Gemini 2.5 Flash model. Automatically saves generated images to a local directory.492 npm11MIT
- AlicenseAqualityDmaintenanceEnables Claude Code to generate, edit, blend, and create variations of images using BytePlus SeeDream AI models, with streaming and Firebase sync.6MIT