Skip to main content
Glama
my13each

Draw Things MCP Server

by my13each

Draw Things MCP Server

Draw Things アプリを活用して、Claude Desktop / Claude Codeから無料でローカル画像生成ができるMCPサーバーです。

Google Gemini、OpenAIなどの有料APIなしで、MacのApple Silicon(M1〜M4)で直接画像を生成します。

元リポジトリ: james-see/mcp-drawthings このリポジトリは、FLUX.1 Schnellモデル向けにデフォルト設定を最適化したバージョンです。

変更点(元リポジトリとの違い)

  • デフォルトsteps: 20 → 4 — FLUX.1 Schnellモデルは4 stepsで十分な品質の画像を生成します。20〜30 stepsは不要に遅くなります。


Related MCP server: Draw Things MCP Server

必要要件

  • macOS(Apple Silicon M1/M2/M3/M4)

  • Draw Things アプリ(App Store

  • Node.js 18以上


ステップ1: Draw Thingsアプリのインストールと設定

1-1. アプリのインストール

App Storeで Draw Things を検索してインストールします。

1-2. 初期設定(モデルのダウンロード)

アプリを初めて起動すると、モデルソースの選択画面が表示されます。

ステップ

選択内容

ステップ 1/3: モデルソース

「Draw Things経由でモデルをダウンロード」 を選択

Cloud Compute

オフ(ローカル実行のみなので不要)

ステップ 2/3: モデル選択

検索欄に flux と入力 → 「FLUX.1 [schnell]」 を選択

ステップ 3/3: 保存先フォルダ

デフォルトまたは任意のフォルダを選択

補足: FLUX.1 [schnell] 通常版は約11.7GBです。Mac RAMが8GBの場合は5-bit版を選択してください。

ダウンロードが完了するまで待ちます。

1-3. APIサーバーを有効にする

  1. 左サイドバーの 「設定」(歯車アイコン)をクリック

  2. 上部タブの 「詳細」 をクリック

  3. 「APIサーバー」 セクションを探す

  4. 以下のように設定:

設定項目

サーバーオンライン

オン(緑色)

プロトコル

HTTP

ポート

7860(デフォルト)

設定後、ターミナルで確認:

curl http://localhost:7860

JSONレスポンスが返ってくれば成功です。

重要: Draw Thingsアプリは常に起動している必要があります。アプリを閉じるとAPIサーバーも停止します。最小化しておけばOKです。


ステップ2: MCPサーバーのインストール

git clone https://github.com/my13each/Drawthings_MCP.git
cd Drawthings_MCP
npm install
npm run build

ステップ3: MCPクライアントの設定

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json に追加:

{
  "mcpServers": {
    "drawthings": {
      "command": "node",
      "args": ["/絶対パス/Drawthings_MCP/dist/index.js"]
    }
  }
}

設定後、Claude Desktopを再起動します。

Claude Code

claude mcp add --scope user drawthings -- node /絶対パス/Drawthings_MCP/dist/index.js

新しい会話を開始すると反映されます。


使い方

Claudeにこう言うだけです:

「宇宙でピザを食べるかわいい猫を描いて」

"Generate an image of a futuristic city at sunset"

英語のプロンプトの方が高品質な画像が生成されます。


MCPツール一覧

ツール

説明

check_status

Draw Things APIの接続状態を確認

get_config

現在ロードされているモデルと設定を取得

generate_image

テキストから画像を生成

transform_image

既存の画像をテキストプロンプトで変換

generate_image パラメータ

パラメータ

必須

説明

prompt

string

O

画像の説明テキスト

negative_prompt

string

X

除外する要素

width

number

X

画像の幅(デフォルト: 512)

height

number

X

画像の高さ(デフォルト: 512)

steps

number

X

推論ステップ数(デフォルト: 4)

cfg_scale

number

X

ガイダンススケール(デフォルト: 7.5)

seed

number

X

シード値(-1 = ランダム)

output_path

string

X

保存先パス

transform_image パラメータ

パラメータ

必須

説明

prompt

string

O

変換の説明テキスト

image_path

string

*

元画像のファイルパス

image_base64

string

*

Base64エンコードされた元画像

denoising_strength

number

X

変換強度 0.0〜1.0(デフォルト: 0.75)

steps

number

X

推論ステップ数(デフォルト: 4)

* image_path または image_base64 のどちらか必須


環境変数

変数

デフォルト

説明

DRAWTHINGS_HOST

localhost

APIサーバーのホスト

DRAWTHINGS_PORT

7860

APIサーバーのポート

DRAWTHINGS_OUTPUT_DIR

~/Pictures/drawthings-mcp

画像の保存ディレクトリ


アーキテクチャ

┌─────────────────┐     stdio      ┌──────────────────┐     HTTP      ┌─────────────┐
│   MCP Client    │◄──────────────►│  mcp-drawthings  │◄────────────►│ Draw Things │
│ (Claude/Cursor) │   JSON-RPC     │                  │  localhost    │    App      │
└─────────────────┘                └──────────────────┘   :7860       └─────────────┘
                                           │
                                           ▼
                                   ┌──────────────┐
                                   │  File System │
                                   │   (images)   │
                                   └──────────────┘

トラブルシューティング

「Draw Things APIに接続できません」

  1. Draw Thingsアプリが起動中か確認

  2. 設定で「サーバーオンライン」がオンになっているか確認

  3. curl http://localhost:7860 でレスポンスを確認

  4. アプリを閉じて再起動した場合、APIサーバーを再度オンにする必要がある場合があります

画像が生成されない

  1. Draw Thingsにモデルがロードされているか確認

  2. Draw Thingsアプリで直接画像生成をテスト

  3. アプリのエラーメッセージを確認


ライセンス

MIT

クレジット

Available Tools

4 tools
check_statusA

Check if the Draw Things API server is running and accessible

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool checks both 'running' and 'accessible' status, which is helpful. However, it doesn't describe the return value format, what indicates success or failure, or any side effects, though for a simple health check this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence with no filler. It communicates the essential purpose immediately and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, minimal risk), the description is nearly complete. It could ideally state what the response looks like, but for a health-check tool the meaning of 'running and accessible' is largely self-explanatory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema confirms this with 100% coverage. The description adds no parameter details because none are needed. Baseline 4 is appropriate for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear, specific action ('Check if the Draw Things API server is running and accessible') with a clear resource (the API server). This distinguishes it from siblings like get_config, generate_image, and transform_image, which perform different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys a clear diagnostic/health-check purpose, which implies it should be used to verify server availability before other API calls. It doesn't explicitly mention when not to use it or name alternatives, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate an image from a text prompt using the Draw Things app. The image will be saved to disk and the file path returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoRandom seed for reproducibility (-1 for random)
modelNoModel filename to use for generation (use list_models to see available models)
stepsNoNumber of inference steps (default: 20)
widthNoWidth of the generated image in pixels (default: 512)
heightNoHeight of the generated image in pixels (default: 512)
promptYesText description of the image to generate
cfg_scaleNoClassifier-free guidance scale (default: 7.5)
output_pathNoCustom file path to save the generated image
negative_promptNoElements to exclude from the generated image

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does state the main side effect and return behavior: an image is generated, saved to disk, and its path is returned. It does not discuss overwrite behavior, blocking behavior, or app availability, but the most important behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the primary action and immediately state the side effect and return value. There is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich, fully documented input schema and no output schema, the description appropriately explains the high-level operation, the external app dependency, and what the caller receives. Minor gaps such as execution time or overwrite semantics are not critical for correct invocation because the schema covers the parameter details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all nine parameters in adequate detail. The description adds no parameter-specific meaning beyond implicitly referencing the prompt and the resulting file path, which meets the baseline but does not go beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair: 'Generate an image from a text prompt using the Draw Things app.' It also declares the key output behavior—saved to disk with the file path returned—which clearly distinguishes it from siblings like check_status, get_config, and transform_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied clearly: use this when you need a new image created from a text prompt. However, the description gives no explicit guidance about when to prefer this tool over transform_image or any other sibling, and it mentions no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_configA

Get the current Draw Things configuration including the loaded model and settings

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Get' and the resource 'configuration' imply a non-mutating read, and 'current' conveys that it reflects live state. However, it does not explicitly state that there are no side effects, nor does it describe output shape or error behavior, which would strengthen transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no filler. The core action and object are front-loaded, and the specific contents of the configuration are appended compactly. Every word contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description provides sufficient context: it states what is returned (current configuration) and what it includes (loaded model and settings). It could add a note about the exact output format or that calling it has no side effects, but those are not strictly necessary for an agent to invoke such a simple tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema declares zero parameters with 100% coverage, so there are no parameter semantics for the description to clarify. The baseline for a zero-parameter tool is 4, and the description appropriately avoids inventing parameter-related details that do not exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a read operation targeting the current Draw Things configuration and names concrete contents ('loaded model and settings'). It is distinguishable from generate_image and transform_image as those are action-oriented, and from check_status as that targets status rather than configuration. However, it never explicitly contrasts itself with the check_status sibling, so differentiation is implicit rather than direct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus check_status, generate_image, or transform_image. No conditions, exclusions, or alternative routing hints are provided. An agent must infer usage purely from the tool name and resource description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_imageA

Transform an existing image using a text prompt (img2img). Either image_path or image_base64 must be provided.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoRandom seed for reproducibility (-1 for random)
stepsNoNumber of inference steps (default: 20)
promptYesText description of the desired transformation
cfg_scaleNoClassifier-free guidance scale (default: 7.5)
image_pathNoPath to the source image file to transform
output_pathNoCustom file path to save the transformed image
image_base64NoBase64-encoded source image (alternative to image_path)
negative_promptNoElements to exclude from the transformed image
denoising_strengthNoStrength of the transformation (0.0-1.0, default: 0.75). Lower values keep more of the original image.

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not mention whether the original image is preserved, whether the tool writes to output_path by default, what side effects occur, or what the return value looks like. For a transformation tool with no annotations, this is a meaningful transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One focused sentence communicates the operation, the input requirement, and the key constraint. Every word earns its place, and the critical either/or input requirement is stated up front.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema covers all parameter details, so the description does not need to repeat them. However, with no annotations and no output schema, the description should provide more context about output behavior, side effects, or when to choose this tool over generate_image. It is adequate but not fully complete for a 9-parameter tool with no behavioral metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all 9 parameters with 100% coverage, establishing a baseline of 3. The description adds value by explicitly calling out that either image_path or image_base64 must be provided, which is not captured by the schema's required list (only prompt is marked required). This helps agents avoid invalid calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Transform'), a specific resource ('an existing image'), and the method ('using a text prompt (img2img)'). This clearly distinguishes it from the sibling generate_image, which is for creating new images rather than modifying existing ones.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the core use case clear: transform an already-existing image rather than generate a new one. However, it does not explicitly state 'use generate_image for new images' or provide explicit exclusion criteria, so it stops short of fully explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.1.2
    • First observedcheck_status
    • First observedgenerate_image
    • First observedget_config
    • First observedtransform_image

TDQS

A4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: status checking, configuration retrieval, text-to-image generation, and image-to-image transformation. There is no meaningful overlap between any pair of tools.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern: check_status, get_config, generate_image, transform_image. The naming style is uniform and predictable.

Tool Count5/5

Four tools is a well-scoped count for an image generation server. Each tool covers an essential operation without unnecessary redundancy or bloat.

Completeness4/5

The core workflow of checking server status, viewing configuration, generating images, and transforming images is covered. A possible minor gap is the lack of a tool to update configuration or switch models, but this does not block the primary use case.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers