Skip to main content
Glama
resccrew

affinity-mcp

by resccrew

affinity-mcp

An MCP server that lets Claude (Claude Code, Claude Desktop or any MCP client) use the Affinity design app on macOS the way a person would: it looks at screenshots, then clicks, drags, types and presses shortcuts.

Affinity has no scripting API, so this works with any version and any studio (Vector, Pixel, Layout). You don't need API keys or plugins.

Tools

Tool

What it does

open_affinity

Launch Affinity or bring it to the front

screenshot

Screenshot of the screen (downscaled to 1280 px); all coordinates refer to the last one

click

Left/right/middle click at (x, y)

double_click

Double-click at (x, y)

drag

Drag from (x1, y1) to (x2, y2): draw shapes, move objects, drag sliders

type_text

Type text (via clipboard, so Cyrillic and emoji work)

press_key

One key: enter, esc, tab, v (Move tool), m (Rectangle)...

hotkey

Key combo, e.g. ["command", "n"] new document, ["command", "z"] undo

scroll

Scroll, optionally over a given point

set_field

Set a panel field (X/Y/W/H, colour hex...) and commit it

menu

Choose any menu item by path, e.g. ["Vector","Geometry","Add"], with no clicks

list_menu

List menu names, so the model can find the exact path

batch

Run many steps in one call and return one screenshot at the end, which is the fast path

click also accepts modifiers (e.g. ["command"] for multi-select).

Every action returns a fresh screenshot by default (screenshot_after=false turns it off). Before every action the server brings Affinity back to the front, so the terminal can't take the clicks.

Related MCP server: Affinity MCP Server

Requirements

  • macOS, Python 3.13+, uv

  • Affinity installed (/Applications/Affinity.app)

  • Permissions for the app that runs your MCP client (Terminal, iTerm, Ghostty, VS Code, Claude): System Settings → Privacy & Security → Accessibility and Screen Recording. Without them, macOS silently ignores clicks and returns an empty screenshot. The server detects this and tells you what's missing.

Install

git clone https://github.com/resccrew/affinity-mcp.git
cd affinity-mcp
uv sync

Claude Code

claude mcp add -s user affinity -- uv --directory /path/to/affinity-mcp run affinity-mcp

Claude Desktop

Add this to ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "affinity": {
      "command": "uv",
      "args": ["--directory", "/path/to/affinity-mcp", "run", "affinity-mcp"]
    }
  }
}

Restart the client, then ask e.g. "open Affinity, create an A4 document and draw a red circle".

Tips

  • Keyboard shortcuts are more reliable than clicking small icons: Cmd+N creates a document (tiles on the welcome screen ignore single synthetic clicks).

  • Coordinates are always relative to the last screenshot. The server maps them to real screen points, including Retina scaling.

Safety

  • The server moves your real mouse and keyboard, so don't use the computer while it works. To abort, move the mouse into any screen corner (pyautogui failsafe).

  • Screenshots of the whole screen go to the model, so close windows with private data first.

Development

uv run pytest

Tests use a fake screen and need no display or permissions.

License

MIT

Available Tools

13 tools
batchA

Run many GUI steps in one call, one screenshot at the end. Stops at the first failing step.

Each step: {"action": ..., ...}. Actions: click{x,y,button?,modifiers?}, double_click{x,y}, drag{x1,y1,x2,y2,duration?}, type{text,press_enter?}, key{key}, hotkey{keys}, scroll{amount,x?,y?}, field{x,y,value}, menu{path}, wait{seconds}. Coordinates refer to the last screenshot taken before this call.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYes
screenshot_afterNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the abort-on-first-failure semantics and the critical coordinate contract ("Coordinates refer to the last screenshot taken before this call"). It still omits the return payload shape, timing behavior, and any permission/state requirements, so it falls short of complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, then failure/coordinate semantics, then the dense action reference. The action list is long but every entry is load-bearing; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description covers the essentials for a batch execution tool: step vocabulary, coordinate reference frame, and error-stop behavior. The main gap is the undocumented screenshot_after parameter and no hint of what the call returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it fully enumerates the per-step action vocabulary (click, drag, type, key, hotkey, scroll, field, menu, wait) with their fields, which is exactly the structure the opaque "steps" array hides. It does not, however, explain the screenshot_after parameter beyond implying "one screenshot at the end".

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Run many GUI steps in one call") and immediately differentiates from the single-action siblings (click, type_text, scroll, etc.) by framing itself as the batched alternative. An agent can tell what this does and why it exists alongside individual action tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Batching intent is implied by "many GUI steps in one call" and the "one screenshot at the end" note, but there is no explicit when-to-use versus calling the individual tools, nor guidance on when this is inappropriate. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickB

Click at (x, y) in the last screenshot. button: left/right/middle; modifiers e.g. ["command"].

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
buttonNoleft
modifiersNo
screenshot_afterNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not disclose whether the click blocks, what happens on failure, whether a new screenshot is captured (the schema's screenshot_after default of true is unexplained), or whether coordinates are absolute pixels in the prior screenshot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the action and coordinate target front-loaded, followed by the optional modifier semantics. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A simple 5-parameter tool with no annotations and no output schema needs the description to cover coordinate semantics and the side effect of screenshot_after. Two of the five parameters are well covered, but the unstated coordinate origin and unconditional post-click screenshot behavior leave real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully documents the button enum (left/right/middle) and modifiers with a concrete example ['command'], but leaves x/y coordinate space and the screenshot_after parameter (default true) entirely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and the exact target (x, y coordinates in the last screenshot), which is precise enough to invoke. It implicitly differentiates itself from sibling double_click and drag by naming a single click, but does not explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'in the last screenshot' implies a prerequisite (a screenshot must exist) but gives no explicit when-to-use guidance or routing against siblings like double_click or menu. Usage is inferable from the mouse-action family but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

double_clickC

Double-click at (x, y) in the last screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
screenshot_afterNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full behavioral burden, and it is thin. It does not disclose the coordinate origin (screen vs. last screenshot), timing between the two clicks, or that 'screenshot_after' defaults to true and triggers a capture. No mention of side effects such as focus/selection changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and no filler. It is efficient, but the brevity shades into under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage across 3 parameters, the description leaves key details (coordinate frame, screenshot_after behavior, interaction with the current UI state) unspecified. Inadequate for a GUI-automation primitive where ambiguous coordinates cause wrong clicks.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate but only partially does. 'in the last screenshot' loosely implies x/y are in screenshot-relative pixel space, but the third parameter 'screenshot_after' is never mentioned and the coordinate frame remains ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete action (double-click) plus the target coordinates and the surface it acts on ('last screenshot'). This is clearly distinguishable from the sibling 'click' by the double-click semantics, though the description never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to prefer this over 'click', 'menu', or 'set_field', and no prerequisites given. The agent must infer from the tool name alone that this is the two-click variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dragC

Drag with the left button from (x1, y1) to (x2, y2) in the last screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
x1Yes
x2Yes
y1Yes
y2Yes
durationNo
screenshot_afterNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses the left mouse button and that coordinates refer to the last screenshot, but omits side effects, reversibility, required permissions, and the meaning of the duration and screenshot_after parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded with the core action and coordinate reference, with no wasted words. It is appropriately sized for a simple pointer operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter action tool with no annotations, no output schema, and 0% schema description coverage, the description is incomplete. It does not explain optional parameters, coordinate semantics, side effects, or when to use drag over other interaction tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names the four required coordinate parameters in sequence but does not explain their units, reference frame, or semantics, and it completely omits the duration and screenshot_after optional parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Drag') and resource (mouse left button from (x1,y1) to (x2,y2) in the last screenshot). It is distinguishable from siblings like click and double_click, though it does not explicitly name an alternative or clarify when a drag is preferred over other interactions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like click, scroll, or set_field. The phrase 'in the last screenshot' implies a prerequisite but does not explain the context or conditions for choosing a drag operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hotkeyC

Press a key combination, e.g. ["command", "n"] for a new document.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes
screenshot_afterNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it omits important traits: the default screenshot_after=true side effect (which itself produces a screenshot) is never mentioned, nor is any note about focus requirements or what the tool returns. It only restates the obvious press action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with an inline example, no wasted words. It is appropriately sized for the tool's simplicity, though brevity comes at the cost of the missing behavioral detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description should disclose the screenshot_after side effect and differentiate from press_key. Neither is present, so an agent cannot confidently determine the tool's full behavior or when to prefer it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are two parameters. The example array format clarifies the 'keys' parameter somewhat, but the screenshot_after parameter and its default-true side effect are entirely undocumented in both schema and description, leaving half the parameters unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Press a key combination') with a concrete example, so the core action is unambiguous. However, it does not distinguish itself from the sibling tool press_key, which appears to overlap in intent, leaving the agent to guess which one to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the sibling press_key or menu/list_menu alternatives. The example implies a single key-combination use case but no conditions, prerequisites, or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_menuB

List menu names: [] = menu bar, ["Vector"] = its items, ["Vector", "Geometry"] = submenu.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the recursive path semantics (menu bar vs. items vs. submenu), which is genuine behavioral context, but it says nothing about error behavior for invalid paths, whether the result is a flat name list or a structured tree, or pagination/size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense line with the core semantics front-loaded and zero filler. It is efficient, though the telegraphic style ('its items') is borderline cryptic for an agent unfamiliar with the domain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should ideally state the return shape (list of menu names, and whether nested menus are included) and invalid-path behavior. The parameter is well covered, but the calling contract for results is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (the single 'path' param is only typed as array-of-string-or-null), so the description must compensate — and it does, by concretely mapping [] / ["Vector"] / ["Vector","Geometry"] to three distinct return scopes. It stops short of clarifying case sensitivity, matching rules, or behavior for a nonexistent path.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List menu names') and immediately clarifies with three path examples showing what each path depth returns. It does not distinguish itself from the sibling 'menu' tool, which is likely a related menu-operation tool, so an agent must infer the boundary from context rather than the text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The examples imply a drill-down pattern (start with [] then descend), but there is no explicit statement of when to use this versus siblings like 'menu', 'click', or 'set_field'. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_affinityC

Launch or bring the Affinity app to the front.

ParametersJSON Schema
NameRequiredDescriptionDefault
screenshot_afterNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but discloses almost nothing beyond the action: whether launching is conditional, what happens on failure, what is returned, or what the screenshot_after side effect does. The 'or bring to the front' clause is the only behavioral detail offered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the action and target both stated immediately; no filler, nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and an undocumented parameter, one sentence is insufficient. The undefined screenshot_after behavior and lack of return/error semantics leave real gaps an agent cannot fill from structured data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (screenshot_after, default true) with 0% schema description coverage, and the description never mentions it. The agent has to infer from the field name alone that a screenshot is captured after opening, which is exactly the kind of side effect the description should clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (launch/bring to front) and a specific resource (the Affinity app), so the agent knows exactly what happens. The siblings are all generic UI primitives (click, type_text, scroll), so this app-lifecycle tool is evidently distinct without needing to name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when this should be used versus the generic interaction tools, nor any prerequisite (e.g. that it's needed before other Affinity interactions if the app isn't running). The one clause 'or bring to the front' hints at idempotence but is not framed as usage advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyC

Press one key, e.g. enter, esc, tab, up, down, backspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
screenshot_afterNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say where the key is sent, what happens when no element is focused, or that screenshot_after defaults to true, leaving the side-effect profile undisclosed for a tool with an implicit capture behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, appropriately sized for a simple action tool. It errs toward under-specification rather than verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter UI action tool with no annotations and no output schema, the description omits the screenshot side effect, focus requirements, and any return information. A caller cannot predict the observable outcome of the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does exemplify valid key values for the required key parameter, but the second parameter, screenshot_after (default true), is never mentioned in the description or the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (press) and resource (one key) and gives concrete example keys (enter, esc, tab, up, down, backspace). It implies a single-key scope that distinguishes it from type_text and hotkey, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use press_key versus type_text (for text) or hotkey (for combos), and no stated prerequisites such as the target needing focus. The examples hint at valid values but the agent must infer the single-key boundary on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotA

Take a screenshot of the whole screen. Coordinates for other tools refer to this image.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses the screen scope and the coordinate reference frame for other tools, but says nothing about the return value (image object, base64, saved path), multi-monitor behavior, or how the result is consumed downstream.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both load-bearing: the first defines the operation and scope, the second delivers the critical interoperability fact about the coordinate system. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter capture tool the description covers what an agent most needs to know to use it in a workflow. The only real gap is the return format, which is not covered by an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Take a screenshot") and adds scope ("of the whole screen"), so an agent knows this is a full-screen capture rather than a window or region grab. It does not explicitly differentiate from siblings, but no sibling performs capture, so ambiguity is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the note that other tools' coordinates refer to this image hints that this should be called first to establish the coordinate frame, but the description never says when to call it, in what order relative to click/type_text, or whether it can be re-invoked mid-task.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollB

Scroll: positive = up, negative = down. Pass x, y (last screenshot) to scroll over a specific area.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
amountYes
screenshot_afterNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It discloses scroll direction semantics but says nothing about side effects such as the automatic screenshot_after behavior (default true in schema), whether scrolling requires a focused window, or what the tool returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the important sign convention front-loaded and no filler. It is terse to the point of being telegraphic ('Pass x, y (last screenshot)'), but every phrase carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with no annotations and no output schema, the description leaves too much uncovered: the screenshot_after parameter, magnitude expectations for amount, and the effect of omitting x/y are all unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does add real meaning for amount (positive = up, negative = down) and x/y (scroll over a specific area). It omits screenshot_after entirely and gives no units or magnitude guidance for amount, leaving one of four parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Scroll') and immediately defines the sign convention of the amount, so an agent knows exactly what the tool does. It does not explicitly distinguish itself from siblings like drag or press_key, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one conditional usage rule — pass x, y (last screenshot) to scroll within a specific area — which implies the default is scrolling the whole viewport. However, it never says when to prefer scroll over alternatives such as press_key or drag, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_fieldB

Replace the value of a panel field at (x, y) (X/Y/W/H, colour hex, ...) and commit with Enter.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
valueYes
screenshot_afterNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden, and it does add one valuable non-schema fact: the value is committed with Enter, so no separate confirm is needed. However, it omits failure handling (what happens if no field is at those coordinates) and the coordinate basis, leaving meaningful gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero waste; the coordinate anchor and commit behavior are stated compactly. Minor loss because the parenthetical examples crowd the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a coordinate-based GUI mutation tool with no annotations, no output schema, and 0% param coverage, the description covers the essentia action but omits failure behavior and the screenshot_after parameter, so it is adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and only partially does: (x, y) implies coordinates and the examples (X/Y/W/H, colour hex) hint at what 'value' holds. The fourth parameter, screenshot_after, is never mentioned, and there is no format/precision guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Replace') and resource ('value of a panel field at (x, y)') with concrete examples of field content. It is distinguishable from type_text/click by targeting a panel field and auto-committing. It does not explicitly name a sibling it is not, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through 'panel field at (x, y)' but gives no explicit when-to-use, when-not-to-use, or guidance on choosing this over type_text/press_key/click for editing a field. An agent is left to infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textC

Type text into the focused field (via clipboard, so any language works). press_enter sends it.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
press_enterNo
screenshot_afterNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the clipboard mechanism and the 'any language works' consequence, which is real behavioral context, but says nothing about the default-on screenshot_after side effect, focus preconditions, or clipboard overwrites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, front-loaded with the action and target, no filler. The parenthetical about the clipboard earns its place by explaining the language-agnostic behavior, though the overall brevity leaves gaps.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter GUI-automation tool with no annotations and no output schema, the description is thin: it omits the default screenshot side effect, the focus precondition, and any indication of failure behavior, all of which an agent needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains press_enter's effect but leaves text and especially screenshot_after (default true, a side effect happening without opt-in) completely undocumented, so it only partially covers the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Type) and target (the focused field), plus the implementation detail that it uses the clipboard. It is clear what the tool does, but it never contrasts itself with the sibling set_field, which an agent would plausibly confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. The mention that 'press_enter sends it' hints at a submit flow, but nothing tells the agent to focus a field first, or when set_field would be the better choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.0
    • First observedbatch
    • First observedclick
    • First observeddouble_click
    • First observeddrag
    • First observedhotkey
    • First observedlist_menu
    • First observedmenu
    • First observedopen_affinity
    • First observedpress_key
    • First observedscreenshot
    • First observedscroll
    • First observedset_field
    • First observedtype_text

TDQS

B3.3/5.0

Scored across 13 tools

Disambiguation4/5

Most tools have clearly distinct purposes: click vs double_click, type_text vs press_key vs hotkey are all separable. The only mild overlaps are between menu and list_menu, and between direct click/drag actions and the batch wrapper, but descriptions make the intended use clear.

Naming Consistency4/5

Nearly all names follow snake_case with a verb_noun or clear action pattern (type_text, press_key, list_menu, open_affinity). Minor deviations like the bare nouns 'menu' and 'batch' are still readable and consistent in casing.

Tool Count5/5

13 tools is well-scoped for a GUI-automation server, covering input primitives, menu access, screenshotting, and batch execution without bloat. Each tool earns its place.

Completeness4/5

The surface covers the full input lifecycle (screenshot, click, double-click, drag, type, keys, hotkey, scroll, field, menu, batch, launch). Minor gaps like an explicit hover/move or context-menu helper are workable via click with right button.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to control the Affinity creative suite on macOS through natural language, allowing for automated design tasks, UI interaction, and file operations. It leverages AppleScript and System Events to bridge AI commands with professional creative software.
    25
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables high-speed native macOS automation for Claude by interacting directly with the Accessibility API, AppleScript, and UI trees instead of using screenshots. It allows users to read app states, click elements, and type text semantically across any macOS application.
    8 npm
    11
    MIT
  • F
    license
    B
    quality
    B
    maintenance
    Drives the real Claude Design web app from your editor or agent, enabling creation, iteration, and retrieval of designs generated on your own claude.ai account via automated browser interaction.
    10
    4
    -