Skip to main content
Glama
w01fgang

MCP Design Comparison Server

by w01fgang

MCP Design Comparison Server

An MCP (Model Context Protocol) server that allows LLMs to compare design screenshots with implementation screenshots using pixelmatch. Supports multiple image formats including PNG, JPEG, WebP, GIF, TIFF, and SVG (rasterized via librsvg). This tool helps identify visual discrepancies between design mockups and actual implementation.

Features

  • Multi-Format Support: Works with PNG, JPEG, WebP, GIF, TIFF, and SVG images

  • Screenshot Comparison: Compare two images pixel-by-pixel

  • SSIM Score: Structural-similarity metric (0–1) that tracks perceived difference, not just raw pixel deltas

  • Auto-Resize: Mismatched resolutions are reconciled automatically — the implementation is scaled to the design's dimensions. resize_fit controls how (contain preserves aspect ratio by default; fill/cover available); toggle off with auto_resize

  • Ignore Regions: Exclude rectangles (e.g. timestamps, avatars) from the comparison so dynamic content doesn't count

  • Vector (SVG) Input: Compare a design PNG against an SVG icon (or SVG vs SVG) — SVGs are rasterized at a configurable density (svg_density, default 288 dpi) with no pre-render step

  • Diff Localization: A diff bounding box and a 3x3 heat grid report where differences cluster, not just how many pixels differ

  • Assertion Gate: Optional max_difference_percentage flips the result to an error when exceeded — a CI regression gate without parsing text

  • Visual Diff Output: Generate a highlighted diff image showing differences

  • Detailed Metrics: Get total pixels, different pixels, and percentage difference

  • Configurable Threshold: Adjust sensitivity of the comparison

  • Base64 Support: Return diff images as base64 or save to file

Related MCP server: MCP Component Review

Installation

For End Users

Install globally via npm:

npm install -g mcp-design-comparison

Or using npx (no installation required):

npx mcp-design-comparison

For Development

git clone https://github.com/w01fgang/mcp-design-comparison.git
cd mcp-design-comparison
npm install
npm run build

Usage

As an MCP Server

Add this server to your MCP client configuration. The server runs on stdio and provides a single tool:

Tool: compare_design

Compare a design screenshot with an implementation screenshot.

Parameters:

  • design_path (string, required): Path to the design screenshot (supports PNG, JPEG, WebP, GIF, TIFF, SVG)

  • implementation_path (string, required): Path to the implementation screenshot (supports PNG, JPEG, WebP, GIF, TIFF, SVG)

  • output_diff_path (string, optional): Path to save the diff image (always saved as PNG). If not provided, the diff image will be returned as base64

  • threshold (number, optional): Matching threshold (0-1). Smaller values make the comparison more sensitive. Default is 0.1

  • auto_resize (boolean, optional): When the two screenshots differ in resolution, scale the implementation to the design's dimensions instead of erroring. Default is true. Set false to require identical dimensions.

  • resize_fit (string, optional): How to scale the implementation when dimensions differ — contain (default, preserves aspect ratio and letterboxes), fill (stretches to exact dimensions), or cover (preserves aspect ratio and crops overflow).

  • ignore_regions (array, optional): Rectangles { x, y, width, height } in design-space coordinates to exclude from the comparison (e.g. dynamic content). Excluded pixels count toward neither the diff nor the percentage denominator. For an SVG design, design space is the rendered raster (intrinsic size × svg_density / 72 — e.g. a 36×36 SVG at the default density 288 is 144×144), not the intrinsic SVG size; the same applies to the returned diff bounds and heat grid.

  • svg_density (number, optional): Rasterization density (DPI) for SVG inputs. Higher = crisper vector render before comparison. Default 288 (4x the 72dpi baseline). Aimed at small assets (icons/logos); lower it for large vector art.

  • localize (boolean, optional): If true (default), include a diff bounding box and a coarse 3x3 heat grid in the result, showing where differences cluster. Set false to skip the extra pass.

  • max_difference_percentage (number, optional): If set and the difference percentage exceeds it (strictly greater), the call returns an error (isError: true) while still producing the diff artifact. Use as a CI gate against a golden image. Omit for report-only.

Returns:

  • Total number of pixels (excluding any ignored regions)

  • Number of different pixels

  • Percentage difference

  • SSIM score (0–1, where 1.0 means identical)

  • Diff bounding box and 3x3 heat grid (when localize is on and differences exist)

  • Diff image (as file or base64)

Configuration Example

Add to your MCP settings file:

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):

{
  "mcpServers": {
    "design-comparison": {
      "command": "npx",
      "args": ["-y", "mcp-design-comparison"]
    }
  }
}

Cursor (~/Library/Application Support/Cursor/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json on macOS):

{
  "mcpServers": {
    "design-comparison": {
      "command": "npx",
      "args": ["-y", "mcp-design-comparison"]
    }
  }
}

After adding the configuration, restart Claude Desktop or Cursor.

How It Works

  1. Loads both images into memory (any supported format → raw RGBA)

  2. If dimensions differ, scales the implementation to the design's dimensions using resize_fit (default contain, aspect-preserving; unless auto_resize is false, which errors instead)

  3. Masks out any ignore_regions in both images so excluded pixels don't count

  4. Uses pixelmatch to compare pixel-by-pixel, computes an SSIM score, and (when localize is on) locates where differences cluster (bounding box + 3x3 heat grid)

  5. Generates a diff image highlighting differences in pink

  6. Returns statistics, the SSIM score, and the diff image

Example Use Cases

  • Design QA: Verify that implementation matches design mockups

  • Regression Testing: Compare screenshots before and after changes

  • Cross-browser Testing: Compare renders across different browsers

  • Responsive Design: Compare layouts at different breakpoints

Testing

Run Automated Tests

npm test

Manual Testing with Your Screenshots

Use the included test script:

node test-manual.mjs design.png implementation.png [output-diff.png]

Requirements

  • Node.js 18+

  • Images in a supported format (PNG, JPEG, WebP, GIF, TIFF, or SVG); differing resolutions are auto-resized unless auto_resize is disabled

License

MIT

Available Tools

1 tool
compare_designA

Compare a design screenshot with an implementation screenshot using pixelmatch and SSIM. Supports PNG, JPEG, WebP, GIF, TIFF, and SVG inputs (SVGs are rasterized at svg_density). Returns the number and percentage of different pixels, a structural-similarity (SSIM) score, a diff bounding box and heat grid showing where differences cluster, and optionally outputs a diff image highlighting the differences.

ParametersJSON Schema
NameRequiredDescriptionDefault
localizeNoIf true (default), include a diff bounding box and a coarse per-cell heat grid in the result, showing where differences cluster. Set false to skip the extra pass.
thresholdNoMatching threshold (0-1). Smaller values make the comparison more sensitive. Default is 0.1.
resize_fitNoHow to scale the implementation when dimensions differ. 'contain' (default) preserves aspect ratio and letterboxes; 'fill' stretches to the exact dimensions; 'cover' preserves aspect ratio and crops the overflow.contain
auto_resizeNoIf true (default), the implementation screenshot is scaled to the design's dimensions when they differ, instead of failing. Set false to require identical dimensions.
design_pathYesPath to the design screenshot (supports PNG, JPEG, WebP, GIF, TIFF, SVG)
svg_densityNoRasterization density (DPI) for SVG inputs. Higher = crisper vector render before comparison. Default 288 (4x the 72dpi baseline). Aimed at small assets (icons/logos); lower it for large vector art.
ignore_regionsNoRectangles (design-space coordinates) to exclude from the comparison, e.g. dynamic content like timestamps or avatars. Excluded pixels count toward neither the diff nor the percentage denominator.
output_diff_pathNoOptional path to save the diff image. If not provided, the diff image will be returned as base64.
implementation_pathYesPath to the implementation screenshot (supports PNG, JPEG, WebP, GIF, TIFF, SVG)
max_difference_percentageNoIf set and the difference percentage exceeds it, the call returns an error (isError: true). Use as a CI gate against a golden image. Omit for report-only.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It reveals the comparison algorithms, SVG rasterization behavior, and diff image option. However, it omits potential side effects (e.g., file system reads only), error conditions, or performance impact, leaving some behavioral aspects implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, somewhat long sentence. It is clear and contains no redundancy, but could be more scannable with bullet points or shorter sentences. It is adequate but not optimally structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description reasonably covers return values (diff counts, SSIM, bounding box, heat grid, optional diff image) and key parameter behaviors. It adequately informs the user of what to expect, though some minor details (e.g., color handling) are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds valuable context beyond schema field descriptions, such as explaining how ignore_regions affects counts and that svg_density is aimed at small assets. It enhances understanding of parameter behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares design and implementation screenshots using pixelmatch and SSIM, lists supported formats, and details return metrics. It uses a specific verb and resource, distinguishing the tool's purpose effectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but does not provide explicit guidance on when to use it versus alternatives or when not to use it. Since no sibling tools are given, this is acceptable but lacks proactive context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.7.0
    • First observedcompare_design

TDQS

A3.9/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no risk of confusion. The tool has a clear, singular purpose.

Naming Consistency5/5

With a single tool, naming is trivially consistent. The name 'compare_design' follows a clear verb_noun pattern.

Tool Count3/5

One tool is borderline for typical server scopes. While it fits a narrow purpose, the server could potentially benefit from a few more tools (e.g., for configuration or batch comparison).

Completeness4/5

The tool covers the core comparison task well, supporting multiple formats and providing useful outputs. Minor gaps exist, such as no ability to adjust comparison thresholds or handle folders of images.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers