Skip to main content
Glama

Layout Detector MCP

An MCP (Model Context Protocol) server that analyzes webpage screenshots to extract layout information. Given a screenshot and image assets, it finds where each asset appears and calculates spatial relationships - enabling AI assistants to rebuild layouts with proper semantic structure.

Quick Start

Install from GitHub

pip install git+https://github.com/katlis/layout-detector-mcp.git

Or clone and install locally

git clone https://github.com/katlis/layout-detector-mcp.git
cd layout-detector-mcp
pip install .

Verify installation:

python3 -c "from layout_detector import server; print('OK')"

Related MCP server: mcp-ux-vision

Configuration

Add to your Claude Code MCP settings (~/.claude.json or project .claude/settings.json):

{
  "mcpServers": {
    "layout-detector": {
      "command": "layout-detector-mcp"
    }
  }
}

After adding the configuration, restart Claude Code and run /mcp to verify the server is connected.

The Problem

When an AI assistant looks at a screenshot, it can describe what it sees but cannot extract precise pixel measurements. This makes it difficult to accurately recreate layouts without human intervention or extensive trial-and-error.

The Solution

This MCP server uses computer vision (OpenCV template matching) to:

  1. Find known assets - Locate images within a screenshot with pixel-perfect coordinates

  2. Analyze relationships - Calculate angles, distances, and relative positions

  3. Detect patterns - Identify radial, grid, stacked, sidebar, or freeform layouts

  4. Enable semantic rebuilds - Provide structured data for modern CSS implementation

Tools

analyze_layout

Performs full layout analysis including pattern detection. This is the main tool you'll use.

Parameters:

  • screenshot_path (string, required): Absolute path to the screenshot image

  • asset_paths (array of strings, required): Absolute paths to asset images to find

  • threshold (number, optional): Match confidence 0-1, default 0.8

Returns:

{
  "viewport": { "width": 900, "height": 650 },
  "pattern": {
    "type": "radial",
    "confidence": 0.90
  },
  "radial": {
    "center_x": 450,
    "center_y": 250,
    "center_element": "logo.gif",
    "average_radius": 196
  },
  "elements": [
    {
      "asset_name": "planet1.gif",
      "x": 628,
      "y": 89,
      "width": 62,
      "height": 62,
      "angle_degrees": 45.0,
      "distance_from_center": 240
    }
  ]
}

find_assets_in_screenshot

Locates image assets within a screenshot without layout analysis.

Parameters:

  • screenshot_path (string, required): Path to the screenshot image

  • asset_paths (array of strings, required): Paths to asset images to find

  • threshold (number, optional): Match confidence 0-1, default 0.8

Returns:

{
  "found": 5,
  "total_assets": 6,
  "matches": [
    {
      "asset_path": "/path/to/logo.png",
      "asset_name": "logo.png",
      "x": 350,
      "y": 200,
      "width": 200,
      "height": 100,
      "center_x": 450,
      "center_y": 250,
      "confidence": 0.95
    }
  ]
}

get_screenshot_info

Get basic screenshot dimensions.

Parameters:

  • screenshot_path (string, required): Path to the screenshot image

Returns:

{
  "path": "/path/to/screenshot.png",
  "width": 900,
  "height": 650
}

Supported Layout Patterns

Pattern

Description

Key Data Returned

Radial

Elements arranged around a center point

Center element, angles, distances

Grid

Elements in rows and columns

Row/column positions, gaps

Stacked

Vertical sections (header/main/footer)

Section names, Y positions

Sidebar

Two-column with narrow sidebar

Sidebar side, widths

Freeform

No clear pattern

Raw X/Y coordinates

Example Usage

Once configured, Claude Code can use these tools:

User: Rebuild this webpage screenshot using the images in /assets

Claude: I'll analyze the layout first using the layout detector.

[Calls analyze_layout tool]

The analysis shows:
- Viewport: 900x650px
- Pattern: Radial (90% confidence)
- Center element: logo.gif at (450, 250)
- 8 elements arranged around the center
- Average distance from center: 196px

I'll implement this using CSS with the logo centered and
other elements positioned using absolute positioning...

Supported Image Formats

  • PNG

  • JPEG

  • GIF (including animated - uses first frame)

  • WebP

  • BMP

Troubleshooting

"No module named 'cv2'"

OpenCV isn't installed. Run:

pip install opencv-python-headless

MCP server not showing in /mcp

  1. Check your settings file path is correct

  2. Ensure the command path is absolute (for source installs)

  3. Restart Claude Code after changing settings

  4. Run python3 test_install.py to verify the package works

Low confidence matches

Try lowering the threshold parameter (default 0.8). Values between 0.6-0.7 may help with compressed or scaled images.

Development

# Install in editable mode with dev dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Test installation
python3 test_install.py

Requirements

  • Python 3.11+

  • OpenCV (opencv-python-headless)

  • NumPy

  • Pillow

  • MCP SDK

License

MIT

Available Tools

3 tools
analyze_layoutA

Analyze the layout of assets in a screenshot. Finds all assets, identifies the center/hero element, calculates relative positions (angle and distance from center), and detects the layout pattern (radial, grid, or freeform). Returns structured data for rebuilding the layout with semantic CSS.

ParametersJSON Schema
NameRequiredDescriptionDefault
screenshot_pathYesAbsolute path to the screenshot image file
asset_pathsYesList of absolute paths to asset images to find
thresholdNoMatch confidence threshold (0-1). Default 0.8

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It describes what the tool does (analysis tasks) and the output format ('structured data for rebuilding the layout with semantic CSS'), but lacks details on performance characteristics (e.g., processing time, error handling), resource requirements, or limitations (e.g., supported image formats, size constraints). It does not contradict any annotations, as none are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and efficiently lists analysis steps in a single sentence, followed by the output use case. It avoids redundancy and each part adds value, though it could be slightly more concise by integrating the output mention with the analysis tasks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (layout analysis with multiple outputs) and lack of annotations or output schema, the description is moderately complete. It covers the analysis tasks and output purpose but omits details on return structure, error conditions, or example outputs. Without an output schema, the agent must rely on the description's vague 'structured data' mention, which is insufficient for full understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents parameters ('screenshot_path', 'asset_paths', 'threshold'). The description does not add any parameter-specific semantics beyond what the schema provides (e.g., it doesn't explain how 'threshold' affects layout detection or what formats 'asset_paths' should be in). The baseline score of 3 is appropriate as the schema handles parameter documentation adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('analyze', 'finds', 'identifies', 'calculates', 'detects') and resources ('layout of assets in a screenshot'), distinguishing it from sibling tools like 'find_assets_in_screenshot' (which likely only locates assets) and 'get_screenshot_info' (which likely provides basic metadata). It explicitly lists the analysis outputs: finding all assets, identifying the center element, calculating relative positions, and detecting layout patterns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for analyzing asset layouts in screenshots to generate semantic CSS data, but it does not explicitly state when to use this tool versus alternatives like 'find_assets_in_screenshot' or 'get_screenshot_info'. No exclusions or prerequisites are mentioned, leaving the agent to infer context from the tool's name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_assets_in_screenshotA

Find known image assets within a screenshot. Uses template matching to locate each asset and return its position (x, y, width, height). Useful for determining where specific images appear in a webpage screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
screenshot_pathYesAbsolute path to the screenshot image file
asset_pathsYesList of absolute paths to asset images to find
thresholdNoMatch confidence threshold (0-1). Default 0.8

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the method ('template matching') and output format, but lacks details on performance characteristics (e.g., speed, accuracy), error handling, or limitations (e.g., image format support, size constraints). It doesn't contradict annotations, but could be more comprehensive for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, followed by method and output details, and ends with a usage context. All three sentences are essential and waste no words, making it highly efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is moderately complete. It covers purpose, method, and usage context, but lacks details on behavioral traits (e.g., what happens if no assets are found) and output specifics beyond position data. For a tool with 3 parameters and no structured output documentation, it could be more thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional parameter semantics beyond what's in the schema (e.g., it doesn't explain 'threshold' beyond the schema's 'Match confidence threshold'). Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Find known image assets within a screenshot'), the method ('Uses template matching'), and the output ('return its position (x, y, width, height)'). It distinguishes from sibling tools like 'analyze_layout' and 'get_screenshot_info' by focusing on asset detection rather than layout analysis or metadata retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('Useful for determining where specific images appear in a webpage screenshot'), which implies it's for image-based asset detection in screenshots. However, it doesn't explicitly state when not to use it or name alternatives among siblings, though the context helps differentiate from tools like 'analyze_layout'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screenshot_infoB

Get basic information about a screenshot image, including its dimensions (width and height in pixels).

ParametersJSON Schema
NameRequiredDescriptionDefault
screenshot_pathYesAbsolute path to the screenshot image file

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool retrieves 'basic information' and 'dimensions', which implies a read-only operation, but doesn't clarify if it requires specific file permissions, handles errors (e.g., invalid paths), or has performance constraints. This leaves gaps in understanding the tool's behavior beyond its core function.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that efficiently conveys the tool's purpose and key output ('dimensions (width and height in pixels)'). It is front-loaded with the main action and avoids unnecessary details, making it easy for an agent to parse and understand quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (single parameter, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose but lacks details on usage context, error handling, or output format (beyond mentioning dimensions), which could help the agent use it more effectively. Without annotations or output schema, more behavioral context would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting the single parameter 'screenshot_path' as an absolute path. The description adds no additional semantic details about the parameter beyond what the schema provides, such as supported file formats or path validation. With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but doesn't need to heavily.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Get basic information') and resources ('screenshot image'), including what information is retrieved ('dimensions (width and height in pixels)'). It distinguishes from siblings like 'analyze_layout' and 'find_assets_in_screenshot' by focusing on basic metadata rather than layout analysis or asset detection, though it doesn't explicitly name these alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings ('analyze_layout', 'find_assets_in_screenshot'), such as for quick metadata checks versus detailed analysis. It implies usage for getting dimensions but lacks explicit context, prerequisites, or exclusions, leaving the agent to infer based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.7/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no ambiguity. analyze_layout focuses on layout analysis and pattern detection, find_assets_in_screenshot handles asset location via template matching, and get_screenshot_info provides basic image metadata. The descriptions clearly differentiate their functions, preventing misselection.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case throughout (analyze_layout, find_assets_in_screenshot, get_screenshot_info). The naming is predictable and readable, using descriptive verbs that align with each tool's action.

Tool Count4/5

With 3 tools, the count is slightly low but reasonable for the server's purpose of layout detection and screenshot analysis. Each tool earns its place by covering distinct aspects (layout analysis, asset finding, metadata retrieval), though a few more tools might enhance coverage without feeling heavy.

Completeness4/5

The tool set covers core operations for layout detection and screenshot analysis, including analysis, asset location, and metadata retrieval. Minor gaps exist, such as no tools for modifying or generating layouts, but agents can work around this with the provided structured data for rebuilding layouts.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/katlis/layout-detector-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server