Layout Detector MCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Layout Detector MCPanalyze this screenshot with the images in /assets"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Layout Detector MCP
An MCP (Model Context Protocol) server that analyzes webpage screenshots to extract layout information. Given a screenshot and image assets, it finds where each asset appears and calculates spatial relationships - enabling AI assistants to rebuild layouts with proper semantic structure.
Quick Start
Install from GitHub
pip install git+https://github.com/katlis/layout-detector-mcp.gitOr clone and install locally
git clone https://github.com/katlis/layout-detector-mcp.git
cd layout-detector-mcp
pip install .Verify installation:
python3 -c "from layout_detector import server; print('OK')"Related MCP server: mcp-ux-vision
Configuration
Add to your Claude Code MCP settings (~/.claude.json or project .claude/settings.json):
{
"mcpServers": {
"layout-detector": {
"command": "layout-detector-mcp"
}
}
}After adding the configuration, restart Claude Code and run /mcp to verify the server is connected.
The Problem
When an AI assistant looks at a screenshot, it can describe what it sees but cannot extract precise pixel measurements. This makes it difficult to accurately recreate layouts without human intervention or extensive trial-and-error.
The Solution
This MCP server uses computer vision (OpenCV template matching) to:
Find known assets - Locate images within a screenshot with pixel-perfect coordinates
Analyze relationships - Calculate angles, distances, and relative positions
Detect patterns - Identify radial, grid, stacked, sidebar, or freeform layouts
Enable semantic rebuilds - Provide structured data for modern CSS implementation
Tools
analyze_layout
Performs full layout analysis including pattern detection. This is the main tool you'll use.
Parameters:
screenshot_path(string, required): Absolute path to the screenshot imageasset_paths(array of strings, required): Absolute paths to asset images to findthreshold(number, optional): Match confidence 0-1, default 0.8
Returns:
{
"viewport": { "width": 900, "height": 650 },
"pattern": {
"type": "radial",
"confidence": 0.90
},
"radial": {
"center_x": 450,
"center_y": 250,
"center_element": "logo.gif",
"average_radius": 196
},
"elements": [
{
"asset_name": "planet1.gif",
"x": 628,
"y": 89,
"width": 62,
"height": 62,
"angle_degrees": 45.0,
"distance_from_center": 240
}
]
}find_assets_in_screenshot
Locates image assets within a screenshot without layout analysis.
Parameters:
screenshot_path(string, required): Path to the screenshot imageasset_paths(array of strings, required): Paths to asset images to findthreshold(number, optional): Match confidence 0-1, default 0.8
Returns:
{
"found": 5,
"total_assets": 6,
"matches": [
{
"asset_path": "/path/to/logo.png",
"asset_name": "logo.png",
"x": 350,
"y": 200,
"width": 200,
"height": 100,
"center_x": 450,
"center_y": 250,
"confidence": 0.95
}
]
}get_screenshot_info
Get basic screenshot dimensions.
Parameters:
screenshot_path(string, required): Path to the screenshot image
Returns:
{
"path": "/path/to/screenshot.png",
"width": 900,
"height": 650
}Supported Layout Patterns
Pattern | Description | Key Data Returned |
Radial | Elements arranged around a center point | Center element, angles, distances |
Grid | Elements in rows and columns | Row/column positions, gaps |
Stacked | Vertical sections (header/main/footer) | Section names, Y positions |
Sidebar | Two-column with narrow sidebar | Sidebar side, widths |
Freeform | No clear pattern | Raw X/Y coordinates |
Example Usage
Once configured, Claude Code can use these tools:
User: Rebuild this webpage screenshot using the images in /assets
Claude: I'll analyze the layout first using the layout detector.
[Calls analyze_layout tool]
The analysis shows:
- Viewport: 900x650px
- Pattern: Radial (90% confidence)
- Center element: logo.gif at (450, 250)
- 8 elements arranged around the center
- Average distance from center: 196px
I'll implement this using CSS with the logo centered and
other elements positioned using absolute positioning...Supported Image Formats
PNG
JPEG
GIF (including animated - uses first frame)
WebP
BMP
Troubleshooting
"No module named 'cv2'"
OpenCV isn't installed. Run:
pip install opencv-python-headlessMCP server not showing in /mcp
Check your settings file path is correct
Ensure the command path is absolute (for source installs)
Restart Claude Code after changing settings
Run
python3 test_install.pyto verify the package works
Low confidence matches
Try lowering the threshold parameter (default 0.8). Values between 0.6-0.7 may help with compressed or scaled images.
Development
# Install in editable mode with dev dependencies
pip install -e ".[dev]"
# Run tests
pytest
# Test installation
python3 test_install.pyRequirements
Python 3.11+
OpenCV (opencv-python-headless)
NumPy
Pillow
MCP SDK
License
MIT
Available Tools
3 toolsanalyze_layoutA
Analyze the layout of assets in a screenshot. Finds all assets, identifies the center/hero element, calculates relative positions (angle and distance from center), and detects the layout pattern (radial, grid, or freeform). Returns structured data for rebuilding the layout with semantic CSS.
| Name | Required | Description | Default |
|---|---|---|---|
| screenshot_path | Yes | Absolute path to the screenshot image file | |
| asset_paths | Yes | List of absolute paths to asset images to find | |
| threshold | No | Match confidence threshold (0-1). Default 0.8 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes what the tool does (analysis tasks) and the output format ('structured data for rebuilding the layout with semantic CSS'), but lacks details on performance characteristics (e.g., processing time, error handling), resource requirements, or limitations (e.g., supported image formats, size constraints). It does not contradict any annotations, as none are given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and efficiently lists analysis steps in a single sentence, followed by the output use case. It avoids redundancy and each part adds value, though it could be slightly more concise by integrating the output mention with the analysis tasks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (layout analysis with multiple outputs) and lack of annotations or output schema, the description is moderately complete. It covers the analysis tasks and output purpose but omits details on return structure, error conditions, or example outputs. Without an output schema, the agent must rely on the description's vague 'structured data' mention, which is insufficient for full understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema fully documents parameters ('screenshot_path', 'asset_paths', 'threshold'). The description does not add any parameter-specific semantics beyond what the schema provides (e.g., it doesn't explain how 'threshold' affects layout detection or what formats 'asset_paths' should be in). The baseline score of 3 is appropriate as the schema handles parameter documentation adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('analyze', 'finds', 'identifies', 'calculates', 'detects') and resources ('layout of assets in a screenshot'), distinguishing it from sibling tools like 'find_assets_in_screenshot' (which likely only locates assets) and 'get_screenshot_info' (which likely provides basic metadata). It explicitly lists the analysis outputs: finding all assets, identifying the center element, calculating relative positions, and detecting layout patterns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for analyzing asset layouts in screenshots to generate semantic CSS data, but it does not explicitly state when to use this tool versus alternatives like 'find_assets_in_screenshot' or 'get_screenshot_info'. No exclusions or prerequisites are mentioned, leaving the agent to infer context from the tool's name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_assets_in_screenshotA
Find known image assets within a screenshot. Uses template matching to locate each asset and return its position (x, y, width, height). Useful for determining where specific images appear in a webpage screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| screenshot_path | Yes | Absolute path to the screenshot image file | |
| asset_paths | Yes | List of absolute paths to asset images to find | |
| threshold | No | Match confidence threshold (0-1). Default 0.8 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the method ('template matching') and output format, but lacks details on performance characteristics (e.g., speed, accuracy), error handling, or limitations (e.g., image format support, size constraints). It doesn't contradict annotations, but could be more comprehensive for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, followed by method and output details, and ends with a usage context. All three sentences are essential and waste no words, making it highly efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is moderately complete. It covers purpose, method, and usage context, but lacks details on behavioral traits (e.g., what happens if no assets are found) and output specifics beyond position data. For a tool with 3 parameters and no structured output documentation, it could be more thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional parameter semantics beyond what's in the schema (e.g., it doesn't explain 'threshold' beyond the schema's 'Match confidence threshold'). Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Find known image assets within a screenshot'), the method ('Uses template matching'), and the output ('return its position (x, y, width, height)'). It distinguishes from sibling tools like 'analyze_layout' and 'get_screenshot_info' by focusing on asset detection rather than layout analysis or metadata retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('Useful for determining where specific images appear in a webpage screenshot'), which implies it's for image-based asset detection in screenshots. However, it doesn't explicitly state when not to use it or name alternatives among siblings, though the context helps differentiate from tools like 'analyze_layout'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screenshot_infoB
Get basic information about a screenshot image, including its dimensions (width and height in pixels).
| Name | Required | Description | Default |
|---|---|---|---|
| screenshot_path | Yes | Absolute path to the screenshot image file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool retrieves 'basic information' and 'dimensions', which implies a read-only operation, but doesn't clarify if it requires specific file permissions, handles errors (e.g., invalid paths), or has performance constraints. This leaves gaps in understanding the tool's behavior beyond its core function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the tool's purpose and key output ('dimensions (width and height in pixels)'). It is front-loaded with the main action and avoids unnecessary details, making it easy for an agent to parse and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (single parameter, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose but lacks details on usage context, error handling, or output format (beyond mentioning dimensions), which could help the agent use it more effectively. Without annotations or output schema, more behavioral context would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the single parameter 'screenshot_path' as an absolute path. The description adds no additional semantic details about the parameter beyond what the schema provides, such as supported file formats or path validation. With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but doesn't need to heavily.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('Get basic information') and resources ('screenshot image'), including what information is retrieved ('dimensions (width and height in pixels)'). It distinguishes from siblings like 'analyze_layout' and 'find_assets_in_screenshot' by focusing on basic metadata rather than layout analysis or asset detection, though it doesn't explicitly name these alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings ('analyze_layout', 'find_assets_in_screenshot'), such as for quick metadata checks versus detailed analysis. It implies usage for getting dimensions but lacks explicit context, prerequisites, or exclusions, leaving the agent to infer based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose with no ambiguity. analyze_layout focuses on layout analysis and pattern detection, find_assets_in_screenshot handles asset location via template matching, and get_screenshot_info provides basic image metadata. The descriptions clearly differentiate their functions, preventing misselection.
All tool names follow a consistent verb_noun pattern with snake_case throughout (analyze_layout, find_assets_in_screenshot, get_screenshot_info). The naming is predictable and readable, using descriptive verbs that align with each tool's action.
With 3 tools, the count is slightly low but reasonable for the server's purpose of layout detection and screenshot analysis. Each tool earns its place by covering distinct aspects (layout analysis, asset finding, metadata retrieval), though a few more tools might enhance coverage without feeling heavy.
The tool set covers core operations for layout detection and screenshot analysis, including analysis, asset location, and metadata retrieval. Minor gaps exist, such as no tools for modifying or generating layouts, but agents can work around this with the provided structured data for rebuilding layouts.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Desktop and mobile website screenshots plus page context for AI agents and automation workflows.
Capture screenshots, detect visual regressions between page versions, and analyze with AI.
Analyze images from multiple angles to extract detailed insights or quick summaries. Describe visu…
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI assistants to capture screenshots of web pages using automated browser sessions. Supports full-page and element-specific screenshots, device simulation, and JavaScript execution for comprehensive web testing and monitoring.617MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI-powered visual analysis of webpages for UI/UX assessment, including screenshot capture, element detection, accessibility auditing, and comprehensive JSON reporting.1
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.8MIT
- AlicenseNot gradedqualityBmaintenanceMCP-first UI/UX review layer for AI-generated frontends. Enables reviewing web pages via URL, capturing screenshots, extracting layout metrics, and generating structured repair plans for agents.22Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/katlis/layout-detector-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server