viewpo-mcp
Provides tools for extracting computed CSS styles and comparing layout properties across different viewport widths to identify and resolve responsive design issues.
Integrates with the Viewpo macOS app to enable AI assistants to capture multi-viewport screenshots and extract structured DOM layout data from web pages.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@viewpo-mcpTake a screenshot of localhost:3000 at mobile and desktop widths"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
viewpo-mcp
MCP server that gives AI coding assistants eyes into responsive viewport rendering via the Viewpo macOS app.
Capture multi-viewport screenshots, extract DOM layout trees, and compare responsive behaviour across breakpoints — all from your AI assistant.
What it does
AI coding assistants (Claude Code, Cursor, Windsurf, etc.) are blind when working on frontend UI. They write CSS and HTML but can't see the visual result. This creates a slow feedback loop:
AI writes code → you check → you describe what's wrong → AI fixes → repeat
viewpo-mcp closes that loop. The AI can now:
Screenshot a URL at multiple viewport widths simultaneously
Extract the DOM layout tree with bounding rects and computed styles
Compare layouts at two different viewport widths to find responsive issues
Requirements
Viewpo macOS app with MCP Bridge enabled (Settings → MCP Bridge → Start)
Node.js 20+
Setup
1. Install
npm install -g viewpo-mcpOr run directly with npx (no install needed):
npx viewpo-mcp2. Enable the MCP Bridge in Viewpo
Open Viewpo on your Mac
Go to Settings → MCP Bridge
Click Start to enable the bridge server
Copy the auth token (click the copy button next to the token)
3. Configure your AI assistant
Claude Code
Add to your MCP config (~/.claude/mcp.json or project .mcp.json):
{
"mcpServers": {
"viewpo": {
"command": "npx",
"args": ["-y", "viewpo-mcp"],
"env": {
"VIEWPO_AUTH_TOKEN": "<paste token from step 2>"
}
}
}
}Cursor
Add to your Cursor MCP settings (.cursor/mcp.json):
{
"mcpServers": {
"viewpo": {
"command": "npx",
"args": ["-y", "viewpo-mcp"],
"env": {
"VIEWPO_AUTH_TOKEN": "<paste token from step 2>"
}
}
}
}Other MCP-compatible assistants
Any assistant supporting the Model Context Protocol can use viewpo-mcp. Set the command to npx -y viewpo-mcp and provide the VIEWPO_AUTH_TOKEN environment variable.
Tools
viewpo_screenshot
Capture screenshots of a URL at one or more viewport widths. Returns base64 JPEG images.
url: "https://example.com" (required)
viewports: [{ width: 375, name: "phone" }, (optional, defaults to 1920px desktop)
{ width: 1920, name: "desktop" }]Common widths: 375 (phone), 820 (tablet), 1920 (desktop).
viewpo_get_layout_map
Extract the DOM layout tree at a given viewport width. Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles.
url: "https://example.com" (required)
viewport: 1920 (optional, default 1920)
selector: ".main-content" (optional, scope to subtree)viewpo_compare_viewports
Compare the layout of a URL at two different viewport widths. Returns elements whose size or CSS styles differ between the two viewports.
url: "https://example.com" (required)
viewport_a: 375 (required)
viewport_b: 1920 (required)
selector: ".hero" (optional, scope to subtree)Environment variables
Variable | Required | Default | Description |
| Yes | — | Bearer token from Viewpo Settings → MCP Bridge |
| No |
| Port the Viewpo bridge server listens on |
How it works
AI Assistant (Claude Code / Cursor / Windsurf)
| stdio (MCP protocol)
v
viewpo-mcp (this package)
| HTTP localhost:9847
v
Viewpo macOS app (NWListener)
| WKWebView rendering
v
Headless browser pool (off-screen, up to 4 concurrent pages)The Viewpo app runs a local HTTP server on localhost:9847. This MCP server translates tool calls into HTTP requests to that bridge. The app loads pages in headless WKWebView instances with real CSS viewport simulation — media queries fire at the target width, not the physical screen width.
Key advantage: WKWebView bypasses X-Frame-Options entirely. Any website loads, regardless of security headers that would block iframe-based tools.
Example workflow
You: "The hero section looks broken on mobile"
AI: 1. viewpo_screenshot("https://preview.example.com", [{width: 375}, {width: 1920}])
→ Sees the layout at both sizes
2. viewpo_compare_viewports("https://preview.example.com", 375, 1920)
→ Gets structured diff: ".hero img" is 1200px wide on mobile (overflowing)
3. Fixes the CSS
4. viewpo_screenshot("https://preview.example.com", [{width: 375}])
→ Verifies the fixDevelopment
git clone https://github.com/littlebearapps/viewpo-mcp.git
cd viewpo-mcp
npm install
npm run build # compile TypeScript
npm run dev # watch modeLicence
MIT - see LICENSE.
Links
Viewpo app - The macOS/iOS app
Model Context Protocol - MCP specification
GitHub - Source code
Available Tools
3 toolsviewpo_compare_viewportsA
Compare the layout of a URL at two different viewport widths. Returns a list of elements whose size or CSS styles differ between the two viewports. Use this to find responsive design issues.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to compare | |
| viewport_a | Yes | First viewport width in CSS pixels (e.g. 375 for phone) | |
| viewport_b | Yes | Second viewport width in CSS pixels (e.g. 1920 for desktop) | |
| selector | No | CSS selector to scope comparison to a subtree |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the core behavior (comparing layouts and returning differences) and the purpose (finding responsive design issues), but lacks details on potential side effects, performance characteristics, error handling, or authentication needs. For a tool with no annotations, this is adequate but leaves gaps in behavioral understanding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise and front-loaded, consisting of two sentences that efficiently convey the tool's purpose, output, and usage. Every sentence earns its place: the first explains what the tool does and returns, while the second provides the usage context. There is zero waste or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (comparing layouts across viewports) and the absence of both annotations and an output schema, the description is minimally complete. It covers the core functionality and use case but lacks details on output format, error conditions, or limitations. Without structured fields to rely on, the description should ideally provide more context, but it meets basic adequacy.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, providing clear documentation for all parameters (url, viewport_a, viewport_b, selector). The description adds minimal value beyond the schema, only implicitly reinforcing the purpose of viewport comparison. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description doesn't significantly enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Compare the layout of a URL at two different viewport widths'), identifies the resource (URL layout), and distinguishes from siblings by focusing on responsive design comparison rather than layout mapping or screenshot capture. It explicitly mentions the outcome ('Returns a list of elements whose size or CSS styles differ') and use case ('find responsive design issues').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('Use this to find responsive design issues'), which implicitly differentiates it from siblings like viewpo_get_layout_map (likely for mapping layouts) and viewpo_screenshot (for capturing visuals). However, it does not explicitly state when NOT to use it or name specific alternatives, keeping it at a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
viewpo_get_layout_mapA
Extract the DOM layout tree of a URL at a given viewport width. Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles. Use this to understand page structure and find layout issues.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to inspect | |
| viewport | No | Viewport width in CSS pixels (default: 1920) | |
| selector | No | CSS selector to scope the layout map to a subtree (e.g. ".main-content", "#hero") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the output ('Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles') and purpose, but lacks details on performance, error handling, or authentication needs, which are relevant for a tool that likely involves network requests.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by a concise usage guideline. Every sentence earns its place by adding value, with no redundant or verbose language, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is mostly complete. It explains what the tool does and its output format, but could benefit from mentioning potential limitations (e.g., timeouts, JavaScript-rendered content) to fully guide an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional parameter semantics beyond what the schema provides, such as examples or edge cases, resulting in a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Extract the DOM layout tree'), target resource ('of a URL'), and scope ('at a given viewport width'), distinguishing it from sibling tools like viewpo_compare_viewports and viewpo_screenshot by focusing on structural analysis rather than comparison or visual capture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('to understand page structure and find layout issues'), but does not explicitly mention when not to use it or name alternatives among the sibling tools, such as using viewpo_screenshot for visual issues instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
viewpo_screenshotA
Capture screenshots of a URL at one or more viewport widths. Returns base64 JPEG images. Use this to SEE what a webpage looks like at different screen sizes.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to screenshot | |
| viewports | No | Viewports to capture. Defaults to desktop (1920px) if omitted. Common widths: 375 (phone), 820 (tablet), 1920 (desktop). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the core behavior (capturing screenshots and returning base64 JPEGs) but lacks details about potential side effects (e.g., network requests, rate limits, authentication needs, or error handling). The description doesn't contradict annotations, but it's incomplete for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, with two sentences that efficiently convey the tool's purpose and usage. Every word earns its place, and there's no redundant or verbose language, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (capturing screenshots with configurable viewports), no annotations, and no output schema, the description is somewhat incomplete. It covers the basic purpose and output format but lacks details on behavioral aspects like error cases, performance implications, or how the output is structured. This is adequate but leaves clear gaps for an agent to fully understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already fully documents the parameters (url and viewports). The description adds minimal value beyond the schema by implying the tool's purpose relates to viewport widths, but it doesn't provide additional semantic context or usage examples for the parameters. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Capture screenshots'), the resource ('a URL at one or more viewport widths'), and the output format ('base64 JPEG images'). It distinguishes from siblings by focusing on visual capture rather than comparison or layout analysis, making the purpose immediately understandable and distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('to SEE what a webpage looks like at different screen sizes'), which implicitly suggests it's for visual inspection. However, it doesn't explicitly state when NOT to use it or name alternatives like the sibling tools, leaving some guidance gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: compare_viewports for detecting differences, get_layout_map for structural analysis, and screenshot for visual capture. There is no overlap in functionality, making it easy for an agent to select the right tool based on the task.
All tool names follow a consistent 'viewpo_' prefix with descriptive snake_case suffixes (compare_viewports, get_layout_map, screenshot). This predictable pattern enhances readability and reduces confusion.
With 3 tools, the server is well-scoped for its viewport testing domain, but it feels slightly thin. While each tool is useful, additional tools like viewpo_get_metrics or viewpo_simulate_device might enhance completeness without being overwhelming.
The tools cover core viewport analysis tasks: layout comparison, structure extraction, and visual capture. However, there are minor gaps, such as missing tools for performance metrics or device simulation, which agents might need to work around for advanced testing scenarios.
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Capture screenshots, detect visual regressions between page versions, and analyze with AI.
Desktop and mobile website screenshots plus page context for AI agents and automation workflows.
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Validate HTML/CSS, audit SEO and JSON-LD, check links, and capture responsive screenshots.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/littlebearapps/viewpo-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server