Skip to main content
Glama

viewpo-mcp

MCP server that gives AI coding assistants eyes into responsive viewport rendering via the Viewpo macOS app.

Capture multi-viewport screenshots, extract DOM layout trees, and compare responsive behaviour across breakpoints — all from your AI assistant.

What it does

AI coding assistants (Claude Code, Cursor, Windsurf, etc.) are blind when working on frontend UI. They write CSS and HTML but can't see the visual result. This creates a slow feedback loop:

AI writes code → you check → you describe what's wrong → AI fixes → repeat

viewpo-mcp closes that loop. The AI can now:

  1. Screenshot a URL at multiple viewport widths simultaneously

  2. Extract the DOM layout tree with bounding rects and computed styles

  3. Compare layouts at two different viewport widths to find responsive issues

Requirements

  • Viewpo macOS app with MCP Bridge enabled (Settings → MCP Bridge → Start)

  • Node.js 20+

Setup

1. Install

npm install -g viewpo-mcp

Or run directly with npx (no install needed):

npx viewpo-mcp

2. Enable the MCP Bridge in Viewpo

  1. Open Viewpo on your Mac

  2. Go to Settings → MCP Bridge

  3. Click Start to enable the bridge server

  4. Copy the auth token (click the copy button next to the token)

3. Configure your AI assistant

Claude Code

Add to your MCP config (~/.claude/mcp.json or project .mcp.json):

{
  "mcpServers": {
    "viewpo": {
      "command": "npx",
      "args": ["-y", "viewpo-mcp"],
      "env": {
        "VIEWPO_AUTH_TOKEN": "<paste token from step 2>"
      }
    }
  }
}

Cursor

Add to your Cursor MCP settings (.cursor/mcp.json):

{
  "mcpServers": {
    "viewpo": {
      "command": "npx",
      "args": ["-y", "viewpo-mcp"],
      "env": {
        "VIEWPO_AUTH_TOKEN": "<paste token from step 2>"
      }
    }
  }
}

Other MCP-compatible assistants

Any assistant supporting the Model Context Protocol can use viewpo-mcp. Set the command to npx -y viewpo-mcp and provide the VIEWPO_AUTH_TOKEN environment variable.

Tools

viewpo_screenshot

Capture screenshots of a URL at one or more viewport widths. Returns base64 JPEG images.

url:       "https://example.com"           (required)
viewports: [{ width: 375, name: "phone" }, (optional, defaults to 1920px desktop)
            { width: 1920, name: "desktop" }]

Common widths: 375 (phone), 820 (tablet), 1920 (desktop).

viewpo_get_layout_map

Extract the DOM layout tree at a given viewport width. Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles.

url:      "https://example.com"    (required)
viewport: 1920                     (optional, default 1920)
selector: ".main-content"          (optional, scope to subtree)

viewpo_compare_viewports

Compare the layout of a URL at two different viewport widths. Returns elements whose size or CSS styles differ between the two viewports.

url:        "https://example.com"  (required)
viewport_a: 375                    (required)
viewport_b: 1920                   (required)
selector:   ".hero"                (optional, scope to subtree)

Environment variables

Variable

Required

Default

Description

VIEWPO_AUTH_TOKEN

Yes

Bearer token from Viewpo Settings &rarr; MCP Bridge

VIEWPO_PORT

No

9847

Port the Viewpo bridge server listens on

How it works

AI Assistant (Claude Code / Cursor / Windsurf)
    | stdio (MCP protocol)
    v
viewpo-mcp (this package)
    | HTTP localhost:9847
    v
Viewpo macOS app (NWListener)
    | WKWebView rendering
    v
Headless browser pool (off-screen, up to 4 concurrent pages)

The Viewpo app runs a local HTTP server on localhost:9847. This MCP server translates tool calls into HTTP requests to that bridge. The app loads pages in headless WKWebView instances with real CSS viewport simulation — media queries fire at the target width, not the physical screen width.

Key advantage: WKWebView bypasses X-Frame-Options entirely. Any website loads, regardless of security headers that would block iframe-based tools.

Example workflow

You: "The hero section looks broken on mobile"

AI:  1. viewpo_screenshot("https://preview.example.com", [{width: 375}, {width: 1920}])
        → Sees the layout at both sizes
     2. viewpo_compare_viewports("https://preview.example.com", 375, 1920)
        → Gets structured diff: ".hero img" is 1200px wide on mobile (overflowing)
     3. Fixes the CSS
     4. viewpo_screenshot("https://preview.example.com", [{width: 375}])
        → Verifies the fix

Development

git clone https://github.com/littlebearapps/viewpo-mcp.git
cd viewpo-mcp
npm install
npm run build     # compile TypeScript
npm run dev       # watch mode

Licence

MIT - see LICENSE.

Available Tools

3 tools
viewpo_compare_viewportsA

Compare the layout of a URL at two different viewport widths. Returns a list of elements whose size or CSS styles differ between the two viewports. Use this to find responsive design issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to compare
viewport_aYesFirst viewport width in CSS pixels (e.g. 375 for phone)
viewport_bYesSecond viewport width in CSS pixels (e.g. 1920 for desktop)
selectorNoCSS selector to scope comparison to a subtree

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the core behavior (comparing layouts and returning differences) and the purpose (finding responsive design issues), but lacks details on potential side effects, performance characteristics, error handling, or authentication needs. For a tool with no annotations, this is adequate but leaves gaps in behavioral understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly concise and front-loaded, consisting of two sentences that efficiently convey the tool's purpose, output, and usage. Every sentence earns its place: the first explains what the tool does and returns, while the second provides the usage context. There is zero waste or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (comparing layouts across viewports) and the absence of both annotations and an output schema, the description is minimally complete. It covers the core functionality and use case but lacks details on output format, error conditions, or limitations. Without structured fields to rely on, the description should ideally provide more context, but it meets basic adequacy.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, providing clear documentation for all parameters (url, viewport_a, viewport_b, selector). The description adds minimal value beyond the schema, only implicitly reinforcing the purpose of viewport comparison. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description doesn't significantly enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Compare the layout of a URL at two different viewport widths'), identifies the resource (URL layout), and distinguishes from siblings by focusing on responsive design comparison rather than layout mapping or screenshot capture. It explicitly mentions the outcome ('Returns a list of elements whose size or CSS styles differ') and use case ('find responsive design issues').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('Use this to find responsive design issues'), which implicitly differentiates it from siblings like viewpo_get_layout_map (likely for mapping layouts) and viewpo_screenshot (for capturing visuals). However, it does not explicitly state when NOT to use it or name specific alternatives, keeping it at a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

viewpo_get_layout_mapA

Extract the DOM layout tree of a URL at a given viewport width. Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles. Use this to understand page structure and find layout issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to inspect
viewportNoViewport width in CSS pixels (default: 1920)
selectorNoCSS selector to scope the layout map to a subtree (e.g. ".main-content", "#hero")

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the output ('Returns element hierarchy with tags, classes, bounding rects, and computed CSS styles') and purpose, but lacks details on performance, error handling, or authentication needs, which are relevant for a tool that likely involves network requests.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by a concise usage guideline. Every sentence earns its place by adding value, with no redundant or verbose language, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is mostly complete. It explains what the tool does and its output format, but could benefit from mentioning potential limitations (e.g., timeouts, JavaScript-rendered content) to fully guide an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional parameter semantics beyond what the schema provides, such as examples or edge cases, resulting in a baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Extract the DOM layout tree'), target resource ('of a URL'), and scope ('at a given viewport width'), distinguishing it from sibling tools like viewpo_compare_viewports and viewpo_screenshot by focusing on structural analysis rather than comparison or visual capture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('to understand page structure and find layout issues'), but does not explicitly mention when not to use it or name alternatives among the sibling tools, such as using viewpo_screenshot for visual issues instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

viewpo_screenshotA

Capture screenshots of a URL at one or more viewport widths. Returns base64 JPEG images. Use this to SEE what a webpage looks like at different screen sizes.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to screenshot
viewportsNoViewports to capture. Defaults to desktop (1920px) if omitted. Common widths: 375 (phone), 820 (tablet), 1920 (desktop).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the core behavior (capturing screenshots and returning base64 JPEGs) but lacks details about potential side effects (e.g., network requests, rate limits, authentication needs, or error handling). The description doesn't contradict annotations, but it's incomplete for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, with two sentences that efficiently convey the tool's purpose and usage. Every word earns its place, and there's no redundant or verbose language, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (capturing screenshots with configurable viewports), no annotations, and no output schema, the description is somewhat incomplete. It covers the basic purpose and output format but lacks details on behavioral aspects like error cases, performance implications, or how the output is structured. This is adequate but leaves clear gaps for an agent to fully understand the tool's behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already fully documents the parameters (url and viewports). The description adds minimal value beyond the schema by implying the tool's purpose relates to viewport widths, but it doesn't provide additional semantic context or usage examples for the parameters. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Capture screenshots'), the resource ('a URL at one or more viewport widths'), and the output format ('base64 JPEG images'). It distinguishes from siblings by focusing on visual capture rather than comparison or layout analysis, making the purpose immediately understandable and distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('to SEE what a webpage looks like at different screen sizes'), which implicitly suggests it's for visual inspection. However, it doesn't explicitly state when NOT to use it or name alternatives like the sibling tools, leaving some guidance gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: compare_viewports for detecting differences, get_layout_map for structural analysis, and screenshot for visual capture. There is no overlap in functionality, making it easy for an agent to select the right tool based on the task.

Naming Consistency5/5

All tool names follow a consistent 'viewpo_' prefix with descriptive snake_case suffixes (compare_viewports, get_layout_map, screenshot). This predictable pattern enhances readability and reduces confusion.

Tool Count4/5

With 3 tools, the server is well-scoped for its viewport testing domain, but it feels slightly thin. While each tool is useful, additional tools like viewpo_get_metrics or viewpo_simulate_device might enhance completeness without being overwhelming.

Completeness4/5

The tools cover core viewport analysis tasks: layout comparison, structure extraction, and visual capture. However, there are minor gaps, such as missing tools for performance metrics or device simulation, which agents might need to work around for advanced testing scenarios.

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/littlebearapps/viewpo-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server