visual-inspector-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| navigateA | Load a URL in a persistent headless browser page (dev server, file://, or public site). Stays open for later click/screenshot calls. |
| screenshotA | Screenshot the current page and return it as a viewable image — see the actual rendered UI instead of guessing from code. Prefer selector (one element, cheapest) over the default viewport, and viewport over fullPage (priciest); use the smallest capture that answers the question. Requires navigate first. |
| clickA | Click an element by selector to reach a UI state (open a menu/modal/tab) before screenshotting it. |
| fillA | Set the value of an input/textarea/select by selector (Playwright locator.fill — clears then types, and dispatches the input/change events React controlled components need). Pass |
| typeA | Type into a field one key at a time (Playwright locator.pressSequentially), firing a keydown/keypress/keyup per character — for inputs whose handlers need real keystrokes and don't react to |
| pressA | Press a keyboard key such as 'Enter' (submit a form), 'Tab', or 'Escape'. With |
| resizeA | Resize the viewport, e.g. to check a responsive breakpoint before screenshotting. |
| console_logsA | Recent browser console messages and page errors — correlate a visual issue with a JS error. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Each tool has a distinct purpose: clicking, filling inputs, navigating, taking screenshots, etc. There is no functional overlap or ambiguity between tools.
All tool names follow a consistent lowercase_snake_case verb pattern (click, fill, navigate, press, resize, screenshot, type), except console_logs which is a noun but still clear. Overall very consistent.
With 8 tools, the server is well-scoped for browser automation and visual inspection. Each tool serves a necessary function without excessive granularity.
The tool set covers core browser interactions (navigation, clicking, input, keyboard, screenshots) and console logs. Minor gaps like explicit wait or scrolling are absent, but the surface is sufficient for most visual inspection tasks.