ui-loop
Provides Next.js-specific integration for route detection and dev server detection. The PostToolUse hook infers routes from Next.js App Router files under app/, strips route groups, maps page/layout/loading/error to their directories, and detects the Next.js dev server via .next/dev/lock. Dynamic segments are skipped with a note.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ui-loopdiff the settings page after my edit and show me what changed"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ui-loop
Context-budgeted visual feedback for coding agents.
ui-loop is an MCP server, a CLI and a Claude Code plugin that lets a coding agent (Claude Code, Cursor, Codex) see the UI it just changed without flooding its context window. Instead of a full screenshot every turn, the agent gets a short structured text summary (console errors, failed requests, overflow, layout shifts, accessibility violations) plus small crops of only the regions that changed since the last capture, all fitted into a token budget you choose. Captures live on disk under .ui-loop/ with automatic eviction, so nothing accumulates in the conversation.
Why
General-purpose browser automation servers are great at driving a browser. They are not designed to be cheap verification primitives:
Playwright MCP accessibility snapshots run 50–540 KB each, and the maintainers' position is that pruning history is the agent's job, not the server's (microsoft/playwright-mcp #1233).
Claude in Chrome screenshots are re-sent every turn. One report measured 18 screenshots at roughly 279K tokens per call, consuming 17% of a Max plan context window in 5 turns (anthropics/claude-code #27869).
Chrome DevTools MCP tool definitions alone cost about 17k tokens before the agent does anything.
None of them ship the thing a coding agent actually needs after an edit: diff-only crops + structured text + eviction as a single primitive. That is all ui-loop does.
Related MCP server: pagelens
Install in 60 seconds
Claude Code (plugin, recommended)
claude plugin marketplace add djig/ui-loop
claude plugin install ui-loop@djig-ui-loopThis registers the MCP server, a PostToolUse hook that diffs the relevant route after you edit a UI file, and a skill that teaches Claude the capture → edit → diff loop.
Plain MCP server without the hook and skill:
claude mcp add ui-loop -- npx -y @djignesh21/ui-loop serveCursor
Add to .cursor/mcp.json:
{ "mcpServers": { "ui-loop": { "command": "npx", "args": ["-y", "@djignesh21/ui-loop", "serve"] } } }Optional best-effort hook: copy examples/cursor/hooks.json to .cursor/hooks.json.
Codex
codex mcp add ui-loop -- npx -y @djignesh21/ui-loop serveAny agent that reads mcp.json / plugin.json
The repo root carries an agent-plugins.org plugin.json and mcp.json, so Copilot, Codex and Cursor can install it from the GitHub URL.
In your project
npx -y @djignesh21/ui-loop init # adds .ui-loop/ to .gitignore, prints config snippetsChromium is resolved from UI_LOOP_CHROMIUM, then Playwright's bundled Chromium, then common system paths (Google Chrome, Chromium). If none is found: npx playwright install chromium or set UI_LOOP_CHROMIUM=/path/to/chrome.
The loop
ui_detect_dev_server→http://localhost:3000ui_capture { url, label }before editing (baseline; returns one downscaled full image the first time)Edit the code
ui_diff { url, label }→ text summary + crops of changed regions within budgetui_region { label, region: 0 }only if a crop needs a closer lookFix,
ui_diffagain (each diff becomes the new baseline)ui_assert { url, checks }to close out with text-only checks
Example transcript (abridged):
> ui_capture { url: "http://localhost:3000/settings", label: "settings" }
ui_capture settings (2026-10-03T17-07-36-866Z) — first capture, saved as baseline
title: Settings
url: http://localhost:3000/settings
viewport: 1280×800
console: clean
a11y: no violations at impact ≥ serious
budget: ~1463/4000 tokens (text 97, images 1366)
[image 1280×800 → 960×600]
… agent edits SettingsForm.tsx …
> ui_diff { url: "http://localhost:3000/settings", label: "settings", maxTokens: 1500 }
ui_diff settings: 2026-10-03T17-07-36-866Z → 2026-10-03T17-07-38-786Z
pixels changed: 2.44% (25002/1024000) in 2 region(s)
#0 [20,360 1260×48] 18001px near <button> "Save changes"
#1 [40,140 160×48] 7001px near <main>
layout shifts: 2
main: 0,80 1280×324 → 0,80 1280×334
#cta: 40,140 160×48 → 640,360 160×48
console: 2 error(s), 0 warning(s), 0 uncaught exception(s)
✖ Boom: failed to hydrate widget (×2)
horizontal overflow: YES (scrollWidth 2420 > innerWidth 1280)
a11y: 1 violation(s) at impact ≥ serious
[critical] image-alt: Images must have alternative text (1 node) e.g. <img src="…">
budget: ~269/1500 tokens (text 226, images 43)
[crop #0 512×33] [crop #1 192×80]With the plugin installed, the PostToolUse hook runs this diff automatically after Edit/Write/MultiEdit on *.tsx|jsx|css|scss|mdx when a dev server is detected, and injects a bounded (≤ 1500 chars, no images) summary as additional context. It never blocks the tool call and stays silent when nothing applies.
Tool reference
Tool | Purpose | Returns images? |
| Screenshot + summary; becomes the baseline for | First capture of a label, or |
| Re-capture, pixel-diff against baseline, return summary + region crops within budget. No baseline → acts like | Region crops (largest first) |
| One crop at higher resolution | One image |
|
| Never |
| Labels and stored captures with sizes | Never |
| Delete a label's captures, or all | Never |
|
| Never |
Every tool also returns structuredContent (JSON) alongside the text for clients that prefer it.
The text summary always includes: page title, final URL, viewport, deduped console errors/warnings and uncaught exceptions (capped), failed requests with status ≥ 400 (capped), horizontal overflow (document.scrollWidth > innerWidth), layout shifts for watchSelectors (default h1,h2,nav,main,button,[role=dialog]), axe-core violations at impact ≥ serious by default (capped), and for diffs the percent of pixels changed plus each region's box and the nearest element (tag, role, text snippet) at its center.
Budget model
Every tool accepts maxTokens (default 4000). Allocation order:
Text summary — always included. Typically 100–300 tokens.
Region crops — by area, largest first, each padded by 16px and downscaled so the longest side ≤ 512px, included while the running total fits.
Omissions are stated —
3 more regions omitted (2, 3, 4); call ui_region …, so the agent knows what it did not see.Full screenshot — only on the first capture of a label or
includeFull: true, downscaled to whatever budget remains.
Image tokens are estimated with Anthropic's published formula, ceil(width × height / 750); text as ceil(chars / 4). Other models bill images differently (OpenAI and Gemini tile at 512px/768px), so treat the estimate as an upper-bound heuristic. The reported budget: line tells you what was actually spent.
Configuration
Environment variables:
Variable | Effect |
| Path to a Chrome/Chromium executable |
| Dev server base URL; used first by detection and by the hook |
| Route the hook should diff, overriding Next.js inference (needed for dynamic segments like |
| Default token budget (4000) |
| Captures kept per label (5) |
| Longest side of region crops in px (512) |
| Navigation timeout (30000) |
| Make the hook a no-op |
Optional ui-loop.config.json in the project root (all keys optional):
{
"keepPerLabel": 5,
"maxTokens": 4000,
"maxCropSide": 512,
"cropPadding": 16,
"maxRegions": 8,
"viewport": { "width": 1280, "height": 800 },
"watchSelectors": ["h1", "h2", "nav", "main", "button", "[role=dialog]"],
"a11yImpact": "serious",
"diffThreshold": 0.1,
"timeoutMs": 30000
}Storage: .ui-loop/captures/<label>/<timestamp>.png + .json. ui-loop init adds .ui-loop/ to .gitignore.
Hook route mapping
For an edited file, the hook picks the route in this order: UI_LOOP_ROUTE → Next.js App Router inference (files under app/; route groups (x) stripped; page/layout/loading/error map to their directory; _private and @slot folders fall back to the parent) → /. Dynamic segments ([id], [...slug]) are skipped with a one-line note instead of guessing.
CLI
ui-loop init
ui-loop capture <url> [--label L] [--json] [--full-page] [--include-full] [--max-tokens N] [--dark]
ui-loop diff <url> [--label L] [--json] [--max-tokens N] [--include-full] [--baseline ID]
ui-loop assert <url> --check noConsoleErrors --check "text=Save@button" --check a11y=serious
ui-loop list | forget [label] | detect
ui-loop hook [--cursor] # stdin: hook payload → stdout: bounded context JSON, exit 0 always
ui-loop serve # MCP over stdio (also the default when stdin is piped)Programmatic API: import { capture, diff, assert, detectDevServer } from '@djignesh21/ui-loop'.
Comparison
ui-loop | Playwright MCP | Chrome DevTools MCP | agent-browser | Claude in Chrome | |
Primary job | Verify a UI change cheaply | General browser automation | Debugging/perf via DevTools | Browser automation for agents | Drive the user's real Chrome |
Diff vs previous state | Yes, region crops | No | No | No | No |
Token budget per call | Yes, explicit | No | No | No | No |
Text health summary (console, network, a11y, overflow) | Yes, always | On request, separate tools | On request, separate tools | Partial | Partial |
Images per turn | Only changed regions; full only on request | Full screenshot on request | Full screenshot on request | Full screenshot | Full screenshot each turn |
Persistent store + eviction | Yes ( | No | No | No | No |
Click/type/navigate flows | No | Yes | Yes | Yes | Yes |
Tool surface | 7 small tools | ~25 | ~26 | Many | Many |
Honest framing: those projects are general browser automation and do far more than ui-loop. ui-loop is a narrow verification primitive meant to run alongside them (or alone, when all you need is "did my edit render correctly?").
Limitations
Static capture of a URL: no clicking, typing or auth flows. Use a dev-only route, query param or mocked state to reach the UI you care about.
Pixel diffs are sensitive to animations, carousels, timestamps and non-deterministic data. ui-loop pauses CSS animations/transitions, but content that changes on every load will show as changed.
fullPagediffs where the page height changes produce large regions; viewport captures are more stable.Route inference covers the Next.js App Router only. Other frameworks get
/unlessUI_LOOP_ROUTEorUI_LOOP_URLis set.Token estimates follow Anthropic's formula; actual billing varies by model and provider.
axe-core runs on the rendered DOM only (no keyboard-navigation checks).
The Cursor hook is best effort; Cursor's hook contract has shifted between versions.
Roadmap
ui_interactwith a strictly bounded action list (click/type/scroll) before capture.Vite/Remix/SvelteKit route inference.
Perceptual diff (ignore anti-aliasing and sub-pixel jitter) and ignore-regions config.
Mobile viewport presets and multi-viewport diffs in one call.
Optional HTML report of a session's captures for humans.
Contributing
git clone https://github.com/djig/ui-loop && cd ui-loop
npm install
npx playwright install chromium # or set UI_LOOP_CHROMIUM
npm run build && npm testUnit tests cover region clustering, budget allocation, token estimation, Next.js route inference, dev-server lock parsing, hook contracts and the summary formatter. The integration test spins up a local HTTP server with before/after fixtures and runs capture → diff → assert in a real Chromium; it skips itself when no Chromium is found.
Issues and PRs welcome. Keep dependencies small, keep tool output bounded, and add a test for behaviour that an agent will rely on.
License
MIT © 2026 Jignesh
Available Tools
7 toolsui_assertAssert page stateARead-onlyIdempotent
Run cheap text-only checks against a live page: text present, element visible/hidden, element count, no console errors, no horizontal overflow, no a11y violations. Returns pass/fail with short evidence. Never returns images.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| checks | Yes | ||
| waitFor | No | CSS selector to wait for, or ms to wait | |
| darkMode | No | ||
| viewport | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), and the description adds a genuine behavioral contract: it returns pass/fail with short evidence and never returns images. It stops short of describing timing behavior (waitFor), timeouts, or what happens on partial failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences, front-loaded with the capability list and closed with the distinguishing return constraint. No filler, no repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, nested check objects, and no output schema, the description adequately covers purpose and return shape but leaves the caller without guidance on timing (waitFor), viewport/darkMode, or how min/max/impact interact with the check types. Enough to call it, not enough to call it well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, so the description should compensate. It partially does by enumerating the check kinds ('text present, element visible/hidden, element count, no console errors, no horizontal overflow, no a11y violations'), which map to the checks[].type enum, but it says nothing about waitFor, darkMode, viewport, or the min/max/impact/selector fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (assert/check) and resource (live page) and enumerates the exact check categories: text, visibility, count, console errors, overflow, a11y. The clause 'Never returns images' implicitly separates it from ui_capture, so an agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'cheap text-only' implies this is the lightweight alternative to screenshot-based siblings like ui_capture, but no sibling is named and there is no explicit when-not condition. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_captureCapture UI baselineA
Screenshot a page and save it as the baseline for its label. Returns a compact text summary (title, console errors, failed requests, overflow, a11y violations). The image is returned only on the first capture of a label or when includeFull=true. Call this BEFORE editing so ui_diff has something to compare against.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Page URL, e.g. http://localhost:3000/settings | |
| label | No | Baseline label. Defaults to a slug of the URL path. Keep one label per route. | |
| waitFor | No | CSS selector to wait for, or ms to wait | |
| darkMode | No | Emulate prefers-color-scheme: dark. | |
| fullPage | No | Capture the full scrollable page (bigger, costlier). Default false. | |
| viewport | No | ||
| maxTokens | No | Token budget for this response (text + images). Default 4000. | |
| includeFull | No | Also return the full screenshot (downscaled to budget). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true), and the description adds real behavior: the return payload is a text summary listing title, console errors, failed requests, overflow and a11y violations, and the image only comes back on the first capture of a label or with includeFull=true. It stops short of saying whether re-capturing a label overwrites the stored baseline, which is the one mutation detail still missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what it does, then the return shape, then the call-ordering rule. No filler and no repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, 88% schema coverage and no output schema, the description usefully fills the return-value gap by naming the summary fields. Everything an agent needs to call it correctly at the right moment is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 88%, so the schema already documents most parameters and the baseline of 3 applies. The description earns the bump by clarifying includeFull's conditional semantics (image returned only on first capture of a label or when includeFull=true) beyond the schema's bare 'Also return the full screenshot'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: screenshots a page and persists it as the baseline for a label. It also implicitly distinguishes itself from ui_diff by framing itself as the producer of the artifact ui_diff consumes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sequencing guidance ('Call this BEFORE editing so ui_diff has something to compare against') names the sibling alternative and the condition that requires this tool first. That is the exact routing information an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_detect_dev_serverDetect dev serverARead-onlyIdempotent
Find running local dev servers (UI_LOOP_URL env, Next.js .next/dev/lock, common ports 3000–3005/5173/4200/8080). Returns candidate base URLs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context beyond those flags: the concrete detection heuristics and the fact that results are unverified 'candidate' URLs rather than confirmed servers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler: the first gives the action plus detection sources, the second gives the return shape. Front-loaded and appropriately sized for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly supplies the return contract ('candidate base URLs') and the detection sources. It stops short of saying what an empty result means or whether multiple candidates are ranked, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 — there is nothing for the description to clarify and no ambiguity an agent could stumble into when invoking it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find running local dev servers') and enumerates exactly how detection is performed (UI_LOOP_URL env, Next.js .next/dev/lock, specific ports). It also states the return type. No sibling (ui_capture, ui_diff, ui_list, etc.) overlaps this discovery function, so an agent can select it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied — an agent would call this to discover a base URL before capture/assert operations — but the description never states when to use it or what to do if detection fails. There is no explicit alternative to route against, so the gap is moderate rather than severe.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_diffDiff UI against baselineA
Re-capture a page, pixel-diff it against the baseline for its label, and return ONLY what changed: a text summary plus cropped images of changed regions (largest first) that fit within maxTokens. Omitted regions are listed so you can fetch them with ui_region. The new capture becomes the baseline. If no baseline exists, behaves like ui_capture.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Page URL, e.g. http://localhost:3000/settings | |
| label | No | Baseline label. Defaults to a slug of the URL path. Keep one label per route. | |
| waitFor | No | CSS selector to wait for, or ms to wait | |
| baseline | No | 'previous' (default) or a captureId from ui_list | |
| darkMode | No | Emulate prefers-color-scheme: dark. | |
| fullPage | No | Capture the full scrollable page (bigger, costlier). Default false. | |
| viewport | No | ||
| maxTokens | No | Token budget for this response (text + images). Default 4000. | |
| includeFull | No | Also return the full screenshot (downscaled to budget). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and idempotentHint=false; the description goes well beyond that by disclosing the key side effect ('The new capture becomes the baseline'), the truncation policy (only changes that 'fit within maxTokens', largest first, omitted regions enumerated), and the degraded no-baseline mode. These are exactly the traits structured fields cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences with no filler: action first, then return payload, then truncation handling, then side effect and fallback. Every clause earns its place and the most decision-relevant facts are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 9 parameters, the description still fully describes the return contract (text summary plus cropped changed-region images, largest first, budget-limited, omissions listed) and the state-mutating side effect, which is the maximum an agent needs before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 89%, so the baseline is 3, but the description adds real meaning: it ties 'label' to the baseline selection ('the baseline for its label'), explains that maxTokens gates which regions are returned at all, and clarifies the default/override relationship for 'baseline'. It does not cover the viewport or waitFor semantics, leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb chain and resource ('Re-capture a page, pixel-diff it against the baseline for its label') and explicitly contrasts itself with siblings ui_capture (fallback when no baseline exists) and ui_region (fetching omitted regions). An agent can distinguish it from all six sibling tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this to see what changed against an existing baseline, and it names the fallback behavior ('If no baseline exists, behaves like ui_capture') plus the follow-up path via ui_region for regions dropped by the token budget. It stops short of an explicit 'prefer this over ui_capture when...' rule, but the routing is inferable and mostly spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_forgetForget capturesADestructiveIdempotent
Delete stored captures for a label, or all labels when omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds the critical behavioral fact that omitting the parameter widens the blast radius to every label, which is exactly the disclosure that matters for a destructive default-scope operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and resource, with the optional-parameter semantics attached directly. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with no output schema, the description plus annotations cover action, scope, default behavior, and safety. It leaves unstated whether deletion is permanent/recoverable, which is the only notable omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the lone parameter carries no description, so the description must compensate — and it does, by defining what 'label' selects and what its absence means. It stops short of naming the label format or where labels come from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Delete') and resource ('stored captures') plus the scoping axis ('for a label, or all labels'). It is clearly distinguishable from siblings like ui_capture and ui_list, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the cleanup path for captures created by ui_capture, but there is no explicit when-to-use, when-not-to-use, or named alternative. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_listList stored capturesARead-onlyIdempotent
List labels and their stored captures (ids, sizes) under .ui-loop/captures.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds value beyond that by disclosing the scope of the listing (labels plus ids and sizes) and the backing location .ui-loop/captures, which is useful context since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the verb, the resource, the returned fields, and the storage path all appear immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and annotations covering the safety profile, the description supplies the remaining essentials: what is listed and where it lives. It stops short of describing ordering, pagination, or behavior on an empty store, but nothing critical to invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description does not need to explain argument semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (labels and their stored captures), and even names the returned fields (ids, sizes) and the storage location (.ui-loop/captures). It is clearly distinguishable from siblings like ui_capture or ui_forget, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent can infer this is the way to inspect what is currently stored before diffing, asserting, or forgetting. There is no explicit when-to-use, when-not-to-use, or pointer to a sibling such as ui_forget for cleanup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_regionFetch one region cropARead-onlyIdempotent
Fetch a single crop from a stored capture at higher resolution. region is either a diff region index (from ui_diff output) or a box {x,y,width,height} in page pixels.
| Name | Required | Description | Default |
|---|---|---|---|
| label | Yes | ||
| region | Yes | ||
| captureId | No | Defaults to the latest capture for the label | |
| maxTokens | No | Default 2000 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds that the crop is returned at higher resolution, which is useful context, but says nothing about the maxTokens-capped return or what happens when captureId is omitted beyond the schema note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the core action is front-loaded before the region-form detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description never indicates what comes back (an image crop consumed as tokens, presumably governed by maxTokens). Combined with the un-explained label parameter, an agent has to infer the return contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, so the description must compensate. It does clarify the dual form of region and its provenance from ui_diff, which the schema's anyOf does not express, but label and the defaulting behavior of captureId/maxTokens are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch a single crop from a stored capture') plus a key differentiator ('at higher resolution'), which separates it from a full capture. It does not explicitly contrast with ui_capture, so sibling differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the region provenance note — 'a diff region index (from ui_diff output)' — which tells the agent this follows a diff workflow. There is no explicit when-to-use vs when-not guidance or statement of alternatives such as ui_capture.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.1- First observed
ui_assert - First observed
ui_capture - First observed
ui_detect_dev_server - First observed
ui_diff - First observed
ui_forget - First observed
ui_list - First observed
ui_region
TDQS
Scored across 7 tools
Each tool has a clearly distinct role: ui_capture establishes a baseline, ui_diff compares against it, ui_region fetches crops, ui_assert runs text checks, ui_list/ui_forget manage stored state, and ui_detect_dev_server discovers servers. The capture-vs-diff distinction is explicitly documented (capture=before editing, diff=after), so misselection is unlikely.
All seven tools use a uniform ui_ prefix with snake_case verb_noun structure (ui_capture, ui_diff, ui_region, ui_assert, ui_list, ui_forget, ui_detect_dev_server). The only slightly longer name follows the same convention, so the pattern is fully predictable.
Seven tools is well-scoped for a visual-regression/diffing workflow: one tool each for capture, diff, inspection, assertion, listing, deletion, and server discovery. No tool feels redundant or missing at this count.
The lifecycle is well covered: capture a baseline, diff after edits, drill into regions, assert cheap checks, and manage labels via list/forget plus dev-server discovery. Minor gaps exist (e.g., no explicit label rename/config or batch operation), but agents can work around these easily.
Maintenance
Related MCP Connectors
Capture screenshots, detect visual regressions between page versions, and analyze with AI.
Give agents eyes on any web page: structured context, and changes explained in plain language.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Grabbit gives AI agents eyes on the web through a hosted MCP server. Send a public URL and get a pixel-perfect hosted image back, without maintaining Chromium, Playwright, or a browser fleet. Capture a full page, exact viewport, or single CSS selector as PNG, JPEG, or WebP. Grabbit handles cookie and consent banners, waits for JavaScript-heavy pages, blocks private and internal URLs, supports safe retries with idempotency keys, and delivers async results through signed webhooks. Completed captures include a CDN URL. Connect with OAuth 2.1 or an API key. Grabbit works with Claude, Cursor, Codex, and any MCP client. Live captures cost $0.002 each. The $50 annual plan includes 25,000 prepaid credits that never reset or expire. Free test keys return placeholder images, so you can wire up the integration before paying. Home: https://grabbit.live Docs: https://grabbit.live/screenshot-api Built by BrainGrid.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.2 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables AI coding agents to visually interact with frontend apps by taking screenshots, clicking elements, reading console logs, and performing visual diffs.12 npm3MIT
- AlicenseAqualityCmaintenanceEnables coding agents to visually verify front-end changes by capturing screenshots at real breakpoints, identifying the single element responsible for overflow, and surfacing console and network errors.22MIT
- AlicenseAqualityBmaintenanceEnables AI coding agents to see, measure, and verify web pages through a real Chrome browser, including screenshots, responsive layout and accessibility audits, pixel diffing against baselines, secure logins, and deployed-fix verification.271MIT