Skip to main content
Glama

ui-loop

Context-budgeted visual feedback for coding agents.

ui-loop is an MCP server, a CLI and a Claude Code plugin that lets a coding agent (Claude Code, Cursor, Codex) see the UI it just changed without flooding its context window. Instead of a full screenshot every turn, the agent gets a short structured text summary (console errors, failed requests, overflow, layout shifts, accessibility violations) plus small crops of only the regions that changed since the last capture, all fitted into a token budget you choose. Captures live on disk under .ui-loop/ with automatic eviction, so nothing accumulates in the conversation.

Why

General-purpose browser automation servers are great at driving a browser. They are not designed to be cheap verification primitives:

  • Playwright MCP accessibility snapshots run 50–540 KB each, and the maintainers' position is that pruning history is the agent's job, not the server's (microsoft/playwright-mcp #1233).

  • Claude in Chrome screenshots are re-sent every turn. One report measured 18 screenshots at roughly 279K tokens per call, consuming 17% of a Max plan context window in 5 turns (anthropics/claude-code #27869).

  • Chrome DevTools MCP tool definitions alone cost about 17k tokens before the agent does anything.

None of them ship the thing a coding agent actually needs after an edit: diff-only crops + structured text + eviction as a single primitive. That is all ui-loop does.

Related MCP server: pagelens

Install in 60 seconds

claude plugin marketplace add djig/ui-loop
claude plugin install ui-loop@djig-ui-loop

This registers the MCP server, a PostToolUse hook that diffs the relevant route after you edit a UI file, and a skill that teaches Claude the capture → edit → diff loop.

Plain MCP server without the hook and skill:

claude mcp add ui-loop -- npx -y @djignesh21/ui-loop serve

Cursor

Add to .cursor/mcp.json:

{ "mcpServers": { "ui-loop": { "command": "npx", "args": ["-y", "@djignesh21/ui-loop", "serve"] } } }

Optional best-effort hook: copy examples/cursor/hooks.json to .cursor/hooks.json.

Codex

codex mcp add ui-loop -- npx -y @djignesh21/ui-loop serve

Any agent that reads mcp.json / plugin.json

The repo root carries an agent-plugins.org plugin.json and mcp.json, so Copilot, Codex and Cursor can install it from the GitHub URL.

In your project

npx -y @djignesh21/ui-loop init   # adds .ui-loop/ to .gitignore, prints config snippets

Chromium is resolved from UI_LOOP_CHROMIUM, then Playwright's bundled Chromium, then common system paths (Google Chrome, Chromium). If none is found: npx playwright install chromium or set UI_LOOP_CHROMIUM=/path/to/chrome.

The loop

  1. ui_detect_dev_server → http://localhost:3000

  2. ui_capture { url, label } before editing (baseline; returns one downscaled full image the first time)

  3. Edit the code

  4. ui_diff { url, label } → text summary + crops of changed regions within budget

  5. ui_region { label, region: 0 } only if a crop needs a closer look

  6. Fix, ui_diff again (each diff becomes the new baseline)

  7. ui_assert { url, checks } to close out with text-only checks

Example transcript (abridged):

> ui_capture { url: "http://localhost:3000/settings", label: "settings" }
ui_capture settings (2026-10-03T17-07-36-866Z) — first capture, saved as baseline
title: Settings
url: http://localhost:3000/settings
viewport: 1280×800
console: clean
a11y: no violations at impact ≥ serious
budget: ~1463/4000 tokens (text 97, images 1366)
[image 1280×800 → 960×600]

… agent edits SettingsForm.tsx …

> ui_diff { url: "http://localhost:3000/settings", label: "settings", maxTokens: 1500 }
ui_diff settings: 2026-10-03T17-07-36-866Z → 2026-10-03T17-07-38-786Z
pixels changed: 2.44% (25002/1024000) in 2 region(s)
  #0 [20,360 1260×48] 18001px near <button> "Save changes"
  #1 [40,140 160×48] 7001px near <main>
layout shifts: 2
  main: 0,80 1280×324 → 0,80 1280×334
  #cta: 40,140 160×48 → 640,360 160×48
console: 2 error(s), 0 warning(s), 0 uncaught exception(s)
  ✖ Boom: failed to hydrate widget (×2)
horizontal overflow: YES (scrollWidth 2420 > innerWidth 1280)
a11y: 1 violation(s) at impact ≥ serious
  [critical] image-alt: Images must have alternative text (1 node) e.g. <img src="…">
budget: ~269/1500 tokens (text 226, images 43)
[crop #0 512×33] [crop #1 192×80]

With the plugin installed, the PostToolUse hook runs this diff automatically after Edit/Write/MultiEdit on *.tsx|jsx|css|scss|mdx when a dev server is detected, and injects a bounded (≤ 1500 chars, no images) summary as additional context. It never blocks the tool call and stays silent when nothing applies.

Tool reference

Tool

Purpose

Returns images?

ui_capture { url, label?, viewport?, waitFor?, fullPage?, includeFull?, maxTokens?, darkMode? }

Screenshot + summary; becomes the baseline for label (default: slug of URL path)

First capture of a label, or includeFull

ui_diff { …same, baseline? }

Re-capture, pixel-diff against baseline, return summary + region crops within budget. No baseline → acts like ui_capture and says so

Region crops (largest first)

ui_region { label, captureId?, region: index | {x,y,width,height}, maxTokens? }

One crop at higher resolution

One image

ui_assert { url, checks: [{type, selector?, text?, min?, max?, impact?}] }

text, visible, hidden, count, noConsoleErrors, noOverflow, a11y → pass/fail with evidence

Never

ui_list

Labels and stored captures with sizes

Never

ui_forget { label? }

Delete a label's captures, or all

Never

ui_detect_dev_server

UI_LOOP_URL, Next.js .next/dev/lock, ports 3000–3005/5173/4200/8080

Never

Every tool also returns structuredContent (JSON) alongside the text for clients that prefer it.

The text summary always includes: page title, final URL, viewport, deduped console errors/warnings and uncaught exceptions (capped), failed requests with status ≥ 400 (capped), horizontal overflow (document.scrollWidth > innerWidth), layout shifts for watchSelectors (default h1,h2,nav,main,button,[role=dialog]), axe-core violations at impact ≥ serious by default (capped), and for diffs the percent of pixels changed plus each region's box and the nearest element (tag, role, text snippet) at its center.

Budget model

Every tool accepts maxTokens (default 4000). Allocation order:

  1. Text summary — always included. Typically 100–300 tokens.

  2. Region crops — by area, largest first, each padded by 16px and downscaled so the longest side ≤ 512px, included while the running total fits.

  3. Omissions are stated — 3 more regions omitted (2, 3, 4); call ui_region …, so the agent knows what it did not see.

  4. Full screenshot — only on the first capture of a label or includeFull: true, downscaled to whatever budget remains.

Image tokens are estimated with Anthropic's published formula, ceil(width × height / 750); text as ceil(chars / 4). Other models bill images differently (OpenAI and Gemini tile at 512px/768px), so treat the estimate as an upper-bound heuristic. The reported budget: line tells you what was actually spent.

Configuration

Environment variables:

Variable

Effect

UI_LOOP_CHROMIUM

Path to a Chrome/Chromium executable

UI_LOOP_URL

Dev server base URL; used first by detection and by the hook

UI_LOOP_ROUTE

Route the hook should diff, overriding Next.js inference (needed for dynamic segments like [id])

UI_LOOP_MAX_TOKENS

Default token budget (4000)

UI_LOOP_KEEP

Captures kept per label (5)

UI_LOOP_MAX_CROP_SIDE

Longest side of region crops in px (512)

UI_LOOP_TIMEOUT_MS

Navigation timeout (30000)

UI_LOOP_HOOK_DISABLED=1

Make the hook a no-op

Optional ui-loop.config.json in the project root (all keys optional):

{
  "keepPerLabel": 5,
  "maxTokens": 4000,
  "maxCropSide": 512,
  "cropPadding": 16,
  "maxRegions": 8,
  "viewport": { "width": 1280, "height": 800 },
  "watchSelectors": ["h1", "h2", "nav", "main", "button", "[role=dialog]"],
  "a11yImpact": "serious",
  "diffThreshold": 0.1,
  "timeoutMs": 30000
}

Storage: .ui-loop/captures/<label>/<timestamp>.png + .json. ui-loop init adds .ui-loop/ to .gitignore.

Hook route mapping

For an edited file, the hook picks the route in this order: UI_LOOP_ROUTE → Next.js App Router inference (files under app/; route groups (x) stripped; page/layout/loading/error map to their directory; _private and @slot folders fall back to the parent) → /. Dynamic segments ([id], [...slug]) are skipped with a one-line note instead of guessing.

CLI

ui-loop init
ui-loop capture <url> [--label L] [--json] [--full-page] [--include-full] [--max-tokens N] [--dark]
ui-loop diff <url>    [--label L] [--json] [--max-tokens N] [--include-full] [--baseline ID]
ui-loop assert <url> --check noConsoleErrors --check "text=Save@button" --check a11y=serious
ui-loop list | forget [label] | detect
ui-loop hook [--cursor]     # stdin: hook payload → stdout: bounded context JSON, exit 0 always
ui-loop serve               # MCP over stdio (also the default when stdin is piped)

Programmatic API: import { capture, diff, assert, detectDevServer } from '@djignesh21/ui-loop'.

Comparison

ui-loop

Playwright MCP

Chrome DevTools MCP

agent-browser

Claude in Chrome

Primary job

Verify a UI change cheaply

General browser automation

Debugging/perf via DevTools

Browser automation for agents

Drive the user's real Chrome

Diff vs previous state

Yes, region crops

No

No

No

No

Token budget per call

Yes, explicit maxTokens

No

No

No

No

Text health summary (console, network, a11y, overflow)

Yes, always

On request, separate tools

On request, separate tools

Partial

Partial

Images per turn

Only changed regions; full only on request

Full screenshot on request

Full screenshot on request

Full screenshot

Full screenshot each turn

Persistent store + eviction

Yes (.ui-loop/)

No

No

No

No

Click/type/navigate flows

No

Yes

Yes

Yes

Yes

Tool surface

7 small tools

~25

~26

Many

Many

Honest framing: those projects are general browser automation and do far more than ui-loop. ui-loop is a narrow verification primitive meant to run alongside them (or alone, when all you need is "did my edit render correctly?").

Limitations

  • Static capture of a URL: no clicking, typing or auth flows. Use a dev-only route, query param or mocked state to reach the UI you care about.

  • Pixel diffs are sensitive to animations, carousels, timestamps and non-deterministic data. ui-loop pauses CSS animations/transitions, but content that changes on every load will show as changed.

  • fullPage diffs where the page height changes produce large regions; viewport captures are more stable.

  • Route inference covers the Next.js App Router only. Other frameworks get / unless UI_LOOP_ROUTE or UI_LOOP_URL is set.

  • Token estimates follow Anthropic's formula; actual billing varies by model and provider.

  • axe-core runs on the rendered DOM only (no keyboard-navigation checks).

  • The Cursor hook is best effort; Cursor's hook contract has shifted between versions.

Roadmap

  • ui_interact with a strictly bounded action list (click/type/scroll) before capture.

  • Vite/Remix/SvelteKit route inference.

  • Perceptual diff (ignore anti-aliasing and sub-pixel jitter) and ignore-regions config.

  • Mobile viewport presets and multi-viewport diffs in one call.

  • Optional HTML report of a session's captures for humans.

Contributing

git clone https://github.com/djig/ui-loop && cd ui-loop
npm install
npx playwright install chromium     # or set UI_LOOP_CHROMIUM
npm run build && npm test

Unit tests cover region clustering, budget allocation, token estimation, Next.js route inference, dev-server lock parsing, hook contracts and the summary formatter. The integration test spins up a local HTTP server with before/after fixtures and runs capture → diff → assert in a real Chromium; it skips itself when no Chromium is found.

Issues and PRs welcome. Keep dependencies small, keep tool output bounded, and add a test for behaviour that an agent will rely on.

License

MIT © 2026 Jignesh

Available Tools

7 tools
ui_assertAssert page stateA
Read-onlyIdempotent

Run cheap text-only checks against a live page: text present, element visible/hidden, element count, no console errors, no horizontal overflow, no a11y violations. Returns pass/fail with short evidence. Never returns images.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
checksYes
waitForNoCSS selector to wait for, or ms to wait
darkModeNo
viewportNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), and the description adds a genuine behavioral contract: it returns pass/fail with short evidence and never returns images. It stops short of describing timing behavior (waitFor), timeouts, or what happens on partial failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the capability list and closed with the distinguishing return constraint. No filler, no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, nested check objects, and no output schema, the description adequately covers purpose and return shape but leaves the caller without guidance on timing (waitFor), viewport/darkMode, or how min/max/impact interact with the check types. Enough to call it, not enough to call it well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, so the description should compensate. It partially does by enumerating the check kinds ('text present, element visible/hidden, element count, no console errors, no horizontal overflow, no a11y violations'), which map to the checks[].type enum, but it says nothing about waitFor, darkMode, viewport, or the min/max/impact/selector fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (assert/check) and resource (live page) and enumerates the exact check categories: text, visibility, count, console errors, overflow, a11y. The clause 'Never returns images' implicitly separates it from ui_capture, so an agent can route between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'cheap text-only' implies this is the lightweight alternative to screenshot-based siblings like ui_capture, but no sibling is named and there is no explicit when-not condition. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_captureCapture UI baselineA

Screenshot a page and save it as the baseline for its label. Returns a compact text summary (title, console errors, failed requests, overflow, a11y violations). The image is returned only on the first capture of a label or when includeFull=true. Call this BEFORE editing so ui_diff has something to compare against.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPage URL, e.g. http://localhost:3000/settings
labelNoBaseline label. Defaults to a slug of the URL path. Keep one label per route.
waitForNoCSS selector to wait for, or ms to wait
darkModeNoEmulate prefers-color-scheme: dark.
fullPageNoCapture the full scrollable page (bigger, costlier). Default false.
viewportNo
maxTokensNoToken budget for this response (text + images). Default 4000.
includeFullNoAlso return the full screenshot (downscaled to budget). Default false.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true), and the description adds real behavior: the return payload is a text summary listing title, console errors, failed requests, overflow and a11y violations, and the image only comes back on the first capture of a label or with includeFull=true. It stops short of saying whether re-capturing a label overwrites the stored baseline, which is the one mutation detail still missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with what it does, then the return shape, then the call-ordering rule. No filler and no repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters, 88% schema coverage and no output schema, the description usefully fills the return-value gap by naming the summary fields. Everything an agent needs to call it correctly at the right moment is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 88%, so the schema already documents most parameters and the baseline of 3 applies. The description earns the bump by clarifying includeFull's conditional semantics (image returned only on first capture of a label or when includeFull=true) beyond the schema's bare 'Also return the full screenshot'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: screenshots a page and persists it as the baseline for a label. It also implicitly distinguishes itself from ui_diff by framing itself as the producer of the artifact ui_diff consumes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sequencing guidance ('Call this BEFORE editing so ui_diff has something to compare against') names the sibling alternative and the condition that requires this tool first. That is the exact routing information an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_detect_dev_serverDetect dev serverA
Read-onlyIdempotent

Find running local dev servers (UI_LOOP_URL env, Next.js .next/dev/lock, common ports 3000–3005/5173/4200/8080). Returns candidate base URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context beyond those flags: the concrete detection heuristics and the fact that results are unverified 'candidate' URLs rather than confirmed servers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler: the first gives the action plus detection sources, the second gives the return shape. Front-loaded and appropriately sized for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return contract ('candidate base URLs') and the detection sources. It stops short of saying what an empty result means or whether multiple candidates are ranked, but nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 — there is nothing for the description to clarify and no ambiguity an agent could stumble into when invoking it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find running local dev servers') and enumerates exactly how detection is performed (UI_LOOP_URL env, Next.js .next/dev/lock, specific ports). It also states the return type. No sibling (ui_capture, ui_diff, ui_list, etc.) overlaps this discovery function, so an agent can select it unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied — an agent would call this to discover a base URL before capture/assert operations — but the description never states when to use it or what to do if detection fails. There is no explicit alternative to route against, so the gap is moderate rather than severe.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_diffDiff UI against baselineA

Re-capture a page, pixel-diff it against the baseline for its label, and return ONLY what changed: a text summary plus cropped images of changed regions (largest first) that fit within maxTokens. Omitted regions are listed so you can fetch them with ui_region. The new capture becomes the baseline. If no baseline exists, behaves like ui_capture.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPage URL, e.g. http://localhost:3000/settings
labelNoBaseline label. Defaults to a slug of the URL path. Keep one label per route.
waitForNoCSS selector to wait for, or ms to wait
baselineNo'previous' (default) or a captureId from ui_list
darkModeNoEmulate prefers-color-scheme: dark.
fullPageNoCapture the full scrollable page (bigger, costlier). Default false.
viewportNo
maxTokensNoToken budget for this response (text + images). Default 4000.
includeFullNoAlso return the full screenshot (downscaled to budget). Default false.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and idempotentHint=false; the description goes well beyond that by disclosing the key side effect ('The new capture becomes the baseline'), the truncation policy (only changes that 'fit within maxTokens', largest first, omitted regions enumerated), and the degraded no-baseline mode. These are exactly the traits structured fields cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences with no filler: action first, then return payload, then truncation handling, then side effect and fallback. Every clause earns its place and the most decision-relevant facts are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 9 parameters, the description still fully describes the return contract (text summary plus cropped changed-region images, largest first, budget-limited, omissions listed) and the state-mutating side effect, which is the maximum an agent needs before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, so the baseline is 3, but the description adds real meaning: it ties 'label' to the baseline selection ('the baseline for its label'), explains that maxTokens gates which regions are returned at all, and clarifies the default/override relationship for 'baseline'. It does not cover the viewport or waitFor semantics, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb chain and resource ('Re-capture a page, pixel-diff it against the baseline for its label') and explicitly contrasts itself with siblings ui_capture (fallback when no baseline exists) and ui_region (fetching omitted regions). An agent can distinguish it from all six sibling tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: use this to see what changed against an existing baseline, and it names the fallback behavior ('If no baseline exists, behaves like ui_capture') plus the follow-up path via ui_region for regions dropped by the token budget. It stops short of an explicit 'prefer this over ui_capture when...' rule, but the routing is inferable and mostly spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_forgetForget capturesA
DestructiveIdempotent

Delete stored captures for a label, or all labels when omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds the critical behavioral fact that omitting the parameter widens the blast radius to every label, which is exactly the disclosure that matters for a destructive default-scope operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action and resource, with the optional-parameter semantics attached directly. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool with no output schema, the description plus annotations cover action, scope, default behavior, and safety. It leaves unstated whether deletion is permanent/recoverable, which is the only notable omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the lone parameter carries no description, so the description must compensate — and it does, by defining what 'label' selects and what its absence means. It stops short of naming the label format or where labels come from.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Delete') and resource ('stored captures') plus the scoping axis ('for a label, or all labels'). It is clearly distinguishable from siblings like ui_capture and ui_list, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the agent can infer this is the cleanup path for captures created by ui_capture, but there is no explicit when-to-use, when-not-to-use, or named alternative. Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_listList stored capturesA
Read-onlyIdempotent

List labels and their stored captures (ids, sizes) under .ui-loop/captures.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds value beyond that by disclosing the scope of the listing (labels plus ids and sizes) and the backing location .ui-loop/captures, which is useful context since no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the verb, the resource, the returned fields, and the storage path all appear immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and annotations covering the safety profile, the description supplies the remaining essentials: what is listed and where it lives. It stops short of describing ordering, pagination, or behavior on an empty store, but nothing critical to invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description does not need to explain argument semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (labels and their stored captures), and even names the returned fields (ids, sizes) and the storage location (.ui-loop/captures). It is clearly distinguishable from siblings like ui_capture or ui_forget, though it does not explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent can infer this is the way to inspect what is currently stored before diffing, asserting, or forgetting. There is no explicit when-to-use, when-not-to-use, or pointer to a sibling such as ui_forget for cleanup.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_regionFetch one region cropA
Read-onlyIdempotent

Fetch a single crop from a stored capture at higher resolution. region is either a diff region index (from ui_diff output) or a box {x,y,width,height} in page pixels.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes
regionYes
captureIdNoDefaults to the latest capture for the label
maxTokensNoDefault 2000

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds that the crop is returned at higher resolution, which is useful context, but says nothing about the maxTokens-capped return or what happens when captureId is omitted beyond the schema note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core action is front-loaded before the region-form detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description never indicates what comes back (an image crop consumed as tokens, presumably governed by maxTokens). Combined with the un-explained label parameter, an agent has to infer the return contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, so the description must compensate. It does clarify the dual form of region and its provenance from ui_diff, which the schema's anyOf does not express, but label and the defaulting behavior of captureId/maxTokens are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch a single crop from a stored capture') plus a key differentiator ('at higher resolution'), which separates it from a full capture. It does not explicitly contrast with ui_capture, so sibling differentiation is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the region provenance note — 'a diff region index (from ui_diff output)' — which tells the agent this follows a diff workflow. There is no explicit when-to-use vs when-not guidance or statement of alternatives such as ui_capture.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.1
    • First observedui_assert
    • First observedui_capture
    • First observedui_detect_dev_server
    • First observedui_diff
    • First observedui_forget
    • First observedui_list
    • First observedui_region

TDQS

A4.2/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a clearly distinct role: ui_capture establishes a baseline, ui_diff compares against it, ui_region fetches crops, ui_assert runs text checks, ui_list/ui_forget manage stored state, and ui_detect_dev_server discovers servers. The capture-vs-diff distinction is explicitly documented (capture=before editing, diff=after), so misselection is unlikely.

Naming Consistency5/5

All seven tools use a uniform ui_ prefix with snake_case verb_noun structure (ui_capture, ui_diff, ui_region, ui_assert, ui_list, ui_forget, ui_detect_dev_server). The only slightly longer name follows the same convention, so the pattern is fully predictable.

Tool Count5/5

Seven tools is well-scoped for a visual-regression/diffing workflow: one tool each for capture, diff, inspection, assertion, listing, deletion, and server discovery. No tool feels redundant or missing at this count.

Completeness4/5

The lifecycle is well covered: capture a baseline, diff after edits, drill into regions, assert cheap checks, and manage labels via list/forget plus dev-server discovery. Minor gaps exist (e.g., no explicit label rename/config or batch operation), but agents can work around these easily.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

  • Capture screenshots, detect visual regressions between page versions, and analyze with AI.

  • Give agents eyes on any web page: structured context, and changes explained in plain language.

  • Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.

  • Grabbit gives AI agents eyes on the web through a hosted MCP server. Send a public URL and get a pixel-perfect hosted image back, without maintaining Chromium, Playwright, or a browser fleet. Capture a full page, exact viewport, or single CSS selector as PNG, JPEG, or WebP. Grabbit handles cookie and consent banners, waits for JavaScript-heavy pages, blocks private and internal URLs, supports safe retries with idempotency keys, and delivers async results through signed webhooks. Completed captures include a CDN URL. Connect with OAuth 2.1 or an API key. Grabbit works with Claude, Cursor, Codex, and any MCP client. Live captures cost $0.002 each. The $50 annual plan includes 25,000 prepaid credits that never reset or expire. Free test keys return placeholder images, so you can wire up the integration before paying. Home: https://grabbit.live Docs: https://grabbit.live/screenshot-api Built by BrainGrid.

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.
    2 npm
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents to see, measure, and verify web pages through a real Chrome browser, including screenshots, responsive layout and accessibility audits, pixel diffing against baselines, secure logins, and deployed-fix verification.
    27
    1
    MIT