Skip to main content
Glama

Your coding agent cannot see the page it just built. pixelpact measures it against the reference and hands back numbers.

CI npm node license

A coding agent writes the CSS and says it is done. It never sees the result. Nothing hands it a number, so the loop cannot close: the agent guesses, you open the page, you send it back, it guesses again. Every round costs you the one thing the agent was supposed to save.

The same gap exists without an agent, when a teammate says a section is finished and nobody in the room has a measurement. The difference is that a person can at least look.

pixelpact reads the reference page and writes down what it actually renders: sizes, colors, spacing, typography, hover and focus states, animation keyframes, design tokens. That file is the contract. Point pixelpact at the implementation and it answers with one line per property that drifted, which is precisely what an agent can act on.

npx pixelpact extract https://reference.example.com -o contract.json
npx pixelpact check contract.json http://localhost:3000

Works with

MCP clients

Playwright

GitHub Actions

Any framework

Figma

Claude Code, Cursor, anything that speaks stdio MCP

the engine it drives

one comment on the pull request

it reads the DOM, not your stack

a frame as the reference, when you want one

If a browser can render it, pixelpact can measure it.

Related MCP server: websight

For coding agents

An agent that writes UI code cannot tell whether it succeeded. pixelpact-mcp gives it the measurement, so the loop closes without a person in the middle: extract the contract once, then let the agent check its own work, read the deviation list, fix, and check again.

// .mcp.json
{
  "mcpServers": {
    "pixelpact": { "command": "npx", "args": ["-y", "pixelpact-mcp"] }
  }
}

Tools exposed: extract_contract, check_implementation, diff_pixels, read_contract_summary.

Why this is not visual regression testing

Percy, Chromatic, Applitools, BackstopJS and Pixeleye all compare your page against a baseline that you approved earlier. That model works once the page already looks right and you want to keep it that way. While you are still building toward a design, there is no baseline to compare with, and the first run of a regression tool simply records whatever you produced.

Visual regression tools

pixelpact

Compares against

a snapshot you approved earlier

the reference design itself

Useful when

the UI is already correct

the UI is being built

First run on a new page

records, cannot judge

measures against the reference

Answer you get

an image diff to inspect by eye

a value per property, with a delta

Fits an autonomous agent

needs a human to approve the diff

the numbers close the loop

The two models are complementary. Use a regression tool to keep a finished page finished, and pixelpact to get it finished in the first place.

Problems pixelpact solves

Without pixelpact

With pixelpact

❌ A coding agent writes CSS, declares it done, and a person has to open the page to find out that it is not.

✅ The agent calls the MCP server, reads the deviations, fixes them and checks again. No person in the middle.

❌ Someone says the section is finished, you disagree, and neither of you has a number. The louder opinion wins.

✅ Every value in the reference is asserted against your page. What failed is listed with expected, actual and the difference.

❌ A regression tool has nothing to compare a brand new page against, so its first run records whatever you happened to build.

✅ The reference is the baseline from the first minute. Nothing has to be approved before the tool is useful.

❌ Hover, focus and animation values are almost never reviewed, because checking them by hand is slow and boring.

✅ They are part of the contract, so they are measured on every run like any other value.

❌ An image diff tells you that something changed and leaves you to hunt for what.

✅ Deviations name the element, the property, the expected value and the measured one.

❌ The design lives in a file nobody opens during code review.

✅ The contract is committed JSON, so a pull request shows exactly which values moved.

Install

pnpm add -D pixelpact playwright
pnpm exec playwright install chromium

playwright is a peer dependency, so the browser download stays under your control and a project that already has Playwright installs nothing extra. Node 22.12 or newer is required.

Quickstart

1. Extract the contract from the reference.

npx pixelpact extract https://reference.example.com \
  --selector "main" \
  --viewport desktop,mobile \
  --screenshots .pixelpact/shots \
  -o contract.json

2. Measure your implementation.

npx pixelpact check contract.json http://localhost:3000 --viewport desktop

3. Fix what it reports, then run it again. The command exits 0 when everything is inside tolerance and 1 when it is not, so it drops straight into a script or a CI job.

What a contract looks like

A contract is plain JSON, readable and diffable, with no proprietary format and no service behind it.

{
  "version": 1,
  "source": { "type": "url", "value": "https://reference.example.com" },
  "root": "main",
  "extractedAt": "2026-09-06T01:20:44.812Z",
  "viewports": [{ "name": "desktop", "width": 1440, "height": 900 }],
  "tokens": { "--brand-600": "rgb(11, 114, 133)" },
  "keyframes": { "fade-up": [{ "offset": "0%", "css": "opacity: 0" }] },
  "byViewport": {
    "desktop": {
      "documentHeight": 4218,
      "elements": [
        {
          "selector": "main > header > a.cta",
          "tag": "a",
          "text": "Get started",
          "box": { "x": 120, "y": 32, "w": 148, "h": 44 },
          "styles": {
            "background-color": "rgb(11, 114, 133)",
            "font-size": "16px",
            "border-radius": "8px"
          },
          "hover": { "background-color": "rgb(8, 90, 105)" },
          "focus": { "outline": "2px solid rgb(11, 114, 133)" }
        }
      ]
    }
  }
}

Because it is a file, you can commit it, review it in a pull request, hand it to another developer, or hand it to an agent.

What a check prints

pixelpact check  FAILED
  target    http://localhost:4173/impl.html
  reference http://localhost:4173/ref.html
  viewport  desktop 1440x900
  elements  14 matched, 0 missing of 14
  checks    1056 passed, 10 failed (99.1% of 1066)

deviations (10)
SELECTOR          PROPERTY                  EXPECTED                ACTUAL                  DIFF
body > main > h1  font-size                 48px                    44px                    4px
body > main > a   box.width                 117.75px                109.75px                8px
body > main > a   padding-right             24px                    20px                    4px
body > main > a   padding-left              24px                    20px                    4px
body > main > a   background-color          rgb(11, 114, 133)       rgb(37, 99, 235)        60.1 (color)
body > main > a   border-top-left-radius    8px                     4px                     4px
body > main > a   border-top-right-radius   8px                     4px                     4px
body > main > a   border-bottom-left-ra...  8px                     4px                     4px
body > main > a   border-bottom-right-r...  8px                     4px                     4px
body > main > a   focus.outline             rgb(11, 114, 133) s...  rgb(37, 99, 235) so...  differs

That is a real run against two copies of one page with four declarations changed. Four edits produce ten measured deviations, because padding moves the box width and one border-radius shorthand sets four corners. Run the same check against the reference itself and all 1066 assertions pass, which is the property that matters: a passing check has to mean something.

Add --json to get the same report as a data structure, which is what CI jobs and agents read.

What a side by side run shows

check says which values moved. diff says how many pixels moved. Neither tells a person where to look. side splits both pages into sections, puts them next to each other, and boxes what differs.

npx pixelpact side https://reference.example.com http://localhost:3000 --widths 1440,390
#   SECTION   WIDTH   VERDICT  DIFF
01  hero      1440px  PASS     0.000%
02  features  1440px  FAIL     0.675%
  .pixelpact/side/1440/02-features.png
03  pricing   1440px  PASS     0.000%
04  foot      1440px  FAIL     1.265%
  .pixelpact/side/1440/04-foot.png

One section compared side by side, differences boxed in red

That is a real run against the two files in examples/side, which are copies of one page where the card gap and the corner radius were changed. Reproduce it with:

npx pixelpact side "file://$PWD/examples/side/reference.html" \
                   "file://$PWD/examples/side/implementation.html" --widths 1440

This is the command to run before telling anyone that a page is finished.

Features

Commands

Command

What it does

pixelpact extract <url>

Reads the reference and writes a contract file

pixelpact check <contract> <url>

Measures an implementation, prints deviations, sets exit code

pixelpact diff <contract> <url>

Pixel comparison against the screenshot stored in the contract

pixelpact side <reference> <url>

Section by section side by side images with the differences boxed

Shared flags cover the browser context (--viewport, --selector, --wait, --timeout, --locale, --timezone, --channel, --headful), the output (--out, --json, --quiet), and the parts of extraction you may want to cap (--max-elements, --max-states, --mask). Run pixelpact <command> --help for the full list.

Exit codes

Code

Meaning

0

Everything inside tolerance

1

Deviations found, or the pixel threshold was exceeded

2

Usage error, for example a bad flag or a contract file that is missing

3

Runtime failure, for example no browser available or the page would not load

Programmatic use

import { extract, check, formatCheckReport, writeContract } from 'pixelpact-core'

const contract = await extract({
  url: 'https://reference.example.com',
  selector: 'main',
  screenshotDir: '.pixelpact/shots',
})
await writeContract('contract.json', contract)

const report = await check(contract, { url: 'http://localhost:3000' })

console.log(formatCheckReport(report, { color: true }))
if (!report.ok) process.exitCode = 1

Everything is typed, and the contract and report shapes are validated at the boundary, so a malformed file fails with a readable message instead of a stack trace.

In CI

The Action measures a preview deployment against the contract committed in the repository and keeps a single pull request comment up to date instead of adding one per push.

- uses: jamalkamaladdin/pixelpact/action@v0
  with:
    contract: contract.json
    url: ${{ steps.preview.outputs.url }}
    viewport: desktop
    tolerance: 1

It needs permissions: pull-requests: write and no secret beyond the automatic GITHUB_TOKEN. Every input, every output and a complete workflow are in action/README.md.

Reading a Figma frame

extract recognises a Figma url and reads the frame through the REST API. No browser is launched for this step.

export FIGMA_TOKEN=figd_...
npx pixelpact extract "https://www.figma.com/design/KEY/Name?node-id=12-345" -o contract.json
npx pixelpact check contract.json http://localhost:3000

A Figma layer has no CSS selector, so a Figma contract binds to your markup through data-contract attributes. Name the element after the layer and matching stops depending on how the design happened to nest its frames:

<a class="btn btn-primary" data-contract="Hero/CTA">Get started</a>

Anything with no match is reported as missing rather than guessed from tag names.

Published styles come across as tokens: a color style as its color, a text style as a css font shorthand such as 600 60px/72px Geist, an effect style as a box shadow. A fill that is a gradient or an image has no single css value, so it is left out and named in the warnings rather than approximated.

How it works

  1. Playwright opens the reference at each requested viewport, waits for fonts, network and paint to settle, and dismisses cookie overlays.

  2. A single function is evaluated inside the page. It walks the DOM under your root selector and records the computed style of every visible element, plus custom properties, keyframes, and the geometry of each box.

  3. Interactive elements are hovered and focused, and only the properties that actually change are stored, so a contract stays small enough to read.

  4. Checking repeats step 2 against your implementation, matches elements by data-contract attribute, then by selector, then by tag and text, and compares property by property. Lengths use a pixel tolerance, colors use a perceptual distance, and animations are compared by name and timing rather than by string equality.

Repository layout

Package

Published as

What it is

packages/core

pixelpact-core

Extraction, checking, reporting

packages/cli

pixelpact

The pixelpact command

packages/mcp

pixelpact-mcp

MCP server for coding agents

action

used from GitHub

Action for pull request checks

FAQ

How does an agent use it? Install pixelpact-mcp, point the MCP client at it, extract the contract once, then let the agent call check_implementation after every edit and read the deviation table it gets back.

How is this different from Percy or Chromatic? They compare your page against a snapshot you approved earlier, which assumes the page is already right. pixelpact compares it against the reference, which is what you actually have while you are still building. The two fit together: pixelpact to get the page correct, a regression tool to keep it that way.

Do I need a baseline? No. The reference is the baseline, and that is the entire point.

My markup does not match the reference structure. Will anything match? Element matching tries the data-contract attribute first, then the selector, then the tag and its text. Put data-contract="hero-cta" on your element and matching stops depending on how the reference happened to nest its divs.

The page has content that changes on every load. Will a check ever pass? Mask it. --mask ".carousel" keeps a region out of the pixel comparison, and --max-elements stops a long page from producing a contract nobody can read.

Which browsers? Chromium through Playwright. Running a matrix across engines is out of scope on purpose: this tool measures agreement with a design, not differences between browsers.

Does a passing check mean the page is correct? It means every value in the contract matched inside tolerance. Elements that exist in your page but not in the reference are not flagged, because the contract only describes what the reference contains. Run diff as well when nothing extra is allowed.

Status

Everything described above is built and every number shown came from a real run: extraction from a live page and from Figma, checking, the pixel diff, the side by side images, the MCP server and the Action.

Version 0.3. Settled enough to use, not settled enough to promise: the contract format will gain fields before 1.0. If a value you need is missing from it, open an issue and say which one, because that is the fastest way for it to appear.

Contributing

Development setup, the commands, and the pull request flow are in CONTRIBUTING.md. Security reports go through SECURITY.md.

License

MIT © Jamal Kamaladdin

Available Tools

4 tools
check_implementationCheck an implementation against a contractA
Read-only

Measures an implementation URL against a previously saved visual contract: box position and size, computed styles, hover and focus states, and pseudo-elements. Call this after building or changing a UI to verify it matches the reference without a human looking at it. Needs contractPath from a prior extract_contract call. Returns a formatted deviation table (capped at 40 rows, with a note on how many were left out) plus the numeric totals, the pass rate, any missing selectors, and ok, which is true only when nothing deviated and nothing is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe implementation URL to measure against the contract.
waitNoExtra settle time in milliseconds after navigation. Defaults to 2000.
timeoutNoNavigation timeout in milliseconds. Defaults to 30000.
headlessNoRun the browser headless. Defaults to true.
selectorNoCSS selector to scope the check to. Defaults to the contract root.
viewportNoViewport name to check, for example desktop. Defaults to the first viewport in the contract.
maxStatesNoMaximum interactive elements to probe. Defaults to 120.
toleranceNoAllowed pixel tolerance for box and length values. Defaults to 1.
contractPathYesPath to a contract JSON file previously written by extract_contract.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds substantial behavioral detail beyond annotations: the exact return shape includes a deviation table capped at 40 rows with an overflow note, numeric totals, pass rate, missing selectors, and the precise meaning of ok. This is rich disclosure for a tool without an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences with no filler. The purpose is front-loaded, followed by usage guidance, prerequisite, and return details. Every sentence contributes distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description compensates by explaining the return contents in detail, including the 40-row cap and ok semantics. Annotations cover safety and open-world behavior, and the prerequisite contractPath origin is stated. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all nine parameters, including defaults and constraints. The description only reiterates that contractPath comes from a prior extract_contract call, which is already stated in the schema description for that parameter. Baseline 3 is appropriate when the schema carries parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: measures an implementation URL against a saved visual contract. It enumerates the visual dimensions checked (box position/size, computed styles, hover/focus states, pseudo-elements) and distinguishes itself from the sibling extract_contract by naming the contract as a required prior artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use guidance: call after building or changing a UI to verify against a reference. It also states the prerequisite that contractPath must come from a prior extract_contract call. It does not explicitly compare against the sibling diff_pixels or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_pixelsPixel-diff an implementation against a contract screenshotA
Read-only

Renders the implementation URL and compares it pixel by pixel against the reference screenshot stored in the contract. Use this for a stricter visual check than check_implementation, after the structural check passes or when a subtle rendering difference is suspected. Requires the contract to have been extracted with screenshotDir set. Returns the percentage of differing pixels, the threshold it was compared against, ok, and the filesystem path of the generated diff image.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe implementation URL to capture and compare pixel by pixel.
waitNoExtra settle time in milliseconds after navigation. Defaults to 2000.
masksNoCSS selectors to blank out before comparing, for example ads or timestamps.
outDirNoDirectory to write the diff image to. Defaults to the OS temp directory.
timeoutNoNavigation timeout in milliseconds. Defaults to 30000.
headlessNoRun the browser headless. Defaults to true.
selectorNoCSS selector to scope the comparison to. Defaults to the contract root.
viewportNoViewport name to compare. Defaults to the first viewport in the contract.
thresholdNoAllowed percent of differing pixels before the comparison fails. Defaults to 0.5.
contractPathYesPath to a contract JSON file previously written by extract_contract.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly and openWorld, so the safety profile is covered. The description adds real value beyond that: the screenshotDir precondition and the concrete return payload (percent differing pixels, threshold, ok, diff image path). It does not mention that a diff image file is written to disk as a side effect, which is a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the core action, then routing guidance, then the precondition, then the return contract. No filler and nothing redundant with structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-param tool with no output schema, the description covers the essential gap: it enumerates the return values so the agent knows what a result contains, plus the precondition for the call to succeed. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each of the 10 parameters carries its own description with defaults, so the schema does the heavy lifting. The description adds no parameter-level syntax or format detail beyond it; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (renders and compares pixel by pixel) and resource (implementation URL vs. contract reference screenshot). It explicitly positions itself against the sibling check_implementation as a 'stricter visual check', so an agent can select between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions ('after the structural check passes or when a subtle rendering difference is suspected') and names the alternative tool. It also states the prerequisite that the contract must have been extracted with screenshotDir set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_contractExtract a visual contractA

Extracts a visual contract from a reference URL and saves it as JSON at outputPath. Call this once per reference design, before check_implementation or diff_pixels can be used, since both need a saved contract to compare against. Pass screenshotDir to also capture reference screenshots, which diff_pixels requires later. Returns a summary: the element count per viewport, any extraction warnings, and the path the contract was written to.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe reference URL to extract a visual contract from.
waitNoExtra settle time in milliseconds after navigation. Defaults to 2000.
masksNoCSS selectors to exclude from the walk, for example ads or timestamps.
timeoutNoNavigation timeout in milliseconds. Defaults to 30000.
fullPageNoCapture the full scrollable page instead of only the viewport. Defaults to true.
headlessNoRun the browser headless. Defaults to true.
selectorNoCSS selector to scope the walk to. Defaults to the document body.
maxStatesNoMaximum interactive elements to probe for hover and focus. Defaults to 120.
viewportsNoViewports to capture. Defaults to desktop 1440x900, tablet 768x1024, mobile 390x844.
outputPathYesFilesystem path to write the contract JSON to, for example ./contracts/home.json.
maxElementsNoMaximum elements to walk, 0 means unbounded. Defaults to 600.
screenshotDirNoDirectory to save reference screenshots to. Required later for diff_pixels to work.
freezeAnimationsNoFreeze CSS animations before measuring. Defaults to true.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true; the description is consistent with that, disclosing the file write to outputPath, the optional screenshot capture, and the shape of the returned summary (element count per viewport, warnings, written path). It stops short of permissions, overwrite behavior, or runtime cost for a browser-driven extraction, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the action and side effect before the workflow constraints. Every sentence contributes routing or dependency information; only the return-summary sentence could arguably be trimmed if an output schema existed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter, no-output-schema tool, the description covers purpose, ordering relative to all three siblings, the downstream screenshot dependency, and the return summary. What remains thin is handling of the many extraction-tuning parameters, though the schema documents them fully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds workflow meaning beyond the schema for two parameters: outputPath is where the contract JSON lands, and screenshotDir is a downstream dependency for diff_pixels rather than merely a directory. The other eleven parameters are only covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Extracts a visual contract from a reference URL') plus the persistence side effect ('saves it as JSON at outputPath'). It names the siblings it precedes (check_implementation, diff_pixels), so an agent can place it in the workflow without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit ordering rule: 'Call this once per reference design, before check_implementation or diff_pixels can be used.' It also gives conditional guidance for a specific parameter ('Pass screenshotDir to also capture reference screenshots, which diff_pixels requires later'), which is exactly the when/when-not information an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_contract_summaryRead a saved contract summaryA
Read-only

Reads a saved contract file and describes it without opening a browser or making any network request. Use this to inspect what a contract covers, for example before deciding whether diff_pixels is possible, since that needs screenshots to already exist. Returns the original source URL, when it was extracted, the viewports it covers, the element count per viewport, whether reference screenshots exist, and any warnings recorded at extraction time.

ParametersJSON Schema
NameRequiredDescriptionDefault
contractPathYesPath to a contract JSON file previously written by extract_contract.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description reinforces this with 'without opening a browser or making any network request' — consistent, not contradictory. It goes further by enumerating the returned fields (source URL, extraction time, viewports, element count per viewport, screenshot existence, extraction warnings), which is genuinely useful since there is no output schema. Minor gap: it does not say what happens when the contract file is missing or malformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and scope, then the usage rationale, then the return shape. Every sentence carries information an agent would otherwise have to guess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a single well-documented parameter, no output schema, and a read-only annotation profile, the description covers the remaining gaps by describing return contents and the no-network behavior. Nothing essential to correct invocation appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and contractPath is fully documented in the schema including its link to extract_contract, so the baseline is 3. The description adds no syntax, format, or path-resolution detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Reads a saved contract file and describes it') and adds a distinguishing negative claim: no browser, no network request. It also names the sibling tool diff_pixels and the condition linking them, so an agent can separate this from extract_contract/check_implementation without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the usage context ('Use this to inspect what a contract covers') and a concrete decision point ('before deciding whether diff_pixels is possible, since that needs screenshots to already exist'). The alternative and its precondition are stated rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcheck_implementation
    • First observeddiff_pixels
    • First observedextract_contract
    • First observedread_contract_summary

TDQS

A4.5/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct role in a clear pipeline: extract_contract (capture), check_implementation (structural compare), diff_pixels (pixel compare), and read_contract_summary (inspect metadata). The descriptions explicitly differentiate the two comparison tools and state their prerequisites, leaving little chance of misselection.

Naming Consistency5/5

All four names follow a consistent snake_case verb_noun pattern (extract_contract, check_implementation, diff_pixels, read_contract_summary). The convention is predictable and readable throughout.

Tool Count4/5

Four tools is a tight, well-scoped set that maps cleanly onto the extract-then-verify workflow, with no redundant operations. It is slightly lean — no management tools (list/delete/update contracts) — but each tool clearly earns its place.

Completeness4/5

The core lifecycle is covered end to end: extract a contract, structurally verify, pixel-verify, and inspect a saved contract. Minor gaps exist (no way to list or delete contracts, no threshold/viewport configuration tool), but these are workarounds an agent can manage.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers