pixelpact
Allows extracting a design contract from a Figma frame via the REST API and using the frame as the reference for implementation checks.
Allows running pixelpact checks in GitHub Actions and posting a comment on the pull request with the results.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pixelpactCompare my checkout page to the Figma design and report the deviations."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Your coding agent cannot see the page it just built. pixelpact measures it against the reference and hands back numbers.
A coding agent writes the CSS and says it is done. It never sees the result. Nothing hands it a number, so the loop cannot close: the agent guesses, you open the page, you send it back, it guesses again. Every round costs you the one thing the agent was supposed to save.
The same gap exists without an agent, when a teammate says a section is finished and nobody in the room has a measurement. The difference is that a person can at least look.
pixelpact reads the reference page and writes down what it actually renders: sizes, colors, spacing, typography, hover and focus states, animation keyframes, design tokens. That file is the contract. Point pixelpact at the implementation and it answers with one line per property that drifted, which is precisely what an agent can act on.
npx pixelpact extract https://reference.example.com -o contract.json
npx pixelpact check contract.json http://localhost:3000Works with
MCP clients | Playwright | GitHub Actions | Any framework | Figma |
Claude Code, Cursor, anything that speaks stdio MCP | the engine it drives | one comment on the pull request | it reads the DOM, not your stack | a frame as the reference, when you want one |
If a browser can render it, pixelpact can measure it.
Related MCP server: websight
For coding agents
An agent that writes UI code cannot tell whether it succeeded. pixelpact-mcp gives it the
measurement, so the loop closes without a person in the middle: extract the contract once,
then let the agent check its own work, read the deviation list, fix, and check again.
// .mcp.json
{
"mcpServers": {
"pixelpact": { "command": "npx", "args": ["-y", "pixelpact-mcp"] }
}
}Tools exposed: extract_contract, check_implementation, diff_pixels,
read_contract_summary.
Why this is not visual regression testing
Percy, Chromatic, Applitools, BackstopJS and Pixeleye all compare your page against a baseline that you approved earlier. That model works once the page already looks right and you want to keep it that way. While you are still building toward a design, there is no baseline to compare with, and the first run of a regression tool simply records whatever you produced.
Visual regression tools | pixelpact | |
Compares against | a snapshot you approved earlier | the reference design itself |
Useful when | the UI is already correct | the UI is being built |
First run on a new page | records, cannot judge | measures against the reference |
Answer you get | an image diff to inspect by eye | a value per property, with a delta |
Fits an autonomous agent | needs a human to approve the diff | the numbers close the loop |
The two models are complementary. Use a regression tool to keep a finished page finished, and pixelpact to get it finished in the first place.
Problems pixelpact solves
Without pixelpact | With pixelpact |
❌ A coding agent writes CSS, declares it done, and a person has to open the page to find out that it is not. | ✅ The agent calls the MCP server, reads the deviations, fixes them and checks again. No person in the middle. |
❌ Someone says the section is finished, you disagree, and neither of you has a number. The louder opinion wins. | ✅ Every value in the reference is asserted against your page. What failed is listed with expected, actual and the difference. |
❌ A regression tool has nothing to compare a brand new page against, so its first run records whatever you happened to build. | ✅ The reference is the baseline from the first minute. Nothing has to be approved before the tool is useful. |
❌ Hover, focus and animation values are almost never reviewed, because checking them by hand is slow and boring. | ✅ They are part of the contract, so they are measured on every run like any other value. |
❌ An image diff tells you that something changed and leaves you to hunt for what. | ✅ Deviations name the element, the property, the expected value and the measured one. |
❌ The design lives in a file nobody opens during code review. | ✅ The contract is committed JSON, so a pull request shows exactly which values moved. |
Install
pnpm add -D pixelpact playwright
pnpm exec playwright install chromiumplaywright is a peer dependency, so the browser download stays under your control and a
project that already has Playwright installs nothing extra. Node 22.12 or newer is required.
Quickstart
1. Extract the contract from the reference.
npx pixelpact extract https://reference.example.com \
--selector "main" \
--viewport desktop,mobile \
--screenshots .pixelpact/shots \
-o contract.json2. Measure your implementation.
npx pixelpact check contract.json http://localhost:3000 --viewport desktop3. Fix what it reports, then run it again. The command exits 0 when everything is inside
tolerance and 1 when it is not, so it drops straight into a script or a CI job.
What a contract looks like
A contract is plain JSON, readable and diffable, with no proprietary format and no service behind it.
{
"version": 1,
"source": { "type": "url", "value": "https://reference.example.com" },
"root": "main",
"extractedAt": "2026-09-06T01:20:44.812Z",
"viewports": [{ "name": "desktop", "width": 1440, "height": 900 }],
"tokens": { "--brand-600": "rgb(11, 114, 133)" },
"keyframes": { "fade-up": [{ "offset": "0%", "css": "opacity: 0" }] },
"byViewport": {
"desktop": {
"documentHeight": 4218,
"elements": [
{
"selector": "main > header > a.cta",
"tag": "a",
"text": "Get started",
"box": { "x": 120, "y": 32, "w": 148, "h": 44 },
"styles": {
"background-color": "rgb(11, 114, 133)",
"font-size": "16px",
"border-radius": "8px"
},
"hover": { "background-color": "rgb(8, 90, 105)" },
"focus": { "outline": "2px solid rgb(11, 114, 133)" }
}
]
}
}
}Because it is a file, you can commit it, review it in a pull request, hand it to another developer, or hand it to an agent.
What a check prints
pixelpact check FAILED
target http://localhost:4173/impl.html
reference http://localhost:4173/ref.html
viewport desktop 1440x900
elements 14 matched, 0 missing of 14
checks 1056 passed, 10 failed (99.1% of 1066)
deviations (10)
SELECTOR PROPERTY EXPECTED ACTUAL DIFF
body > main > h1 font-size 48px 44px 4px
body > main > a box.width 117.75px 109.75px 8px
body > main > a padding-right 24px 20px 4px
body > main > a padding-left 24px 20px 4px
body > main > a background-color rgb(11, 114, 133) rgb(37, 99, 235) 60.1 (color)
body > main > a border-top-left-radius 8px 4px 4px
body > main > a border-top-right-radius 8px 4px 4px
body > main > a border-bottom-left-ra... 8px 4px 4px
body > main > a border-bottom-right-r... 8px 4px 4px
body > main > a focus.outline rgb(11, 114, 133) s... rgb(37, 99, 235) so... differsThat is a real run against two copies of one page with four declarations changed. Four edits
produce ten measured deviations, because padding moves the box width and one border-radius
shorthand sets four corners. Run the same check against the reference itself and all 1066
assertions pass, which is the property that matters: a passing check has to mean something.
Add --json to get the same report as a data structure, which is what CI jobs and agents read.
What a side by side run shows
check says which values moved. diff says how many pixels moved. Neither tells a person
where to look. side splits both pages into sections, puts them next to each other, and boxes
what differs.
npx pixelpact side https://reference.example.com http://localhost:3000 --widths 1440,390# SECTION WIDTH VERDICT DIFF
01 hero 1440px PASS 0.000%
02 features 1440px FAIL 0.675%
.pixelpact/side/1440/02-features.png
03 pricing 1440px PASS 0.000%
04 foot 1440px FAIL 1.265%
.pixelpact/side/1440/04-foot.png![]()
That is a real run against the two files in examples/side, which are
copies of one page where the card gap and the corner radius were changed. Reproduce it with:
npx pixelpact side "file://$PWD/examples/side/reference.html" \
"file://$PWD/examples/side/implementation.html" --widths 1440This is the command to run before telling anyone that a page is finished.
Features
Commands
Command | What it does |
| Reads the reference and writes a contract file |
| Measures an implementation, prints deviations, sets exit code |
| Pixel comparison against the screenshot stored in the contract |
| Section by section side by side images with the differences boxed |
Shared flags cover the browser context (--viewport, --selector, --wait, --timeout,
--locale, --timezone, --channel, --headful), the output (--out, --json,
--quiet), and the parts of extraction you may want to cap (--max-elements, --max-states,
--mask). Run pixelpact <command> --help for the full list.
Exit codes
Code | Meaning |
| Everything inside tolerance |
| Deviations found, or the pixel threshold was exceeded |
| Usage error, for example a bad flag or a contract file that is missing |
| Runtime failure, for example no browser available or the page would not load |
Programmatic use
import { extract, check, formatCheckReport, writeContract } from 'pixelpact-core'
const contract = await extract({
url: 'https://reference.example.com',
selector: 'main',
screenshotDir: '.pixelpact/shots',
})
await writeContract('contract.json', contract)
const report = await check(contract, { url: 'http://localhost:3000' })
console.log(formatCheckReport(report, { color: true }))
if (!report.ok) process.exitCode = 1Everything is typed, and the contract and report shapes are validated at the boundary, so a malformed file fails with a readable message instead of a stack trace.
In CI
The Action measures a preview deployment against the contract committed in the repository and keeps a single pull request comment up to date instead of adding one per push.
- uses: jamalkamaladdin/pixelpact/action@v0
with:
contract: contract.json
url: ${{ steps.preview.outputs.url }}
viewport: desktop
tolerance: 1It needs permissions: pull-requests: write and no secret beyond the automatic
GITHUB_TOKEN. Every input, every output and a complete workflow are in
action/README.md.
Reading a Figma frame
extract recognises a Figma url and reads the frame through the REST API. No browser is
launched for this step.
export FIGMA_TOKEN=figd_...
npx pixelpact extract "https://www.figma.com/design/KEY/Name?node-id=12-345" -o contract.json
npx pixelpact check contract.json http://localhost:3000A Figma layer has no CSS selector, so a Figma contract binds to your markup through
data-contract attributes. Name the element after the layer and matching stops depending on
how the design happened to nest its frames:
<a class="btn btn-primary" data-contract="Hero/CTA">Get started</a>Anything with no match is reported as missing rather than guessed from tag names.
Published styles come across as tokens: a color style as its color, a text style as a css font
shorthand such as 600 60px/72px Geist, an effect style as a box shadow. A fill that is a
gradient or an image has no single css value, so it is left out and named in the warnings
rather than approximated.
How it works
Playwright opens the reference at each requested viewport, waits for fonts, network and paint to settle, and dismisses cookie overlays.
A single function is evaluated inside the page. It walks the DOM under your root selector and records the computed style of every visible element, plus custom properties, keyframes, and the geometry of each box.
Interactive elements are hovered and focused, and only the properties that actually change are stored, so a contract stays small enough to read.
Checking repeats step 2 against your implementation, matches elements by
data-contractattribute, then by selector, then by tag and text, and compares property by property. Lengths use a pixel tolerance, colors use a perceptual distance, and animations are compared by name and timing rather than by string equality.
Repository layout
Package | Published as | What it is |
| Extraction, checking, reporting | |
| The | |
| MCP server for coding agents | |
used from GitHub | Action for pull request checks |
FAQ
How does an agent use it? Install pixelpact-mcp, point the MCP client at it, extract the
contract once, then let the agent call check_implementation after every edit and read the
deviation table it gets back.
How is this different from Percy or Chromatic? They compare your page against a snapshot you approved earlier, which assumes the page is already right. pixelpact compares it against the reference, which is what you actually have while you are still building. The two fit together: pixelpact to get the page correct, a regression tool to keep it that way.
Do I need a baseline? No. The reference is the baseline, and that is the entire point.
My markup does not match the reference structure. Will anything match? Element matching
tries the data-contract attribute first, then the selector, then the tag and its text. Put
data-contract="hero-cta" on your element and matching stops depending on how the reference
happened to nest its divs.
The page has content that changes on every load. Will a check ever pass? Mask it.
--mask ".carousel" keeps a region out of the pixel comparison, and --max-elements stops a
long page from producing a contract nobody can read.
Which browsers? Chromium through Playwright. Running a matrix across engines is out of scope on purpose: this tool measures agreement with a design, not differences between browsers.
Does a passing check mean the page is correct? It means every value in the contract matched
inside tolerance. Elements that exist in your page but not in the reference are not flagged,
because the contract only describes what the reference contains. Run diff as well when
nothing extra is allowed.
Status
Everything described above is built and every number shown came from a real run: extraction from a live page and from Figma, checking, the pixel diff, the side by side images, the MCP server and the Action.
Version 0.3. Settled enough to use, not settled enough to promise: the contract format will gain fields before 1.0. If a value you need is missing from it, open an issue and say which one, because that is the fastest way for it to appear.
Contributing
Development setup, the commands, and the pull request flow are in CONTRIBUTING.md. Security reports go through SECURITY.md.
License
MIT © Jamal Kamaladdin
Available Tools
4 toolscheck_implementationCheck an implementation against a contractARead-only
Measures an implementation URL against a previously saved visual contract: box position and size, computed styles, hover and focus states, and pseudo-elements. Call this after building or changing a UI to verify it matches the reference without a human looking at it. Needs contractPath from a prior extract_contract call. Returns a formatted deviation table (capped at 40 rows, with a note on how many were left out) plus the numeric totals, the pass rate, any missing selectors, and ok, which is true only when nothing deviated and nothing is missing.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The implementation URL to measure against the contract. | |
| wait | No | Extra settle time in milliseconds after navigation. Defaults to 2000. | |
| timeout | No | Navigation timeout in milliseconds. Defaults to 30000. | |
| headless | No | Run the browser headless. Defaults to true. | |
| selector | No | CSS selector to scope the check to. Defaults to the contract root. | |
| viewport | No | Viewport name to check, for example desktop. Defaults to the first viewport in the contract. | |
| maxStates | No | Maximum interactive elements to probe. Defaults to 120. | |
| tolerance | No | Allowed pixel tolerance for box and length values. Defaults to 1. | |
| contractPath | Yes | Path to a contract JSON file previously written by extract_contract. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds substantial behavioral detail beyond annotations: the exact return shape includes a deviation table capped at 40 rows with an overflow note, numeric totals, pass rate, missing selectors, and the precise meaning of ok. This is rich disclosure for a tool without an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences with no filler. The purpose is front-loaded, followed by usage guidance, prerequisite, and return details. Every sentence contributes distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description compensates by explaining the return contents in detail, including the 40-row cap and ok semantics. Annotations cover safety and open-world behavior, and the prerequisite contractPath origin is stated. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters, including defaults and constraints. The description only reiterates that contractPath comes from a prior extract_contract call, which is already stated in the schema description for that parameter. Baseline 3 is appropriate when the schema carries parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: measures an implementation URL against a saved visual contract. It enumerates the visual dimensions checked (box position/size, computed styles, hover/focus states, pseudo-elements) and distinguishes itself from the sibling extract_contract by naming the contract as a required prior artifact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use guidance: call after building or changing a UI to verify against a reference. It also states the prerequisite that contractPath must come from a prior extract_contract call. It does not explicitly compare against the sibling diff_pixels or state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_pixelsPixel-diff an implementation against a contract screenshotARead-only
Renders the implementation URL and compares it pixel by pixel against the reference screenshot stored in the contract. Use this for a stricter visual check than check_implementation, after the structural check passes or when a subtle rendering difference is suspected. Requires the contract to have been extracted with screenshotDir set. Returns the percentage of differing pixels, the threshold it was compared against, ok, and the filesystem path of the generated diff image.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The implementation URL to capture and compare pixel by pixel. | |
| wait | No | Extra settle time in milliseconds after navigation. Defaults to 2000. | |
| masks | No | CSS selectors to blank out before comparing, for example ads or timestamps. | |
| outDir | No | Directory to write the diff image to. Defaults to the OS temp directory. | |
| timeout | No | Navigation timeout in milliseconds. Defaults to 30000. | |
| headless | No | Run the browser headless. Defaults to true. | |
| selector | No | CSS selector to scope the comparison to. Defaults to the contract root. | |
| viewport | No | Viewport name to compare. Defaults to the first viewport in the contract. | |
| threshold | No | Allowed percent of differing pixels before the comparison fails. Defaults to 0.5. | |
| contractPath | Yes | Path to a contract JSON file previously written by extract_contract. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly and openWorld, so the safety profile is covered. The description adds real value beyond that: the screenshotDir precondition and the concrete return payload (percent differing pixels, threshold, ok, diff image path). It does not mention that a diff image file is written to disk as a side effect, which is a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action, then routing guidance, then the precondition, then the return contract. No filler and nothing redundant with structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-param tool with no output schema, the description covers the essential gap: it enumerates the return values so the agent knows what a result contains, plus the precondition for the call to succeed. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each of the 10 parameters carries its own description with defaults, so the schema does the heavy lifting. The description adds no parameter-level syntax or format detail beyond it; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (renders and compares pixel by pixel) and resource (implementation URL vs. contract reference screenshot). It explicitly positions itself against the sibling check_implementation as a 'stricter visual check', so an agent can select between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions ('after the structural check passes or when a subtle rendering difference is suspected') and names the alternative tool. It also states the prerequisite that the contract must have been extracted with screenshotDir set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_contractExtract a visual contractA
Extracts a visual contract from a reference URL and saves it as JSON at outputPath. Call this once per reference design, before check_implementation or diff_pixels can be used, since both need a saved contract to compare against. Pass screenshotDir to also capture reference screenshots, which diff_pixels requires later. Returns a summary: the element count per viewport, any extraction warnings, and the path the contract was written to.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The reference URL to extract a visual contract from. | |
| wait | No | Extra settle time in milliseconds after navigation. Defaults to 2000. | |
| masks | No | CSS selectors to exclude from the walk, for example ads or timestamps. | |
| timeout | No | Navigation timeout in milliseconds. Defaults to 30000. | |
| fullPage | No | Capture the full scrollable page instead of only the viewport. Defaults to true. | |
| headless | No | Run the browser headless. Defaults to true. | |
| selector | No | CSS selector to scope the walk to. Defaults to the document body. | |
| maxStates | No | Maximum interactive elements to probe for hover and focus. Defaults to 120. | |
| viewports | No | Viewports to capture. Defaults to desktop 1440x900, tablet 768x1024, mobile 390x844. | |
| outputPath | Yes | Filesystem path to write the contract JSON to, for example ./contracts/home.json. | |
| maxElements | No | Maximum elements to walk, 0 means unbounded. Defaults to 600. | |
| screenshotDir | No | Directory to save reference screenshots to. Required later for diff_pixels to work. | |
| freezeAnimations | No | Freeze CSS animations before measuring. Defaults to true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true; the description is consistent with that, disclosing the file write to outputPath, the optional screenshot capture, and the shape of the returned summary (element count per viewport, warnings, written path). It stops short of permissions, overwrite behavior, or runtime cost for a browser-driven extraction, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and side effect before the workflow constraints. Every sentence contributes routing or dependency information; only the return-summary sentence could arguably be trimmed if an output schema existed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter, no-output-schema tool, the description covers purpose, ordering relative to all three siblings, the downstream screenshot dependency, and the return summary. What remains thin is handling of the many extraction-tuning parameters, though the schema documents them fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds workflow meaning beyond the schema for two parameters: outputPath is where the contract JSON lands, and screenshotDir is a downstream dependency for diff_pixels rather than merely a directory. The other eleven parameters are only covered by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Extracts a visual contract from a reference URL') plus the persistence side effect ('saves it as JSON at outputPath'). It names the siblings it precedes (check_implementation, diff_pixels), so an agent can place it in the workflow without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit ordering rule: 'Call this once per reference design, before check_implementation or diff_pixels can be used.' It also gives conditional guidance for a specific parameter ('Pass screenshotDir to also capture reference screenshots, which diff_pixels requires later'), which is exactly the when/when-not information an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_contract_summaryRead a saved contract summaryARead-only
Reads a saved contract file and describes it without opening a browser or making any network request. Use this to inspect what a contract covers, for example before deciding whether diff_pixels is possible, since that needs screenshots to already exist. Returns the original source URL, when it was extracted, the viewports it covers, the element count per viewport, whether reference screenshots exist, and any warnings recorded at extraction time.
| Name | Required | Description | Default |
|---|---|---|---|
| contractPath | Yes | Path to a contract JSON file previously written by extract_contract. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, and the description reinforces this with 'without opening a browser or making any network request' — consistent, not contradictory. It goes further by enumerating the returned fields (source URL, extraction time, viewports, element count per viewport, screenshot existence, extraction warnings), which is genuinely useful since there is no output schema. Minor gap: it does not say what happens when the contract file is missing or malformed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope, then the usage rationale, then the return shape. Every sentence carries information an agent would otherwise have to guess.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a single well-documented parameter, no output schema, and a read-only annotation profile, the description covers the remaining gaps by describing return contents and the no-network behavior. Nothing essential to correct invocation appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and contractPath is fully documented in the schema including its link to extract_contract, so the baseline is 3. The description adds no syntax, format, or path-resolution detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reads a saved contract file and describes it') and adds a distinguishing negative claim: no browser, no network request. It also names the sibling tool diff_pixels and the condition linking them, so an agent can separate this from extract_contract/check_implementation without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the usage context ('Use this to inspect what a contract covers') and a concrete decision point ('before deciding whether diff_pixels is possible, since that needs screenshots to already exist'). The alternative and its precondition are stated rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
check_implementation - First observed
diff_pixels - First observed
extract_contract - First observed
read_contract_summary
TDQS
Scored across 4 tools
Each tool has a distinct role in a clear pipeline: extract_contract (capture), check_implementation (structural compare), diff_pixels (pixel compare), and read_contract_summary (inspect metadata). The descriptions explicitly differentiate the two comparison tools and state their prerequisites, leaving little chance of misselection.
All four names follow a consistent snake_case verb_noun pattern (extract_contract, check_implementation, diff_pixels, read_contract_summary). The convention is predictable and readable throughout.
Four tools is a tight, well-scoped set that maps cleanly onto the extract-then-verify workflow, with no redundant operations. It is slightly lean — no management tools (list/delete/update contracts) — but each tool clearly earns its place.
The core lifecycle is covered end to end: extract a contract, structurally verify, pixel-verify, and inspect a saved contract. Minor gaps exist (no way to list or delete contracts, no threshold/viewport configuration tool), but these are workarounds an agent can manage.
Maintenance
Related MCP Connectors
- WhoogyOAuthcom.whoogy
Compare Figma designs against live websites and get visual + content QA reports, from inside Claude.
- miromiroOAuthapp.miromiro
Turn any live website into brand colors, fonts, design tokens, SVGs, Lottie and paste-ready code.
Score any URL against a real design contract — 42 checks, A-F grade, token + motion validation.
On-demand drift checks: declared CSS color, radius, spacing & type vs your own tokens or a pack
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI to capture, compare, and automatically patch frontend code against reference designs, achieving pixel-perfect fidelity without manual CSS tweaking.24 npm10MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.0MIT
- AlicenseNot gradedqualityDmaintenanceVerifies that AI-generated UI code matches Figma design specs by rendering components in a real browser, comparing computed CSS, and returning patch-ready fixes with scored parity reports.MIT
- FlicenseAqualityBmaintenanceCompares Figma frames to live pages, checking colors, fonts, and border radii, and generates a shareable HTML report.6-