Designesy
Server Details
Score any URL against a real design contract — 42 checks, A-F grade, token + motion validation.
- Status
- Healthy
- Uptime
- 97.8% over 43 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- LE-VAI/designesy-org
- GitHub Stars
- 0
- Server Listing
- Designesy
TDQS
Scored across 17 tools
There are many score-type tools (score, drift_score, readiness_score, monitor_score, tokens_score, motion_score, a11y_score, report), which could be confused, but each description includes an explicit 'When NOT to use' clause pointing to the correct sibling (e.g. drift vs monitor, score vs readiness). The boundaries are well-drawn, though the sheer number of *_score variants still requires careful reading to disambiguate.
All tools share the designesy_ snake_case prefix, and the analytic tools consistently use a <subject>_score suffix. A few resource getters deviate to bare nouns (catalog, contract, report, guardrails) but remain readable and follow the same casing convention.
17 tools is slightly heavy but justified for a broad design-intelligence platform spanning a11y, drift, readiness, tokens, motion, guardrails, and discovery. Each tool maps to a distinct artifact or engine, though the score family pushes the upper end of the comfortable range.
The surface covers the full analysis lifecycle: contract retrieval, multi-axis scoring, drift/monitor temporal governance, token and motion validation, readiness probing, guardrail/bundle generation, and agent-facing exports (llms, skill_md, agent_json). It is largely read-only, which is consistent with the domain, so no major gaps are apparent.
Available Tools
17 toolsdesignesy_a11y_scoreAInspect
Get the Designesy WCAG 2.2 AA accessibility verification framework: 11 conformance checks (a01-a11) plus a ready-to-run Playwright + axe-core 4.13.0 script template targeting your URL. Use this to audit a site for accessibility violations. When NOT to use: for a full design-contract score (not just a11y), use designesy_score. Does NOT run the scan: axe-core needs a real browser DOM. Returns the 11 checks + a Playwright script you execute locally (npm i -D @axe-core/playwright). The score comes from your local run, not from this tool. Returns JSON: { checks[{id (a01-a11), name, status: "PENDING_EXECUTION"}], playwright_script, install_command, run_command }. Pass config (JSON string) to customize axe.configure(), for example with branding overrides or rule disables. Omit for standard WCAG 2.2 AA.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to scan for accessibility. The returned script template will target this URL. | |
| config | No | Brand customization JSON for axe.configure(): branding, checks, rules, disableOtherRules. | |
| ruleset | No | Ruleset tag (default: wcag22aa). Options: wcag2a, wcag2aa, wcag21aa, wcag22aa, best-practice. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does: it discloses that the tool does NOT execute the scan ('axe-core needs a real browser DOM'), that the score comes from a local run, the required install step (npm i -D @axe-core/playwright), and the exact return shape including status values of PENDING_EXECUTION. This is exactly the kind of non-obvious behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then exclusion, then the critical non-execution caveat. Slightly padded by both an inline 'Returns JSON: {...}' block and separate sentences covering install_command/run_command, which partially restate each other, but no sentence is truly wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with no annotations and no output schema, the description covers purpose, exclusions, the non-obvious non-execution model, the prerequisite install command, and the full response structure. Nothing an agent needs to invoke or interpret the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds genuine meaning on top by explaining the purpose of config ('customize axe.configure(), for example with branding overrides or rule disables') and the default behavior when omitted ('Omit for standard WCAG 2.2 AA'). It does not add anything for ruleset beyond what the schema already enumerates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Designesy WCAG 2.2 AA accessibility verification framework'), quantifies the deliverable (11 checks a01-a11 plus a Playwright + axe-core script template), and names the sibling it differs from (designesy_score). An agent can distinguish it from designesy_score, designesy_drift_score, and designesy_contract without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit both ways: 'Use this to audit a site for accessibility violations' and 'When NOT to use: for a full design-contract score (not just a11y), use designesy_score.' It also pre-empts the most likely misfire by stating the tool does not run the scan, so an agent cannot expect results from this call alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_agent_jsonAInspect
Get the Designesy agent discovery document (/.well-known/agent.json): the org identity, authority, ingest protocol, package index, machine-export list, permission policy, and citation templates. Use this when you are integrating with or enumerating Designesy as a machine agent and need the canonical discovery/manifest endpoint rather than one specific contract. When NOT to use: for the package list, use designesy_catalog (lighter); for the contract, use designesy_contract. Read-only: no side effects. Returns the /.well-known/agent.json object: { identity, authority, ingest_protocol, package_index, permission_policy, citation_templates }. No parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose 'Read-only: no side effects' plus the exact return shape and keys. It stops short of stating auth requirements, caching, or rate limits, so it is not exhaustive, but it covers the key behavioral facts for a zero-parameter GET.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the purpose, then usage, exclusion, side effects, and return shape in a logical order. There is mild redundancy: the payload keys listed in the opening sentence are repeated verbatim in the 'Returns' sentence, which costs a little length without adding information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single zero-param read endpoint with no output schema, the description supplies the canonical path, the full set of returned keys, the read-only guarantee, and sibling routing. An agent has everything needed to select and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, which is the baseline-4 case; the description confirms 'No parameters.' Nothing further is needed, and no meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Designesy agent discovery document (/.well-known/agent.json)') and enumerates the payload contents. It distinguishes itself from siblings by naming designesy_catalog and designesy_contract as the alternatives for narrower needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the use condition ('when you are integrating with or enumerating Designesy as a machine agent and need the canonical discovery/manifest endpoint rather than one specific contract') and an explicit 'When NOT to use' clause routing to designesy_catalog for the package list and designesy_contract for the contract.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_catalogAInspect
List the 23 published Designesy packages with versions, URLs, and statuses. Use this to discover what Designesy publishes before fetching a specific contract. When NOT to use: if you already know which package you need, skip this and call designesy_contract directly. Read-only: no side effects. Returns JSON: { package_count, packages[{id, kind, title, version, status, human_url, machine_url}], standing_rules[], machine_exports[] }. No parameters: accepts empty input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose the key traits: read-only, no side effects, no parameters accepted, and the exact JSON response shape. It stops short of things like rate limits or pagination/caching behavior, but for a static catalog listing the safety and output profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, then exclusions, safety, and return shape. Each sentence adds distinct information with no restatement of the name or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by enumerating the response fields (package_count, packages[], standing_rules[], machine_exports[]). Combined with the sibling routing and safety disclosure, nothing needed to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4; the description confirms 'No parameters: accepts empty input,' which removes any ambiguity about whether an empty call is valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the 23 published Designesy packages') and enumerates the returned fields (versions, URLs, statuses), which distinguishes it cleanly from the many scoring/report siblings and from designesy_contract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('discover what Designesy publishes before fetching a specific contract') and an explicit when-NOT-to-use with the named alternative ('if you already know which package you need, skip this and call designesy_contract directly'). This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_compareAInspect
Diff two design systems from live URLs, using the only URL-scoped design-token diff engine. Fetches both URLs in parallel, extracts their :root custom properties, and produces a structured diff across 8 dimensions: tokens added (in A not B), removed (in B not A), renamed (heuristic Levenshtein ≤ 2), value-changed (same name, different value), scale-stop-changed (spacing/radius/color scale steps), contrast-drift-per-pair (WCAG contrast ratio change for shared color tokens), structure-delta (token count + category distribution), and score-delta (runs /score on both URLs and diffs). Use this to answer "what actually changed between two design systems" or "how does our design system differ from a reference". When NOT to use: for single-site drift detection, use designesy_drift_score; for continuous monitoring, use designesy_monitor_score. Executable: fetches both URLs, extracts CSS + tokens, computes diff. No browser needed. Returns JSON: { ok, urlA, urlB, score (0-100, diff completeness), grade, pass, warn, fail, total, tokensA, tokensB, added[], removed[], renamed[], valueChanged[], scaleDiff, structureDelta, contrastDrift[], scoreDelta, checks[] }. Results cached ~24h per URL pair.
| Name | Required | Description | Default |
|---|---|---|---|
| urlA | Yes | First URL to compare (e.g. your design system). | |
| urlB | Yes | Second URL to compare (e.g. a reference or competitor). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses parallel fetching of both URLs, :root custom-property extraction, the Levenshtein <= 2 rename heuristic, no-browser execution, and ~24h per-URL-pair caching. It stops short of stating auth requirements, rate limits, or failure behavior when a URL is unreachable, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content is front-loaded with the core purpose before the dimensions, alternatives, and return shape. The long inline JSON return listing is verbose, but since no output schema exists it earns most of its space; a few clauses could still be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description takes on the job of enumerating return fields (score, grade, added[], renamed[], contrastDrift[], etc.), and it does so thoroughly alongside scope, exclusions, and caching semantics. Nothing an agent needs to invoke and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters already carry their own descriptions with examples ('your design system' / 'reference or competitor'), so the baseline is 3. The prose adds role framing ('live URLs') but no format or constraint detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Diff two design systems from live URLs') and immediately differentiates itself by naming the mechanism ('URL-scoped design-token diff engine') and enumerating the 8 diff dimensions. An agent can distinguish this from designesy_drift_score and designesy_monitor_score without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger phrasings ('what actually changed between two design systems') and an explicit 'When NOT to use' clause that routes single-site drift to designesy_drift_score and continuous monitoring to designesy_monitor_score. This is the full when/when-not/alternative pattern.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_contractAInspect
Get the Designesy design-system contract: the canonical tokens, motion, acoustic, takt, cadence, typography, components, and verification rules that define what the Designesy org considers legitimate design. Use this when you need the actual contract values (token names and values, motion timings, accessibility rules) to author, check, or bind a design. When NOT to use: for a pass/fail score of a live site, use designesy_score; for an agent-skill-format export, use designesy_skill_md. Read-only; cached ~24h server-side. Returns the full contract JSON, or a single section when "section" is provided. Pass section to get one slice (e.g. "motion" for just the motion tokens) instead of the full contract, which saves tokens when you only need one dimension.
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | Optional: filter to a specific contract section (colors, motion, acoustic, typography, takt, cadence, verification, open_tensions, components, interaction). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses key traits: read-only, cached ~24h server-side, and the return shape (full contract JSON or a single section). It does not mention auth/permission requirements or rate limits beyond caching, so it falls short of exhaustive disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, exclusions, behavior, and return semantics in a logical order. Every sentence earns its place; the section-parameter advice directly informs invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly covers the return format itself (full contract JSON or a single section). Combined with usage routing and caching notes, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the schema already documents the section parameter, giving a baseline of 3. The description adds genuine value beyond the schema by explaining the motivation ('saves tokens when you only need one dimension') and giving a concrete example ('motion' for just the motion tokens).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Designesy design-system contract') and enumerates what it contains (tokens, motion, acoustic, takt, cadence, typography, components, verification rules). It explicitly names sibling tools it is not (designesy_score, designesy_skill_md), so an agent can distinguish it without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('when you need the actual contract values ... to author, check, or bind a design') and when-NOT-to-use with named alternatives for each condition. Nothing is left to inference for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_design_reviewAInspect
Get the Designesy Design Review framework: an 8-dimension rubric (Purpose, Clarity, Context, Inclusion, System coherence, Durability, Delight, Responsibility) plus the agent prompt, output format, and verification checklist for a qualitative design critique. Use this when you want a structured rubric to critique a design holistically, rather than a numeric compliance score. When NOT to use: for a deterministic numeric score, use designesy_score; this tool gives you a rubric, not a number. Read-only: returns the rubric + prompt. The calling agent performs the actual critique (this tool does not evaluate the design for you). Returns JSON: { rubric, dimensions[8], agent_prompt, output_format, verification_checklist }. Pass artifact/purpose/context/rules to get a pre-filled critique prompt; omit all four to get the blank framework.
| Name | Required | Description | Default |
|---|---|---|---|
| rules | No | Governing rules or contract version (default: designesy design system v0.4.1). | |
| context | No | Audience, device, environment, and constraints. | |
| purpose | No | What the design is trying to make possible. | |
| artifact | No | URL or description of the artifact to review. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden and mostly meets it: it declares the operation read-only, states the return payload, and crucially warns that the tool does not evaluate the design itself — the calling agent performs the critique. It also states the default for `rules` (design system v0.4.1). It stops short of covering error behavior or any failure modes, keeping it just under full marks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource definition, then when-to-use, when-not-to-use, behavior, return shape, and parameter effect — a logical funnel with no filler sentences. The only minor redundancy is restating the rubric-vs-number distinction across two sentences, which could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly compensates by naming the exact return fields ({ rubric, dimensions[8], agent_prompt, output_format, verification_checklist }) and by clarifying the essential non-obvious point that the tool returns scaffolding rather than performing the critique. Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real semantics the schema cannot express: that passing artifact/purpose/context/rules yields a pre-filled critique prompt while omitting all four returns the blank framework, clarifying the optional-parameter interaction. It stops short of describing how each field shapes the resulting prompt.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Designesy Design Review framework') and enumerates its concrete contents: an 8-dimension rubric with all eight named dimensions, plus agent prompt, output format, and verification checklist. It explicitly contrasts itself with the sibling designesy_score, so the agent can distinguish rubric-delivery from numeric scoring without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('when you want a structured rubric to critique a design holistically') and an explicit when-not-to-use naming the alternative tool and the selecting condition ('for a deterministic numeric score, use designesy_score'). Nothing about tool choice is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_drift_scoreAInspect
Score a live URL for AI-generated UI drift. Its 12 checks detect the four documented 2026 drift failure modes: token fabrication (var() to undeclared custom properties), within-session drift (spacing/color/radius value variance), between-session amnesia (inconsistent font stacks, shadows, transitions), and silent breaking changes (z-index chaos, dangling alias chains). Use this when you need to verify whether a site (especially an AI-generated one) is drifting off its own declared token system. When NOT to use: for a full 42-check design-contract score, use designesy_score; for token-file format validation, use designesy_tokens_score. Executable: fetches the URL server-side, extracts all CSS (inline + linked stylesheets), parses :root custom properties and var() references, runs 12 drift checks. No browser needed. Returns JSON: { ok, url, score (0-100), grade (A-F), pass, warn, fail, total, tokensExtracted, checks[{id, item, category, status, detail}] }. Results cached ~24h per URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to scan for drift. Defaults to https://www.designesy.org/ if not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that execution is server-side, that CSS is extracted from inline and linked stylesheets, that :root custom properties and var() references are parsed, that no browser is needed, and that results are cached ~24h. It does not mention authentication or rate limits, but the behavioral surface is otherwise well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then checks, then usage routing, then execution mechanics, then return shape, then caching. Sentences are information-dense and every one earns its place, including the specific parentheticals defining each failure mode.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers everything an agent needs despite having no output schema: it inlines the full return JSON shape (score, grade, pass/warn/fail, checks array), explains execution semantics, and notes the 24h cache. Complete for a single-param scoring tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single url param (with its default) is already fully documented in the schema. The description adds no syntax or format detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope ('Score a live URL for AI-generated UI drift') and immediately enumerates the 12 checks and the four failure modes it detects. It explicitly names the siblings it is not (designesy_score, designesy_tokens_score), so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('verify whether a site is drifting off its own declared token system') and a when-NOT-to-use block that names two alternatives with the exact conditions selecting them (full 42-check contract vs token-file format validation). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_guardrailsAInspect
Generate a frozen build-contract bundle for AI coding agents from any design system URL (the product layer). Ingests a site, extracts its :root tokens, and emits 6 outputs: (1) DTCG-format token file, (2) Stylelint config generated from token values, (3) AGENTS.md-format rules with token allowlist, (4) component contract with allowed prop patterns, (5) anti-pattern documentation, (6) DESIGN.md file (Google open spec, google-labs-code/design.md), the de-facto AI-readable design-context standard: YAML front matter plus a markdown body. Use this when you need to turn a design system into the file AI agents read and the lint that enforces it. When NOT to use: for design-contract scoring, use designesy_score; for token-file validation, use designesy_tokens_score; for drift detection, use designesy_drift_score. Executable: fetches the URL, extracts CSS + :root custom properties, generates the bundle. No browser needed. Returns JSON: { ok, url, score (0-100, emission completeness), grade, pass, warn, fail, total, tokensExtracted, bundle: { tokens, lintConfig, agentRules, componentContract, antiPatterns, designMd }, checks[{id, item, category, status, detail}] }. Results cached ~24h per URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to generate guardrails for. Defaults to https://www.designesy.org/ if not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that it is executable, fetches the URL, extracts CSS and :root custom properties, produces six specific outputs, requires no browser, returns a detailed JSON shape, and caches results ~24h per URL. This is genuinely rich behavioral context beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and outputs before the routing and behavioral details, and each block (outputs, when-to-use, when-not-to-use, execution, return shape) is functional. The return-JSON blob is dense, but it usefully substitutes for the absent output schema rather than being filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-output ingestion tool with no output schema, the description specifies the artifact list, execution model, caching, and full return payload including the score/grade/checks breakdown. Nothing essential to invoking or interpreting it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is only one optional parameter, so the schema already documents the url and its default. The description adds only that ingestion is from 'any design system URL,' which is marginal beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (generate a frozen build-contract bundle) plus scope (from any design system URL) and enumerates the six concrete artifacts it emits. It also names the sibling tools it is not (designesy_score, designesy_tokens_score, designesy_drift_score), so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the selection condition ('when you need to turn a design system into the file AI agents read and the lint that enforces it') and a full when-NOT-to-use section routing to three named sibling tools by task. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_llms_full_txtAInspect
Get the Designesy /llms-full.txt, the complete agent-facing brief: ingest protocol, discovery endpoints, every package, standing rules, anti-patterns, and a paste-ready agent prompt. Use this for comprehensive onboarding to the Designesy ecosystem when the short /llms.txt is not enough. When NOT to use: for a quick orientation, use designesy_llms_txt first (~500 tokens vs ~3000). Read-only: no side effects. Returns text/plain (~3000 tokens, includes a paste-ready agent prompt). No parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely discharges it: it declares read-only behavior, no side effects, no parameters, and the return format and size (text/plain, ~3000 tokens). It does not mention auth requirements or rate limits, but for a static text fetch that is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool fetches, then routes to the alternative, then the behavioral facts. Every sentence adds distinct value with no redundancy, despite the definition being fairly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating the media type and approximate token size and flagging the embedded agent prompt. For a parameterless static-text tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4; the description reinforces this with an explicit 'No parameters' statement. No further parameter meaning is possible or needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get the Designesy /llms-full.txt) and enumerates its contents (ingest protocol, discovery endpoints, packages, rules, anti-patterns, agent prompt). It explicitly distinguishes itself from the sibling designesy_llms_txt, so an agent can select correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('comprehensive onboarding when the short /llms.txt is not enough') and when-NOT-to-use ('for a quick orientation, use designesy_llms_txt first'), with a quantified cost comparison (~500 vs ~3000 tokens). This is exactly the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_llms_txtAInspect
Get the Designesy /llms.txt: a short agent-facing brief with the canonical reference, topic index, ingest steps, package list, and contact. Use this first when you don't know what Designesy is; it's the cheapest orientation path before pulling heavier artifacts. When NOT to use: for the full expanded brief, use designesy_llms_full_txt; for the contract itself, use designesy_contract. Read-only: no side effects. Returns text/plain (~500 tokens). No parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it declares the operation read-only with no side effects, states the return medium and size (text/plain, ~500 tokens), and confirms the absence of parameters. Cost/size disclosure and safety profile are both present, leaving no behavioral surprises.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the payload contents, then the use/when-not-to-use routing, then the operational facts; every clause is informative. The trailing fragment 'No parameters.' is minor redundancy against an already-empty schema, keeping it just short of maximum.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, side-effect-free fetch with no output schema, the description covers what it returns, its approximate size, its safety profile, and how it relates to sibling artifacts. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which the rubric sets at a baseline of 4. The description reinforces this with 'No parameters', consistent with the empty schema and 100% coverage, though it adds no further semantic value beyond confirming the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('the Designesy /llms.txt') and enumerates the payload (canonical reference, topic index, ingest steps, package list, contact). It also names the sibling tools it must not be confused with (designesy_llms_full_txt, designesy_contract), so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use-first rule ('Use this first when you don't know what Designesy is'), a rationale ('cheapest orientation path before pulling heavier artifacts'), and an explicit when-NOT-to-use section that routes to the full brief and the contract tool. Both the selection and exclusion conditions are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_monitor_scoreAInspect
Score a URL for continuous design-drift governance, the temporal layer over the drift radar. Re-runs the 12 drift checks (d01-d12) on the URL and computes 10 monitor checks (m01-m10): schedule registered, last run fresh, drift delta vs baseline, trend slope (3-run trajectory), new violations since last run, resolved since last run (the healing signal), score degradation threshold, token-set mutation, contract version drift, and alert delivered. When alerts fire and an email address is provided, sends an HTML drift-alert email via Resend (requires RESEND_API_KEY env var). Pass a history array of prior snapshots to compute deltas; omit it for a first-run baseline. Use this to watch a design system over time: "weekly audits at cents per report" (Into Design Systems 2026). When NOT to use: for a single point-in-time drift check, use designesy_drift_score; for design-contract scoring, use designesy_score. Executable: fetches the URL, extracts CSS + :root tokens, runs checks, computes deltas. No browser needed. Returns JSON: { ok, url, score (0-100, governance health), grade (A-F), pass, warn, fail, total, currentSnapshot, baseline, previous, driftChecks, monitorChecks, alerts, emailAlert }. Results cached ~24h per URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to monitor for drift. Defaults to https://www.designesy.org/ if not provided. | |
| No | Email address to receive drift alerts. When alerts fire AND this is provided AND RESEND_API_KEY is set, an HTML alert email is sent. Optional: without it, alerts surface in-UI only. | ||
| history | No | Prior snapshots for delta computation. Omit for a first-run baseline. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and meets it: it discloses that an HTML alert email is sent via Resend (requiring RESEND_API_KEY), that it is executable (fetches URL, extracts CSS + :root tokens, computes deltas, no browser), that results are cached ~24h per URL, and it lists the return JSON fields. Effects, dependencies, and side effects are all surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is dense and long, but it is front-loaded with the core purpose and the check inventory before routing rules and return shape. The 'when NOT to use' and return JSON are separated into clear segments. Slightly over-packed with check IDs and marketing quote, but every segment earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param, no-annotation, no-output-schema executable tool with external dependencies, the description is complete: it covers inputs, execution model, alert side effects, caching, and the return field list, which substitutes for the missing output schema. An agent has everything needed to call it correctly or route elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three params, giving a baseline of 3. The description adds genuine meaning beyond the schema: it explains the history array as prior snapshots for delta computation and the omit-for-baseline behavior, and clarifies the email/alert trigger chain (alerts fire AND email provided AND RESEND_API_KEY set). This is useful semantics layered on top of complete schema docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Score a URL for continuous design-drift governance') and frames the tool's role as 'the temporal layer over the drift radar.' It enumerates the concrete work (12 drift checks d01-d12, 10 monitor checks m01-m10) and names the exact sibling alternatives, so an agent can distinguish it from designesy_drift_score and designesy_score without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('watch a design system over time'), an operational rule for the history param ('Pass a history array ... omit it for a first-run baseline'), and an explicit 'When NOT to use' clause naming two alternatives by name with the condition that selects each. Nothing about tool selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_motion_scoreAInspect
Validate a Lottie animation file against the Lottie spec v1.0.1 and the Designesy §16 Ten Non-Negotiable Motion Standards, returning 10 checks (m01-m10) with PASS/FAIL/WARN. The DTCG 2025.10 spec leaves motion tokens as a second-class citizen: there is no standard for motion token structure, reduced-motion markers, or animation accessibility. Designesy's motion validator fills this gap: it checks required fields (v, fr, ip, op, w, h, layers), $version, a markers array for reduced-motion compliance, and no deprecated version. Use this to verify a motion/animation asset is well-formed AND accessible: the only validator that checks both. When NOT to use: for full-site motion scoring (not a single Lottie file), use designesy_score. Executable: fetches the URL or parses the raw Lottie JSON, runs 10 checks server-side. No browser needed. Returns JSON: { checks[{id (m01-m10), name, status (PASS/FAIL/WARN), detail}], valid, score }. Pass url to fetch a remote Lottie file, or lottie_file to validate an inline JSON string. Provide exactly one.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to a Lottie JSON file. The tool fetches and validates it. | |
| lottie_file | No | Raw Lottie JSON string to validate (alternative to url). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden, and it does well: it discloses the execution model (fetches the URL or parses raw JSON, runs checks server-side, no browser needed) and enumerates what is actually checked. It does not mention auth, rate limits, or failure modes on an unreachable URL, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and return shape, which is good, but it spends several sentences on positioning ('fills this gap', 'the only validator that checks both') and DTCG background that inform the agent less than the operational facts. It is longer than it needs to be for a two-parameter validator.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although no output schema exists, the description supplies the return structure ({ checks[{id, name, status, detail}], valid, score }), the check count/range, and the input mode selection. Combined with the usage guidance, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a genuine constraint beyond the schema: the mutual exclusivity of url and lottie_file ('Provide exactly one'). That resolves an ambiguity the schema alone leaves open.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validate) and resource (Lottie animation file) with concrete scope (Lottie spec v1.0.1 plus Designesy §16 Motion Standards, returning checks m01-m10). It explicitly differentiates itself from the sibling designesy_score by naming it in the 'When NOT to use' clause.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-to-use case ('verify a motion/animation asset is well-formed AND accessible') and a when-NOT-to-use case that names the alternative tool ('for full-site motion scoring ... use designesy_score'). This is exactly the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_readiness_scoreAInspect
Score a URL for design-system AI readiness, the 6th maturity axis (zeroheight 2026). 10 checks probe the target origin for machine-readable artifacts: DTCG token files, llms.txt, agent.json, MCP endpoint (tools/list), DESIGN.md, token $description, component schemas, sitemap.xml, robots.txt, and Open Graph/Twitter meta. Use this to verify whether a design system is the default context AI tools build from, or whether AI is silently working around it. When NOT to use: for full design-contract scoring, use designesy_score; for AI-drift detection, use designesy_drift_score. Executable: fetches the URL and probes the origin via HEAD/GET for each artifact. No browser needed. Returns JSON: { ok, url, score (0-100), grade (A-F), pass, warn, fail, total, checks[{id, item, category, status, detail}] }. Results cached ~24h per URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to score for AI readiness. Defaults to https://www.designesy.org/ if not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the execution model ('fetches the URL and probes the origin via HEAD/GET for each artifact'), that no browser is needed, and a caching policy ('~24h per URL'). These are non-obvious operational traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then probes, exclusions, execution, and return shape in a logical order; every sentence earns its place. It is dense but not padded, though the full inline return-shape listing is the only slightly heavy element.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description fully compensates by explaining the JSON return contract (ok, url, score, grade, pass/warn/fail/total, checks[]), the execution model, and caching. An agent has everything needed to call and interpret this tool without opening anything else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'url' parameter is fully documented in the schema, including its default. The description adds no format constraints, validation rules, or edge-case meaning beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Score') and resource ('a URL for design-system AI readiness'), and enumerates the exact 10 artifacts probed, so an agent knows precisely what the tool evaluates. It also positions itself within the designesy family ('the 6th maturity axis') and the sibling set, distinguishing it clearly from adjacent scoring tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('verify whether a design system is the default context AI tools build from') and an explicit 'When NOT to use' clause routing to designesy_score for design-contract scoring and designesy_drift_score for AI-drift detection. Nothing is left to inference about tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_reportAInspect
Generate a unified design-intelligence report for a single URL, the synthesis capstone. Fires /score (42-check audit), /drift (12-check drift radar), and /readiness (10-check AI readiness) in parallel, then computes a weighted composite: score × 0.5 + drift × 0.3 + readiness × 0.2. One input, one output, one composite grade. Use this when you need a single holistic assessment instead of three separate scans, or when sharing a design-intelligence verdict (the report is the most shareable surface). When NOT to use: for just the audit score, use designesy_score; for just drift, use designesy_drift_score; for just AI readiness, use designesy_readiness_score. Executable: fires 3 internal APIs in parallel, each fetches the target URL. No browser needed. Returns JSON: { ok, url, compositeScore (0-100), compositeGrade (A-F), score { sub-result }, drift { sub-result }, readiness { sub-result }, totalChecks, totalPass, totalWarn, totalFail, totalSkip, checks[] (all checks across all engines, tagged with engine), synthesis[] (8 synthesis checks verifying the report ran correctly), appUrl (standalone interactive dashboard URL) }. Results cached ~24h per URL. MCP Apps: hosts that support io.modelcontextprotocol/ui render an interactive dashboard inline; others get the JSON plus an appUrl link.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public URL to generate a design-intelligence report for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations supplied, the description carries the full burden and meets it: it discloses that three internal APIs fire in parallel, that each fetches the target URL, that no browser is required, that results are cached ~24h per URL, and that an appUrl/dashboard is produced for io.modelcontextprotocol/ui hosts versus JSON for others.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, formula, usage, execution, and return shape in a logical order with no filler prose. It runs long, and the full JSON key enumeration plus the 'One input, one output, one composite grade' restatement of the formula are slightly redundant, keeping it short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description fully compensates, spelling out the composite formula, the agents it orchestrates, the composite and sub-result return shape, check tallies, and caching/render behavior. An agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single 'url' parameter, so the schema already documents it fully. The description adds only the notion that the URL is fetched by each internal API, which is marginal beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (generate a unified design-intelligence report for a single URL) and frames it as the 'synthesis capstone' of a family. It explicitly distinguishes itself from designesy_score, designesy_drift_score, and designesy_readiness_score by naming each sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions (need a single holistic assessment instead of three scans, or sharing a verdict) and an explicit 'When NOT to use' block that routes to the correct alternative for each sub-scan. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_scoreAInspect
Score a live URL against the Designesy design contract: a deterministic 42-check verification engine that returns a numeric score, letter grade (A-F), and per-check breakdown. Use this to audit whether a website or AI-generated UI complies with a real design contract (tokens, motion, accessibility, cadence, takt, typography, copywriting). When NOT to use: for token-file validation only, use designesy_tokens_score; for a Lottie file, use designesy_motion_score; for a qualitative critique, use designesy_design_review. Executable: fetches the URL server-side, extracts CSS, runs 42 checks. Results cached ~24h per URL. Checks needing a live browser (Core Web Vitals, sound toggle, overflow) return MANUAL (not FAIL); run the full audit (/api/score/audit) to resolve them. Checks that are not applicable to the site (no tokens, no buttons, no DESIGN.md) return SKIP (N/A). Returns JSON: { url, score (0-100), grade (A-F), pass_count, fail_count, checks[{id, name, status, weight, category}] }. Pass format="canonical" for review-findings.json schema, "review" for markdown, or "google" for design.md-compatible output.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to score. Defaults to https://www.designesy.org/ if not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and meets it: it discloses server-side fetching, ~24h caching per URL, the MANUAL status for browser-dependent checks and the /api/score/audit escape hatch, and the SKIP (N/A) semantics for inapplicable checks. These are non-obvious behaviors an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then exclusions, then execution behavior, then return shape; every section earns its place. The parenthetical category list and status explanations add length that is mostly justified for a 42-check engine, though the format sentence is wasted given it is not a real parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description steps in with a full return JSON sketch (url, score, grade, pass_count, fail_count, checks[]), status semantics, and caching behavior. The only gap is the phantom 'format' parameter, which makes the tool look richer than its schema allows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single 'url' parameter, so the baseline is 3. The description corroborates the default URL behavior but then instructs 'Pass format="canonical"' — a parameter that does not exist in the input schema, which risks an agent emitting an invalid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Score a live URL against the Designesy design contract'), quantifies the engine (42-check deterministic) and enumerates what it audits (tokens, motion, accessibility, cadence, takt, typography, copywriting). It also names the sibling tools it is not, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When NOT to use' block with three named alternatives and the condition selecting each (designesy_tokens_score for token files, designesy_motion_score for Lottie, designesy_design_review for qualitative critique). This is exactly the routing guidance the dimension rewards.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_skill_mdAInspect
Get the Designesy SKILL.md: the agent-skill-format export of the design-system contract, written as behavioral rules an AI coding agent can drop into .agents/skills/ or a system prompt. Use this when you want the contract in a form that steers how an agent builds UI (tokens, anti-patterns, behavioral rules, verification). When NOT to use: for the raw contract JSON, use designesy_contract; for scoring, use designesy_score. Read-only: no side effects. Returns markdown text (SKILL.md format) to drop into .agents/skills/ or paste into a system prompt. No parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does disclose 'Read-only: no side effects' plus the return format (markdown text in SKILL.md format). It does not mention error behavior or limits, but for a zero-param read-only export tool this is nearly complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's identity and purpose, then the when/when-not guidance. Slight redundancy: the '.agents/skills/ or system prompt' destination is stated twice in near-identical wording, which could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only export with no output schema, the description states what is returned (markdown SKILL.md text) and how to use it. An agent has everything needed to call and consume it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description confirms 'No parameters,' consistent with the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get the Designesy SKILL.md') and defines exactly what the resource is: the agent-skill-format export of the design-system contract. It explicitly differentiates from siblings designesy_contract (raw JSON) and designesy_score (scoring).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'Use this when' condition (steering how an agent builds UI) and an explicit 'When NOT to use' clause routing to designesy_contract and designesy_score. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
designesy_tokens_scoreAInspect
Validate a design token file against the W3C Design Tokens Community Group (DTCG) 2025.10 Final Community Group Report (the spec's first stable version, published Oct 28 2025 as a Candidate Recommendation and considered stable). Returns 10 conformance checks (t01-t10) with PASS/FAIL/WARN. Use this to verify a tokens.json (or any DTCG token export) is structurally correct: $type/$value/$description present, structured colors (colorSpace + components rather than bare hex), a valid $schema pointer to designtokens.org, and correct dimension units. With 84% of teams now using design tokens (zeroheight Design Systems Report 2025, up from 56% in 2024) and the spec finally stable, every adopting team needs a validator. When NOT to use: for scoring a whole live site (not just its token file), use designesy_score. Executable: fetches the URL or parses the raw JSON you provide, runs 10 checks server-side. No browser needed. Returns JSON: { checks[{id (t01-t10), name, status (PASS/FAIL/WARN), detail}], valid, score }. Pass url to fetch a remote token file, or dtcg_file to validate an inline JSON string. Provide exactly one.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to a DTCG token file (JSON). The tool fetches and validates it. | |
| dtcg_file | No | Raw DTCG token JSON string to validate (alternative to url). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely meets it: it discloses that the tool fetches a URL or parses inline JSON, runs checks server-side, needs no browser, and returns a defined JSON shape. It stops short of permission/auth or rate-limit behavior, but for a read-only validator this is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the not-use case are front-loaded, which is good, but the description is padded with promotional copy ('84% of teams now using design tokens... every adopting team needs a validator') and a hedged spec-history sentence ('published Oct 28 2025 as a Candidate Recommendation and considered stable') that do not help an agent invoke the tool. Roughly two sentences of pure sales pitch.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, usage, exclusion, execution model, and the return shape (checks with id/name/status/detail, valid, score) even though no output schema exists. Nothing an agent needs to select and call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds a constraint not present in the schema: 'Pass url to fetch a remote token file, or dtcg_file to validate an inline JSON string. Provide exactly one.' This mutual-exclusivity rule is genuine added meaning beyond the two property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Validate) plus a precise resource (design token file) against a named standard (W3C DTCG 2025.10). It explicitly distinguishes itself from the sibling designesy_score by scoping to a token file rather than a whole live site.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a positive trigger (verify a tokens.json or any DTCG export is structurally correct) and an explicit exclusion ('When NOT to use: for scoring a whole live site... use designesy_score'). The alternative is named with the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Changed
designesy_a11y_score1 field changed- changed
Input schema / properties / config / descriptionPrevious value: -"Brand customization JSON for axe.configure() — branding, checks, rules, disableOtherRules."New value: +"Brand customization JSON for axe.configure(): branding, checks, rules, disableOtherRules."
- Changed
designesy_design_review1 field changed- changed
Input schema / properties / rules / descriptionPrevious value: -"Governing rules or contract version (default: designesy design system v0.4.0)."New value: +"Governing rules or contract version (default: designesy design system v0.4.1)."
- Changed
designesy_monitor_score1 field changed- changed
Input schema / properties / email / descriptionPrevious value: -"Email address to receive drift alerts. When alerts fire AND this is provided AND RESEND_API_KEY is set, an HTML alert email is sent. Optional — without it, alerts surface in-UI only."New value: +"Email address to receive drift alerts. When alerts fire AND this is provided AND RESEND_API_KEY is set, an HTML alert email is sent. Optional: without it, alerts surface in-UI only."
1 tool update
- Changed
designesy_monitor_score1 field changed- added
Input schema / properties / emailAdded value: +{ + "description": "Email address to receive drift alerts. When alerts fire AND this is provided AND RESEND_API_KEY is set, an HTML alert email is sent. Optional — without it, alerts surface in-UI only.", + "type": "string" +}
17 tool updates
- Changed
designesy_a11y_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_agent_json1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
designesy_catalog1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
designesy_compare2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_contract2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_design_review2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_drift_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_guardrails2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_llms_full_txt1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
designesy_llms_txt1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
designesy_monitor_score4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / history / items / additionalPropertiesRemoved value: -false - removed
Input schema / properties / history / items / properties / checks / items / additionalPropertiesRemoved value: -false
- Changed
designesy_motion_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_readiness_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_report2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
designesy_skill_md1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
designesy_tokens_score2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
1 tool update
- Added
designesy_report
1 tool update
- Added
designesy_compare
1 tool update
- Added
designesy_monitor_score
3 tool updates
- Added
designesy_drift_score - Added
designesy_guardrails - Added
designesy_readiness_score
1 tool update
- Changed
designesy_design_review1 field changed- changed
Input schema / properties / rules / descriptionPrevious value: -"Governing rules or contract version (default: designesy design system v0.3.0)."New value: +"Governing rules or contract version (default: designesy design system v0.4.0)."
11 tool updates
- First observed
designesy_a11y_score - First observed
designesy_agent_json - First observed
designesy_catalog - First observed
designesy_contract - First observed
designesy_design_review - First observed
designesy_llms_full_txt - First observed
designesy_llms_txt - First observed
designesy_motion_score - First observed
designesy_score - First observed
designesy_skill_md - First observed
designesy_tokens_score
Related MCP Connectors
- miromiroOAuthapp.miromiro
Turn any live website into brand colors, fonts, design tokens, SVGs, Lottie and paste-ready code.
On-demand drift checks: declared CSS color, radius, spacing & type vs your own tokens or a pack
Scores any public website on how usable it is by AI agents, with per-check evidence.
Free website analyzer: score any public URL 0-100 across 8 quality dimensions. No auth.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables coding agents and CI to verify rendered web UIs against design token usage, layout constraints, and accessibility rules, returning concise pass/fail verdicts with actionable findings.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to audit rendered web pages for design and UI defects by measuring geometry, styles, contrast and motion in a local browser and judging them against evidence-tiered rules. Findings come back as short, fixable items expressed in the project's own design tokens, while only sanitized numeric measurements leave the machine.Apache 2.0
- AlicenseAqualityCmaintenancePoint your coding agent at a URL and get a real-browser QA audit: broken signup/login/checkout flows, JS console errors, missing analytics, consent + security headers, mobile tap targets, and accessibility — returned as machine-verified findings graded A-F.442Apache 2.0

uxlintofficial
AlicenseAqualityBmaintenanceAudits any site's UX the way a design-literate reviewer would — contrast, tap targets, type scale, colour discipline, scan patterns, copy — and returns the rule, the source line and the exact fix.6Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.