saglitzdesign-mcp
The SaglitzDesign MCP server is an expert design & marketing knowledge server that gives AI agents prescriptive, sourced guidance on web, iOS, Android, and macOS design — plus UX, copywriting, SEO, GEO, and marketing best practices.
Browse & Search the Knowledge Base: Explore and search 57 curated documents across design languages, components, UX, SEO, GEO, craft, books, process, and marketing — filterable by category or platform.
Fetch Full Documents: Retrieve any knowledge document in full by ID for deep reference.
Component Guidance: Expert specs, states, labels, and anti-patterns for UI components and screens (buttons, forms, paywalls, hero sections, pricing, onboarding, dashboards, etc.).
Design Language References: Full references for Material 3, Apple HIG / Liquid Glass (iOS 26), iOS/Android/macOS app design guides, Fluent 2, 2026 web trends, and design tokens/theming architecture.
Design Roadmaps: Phased, step-by-step expert processes for any project type (website, landing page, iOS/Android/macOS app, SaaS web app) with phase goals, exit criteria, and recommended docs.
Design Review Checklists: Structured audit checklists for mobile apps, websites, landing pages, or dashboards — filterable by focus area (UI, UX, accessibility, SEO, GEO, conversion, or copywriting).
SEO & GEO Guides: Comprehensive guidance for technical/on-page SEO, Core Web Vitals, and Generative Engine Optimization (GEO) for ChatGPT, Perplexity, and Google AI Overviews.
Real-World Design Examples: Annotated screenshots from top apps and websites (Mobbin-curated) for patterns like paywalls, onboarding, auth, pricing, and more.
Generate Design Tokens: Produces ready-to-use CSS variables, Tailwind v4, SwiftUI, Jetpack Compose, and W3C DTCG JSON tokens from color/spacing/type specs.
Accessibility Audits: Deterministic WCAG 2.2 checks with exact contrast ratio calculations and tap-target size validation per platform, with specific fixes.
Knowledge Freshness Reports: Identifies which documents are outdated based on per-category staleness thresholds.
Build Workflows: Orchestrate end-to-end builds of landing pages, multi-page sites, and mobile app UIs with code generation and iterative visual critique loops.
Extensible: Add custom Markdown knowledge files under
knowledge/with frontmatter; the server indexes them automatically.
Provides expert-level guidance on Android app design, including Material 3, Android 16, and M3 Expressive specifications, with prescriptive rules for components, UX, and accessibility.
Offers deep iOS app-design guides covering Apple HIG, Liquid Glass (iOS 26), and platform-specific components, UX patterns, and accessibility standards.
Delivers macOS design guidance with Apple HIG, macOS-specific UI patterns, and platform-appropriate UX and accessibility rules.
Provides Generative Engine Optimization (GEO) strategies to improve visibility in Perplexity's AI search results, including llms.txt setup, citation tactics, and content optimization.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@saglitzdesign-mcpsearch for iOS button design guidelines"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SaglitzDesign MCP
An expert design & marketing brain for your AI coding agent.
A Model Context Protocol server that gives Claude, Cursor, and any MCP client expert‑level guidance on web, iOS, Android and macOS design — plus the UX, copywriting, SEO, GEO and marketing knowledge that makes a product actually convert.
96 curated knowledge documents · 36 tools · 9 build/review/port workflows · MCP resources with id autocomplete · a one‑call design‑system builder · real token/color/type/elevation/motion/a11y generators · design & UX‑copy linters · production component recipes · phased roadmaps · real‑world visual examples
Why · What's inside · Tools · Resources · Install · Usage · Your own rules · Changelog · License
Why
LLMs are confidently wrong about design. They reach for defaults, invent outdated specs, and give generic "make it clean and modern" advice. SaglitzDesign replaces that with a curated, sourced, prescriptive knowledge base your agent can query on demand — the kind of guidance you'd get from a senior product designer, a conversion copywriter, and a technical SEO all at once.
Ask your agent to design a paywall, audit a landing page, or plan an iOS app, and it pulls concrete rules ("44pt minimum touch target", "one primary button per screen", "LCP ≤ 2.5s"), real patterns from top apps, and a phased roadmap — instead of guessing.
Runtime‑independent. The server reads only local files. No external API, no account, nothing to configure. It just works, offline.
Local by design. It speaks MCP over stdio only and runs as a child process of your MCP client. There is no HTTP or SSE transport, so it cannot be hosted remotely or reached over a network — which is also why your prompts and code never leave your machine.
Prescriptive, not vague. Every doc is written as rules an agent applies verbatim — numbers, thresholds, do/don't lists, anti‑patterns.
Grounded. Design‑language specs from official sources; patterns studied from real top apps and sites; classics distilled from the actual books.
Related MCP server: seedflip-mcp
What's inside
96 knowledge documents across 11 categories:
Category | Coverage |
🎨 Design languages | Material 3 & M3 Expressive · Apple HIG + Liquid Glass (iOS 26) · deep iOS, Android (Android 16 / M3 Expressive) and macOS app‑design guides · Apple accessibility (Dynamic Type, VoiceOver, hit regions) and Apple shipping readiness ( |
🧩 Components | Buttons (hierarchy, sizing, states, labels) · forms & inputs · navigation · cards / lists / modals / sheets / empty states |
🧠 UX | Nielsen heuristics & behavioral laws · accessibility (WCAG 2.2) · typography · color & dark mode · spacing & grids · motion · mobile UX · conversion / CRO · data visualization · information architecture · i18n / localization (RTL) · AI product UX (chat, streaming, agentic) · onboarding & permission priming |
✨ Craft | Expert polish standards · typographic craft · animation craft (easing, springs, interruptibility) · UX writing & cognitive load · 0–40 critique rubric · clean/minimal app design · design‑engineering (semantic HTML, CSS architecture, tokens‑in‑code) · ethical design (avoiding dark patterns) · iconography (choosing & using an icon system) · the AI‑default aesthetic (the stock gradient, font, card chrome and copy generated interfaces reach for, cited to each system's own docs) |
📚 Books | Distilled classics — design: Norman, Krug, Refactoring UI, psychology of design, grid/typography, interaction design (Cooper/Tidwell), emotional design (Walter/Norman) · marketing: Cialdini, Positioning, StoryBrand + Ogilvy, Hooked |
🗺️ Process | Product‑design & marketing‑website roadmaps · design‑systems methodology (Atomic Design, component API, governance) · design handoff (Figma Dev Mode, Code Connect, design↔dev) |
📣 Marketing | Branding & identity · email marketing · HTML email development (Outlook, dark mode, bulletproof) · ad creative · paywall benchmarks (RevenueCat 2026) · growth frameworks (loops/AARRR/PLG) · pricing strategy · analytics & experimentation · value proposition & JTBD · content & distribution (topic clusters, community, referral) · App Store Optimization (ASO) |
🔎 SEO | Technical SEO (Core Web Vitals) · on‑page & E‑E‑A‑T · SEO for designers |
🤖 GEO | Generative Engine Optimization — visibility in ChatGPT / Perplexity / AI Overviews, llms.txt, citation tactics |
🔐 Security | Web security headers & CSP — strict nonce/hash policies, |
🖼️ Patterns & examples | Real‑world patterns studied from top apps & sites (incl. e‑commerce & checkout and fintech / trust flows), plus a curated library of real‑world example screens |
Which documents' sources are enforced, and which are not. Of the 96
documents, 11 have their sources: checked by the test suite against a tiered
allowlist — the six Apple design‑language guides (apple-hig-liquid-glass,
ios-app-design, macos-app-design, apple-accessibility,
apple-shipping-readiness, wwdc-design-principles) and the five security
documents. The other 85 are curated but not yet checked, and extending the
assertion to them is its own piece of work rather than a formality: measured
today, 66 of the 85 would fail on 282 sources between them — 227 URL citations
across 133 distinct hosts, plus 55 that are not URLs at all and so have no
host to check. (The 227 and the 55 are the two halves of the 282, not two
figures to add.) Note
that apple-intelligence-design and visionos-spatial-design are Apple‑topic
documents that sit outside the enforced set. get_design_doc prints which
side of this line a document is on, next to its sources, so you never have to
come back here to find out.
Workflows (/ prompts) — "build me a…"
Beyond answering questions, SaglitzDesign ships prompts that orchestrate an
entire build end‑to‑end. In Claude Code they appear in the / menu under a name
that depends on how you installed it. Installed as the Claude Code plugin they
are /saglitzdesign:build_landing_page and so on — plugin commands, one
generated file per workflow in commands/. With the server installed on its own
they are /mcp__saglitzdesign__build_landing_page, which the menu labels
saglitzdesign:build_landing_page (MCP). That label is typeable too, (MCP)
suffix and all — it is just a clumsier way to reach the same prompt. What does
not work is that short form with the suffix dropped: Claude Code hands the bare
server:prompt alias to first‑party Anthropic connectors only — the URL has to
be https, on api.anthropic.com, under /v1/design/ — and no other server gets
one, stdio or remote. (The plugin's /saglitzdesign:build_landing_page above
only looks like that alias — it is a plugin command namespace, a different
mechanism.)
Invoke one and the agent runs the full method — roadmap → positioning & copy → generates
the design system (color, type, layout, elevation, tokens) → real examples →
writes the actual code from the component recipes → runs the deterministic
verify gate (design_lint, audit_accessibility, audit_design_system,
audit_ux_copy) → opens it in a browser, screenshots, scores it against the
critique rubric, and iterates until it passes.
Workflow | What it does |
| Designs & builds a conversion‑focused landing page, copy‑first, with a visual critique loop. |
| Builds a multi‑page marketing site — positioning, IA, SEO/GEO, shared design system. |
| Builds iOS or Android screens on the correct platform baseline (HIG/Liquid Glass or Material 3). |
| Builds a dense SaaS dashboard / app shell — navigation, tables, empty states — not a marketing page with KPI cards glued on. |
| Measures the screenshot, then critiques it against the fixed 0–40 rubric — cites real ratios and colour counts, specific elements, no padding. |
| Scores a paywall / subscription onboarding against real RevenueCat 2026 conversion benchmarks. |
| Audits an existing site/app — runs the deterministic auditors first, so findings lead with measured numbers, then the checklists and the 0–40 rubric, ranked by severity. |
| Improves an existing UI (bolder / quieter / higher‑converting) using the craft standards, with a measured before→after (consistency score, critique score, lint findings, contrast failures). |
| Takes an existing UI to another platform (iOS ↔ Android ↔ macOS ↔ web) surface by surface — porting the intent and IA, never the components. |
Just type, e.g.,
/saglitzdesign:build_landing_page a SaaS invoicing tool for freelancers— everything after the command name becomes the brief, and the workflow asks for anything missing, then builds it. The/mcp__saglitzdesign__…form takes only the first whitespace‑separated word as its brief, so use the plugin command whenever the brief is a phrase. These are yours to type: every command carriesdisable-model-invocation, so the agent never starts one on its own. Asked in prose it reaches for the skills instead, and they cover much of the same ground — design review, landing‑page conversion, Apple platforms — as guidance rather than as this orchestrated build. Thesaglitzdesignskill's own trigger vocabulary names paywalls, so prose about one does reach a skill; what no skill carries isreview_paywall's RevenueCat benchmark scoring, which runs only when you type it.The visual critique loop uses whatever browser tool is connected (Claude in Chrome, Playwright, or chrome‑devtools MCP) to see and refine its own output. Without one, it reviews the code directly.
Tools
Tool | What it does |
| The one‑call foundation. Brand color + vibe + platform → a direction card (named defaults to leave, the type pairing as a decision, one signature move, and a stock‑region fact if the seed sits in indigo/violet/purple) plus WCAG‑verified color (light+dark), matched fonts, an icon library, a modular type scale, an elevation ramp, ready‑to‑paste tokens, the components to build, and a checklist — all generated to work together. |
| Start here for process. A phased, expert process for a project type (website, landing page, iOS / Android / macOS app, SaaS web app) — each phase has a goal, exit criteria, and the exact docs to read. |
| Natural‑language search across everything, returning the most relevant section of the best‑matching docs. |
| Fetch any document in full by id. |
| Deep dive on a component or screen (button, form, paywall, hero, pricing…) — specs + real‑world patterns. |
| Full platform / design‑system references (Material 3, Liquid Glass, iOS/Android/macOS, Fluent 2, web trends, tokens). |
| Building on more than one platform? Side‑by‑side iOS / Android / macOS / web conventions for one surface (navigation, buttons, sheets, motion, forms…), plus the porting rules — and an explicit do NOT port list. |
| Curated real‑world examples of a pattern from top apps/sites, with notes on what each does well. (Screenshots are a local‑only asset — see Visual examples.) |
| An assembled audit checklist per project type and focus (UI, UX, accessibility, SEO, GEO, conversion, copywriting). |
| SEO and GEO guides, optionally narrowed to a topic. |
| Already have a design system? Paste your CSS custom properties, shadcn |
| Real artifacts, not advice — turns a color/spacing/type spec into CSS variables, Tailwind v4, SwiftUI, Jetpack Compose, and W3C DTCG JSON. |
| Deterministic WCAG 2.2 checks — exact contrast ratios for color pairs + tap‑target sizes per platform, with fixes. |
| Production‑ready, accessible reference code for a component (button, input, modal, toast, card, switch, tabs, empty‑state, list‑row, navigation, search, select, table, tooltip, form, pagination, skeleton, badge, breadcrumb) in react‑tailwind, html‑css, SwiftUI, or Compose — all states, ARIA, keyboard, correct motion. Pass your |
| One brand color → a full palette. A 50–950 tonal scale, a cohesive neutral ramp, status colors (danger / success / warning) harmonised with your brand, and light + dark semantic tokens — every text/UI pair WCAG‑verified and auto‑corrected. Feeds straight into |
| Curated, production font pairings for a vibe (SaaS, editorial, bold, native…) — heading + body (+ mono) with paste‑ready CSS stacks, weights, rationale, and a type scale. |
| Repairs a failing color pair: computes the nearest accessible color (hue/saturation preserved) that meets your WCAG target — the corrected value, not just a fail report. |
| Recommends the right icon system for a vibe/platform (Lucide, Phosphor, Solar, SF Symbols, Material Symbols…) — with license, install command, coverage, fit rationale, and usage rules. |
| A modular type scale (base × ratio) → named steps with line‑heights, tracking, and fluid |
| A cohesive layered box‑shadow ramp (flat→modal) as CSS variables + Tailwind, with dark‑mode guidance. |
| Breakpoints (with what changes at each), container widths, a column grid, an intrinsic auto‑fit grid, container queries, and a fluid section‑rhythm scale — CSS variables + Tailwind v4. |
| Easing + duration tokens and ready‑to‑paste keyframe animations (fade/slide/scale/spring/shimmer) in CSS, Framer Motion, or SwiftUI — reduced‑motion included. |
| Lints an HTML/CSS/JSX/Tailwind snippet for design & a11y anti‑patterns (hardcoded values, killed focus, missing alt/labels, clickable divs, unlabelled inputs…) with line numbers and fixes. Tag‑aware, so formatting never changes the verdict. Returns markdown plus structured output — findings, a severity summary, and a machine‑readable |
| Measures your actual screen. Give it a PNG and it reports the real palette and colour count, true WCAG contrast ratios for the pairs on screen, density, structural detections (alignment, rhythm, off‑grid gaps) each with a confidence level, and a stock‑region fact when a significant cluster sits in Tailwind indigo/violet/purple — plus a self‑contained HTML report you can open and share. Pure‑Node PNG decoding, no network, no dependencies. |
| Audits a real codebase, not a snippet. Point it at a directory: it walks your design source, lints every file, and scores the whole project for consistency — cross‑file drift being exactly what a single‑file lint cannot see. Findings ranked worst‑file‑first with file:line, plus an explicit list of what it did not look at. Returns markdown plus structured output — findings, a severity summary, a machine‑readable |
| Audits a web project or snippet for the defects that actually ship — missing or weak Content‑Security‑Policy, absent HSTS, unpinned cross‑origin scripts, mixed content, credentials in |
| Audits the SEO and GEO signals that are actually in your source — a missing |
| Audits the performance signals that are actually in your source — a hero image held back by |
| Audits an iOS or macOS app against Apple's own documentation. Point it at the Xcode project directory: it reads the four surfaces a project declares configuration on — the information property list, |
| Audits an Android app against Material 3 and android-app-design. Point it at the project or module directory: it reads |
| Is there actually a system here? Point it at real CSS/JSX and get a consistency score plus the sprawl behind it: how many distinct colors, sizes, radii, shadows and spacings are in use, which colors are indistinguishable duplicates, what's off the 4pt grid, token adoption — and a consolidation plan. |
| Audits for the specific defaults generated interfaces reach for — the stock Tailwind indigo/violet/purple gradient (as classes, hex, or OKLCH), Inter/Roboto/Open Sans/DM Sans/Plus Jakarta Sans as the only declared typeface on a brand surface, emoji standing in for icons, the |
| Objective copy audit — readability (Flesch), sentence length, passive voice, jargon, filler, user‑focus, weak CTAs — with flagged phrases and fixes. |
| Audits a pasted snippet for named deceptive-pattern tells from |
| Browse the full knowledge index by category / platform. |
| Reports each doc's age vs a per‑category staleness threshold, so the base can be kept current. |
Resources
Beyond tools, the knowledge base is exposed as MCP resources, so clients that
support them (Claude Desktop, Cursor) can @‑mention a document directly —
with id autocompletion — instead of spending a tool call:
URI | What it is |
| The whole index: every document by category, with platform and last‑verified date. |
| One knowledge document in full (96 of them). Autocompletes on |
| A component's spec plus its reference implementation in every available stack. |
Visual examples
get_design_examples serves a curated library of real app/site screens. The
screenshots themselves are third‑party assets and are not redistributed, so
the published npm package ships the annotations and source links without the
images — the tool detects this and says so rather than pretending otherwise. If
you clone the repo you can rebuild the local image library; see
scripts/regenerate-examples.md.
Install
Requirements: Node.js 20+.
From npm (no clone needed):
npx saglitzdesign-mcpFrom source:
git clone https://github.com/HalidSaglam/saglitzdesign-mcp.git
cd saglitzdesign-mcp
npm install
npm run buildClaude Code
Register it once (via npm — no clone), available in every project:
claude mcp add --scope user saglitzdesign -- npx -y saglitzdesign-mcpOr, if you cloned the repo, point it at the built file:
claude mcp add --scope user saglitzdesign node /absolute/path/to/saglitzdesign-mcp/dist/index.jsAs a plugin
One install brings all three pieces — the MCP server, all nine skills (the eight depth skills plus the umbrella that routes into them) and a slash command for every workflow:
claude plugin marketplace add HalidSaglam/saglitzdesign-mcp
claude plugin install saglitzdesign@saglitzThe plugin declares its server as npx -y saglitzdesign-mcp@latest, so the
server itself still comes from npm; the skills and the commands are files inside
the plugin. claude plugin details saglitzdesign@saglitz lists what arrived.
The three ways of installing do not carry the same payload:
Install | Server, knowledge base, recipes | Skills |
|
| yes | all nine — eight depth skills plus the umbrella | one per workflow |
| yes | — | — |
| — | all nine — eight depth skills plus the umbrella | — |
The npm package's files: list covers dist/, knowledge/ and recipes/, and
does not name skills/ or commands/ — so an MCP-only install has every tool
and every document and none of the skill or command files. The skills CLI is the
mirror image: it copies each SKILL.md into your agent and brings no server, so
the tools those skills point at are not there unless you also install one.
Cloning the repository gets the source of all of it — but not a runnable
server: dist/ is gitignored, so a clone needs npm install && npm run build
before node dist/index.js starts. Everything else in the table is a tracked
file and arrives with the clone.
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"saglitzdesign": {
"command": "npx",
"args": ["-y", "saglitzdesign-mcp"]
}
}
}Cursor / other MCP clients
Run npx -y saglitzdesign-mcp over stdio (or node /absolute/path/to/dist/index.js from a clone).
Requires Node 20 or newer. It also runs on Node 18 with a warning; below that the launcher exits with an explanation instead of a syntax error.
Your client launches the server — you don't start it yourself. Run by hand it will print its status and then wait on standard input for a client that never arrives, which looks like a hang but is the server working correctly.
As skills (no MCP server)
Prefer lightweight, self-contained skills? Install the condensed SaglitzDesign skills into any skills-compatible agent:
npx skills@latest add HalidSaglam/saglitzdesign-mcpNine skills. saglitzdesign is the umbrella: it carries no depth guidance
of its own, and routes the work into whichever of the other eight fits. Those
eight — clean-interface-design, landing-page-conversion, design-review,
motion-and-animation, apple-platform-design, android-platform-design,
design-system-audit, ship-quality-gate — are each standalone guidance that also points to the
full MCP for depth. See skills/.
Each skill is copied into your agent and its content hash pinned in
skills-lock.json, so an installed skill does not change when this repository
does. Re-run the command above to pick up a new skill or an edited one.
Dev & debug
npm run dev # run from TypeScript via tsx
npm run inspect # open the MCP Inspector UIUsage
Once connected, just talk to your agent naturally — it decides when to call the tools:
"Using saglitzdesign, plan the design of an iOS fitness app." →
get_design_roadmapreturns a 7‑phase plan with the docs to read at each step.
"Review my landing page for conversion with saglitzdesign." →
design_review_checklist(landing‑page / conversion) +get_design_examples.
"How should a primary button behave on mobile?" →
get_component_guidancereturns specs, states, labels and anti‑patterns.
"Show me real paywall examples." →
get_design_examplesreturns annotated screenshots.
"What's llms.txt and how do I set it up?" →
seo_geo_guide(GEO) returns the tactic with a ready‑to‑use example.
Your own design rules
Point the server at a directory of your own and your documents join the base — searchable, readable, and, if you ask, part of the review checklist:
claude mcp add --scope user saglitzdesign \
--env SAGLITZDESIGN_KNOWLEDGE_DIR=/path/to/our-design-rules \
-- npx -y saglitzdesign-mcpOr in claude_desktop_config.json / any MCP client, alongside command and args:
"env": { "SAGLITZDESIGN_KNOWLEDGE_DIR": "/path/to/our-design-rules" }Several directories are allowed, separated the way PATH is on your platform.
Your rules lead. When a search is genuinely about something you documented, your document comes first — measured by how many of the query's terms it covers, so a short house‑rules file is not buried by a long reference.
Same id replaces ours. A file with
id: buttonstakes over from the built‑in one. That is deliberate, and it is announced at startup rather than happening quietly. Wherever it is served it is marked as your team's document, so nothing of yours is ever quoted as though it were sourced platform guidance.review: [website, saas-web-app]in the frontmatter puts the document into those project types'design_review_checklist, ahead of the curated list — the difference between your rules being findable and your rules being enforced.
Editing files inside the installed package works until
npm updatedeletes them. Use the environment variable.
Adding to this repository
Drop a Markdown file anywhere under knowledge/ with frontmatter:
---
id: my-topic
title: "My Topic"
category: ux # design-language | component | ux | seo | geo | pattern | craft | book | process | marketing | security
platform: both # mobile | web | macos | both
tags: [tag1, tag2]
sources: ["https://…"]
updated: 2026-07-08
---
Content served verbatim to clients…The server indexes every .md on startup — no rebuild needed for content
changes (just restart the server). A /refresh-knowledge command
(.claude/commands/) can re‑research stale docs with agents.
How the knowledge was built
Design‑language and SEO/GEO docs were researched from official documentation
and current sources (cited in each file's sources). Real‑world UI patterns
were studied from top apps and websites; the visual example library was curated
the same way. Classic design and marketing books were distilled into original,
prescriptive syntheses — no source text is reproduced.
On images: screenshot files are a local research asset. They are not
included in this repository or any published package. Without them,
get_design_examples gracefully degrades to descriptions plus source links.
To rebuild the local image library (or add your own examples), see
scripts/regenerate-examples.md.
Find it on
Also listed in the official MCP Registry and on npm.
License
MIT © 2026 Saglitz Design.
The knowledge/ documents are original syntheses with sources cited per file.
Referenced screenshots are not part of this repo and may not be redistributed —
see NOTICE.md.
Available Tools
36 toolsaudit_accessibilityAudit AccessibilityARead-onlyIdempotent
Deterministic design-time accessibility checks: WCAG 2.2 color-contrast ratios for text/UI color pairs, and minimum tap/target sizes per platform (iOS 44pt, Android 48dp, web 24px min / 44 recommended). Returns exact ratios, pass/fail, and fixes — the machine-verifiable slice of a11y you can run before code. For keyboard/screen-reader/Dynamic Type checks, see get_design_doc('accessibility').
| Name | Required | Description | Default |
|---|---|---|---|
| tap_targets | No | Interactive targets to check for minimum size | |
| contrast_pairs | No | Color pairs to check for contrast |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds useful context beyond that: the checks are deterministic, run at design time, and return 'exact ratios, pass/fail, and fixes' instead of just an audit result. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that front-load the tool's purpose and scope, then provide the alternative. Every clause adds information; there is no fluff or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description explains what the tool returns: exact ratios, pass/fail, and fixes. Combined with comprehensive parameter schemas, clear scope boundaries, and an explicit pointer to the alternative tool, this is complete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds meaningful context by specifying platform thresholds (iOS 44pt, Android 48dp, web 24px/44px) and the WCAG 2.2 contrast rules, which directly inform how tap_targets and contrast_pairs are evaluated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Deterministic design-time accessibility checks' for WCAG 2.2 contrast and tap/target sizes. It clearly separates this tool from other design audits and explicitly routes non-mechanical a11y checks elsewhere, making it easy to tell apart from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly defines when to use this tool: the 'machine-verifiable slice of a11y you can run before code.' It also names the alternative for keyboard/screen-reader/Dynamic Type checks with get_design_doc('accessibility'), giving clear exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_android_uiAudit Android UiARead-onlyIdempotent
Audit an Android app's UI against Material 3 and android-app-design: point it at the project (or module) directory and it reads AndroidManifest.xml, resource XML and Kotlin, infers whether the folder is an Android project from those plus Compose/Android imports, and runs six rules. Configuration: windowOptOutEdgeToEdgeEnforcement="true" (deprecated and ignored on Android 16), enableOnBackInvokedCallback="false" (the predictive-back opt-out), and a values/themes.xml or values/colors.xml with no values-night/ file among the surfaces read. Compose: a Color(0x…) literal, fontSize = N.sp, and an androidx.compose.material.Button-style Material 2 import (not material3, not material.icons). Every configuration rule stays silent when no Android signal was found, and the report says whether the platform was inferred and from what, so a silence can be read as the gate rather than as a result. It reads source and does not measure anything: it builds nothing, starts no emulator, takes no screenshot, and no finding is or can be a rendered-output, contrast, TalkBack or Play review result. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line — configuration findings carry no line, since a manifest attribute has no useful source position), a severity summary, a scan block saying how much was actually read, and a machine-readable notVisible list of what it structurally could not check, every entry of it derived from a run. Directory only — there is no snippet mode, because configuration is the backbone of this audit and a snippet carries none of it. A missing path, a path that is a file, or a code argument is returned as an error result, not as an empty audit. Pair with get_design_doc("material-3") and get_design_doc("android-app-design") for the guidance behind the rules.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | Not supported by this tool, and rejected with an explanation if passed. There is no snippet mode: the platform verdict and half the rules are read from AndroidManifest.xml and resource XML, and a pasted snippet carries none of them. Pass `path` instead. | |
| path | No | The Android project or module directory to audit — the folder holding `AndroidManifest.xml`, `res/` and Kotlin sources. Required. Absolute paths are strongly preferred: a relative path is resolved against the server's working directory, which is usually not your project folder. |
Output Schema
| Name | Required | Description |
|---|---|---|
| scan | Yes | What the scan reached. Read it before trusting any absence claim: a capped scan looked at part of the project, and the configuration rules are gated on Android signals this scan may never have opened. |
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the read-only and idempotent annotations, the description discloses that it builds nothing, starts no emulator, takes no screenshot, and cannot produce rendered-output, contrast, TalkBack, or Play review findings. It also explains that silence can mean the Android-project gate rather than a passing result and that the report exposes notVisible items derived from the run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with the purpose and then proceeds logically through configuration, Compose signals, behavioral limits, output shape, error handling, and companion docs. It earns its length for a complex tool, though there is minor redundancy around snippet mode and configuration being the backbone that keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with high complexity, the description covers the input contract, platform inference, configuration-sensitive rules, output format, notVisible semantics, error behavior, and related documentation. There is no critical missing context an agent would need to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters at 100% coverage, so the baseline is 3. The description adds meaning by clarifying that path refers to a project or module directory, that absolute paths are preferred, that code is rejected rather than ignored, and that missing paths or file paths produce an error result rather than an empty audit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and target: 'Audit an Android app's UI against Material 3 and android-app-design' and explains the mechanism by reading AndroidManifest.xml, resource XML, and Kotlin. This clearly distinguishes the tool from sibling audits like audit_apple_ui and audit_generic_design by platform and audit basis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear invocation context: point it at a project or module directory, directory-only, no snippet mode, a code argument is rejected, and a missing path or file path returns an error. It also suggests pairing with get_design_doc for rule guidance, but it does not explicitly name sibling alternatives or state when-not-to-use conditions beyond the snippet limitation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_apple_uiAudit Apple UiARead-onlyIdempotent
Audit an iOS or macOS app's UI against Apple's own documentation: point it at the project directory and it reads the four surfaces an Xcode project declares configuration on — the information property list, INFOPLIST_KEY_* build settings in project.pbxproj, the entitlements plist, and each colorset's Contents.json — infers whether the project targets iOS or macOS from those plus the Swift imports, and runs eight rules. Configuration: a custom colorset with no luminosity: dark appearance, the deprecated UIRequiresFullScreen key (either spelling), a microphone entitlement declared under one capability and not its twin, and — on macOS only — no App Sandbox entitlement, reported as a fact about the Mac App Store channel rather than as a defect. Swift: NavigationView, .font(.system(size:)) on iOS only, a colour written as numbers, and a Button whose whole label is one SF Symbol. Every platform-scoped rule stays silent when the platform signals do not settle the question, and the report says which platform was inferred and from what, so a silence can be read as the gate rather than as a result. It reads source and does not measure anything: it builds nothing, runs no simulator, takes no screenshot, and no finding is or can be a rendered-output, contrast, notarization or App Review result. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line — configuration findings carry no line, since a missing key has no position), a severity summary, a scan block saying how much was actually read, and a long machine-readable notVisible list of what it structurally could not check, every entry of it derived from a run. Directory only — there is no snippet mode, because configuration is the backbone of this audit and a snippet carries none of it. A missing path, a path that is a file, or a code argument is returned as an error result, not as an empty audit. Pair with get_design_doc("apple-hig-liquid-glass"), get_design_doc("apple-accessibility") and get_design_doc("apple-shipping-readiness") for the guidance behind the rules.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | Not supported by this tool, and rejected with an explanation if passed. There is no snippet mode: the platform verdict and half the rules are read from the information property list, the entitlements plist, `project.pbxproj` and the asset catalog, and a pasted snippet carries none of them. Pass `path` instead. | |
| path | No | The Xcode project directory to audit — the folder holding your `.xcodeproj`, `Info.plist`, `Assets.xcassets` and Swift sources. Required. Absolute paths are strongly preferred: a relative path is resolved against the server's working directory, which is usually not your project folder. |
Output Schema
| Name | Required | Description |
|---|---|---|
| scan | Yes | What the scan reached. Read it before trusting any absence claim: a capped scan looked at part of the project, and the platform-scoped rules are gated on configuration this scan may never have opened. |
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description adds substantial behavioral detail: it builds nothing, runs no simulator, takes no screenshot, reports platform inference, and returns a notVisible list of unchecked items. No contradiction exists between description and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and information-dense, but it is a lengthy block with redundancy ('no snippet mode' appears multiple times) and a long enumeration of rules and output fields. It earns its place for a complex tool, but could be tighter and better structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, input requirements, surfaces read, rule categories, platform inference, output structure, error handling, and complementary tools. An output schema exists, but the description still explains what the findings contain, making the tool well-specified for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds meaning by clarifying that a path must be a directory, that missing/file paths return error results rather than empty audits, and why code is rejected. This goes slightly beyond the schema's parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Audit an iOS or macOS app's UI against Apple's own documentation.' It further specifies the project surfaces read and lists the eight rules, making the tool unmistakably distinct from siblings like audit_android_ui, audit_accessibility, and generic design audits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: point it at the project directory, directory only, no snippet mode, and code/missing path/file are errors. It also suggests pairing with get_design_doc calls. However, it does not explicitly name sibling alternatives (e.g., 'use audit_android_ui for Android projects') to provide when-not-to-use guidance against specific alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_design_systemAudit Design SystemARead-onlyIdempotent
Measure how systematic an existing UI really is: paste CSS / SCSS / Tailwind / JSX source and get a consistency score plus the sprawl behind it — how many distinct colors, font sizes, radii, shadows and spacing values it actually uses, which colors are near-duplicates nobody can tell apart, which spacing is off the 4pt grid, token adoption, stray font families, !important and magic z-index. Returns a consolidation plan wired to the generators. Use it before a redesign, on an inherited codebase, or to prove a design system is (or isn't) being followed. Deterministic static analysis; complements design_lint (per-line anti-patterns) with a whole-codebase view.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The stylesheet / token file / component source to audit. Concatenate several files to audit them together. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool read-only, idempotent, and non-destructive; the description reinforces this with 'Deterministic static analysis' and 'paste source,' implying no side effects and no external repo access. It adds context about the return value, including a consolidation plan, though it does not mention limits like input size or exact response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover what the tool measures, when to use it, and how it relates to a sibling tool, with no filler. The output metrics are front-loaded, making the core purpose immediately clear before the use cases and comparison.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only analysis tool, this is nearly complete: input format, output highlights, use cases, and sibling differentiation are all covered. Since there is no output schema, a bit more detail on exact return shape or constraints would make it fully complete, but the description is still sufficient for correct tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single code parameter already has a clear schema description, including concatenating multiple files, so schema coverage is 100%. The tool description adds value beyond the schema by naming accepted formats: 'CSS / SCSS / Tailwind / JSX source,' which helps an agent know exactly what input is valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Measure how systematic an existing UI really is' and enumerates concrete outputs like consistency score, color/spacing/token sprawl, near-duplicates, and off-grid values. It also distinguishes itself from design_lint by explicitly contrasting a 'whole-codebase view' with 'per-line anti-patterns.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger scenarios: 'Use it before a redesign, on an inherited codebase, or to prove a design system is (or isn't) being followed.' It also names the relevant alternative tool, design_lint, and explains how the two complement each other, giving an agent clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_ethical_designAudit Ethical DesignARead-onlyIdempotent
Audit a pasted snippet for named deceptive-pattern tells from ethical-design: confirmshaming decline copy, a pre-checked marketing or newsletter checkbox, a scarcity or deadline phrase with no live binding nearby, and an Accept all with no equally-named reject. It reads source and does not measure anything: no checkout is completed, no countdown is timed, no warehouse is queried, and no consent banner is clicked, so no finding is or can be a verdict on the business. Snippet only — there is no directory mode. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, line), a severity summary, and a machine-readable notVisible list of what it could not check. Pair with get_design_doc('ethical-design') and a human looking at the render.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The HTML/JSX/copy snippet to audit for named deceptive-pattern tells |
Output Schema
| Name | Required | Description |
|---|---|---|
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations' readOnly/idempotent hints, the description discloses concrete behavioral boundaries: no checkout is completed, no countdown is timed, no warehouse is queried, and no consent banner is clicked. It also states that findings are not verdicts on the business, which is critical context for interpreting results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed yet tightly organized: purpose first, then behavioral caveats, then return contents, then usage pairing. Every sentence earns its place; the explicit 'not doing' list prevents harmful assumptions without being padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one required parameter, an output schema, and detailed annotations, the description still adds valuable context: the exact patterns checked, the return shape including the notVisible list, and the limitation that nothing is actually measured. Agents have everything needed to invoke and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the code parameter is already well documented. The description adds marginal context by calling it a 'pasted snippet' and emphasizing 'Snippet only,' but it does not need to add more meaning because the schema handles the definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Audit a pasted snippet for named deceptive-pattern tells from ethical-design.' It then enumerates the exact patterns checked, making the tool's scope unmistakable and clearly distinct from broader design or accessibility audits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly scopes usage with 'Snippet only — there is no directory mode' and advises pairing with get_design_doc('ethical-design') and a human reviewing the render. This gives clear when-to-use and when-not-to-use guidance, though it does not name a specific alternative for full-page audits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_generic_designAudit Generic DesignARead-onlyIdempotent
Audits a web project or snippet for the specific defaults generated interfaces reach for: the stock Tailwind indigo/violet/purple gradient (as classes, hex, or OKLCH), Inter/Roboto/Open Sans/DM Sans/Plus Jakarta Sans as the only declared typeface on a brand surface, emoji standing in for icons, the rounded-2xl + shadow-lg + border card recipe repeated across a page, gradient-filled heading text, an eyebrow label over every heading, the backdrop-blur + white/10 glassmorphism recipe, three or more animate-pulse or animate-shimmer placeholders, stock hype-opener copy ('unlock the power of', 'say goodbye to', …), stacked filler adverbs ('seamlessly', 'effortlessly', …), and a page whose every call to action is drawn from the stock set ('Get Started', 'Learn More'). Every finding is a fact about the source text — a class name, a phrase, a repeated structure — never a judgement about whether the result is good design; it reports facts, not taste, so pair it with design_review_checklist or get_design_doc("design-critique-scoring") for actual critique. It reads source and does not measure anything: it makes no network request, renders nothing, and no finding is or can be a rendered-output or aesthetic judgement. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line), a severity summary, a machine-readable notVisible list of what it could not check, and a 0-100 score itemised to the same rule, weight and evidence the markdown prints — each rule counts once no matter how many times it fires, so a long page never scores higher purely for its length. In directory mode it does not read story, test or fixture files (.stories., .story., .spec., .test., fixtures/, mocks/), whose job is to demonstrate a component rather than ship a surface; it reports how many it skipped. The copy rules match English only, so a page in another language is scored by the visual rules alone. A missing or non-directory path is returned as an error result, not as an empty audit. Pair with audit_project for design drift and design_review_checklist for critique.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | A single snippet to audit instead of a directory. | |
| path | No | Directory to audit. Absolute paths are strongly preferred. | |
| filename | No | Filename for the snippet, e.g. 'page.html' or 'Page.tsx'. Some rules — the typeface check in particular — use it to tell a landing page from a dashboard. |
Output Schema
| Name | Required | Description |
|---|---|---|
| scan | No | Present only in directory mode: what the scan actually reached. The markdown's "Stopped at the file cap" sentence has no other counterpart in structuredContent — `findings`, `summary` and `score` all look identical for a clean project and a truncated one that never opened its worst file — so a caller reading only structuredContent must check this before trusting a clean score or an absent finding. |
| score | Yes | The generic-design score, itemised. Not a quality judgement — it counts documented defaults that were left unchanged. |
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly, idempotent, and non-destructive, and the description reinforces this with concrete behavior: 'it makes no network request, renders nothing,' and no finding is or can be a rendered-output or aesthetic judgement. It discloses non-obvious scoring semantics ('each rule counts once no matter how many times it fires'), directory-mode exclusions of story/test/fixture files, and reports how many files it skipped. This adds substantial behavioral context beyond what the annotations alone convey, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but well-segmented: it moves from the audited pattern list to the reporting stance, output format, edge cases, and finally sibling pairings, with no filler phrases. It runs long (~280 words), and some output-format detail (findings fields) overlaps with the existing output schema, so a small trim was possible; still, nearly every clause carries non-redundant operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param, 0-required tool with annotations and an output schema already present, the description covers everything an agent needs: detailed scope, output shape, scoring semantics, file exclusions, language limitation, error behavior, and sibling pairings. Because the output schema already defines return values, the description correctly focuses on behavior and selection criteria instead of restating structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter (code, path, filename) already documented, including the filename nuance that the typeface check uses it to tell a landing page from a dashboard. The description adds behavioral context around the path parameter (directory-mode exclusions) and the error result, but the schema already carries the parameter definitions, so the description reinforces rather than compensates. Baseline 3 is appropriate since the structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Audits a web project or snippet' for a precisely enumerated set of generic-design defaults (stock Tailwind gradients, default typefaces, emoji icons, card recipes, hype copy). It further distinguishes itself by its factual stance — 'reports facts, not taste' — and explicitly positions itself against siblings like audit_project and design_review_checklist, so an agent can tell it apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit pairing guidance: use design_review_checklist or get_design_doc for actual critique, audit_project for design drift, and this tool for factual generic-default detection. It also documents when to pass a snippet (code) vs a directory (path), states the English-only copy-rule limitation, and defines the error result for missing paths — leaving no ambiguity about when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_performanceAudit PerformanceARead-onlyIdempotent
Audit a page, a component or a whole web project for the performance signals that are actually in the source: a hero image held back by loading="lazy", or contradicting its own fetchpriority; the LCP-candidate image declaring no fetch priority at all; a hero background declared in CSS or an inline style, which the HTML preload scanner never sees; images with no width/height or aspect-ratio to reserve their box; a in the carrying neither defer nor async nor type="module"; @font-face blocks with no font-display; fonts served from a third-party font CDN; and scripts loaded from more distinct remote domains than any reading of "minimise" defends. It reads source and does not measure anything: Core Web Vitals are 75th-percentile field data from real devices, this loads nothing and times nothing, and no finding is or can be an LCP, INP or CLS verdict — so do not call it expecting a vitals report. Its hero rules are deliberately narrow (the first image inside , with the header logo and the mid-article diagram structurally excluded), which means some pages get no hero finding at all; that limitation and the others are returned explicitly rather than left to read as a clean result. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line), a severity summary, and a machine-readable notVisible list of what it could not check. A missing or non-directory path is returned as an error result, not as an empty audit. Pair with audit_seo_geo for the crawl and answer-engine signals, and measure_screenshot for the rendered result.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | A single snippet to audit instead of a directory. | |
| path | No | Directory to audit. Absolute paths are strongly preferred. Every file is audited on its own — a stylesheet in another file does not size an image in this one, even when both are scanned. | |
| filename | No | Filename for the snippet, e.g. 'index.html', 'Page.tsx' or 'styles.css'. Some rules depend on it: a stylesheet and a component are read differently. |
Output Schema
| Name | Required | Description |
|---|---|---|
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond annotations by stating it reads source and does not measure: no page loading, no timing, and no Core Web Vitals verdicts. It also discloses behavioral edge cases such as narrow hero rules, a machine-readable notVisible list of unchecked items, and a missing path returning an error result rather than an empty audit. No contradiction with readOnlyHint, idempotentHint, or destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the tool is genuinely complex and every sentence carries decision-relevant information: scope, examples, limitations, error behavior, return shape, and sibling pairings. The purpose is front-loaded and the added caveats prevent misuse, so the length is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema already documents return values, the description covers everything needed for selection and invocation: what the tool audits, what it does not do, how it handles files and filenames, how errors appear, and which siblings complement it. An agent would be unlikely to misuse or under-use this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters. The description adds meaningful nuance beyond the schema, especially that path audits every file on its own and does not let a stylesheet in another file size an image in this one, and that filename changes how rules read a component versus a stylesheet. This extra context helps invocation correctness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: audit a page, component, or web project for source-level performance signals, with concrete examples. It explicitly contrasts itself with measurement tools like measure_screenshot by stating it loads nothing and times nothing, so it is clearly distinguishable from audit_accessibility, audit_seo_geo, and the other audit siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance on when not to use it: 'do not call it expecting a vitals report.' It also names the complementary sibling tools to pair with for different signal types, audit_seo_geo and measure_screenshot, and surfaces limitations around hero rules and missing paths. This is strong usage-direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_projectAudit ProjectARead-onlyIdempotent
Audit a real codebase instead of a pasted snippet: point it at a directory and it walks the design source, runs the design/accessibility lint over every file, and scores the whole thing for consistency — how many distinct colours, type sizes, radii, shadows and spacings the project actually uses, and which colours are indistinguishable duplicates. Returns findings ranked worst-file-first with file:line, plus an explicit list of what it did not look at. It reads source and does not measure anything: it loads no page, renders nothing, takes no screenshot, and no finding is or can be a rendered-output result. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line), a severity summary, a machine-readable notVisible list of what it could not check, and a scan block saying how many files and bytes were actually read, which files were skipped for size, which could not be opened, and whether the file or byte cap was hit — check that before trusting any absence in findings. It runs design_lint's rules and the consistency count, and nothing else: run audit_security, audit_generic_design, audit_seo_geo and audit_performance on the same directory for theirs. A missing or non-directory path is returned as an error result, not as an empty audit. Cross-file drift is the thing a single-file lint cannot see, which is the point of this tool. Reads only the directory you name; makes no network call. Pair with measure_screenshot for the rendered result and audit_ux_copy for the words.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Directory to audit. Absolute paths are strongly preferred — a relative path is resolved against the server's working directory, which is usually not your project folder. | |
| extensions | No | Override which file extensions are scanned, e.g. ['.tsx','.css','.js']. Replaces the default list rather than adding to it, so name every extension you want read, each with a leading dot — 'js' matches nothing and the audit comes back empty, where '.js' works. Defaults to .css, .scss, .sass, .less, .html, .htm, .jsx, .tsx, .vue, .svelte, .astro — that whole list, read from the array the audit actually uses rather than restated here; .js and .ts are excluded by default because most are logic, not UI. |
Output Schema
| Name | Required | Description |
|---|---|---|
| scan | Yes | What the scan reached. Read it before trusting any absence claim: a capped scan looked at part of the project. |
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds substantial context beyond them: 'it loads no page, renders nothing, takes no screenshot, and no finding is or can be a rendered-output result'; 'Reads only the directory you name; makes no network call'; and the error-handling trait 'A missing or non-directory path is returned as an error result, not as an empty audit' plus the file/byte-cap warning. These are exactly the behavioral traits that prevent misuse, and nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but earns nearly every sentence: it front-loads the core purpose, then covers output shape, safety profile, sibling routing, error behavior, and pairing. Given 35 siblings and a complex multi-part return, the density is justified. Minor redundancy exists ('Cross-file drift... which is the point of this tool' restates the opening's 'instead of a pasted snippet'), which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool of this complexity, nothing essential is missing: inputs (directory + extension override), behavior (read-only, no network), return shape summarized (findings, severity summary, notVisible, scan block) with the output schema covering details, error semantics, exclusions, alternatives, and a cap-related caution that protects against false-negative conclusions. The output schema exists, so the description correctly delegates return-value details to it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the schema's `extensions` documentation is already exhaustive (leading-dot requirement, replacement semantics, default list, .js/.ts exclusion rationale). The description reinforces the `path` meaning ('point it at a directory') and warns to check the `scan` block, but adds little parameter-specific meaning that the schema doesn't already carry. The schema does the heavy lifting, which the rubric treats as acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Audit a real codebase... point it at a directory') and precisely enumerates what the tool does: walk design source, run design/accessibility lint over every file, score consistency (colors, type sizes, radii, shadows, spacings), and flag indistinguishable duplicates. It distinguishes itself from siblings by name (audit_security, audit_generic_design, audit_seo_geo, audit_performance, design_lint, measure_screenshot, audit_ux_copy), so an agent can tell exactly which tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing: 'It runs design_lint's rules and the consistency count, and nothing else: run audit_security, audit_generic_design, audit_seo_geo and audit_performance on the same directory for theirs.' It states when the tool is the right choice ('Cross-file drift is the thing a single-file lint cannot see') and even gives pairing guidance ('Pair with measure_screenshot for the rendered result and audit_ux_copy for the words'). Exclusion is explicit and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_securityAudit SecurityARead-onlyIdempotent
Audit a web project or snippet for security defects a frontend actually ships: missing or weak Content-Security-Policy, absent HSTS, unpinned cross-origin scripts, mixed content, credentials in localStorage, secret-named NEXT_PUBLIC_/VITE_ variables, unsandboxed third-party iframes, wildcard postMessage, raw-HTML sinks with no sanitiser, production source maps and un-ignored .env files. Header state is inferred from wherever your stack declares it — next.config, vercel.json, netlify.toml, _headers, staticwebapp.config.json, firebase.json, Nuxt routeRules, a Remix/React Router headers export, SvelteKit hooks.server.ts and kit.csp, Astro middleware and Next.js middleware or its Next 16 rename proxy.ts, new Response(body, { headers }) and new Headers({…}) on Cloudflare Workers/Deno/Bun — and, as rules rather than places, a quoted header-name property in any JSON or object literal, and any call whose method name is set, setHeader, append or header whatever the object is called (res.set, res.setHeader, headers.set, headers.append, Fastify reply.header, Hono c.header, Koa ctx.set) — and , read as text and never evaluated. It reads source and does not measure anything: it makes no request to your site, tests no live endpoint, and no finding is or can be a penetration-test or vulnerability-scan result — so do not call it expecting one. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line), a severity summary, and a machine-readable notVisible list of what it could not check. A missing or non-directory path is returned as an error result, not as an empty audit. Pair with audit_project for design drift and audit_accessibility for WCAG.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | A single snippet to audit instead of a directory. Source rules only. | |
| path | No | Directory to audit. Absolute paths are strongly preferred. Required for configuration and header rules — a snippet cannot show them. | |
| filename | No | Filename for the snippet, e.g. 'page.html' or 'Page.tsx'. Some rules depend on it: an inline onclick is a defect in HTML and normal JSX in a .tsx file. |
Output Schema
| Name | Required | Description |
|---|---|---|
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and idempotentHint annotations, the description discloses static-only behavior, no network evaluation, header parsing as text rather than execution, error handling for missing or non-directory paths, and the structured output shape including the `notVisible` list. This is exactly the behavioral context an agent needs beyond structured metadata. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence front-loads purpose effectively, and the long enumerations of security rules and config-file locations are dense with actionable detail rather than fluff. However, the description is quite long and could be more scannable with structured lists, so it earns four rather than five.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with this complexity, the description is remarkably complete: it covers input modes, detection scope, framework inference, limitations, return format, error behavior, and related sibling tools. The output schema already handles return-value documentation, so the description does not need to repeat that. Nothing essential for correct selection or invocation appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds genuinely new meaning: `path` is required for configuration and header rules while `code` is source-rules-only, absolute paths are strongly preferred, and `filename` affects rule interpretation by extension. This helps an agent select and fill parameters correctly in a way the schema alone does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb-object pairing: 'Audit a web project or snippet for security defects a frontend actually ships.' It enumerates concrete defect categories, making the tool's scope unmistakable. It also distinguishes itself from siblings by explicitly pairing with audit_project and audit_accessibility for other concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context by scoping to source-based frontend security checks and explicitly stating what it cannot do: 'it makes no request to your site, tests no live endpoint, and no finding is or can be a penetration-test or vulnerability-scan result — so do not call it expecting one.' It also names companion tools for adjacent use cases, helping an agent route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_seo_geoAudit Seo GeoARead-onlyIdempotent
Audit a page, a component or a whole web project for the SEO and GEO signals that are actually in the source: a missing , and one longer or shorter than the width a result gives it; a missing meta description, and one outside the width a snippet gives it; more than one on a page; a heading level skipped in the outline; no canonical link at all; a canonical written as a relative URL in a self-contained document; a canonical left pointing at localhost or a staging host; an hreflang set that never lists the page itself; JSON-LD that does not parse, or that declares no @context or @type; JSON-LD declaring a type whose rich result Google has retired; an image with no alt attribute at all; robots.txt crawl rules, including the AI crawlers behind ChatGPT, Claude, Perplexity and Google's AI surfaces; a robots.txt that names no sitemap; no llms.txt beside it; and content that exists only once a script has run. It reads source and does not measure anything: no request is made to your site, nothing is rendered, and no finding is or can be a Core Web Vitals result, an indexing status or a ranking outcome — so do not call it expecting a vitals or ranking report. Absence is only ever claimed where it can be proven — a self-contained HTML document, or a whole directory — and a scan that hits its cap downgrades every absence claim to an unconfirmed note. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, file, line), a severity summary, and a machine-readable notVisible list of what it could not check. A missing or non-directory path is returned as an error result, not as an empty audit. Pair with audit_performance for the delivery signals, audit_ux_copy for whether the writing earns the click, and seo_geo_guide for the guidance behind the rules.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | A single snippet to audit instead of a directory. Page rules only. | |
| path | No | Directory to audit. Absolute paths are strongly preferred. This is the useful mode — robots.txt, llms.txt, sitemap and project-wide metadata rules all need a directory. | |
| filename | No | Filename for the snippet, e.g. 'index.html' or 'page.tsx'. Load-bearing: a plain HTML file carries its whole <head> and can prove metadata absent, a framework component cannot. |
Output Schema
| Name | Required | Description |
|---|---|---|
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces them with concrete behavior: 'no request is made to your site, nothing is rendered'. It goes well beyond annotations by disclosing the proof policy for absence claims, the cap's downgrading of claims to 'unconfirmed notes', and the error result for missing/non-directory paths — all genuinely useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries information, and the definition is appropriately sized for a complex tool with exclusions, modes, and sibling routing. However, it is delivered as one dense unbroken paragraph — the long enumeration of checks is a sprawling run-on — which makes it harder for an agent to parse quickly than the content justifies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter, two-mode tool with an output schema, the description is complete: it enumerates all rule categories, states the safety profile, explains the absence-proof policy and cap behavior, describes the return shape, documents the error condition, and routes to complementary tools. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real meaning above the schema: that code mode runs 'page rules only', that path mode is needed for robots.txt/llms.txt/sitemap rules, and that filename is 'load-bearing' because a plain HTML file can prove metadata absent while a framework component cannot. This explanation of the code/path/filename interplay is valuable beyond the field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb ('Audit') and resource ('SEO and GEO signals that are actually in the source') and enumerates an exhaustive, concrete list of checks. It further sharpens the definition by stating what it is not ('no request is made to your site... not a vitals or ranking report'), which cleanly distinguishes it from audit_performance and other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-not ('do not call it expecting a vitals or ranking report'), names complementary siblings with their distinct roles ('audit_performance for the delivery signals, audit_ux_copy for whether the writing earns the click, seo_geo_guide for the guidance behind the rules'), and even tells the agent which mode is the 'useful mode' (path over code). This is exemplary routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_ux_copyAudit Ux CopyARead-onlyIdempotent
Audit UI / marketing copy objectively: readability (Flesch reading ease + grade level), average sentence length, passive voice, jargon/hype words, filler, user-focus ('you' vs 'we'), and weak CTAs. Returns metrics plus specific flagged phrases and fixes. The machine-checkable slice of UX writing — pair with get_design_doc('ux-writing') for voice/tone judgment. It reads source and does not measure anything: no usability session is run, no A/B result is read, and no finding here is or can be a statement about whether the copy actually works for a reader. It also has no notion of register: it has no way to tell short UI copy from long-form prose, so a paragraph of accurate technical documentation can draw more jargon/filler hits than a paragraph of real hype. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, line), a severity summary, a machine-readable notVisible list of what it could not check, and a metrics block carrying the same words/sentences/avgSentenceLen/Flesch/grade-level/you-we numbers the markdown table prints.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The copy to audit (a headline, paragraph, button label, error message, or full page copy) |
Output Schema
| Name | Required | Description |
|---|---|---|
| metrics | Yes | The metrics table the markdown half prints, as numbers. Facts about the text's shape, not measurements of how it reads — see `notVisible`. |
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description discloses important non-obvious behaviors: it reads source and does not measure anything, has no concept of register, and may falsely flag technical documentation for jargon/filler. This significantly shapes how an agent should interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loads purpose, but the final sentence enumerating the output structure is verbose and partially redundant with the markdown/metrics repetition. Every sentence does add value, though the length is slightly more than needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description thoroughly covers the return shape, limitations, and relationship to sibling tools. An agent has everything needed to call it correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides full coverage for the single 'text' parameter with concrete examples of acceptable input. The description adds no additional parameter-level semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb ('Audit') and a well-defined resource ('UI / marketing copy') followed by an explicit list of metrics. It also distinguishes itself from the voice/tone-oriented sibling get_design_doc, making its scope clear among many audit_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to pair with get_design_doc('ux-writing') for voice/tone judgment, giving a clear alternative for non-machine-checkable aspects. It also specifies what the tool cannot do (no usability claims, no register awareness), which helps an agent decide when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_design_languagesCompare Design LanguagesARead-onlyIdempotent
Compare how iOS (HIG/Liquid Glass), Android (Material 3), macOS and the web each solve ONE design problem — navigation, buttons, modals/sheets, typography, color, elevation, motion, forms, lists, icons, search or settings. Returns a side-by-side table of the concrete conventions per platform, the rules for porting a design between them, and an explicit 'do NOT port' list. Use when building the same product on more than one platform, or when deciding whether a pattern that works on one platform belongs on another.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | The surface to compare, e.g. 'navigation', 'buttons', 'modals-sheets', 'motion' | |
| platforms | No | Which platforms to include as columns (default all four). e.g. ['ios','android'] for a mobile-only comparison |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint, idempotentHint, and non-destructive behavior. The description goes further by revealing the concrete return shape — a side-by-side table of conventions, porting rules, and an explicit 'do NOT port' list — and by emphasizing the one-problem scope, which is useful contextual behavior beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of three substantive sentences that front-load the core action and then provide output and use-case detail. The long enumeration of design problems mostly mirrors the schema enum, introducing minor redundancy, but the overall structure remains tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates well by explicitly describing the side-by-side table, porting rules, and do-NOT-port list. Schema coverage is 100%, annotations clarify the read-only safety profile, and the use cases are stated. An agent has everything needed to decide whether and how to invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both topic and platforms with examples. The description reinforces that topic is a design problem and platforms become columns, but it adds no new syntactic or semantic detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Compare') and a clearly delimited resource: how iOS, Android, macOS, and web each solve ONE design problem, with the problem categories enumerated. This strongly distinguishes it from sibling tools like get_design_language or search_design_knowledge, making its cross-platform comparative purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit use conditions: 'when building the same product on more than one platform' or 'when deciding whether a pattern that works on one platform belongs on another.' It does not, however, name alternatives or state when not to use it, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_design_systemCreate Design SystemARead-onlyIdempotent
THE one-call foundation. Turn a brand color + product vibe + platform into a complete, coherent design-system starter: a direction card (named defaults to leave, the type pairing as a decision, one signature move, and a stock-region fact if the seed sits in indigo/violet/purple), accessibility-verified color (light+dark), a matched font pairing, an icon library, a modular type scale, an elevation ramp, ready-to-paste design tokens (CSS/Tailwind or SwiftUI/Compose), the components to build, and a build checklist — all generated to work together. Use this FIRST when someone says 'design/build me a website/app' to lay the foundation, then get_component_recipe for each component and get_design_roadmap for the full process.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Brand/token name (default 'Brand') | |
| vibe | Yes | Product vibe / use case, e.g. 'modern SaaS dashboard', 'premium fintech app', 'bold marketing site', 'minimal portfolio' | |
| platform | No | Target platform (default web) — picks icon set and token output | |
| brand_color | Yes | Brand / primary color as hex, e.g. '#4F46E5' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint=true, idempotentHint=true, and destructiveHint=false, lowering the bar. The description adds genuinely useful behavioral context: outputs are platform-dependent ('CSS/Tailwind or SwiftUI/Compose'), one deliverable is conditional ('a stock-region fact if the seed sits in indigo/violet/purple'), and 'ready-to-paste' clarifies the tool returns artifacts for the user to apply rather than persisting state. No contradiction with the readOnly annotation since the generation is computational, not a state mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The usage rule is front-loaded at the end of the description and the headline is punchy, but the opening sentence is a single dense run-on enumerating ten artifacts with nested parentheticals ('named defaults to leave, the type pairing as a decision, one signature move, and a stock-region fact...'). The marketing phrase 'THE one-call foundation' adds emphasis but little functional value, and the wall-of-text structure makes scanning harder than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description bears the return-value burden and largely carries it: it lists every generated artifact, explains the conditional and platform-dependent behaviors, and names the workflow that follows. The only real gap is that the response envelope (structure/format of the returned design system) is unspecified, though the artifact list makes the return content predictable for a 4-parameter generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — all four parameters have meaningful descriptions with examples and defaults, so the baseline is 3. The description reinforces the relationship (brand_color + vibe + platform drive the whole output) and notes platform 'picks icon set and token output,' but adds no syntax or format details beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Turn a brand color + product vibe + platform into a complete, coherent design-system starter' and enumerates concrete deliverables (direction card, tokens, type scale, elevation ramp, components, build checklist). It differentiates from the many design siblings by positioning itself as the 'one-call foundation' and naming get_component_recipe and get_design_roadmap as the follow-up tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs: 'Use this FIRST when someone says design/build me a website/app to lay the foundation, then get_component_recipe for each component and get_design_roadmap for the full process.' This gives an agent a clear decision rule, sequential ordering, and named alternatives without leaving anything to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_lintDesign LintARead-onlyIdempotent
Lint a snippet of HTML / CSS / JSX / Tailwind for design & accessibility anti-patterns: hardcoded colors instead of tokens, px font-sizes, removed focus outlines, images without alt, clickable divs, icon-only buttons without labels, positive tabindex, ad-hoc radii, !important overuse. Returns findings with line numbers, severity, and fixes. It reads source and does not measure anything: nothing is rendered, no contrast ratio is computed and no tap target is sized, so no finding is or can be a visual or an accessibility verdict. Returns markdown plus structured output: findings (rule, severity, message, fix, doc, line), a severity summary, and a machine-readable notVisible list of what it could not check. Fast static design-time check — not a replacement for a full audit. Complements design_review_checklist.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The HTML/CSS/JSX/Tailwind snippet to lint |
Output Schema
| Name | Required | Description |
|---|---|---|
| summary | Yes | Counts by severity. Always agrees with `findings` — it is derived from the same list. |
| findings | Yes | Every finding, in the order the markdown report lists them. |
| notVisible | Yes | What this audit structurally could not check, one limitation per entry. Read it as a peer of `findings`: silence on a subject named here is this tool's reach, not a clean result. Nothing any of these tools reports is measured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description reinforces and extends this with valuable behavioral context: 'It reads source and does not measure anything: nothing is rendered, no contrast ratio is computed and no tap target is sized.' It also discloses the output shape (markdown plus structured findings, severity summary, machine-readable notVisible list), so an agent understands the tool's limits and return value beyond what the annotations convey. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and anti-pattern list, then progressively covers limitations, output structure, and tool relationships. It is longer than average but every sentence adds meaningful information. There is minor redundancy between 'Returns findings with line numbers, severity, and fixes' and the later structured-output breakdown, but this is acceptable given the added specificity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read-only tool with an output schema and rich annotations, this description is complete. It covers accepted input languages, checked anti-patterns, limitations, output format, and the relationship to design_review_checklist. An agent has all the information needed to invoke it correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'code' parameter is described as 'The HTML/CSS/JSX/Tailwind snippet to lint.' The description adds extra semantic value by enumerating the anti-patterns it checks for, which helps an agent decide what kind of snippet to pass and what to expect. It does not add format-specific parsing details, but with one well-documented parameter, the marginal gain is modest yet real.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a precise verb and resource: 'Lint a snippet of HTML / CSS / JSX / Tailwind for design & accessibility anti-patterns.' It lists concrete patterns (hardcoded colors, px font-sizes, removed focus outlines, clickable divs, etc.), making the tool's scope unmistakable. It also differentiates itself from sibling tools by explicitly stating what it does not do: no rendering, no contrast ratio computation, no tap target sizing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use context ('Fast static design-time check') and when-not-to-use context ('not a replacement for a full audit'). It names a complementary sibling (design_review_checklist) and warns that findings are not visual or accessibility verdicts, steering agents away from using it in place of measurement-based tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_review_checklistDesign Review ChecklistARead-onlyIdempotent
Generate a structured design-review checklist for a project type (mobile app, website, landing page, dashboard), assembled from the knowledge base: key rules and anti-patterns per area. Use it to audit an existing design or as acceptance criteria for a new one.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | Narrow the review to one dimension (default: all) | |
| project_type | Yes | What is being reviewed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior. The description adds meaningful context beyond that by stating the checklist is assembled from the knowledge base and includes key rules and anti-patterns per area, implying no live design analysis is performed. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary action and output are front-loaded, and the second sentence clarifies practical use. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two enum-bounded parameters and no output schema, the description is largely complete: it defines the output, how it is assembled, and how it should be used. It does not explicitly mention the focus parameter, but the schema covers that, so no significant gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already well documented. The description adds only general context around 'areas' and project types, but does little to explain the focus parameter beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('generate a structured design-review checklist') on a specific resource (project type), and explains what the checklist contains. It distinguishes itself from audit tools by framing the output as a checklist for auditing rather than an audit itself, though it does not explicitly name sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: use the checklist to audit an existing design or as acceptance criteria for a new one. It does not mention alternatives or exclusions, but the intended use cases are concrete enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fix_contrastFix ContrastARead-onlyIdempotent
Repair a failing color pair: given a foreground and background hex, compute the NEAREST accessible color (hue & saturation preserved, lightness nudged) that meets the WCAG 2.2 target — not just a pass/fail report. Use when audit_accessibility flags a pair and you need the corrected value to ship. For a full pass/fail audit use audit_accessibility; to build a whole palette use generate_color_system.
| Name | Required | Description | Default |
|---|---|---|---|
| adjust | No | Which color to move (default 'foreground' — the text) | |
| target | No | Target contrast ratio (default 4.5 = AA normal text; use 3 for large text/UI, 7 for AAA) | |
| background | Yes | Background hex it sits on, e.g. '#FFFFFF' | |
| foreground | Yes | Foreground/text hex to adjust, e.g. '#9CA3AF' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context: it returns a corrected value rather than a pass/fail report, preserves hue/saturation, and nudges lightness. It could mention whether the function returns both colors or just the adjusted one, but with strong annotations this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The core operation and key constraint are front-loaded, and the sibling routing is compact at the end. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param tool with 100% schema coverage, no output schema, and strong annotations, the description is nearly complete. It explains the algorithm, the scope, and the alternatives. The only missing context is exactly what the return value looks like (single hex? pair? object?), but since no output schema exists, that information would have been valuable. Still, the tool is simple enough that the gap is small.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description adds useful context by explaining the default behavior ('foreground' as the text, target default 4.5, and the meaning of different targets), but it doesn't deeply elaborate on parameter formats beyond what the schema provides. A 3 is appropriate because the schema carries the heavy lifting and the description adds modest interpretive value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Repair'), a precise resource (color pair contrast), and the algorithm's intent (nearest accessible color, hue/saturation preserved, lightness nudged). It explicitly distinguishes itself from audit_accessibility and generate_color_system, so an agent can separate it from siblings without inspecting their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use ('when audit_accessibility flags a pair') and when-not-to-use ('For a full pass/fail audit use audit_accessibility; to build a whole palette use generate_color_system'), naming the alternatives directly. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_color_systemGenerate Color SystemARead-onlyIdempotent
Turn ONE brand color into a complete, accessibility-verified palette: a 50–950 tonal scale, a cohesive brand-tinted neutral ramp, and full light + dark semantic tokens (background, surface, border, text, primary/onPrimary, subtle, focus ring). Every text/UI pair is checked against WCAG 2.2 and auto-adjusted to pass. Deterministic — outputs a real palette, not advice. Feed the result into generate_design_tokens, then audit_accessibility.
| Name | Required | Description | Default |
|---|---|---|---|
| brand_color | Yes | The brand / primary color as hex, e.g. '#4F46E5' or '#e11d48' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/destructive annotations, the description discloses deterministic behavior, WCAG 2.2 checking with auto-adjustment, and that it returns a real generated palette rather than advice. This gives the agent strong expectations about side-effect-free but substantive output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the first defines scope and output, the second explains the accessibility guarantee, and the third conveys determinism and downstream steps. The description is dense but not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining return value composition, and it does so thoroughly: tonal scale, neutral ramp, light/dark semantic tokens, and WCAG verification. For a one-parameter deterministic generator, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter, brand_color, is already described as a hex string with examples. The description adds the 'ONE brand color' emphasis but little new parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Turn ONE brand color into...') and names the concrete output components: tonal scale, neutral ramp, and semantic tokens. It clearly differentiates from siblings like generate_design_tokens and audit_accessibility by framing this as palette generation followed by those downstream tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: a single brand color in, a full system out, plus an explicit downstream workflow ('Feed the result into generate_design_tokens, then audit_accessibility'). It stops short of stating explicit when-not-to-use cases or alternatives like generate_type_scale or generate_elevation_system.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_design_tokensGenerate Design TokensARead-onlyIdempotent
Turn a design-token spec (semantic colors + optional spacing/radius/type scales) into REAL, ready-to-use artifact files: CSS custom properties, Tailwind v4 @theme, SwiftUI, Jetpack Compose, and W3C DTCG JSON. Deterministic — outputs code, not advice. Use it to give a project one source of truth across web, iOS and Android. Pair with audit_accessibility to verify the palette's contrast.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Token set / brand name (default 'Brand') | |
| radii | No | radius name→px (default sm/md/lg/xl/full; use 9999 for pill) | |
| colors | Yes | Semantic color roles → hex. e.g. {"primary":"#4F46E5","onPrimary":"#FFFFFF","surface":"#0A0A0B","textPrimary":"#F5F5F5","danger":"#EF4444"} | |
| format | No | Output format (default 'all') | |
| spacing | No | px spacing scale (default 8pt scale 2..96) | |
| fontSizes | No | type scale name→px (default xs..4xl) | |
| fontFamilies | No | font role→stack (default sans/mono) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint safety traits. The description adds meaningful behavioral context by stating it is 'Deterministic — outputs code, not advice,' which clarifies it does not produce opinions or partial recommendations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences: the first front-loads the core transform and output formats, the second adds a key behavioral trait, and the third gives the primary use case and a related pairing. Every sentence earns its place with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description leaves the return structure unspecified — an agent may not know whether the generated 'artifact files' come back as inline strings, a zip, or a structured mapping. It also does not mention that 'format' defaults to 'all' or supports selective output. For a 7-parameter tool producing multi-file output, this is a notable gap, though the rich input schema and clear purpose mitigate it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema alone documents all parameters with examples and defaults. The description adds modest value by mapping the abstract format enum to real-world outputs ('Tailwind v4 @theme', 'SwiftUI', 'Jetpack Compose', 'W3C DTCG JSON') and grouping the optional scales, but this is marginal beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Turn ... into') and a concrete resource ('a design-token spec'), then enumerates the exact artifact formats produced. It distinguishes itself from sibling generators like generate_color_system or generate_type_scale by emphasizing multi-platform, ready-to-use files and 'one source of truth across web, iOS and Android.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear use case ('give a project one source of truth across web, iOS and Android') and suggests pairing with audit_accessibility to verify contrast. However, it does not explicitly state when to avoid this tool in favor of more targeted generators like generate_type_scale or generate_color_system, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_elevation_systemGenerate Elevation SystemARead-onlyIdempotent
Generate a cohesive elevation / box-shadow ramp (layered ambient + direct light) with semantic level names (flat…modal), as CSS custom properties and Tailwind @theme, plus dark-mode guidance. Deterministic. Use one shadow token per level instead of hand-tuning shadows per component.
| Name | Required | Description | Default |
|---|---|---|---|
| hue | No | Optional shadow tint as 'H S%' e.g. '220 40%' for a cool cast (default neutral black) | |
| levels | No | Number of raised levels (default 5) | |
| strength | No | Opacity multiplier 0.5–1.5 (default 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal read-only, idempotent, and non-destructive behavior. The description adds useful context beyond annotations by stating the tool is deterministic and by describing what the output includes (CSS custom properties, Tailwind @theme, dark-mode guidance). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the primary purpose, and each sentence earns its place. The deterministic note and the one-token-per-level guidance are valuable without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately explains what the tool returns: CSS custom properties, Tailwind @theme tokens, and dark-mode guidance. It also communicates the semantic naming scheme. It could be more explicit about output structure or examples, but it is sufficiently complete for a deterministic, read-only generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already fully documents hue, levels, and strength. The description does not add much parameter-level detail, but it does clarify the overall purpose of the generated ramp and the semantic naming convention, which is enough to maintain the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific deliverable: a cohesive elevation/box-shadow ramp with semantic levels, CSS custom properties, Tailwind @theme integration, and dark-mode guidance. It is distinct from sibling generators like generate_color_system or generate_type_scale because it names its exact output domain and format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need elevation tokens rather than hand-tuning per-component shadows. However, it does not explicitly mention alternatives or state when not to use it, leaving tool-selection guidance mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_layout_systemGenerate Layout SystemARead-onlyIdempotent
Generate the layout foundation the other generators leave out: breakpoints (with what changes at each), container max-widths, edge padding, a column grid, an intrinsic auto-fit card grid, container queries, and a fluid section-rhythm scale — as CSS custom properties and a Tailwind v4 @theme block, plus the rules that matter more than the numbers (design narrow-first, cap the measure at 45–75ch, prefer intrinsic layout to media queries). Deterministic real code. Pair with generate_type_scale and generate_design_tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| gutter | No | Gutter between columns in px (default 16–32 by preset) | |
| preset | No | Layout archetype (default 'marketing-site'): 'web-app' for a dense app shell with a sidebar, 'docs' for a three-zone documentation layout, 'mobile-first' for a 4→12 column phone-first grid | |
| columns | No | Grid columns (default 12, or 4 for mobile-first) | |
| max_width | No | Max content width in px (default depends on preset: 960–1440) | |
| container_queries | No | Include a container-query example so components respond to their own width (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/destructive annotations, the description adds that output is 'Deterministic real code' and states the delivery form: CSS custom properties plus Tailwind v4 @theme. This gives useful behavioral expectations without contradicting any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: the main action and scope lead, followed by a concrete deliverable list, output format, and usage recommendation. Every clause contributes information with no filler; it is slightly long but justified by the tool's breadth.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description serves as the output contract, listing all generated artifacts, the technical format, and the embedded design rules. Combined with complete parameter schemas and safety annotations, an agent has everything needed to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description itself adds no parameter-level detail, but schema description coverage is 100%, with each of the five parameters including defaults, ranges, and enum context. The baseline of 3 applies because the schema carries the full burden for parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Generate the layout foundation' and enumerates exactly what is produced: breakpoints, container max-widths, edge padding, a column grid, an auto-fit card grid, container queries, and a fluid section-rhythm scale. It also states the output format, CSS custom properties plus a Tailwind v4 @theme block, which clearly distinguishes it from sibling generators like generate_color_system or generate_type_scale.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by positioning the tool as the layout foundation other generators leave out, and explicitly recommends pairing it with generate_type_scale and generate_design_tokens. It doesn't provide explicit when-not-to-use conditions or compare against all generate_* siblings, but the context is sufficient for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_motionGenerate MotionARead-onlyIdempotent
Generate a motion system: easing tokens (decelerate/accelerate/standard/spring as cubic-beziers), duration tokens, and ready-to-paste keyframe animations (fade-in, slide-up, scale-in, spring-pop, shimmer) in CSS, Framer Motion, or SwiftUI — grounded in the animation-craft rules (ease-out on enter, small distances, never scale(0), honor reduced-motion). Deterministic real code.
| Name | Required | Description | Default |
|---|---|---|---|
| stack | No | Target stack (default css) | |
| animation | No | Which animation to emit (default all) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description goes beyond this by revealing behavioral traits: outputs are 'grounded in the animation-craft rules', 'deterministic', and 'ready-to-paste real code'. This adds meaningful expectations about output style and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single information-dense sentence that front-loads the core action and resource, then enumerates deliverable components, target stacks, constraints, and output properties. Every phrase contributes meaning; the only minor issue is that the long sentence could have been split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description compensates by detailing what is generated (tokens and animations), which stacks are supported, and the craft rules that shape the output. It doesn't explicitly describe the response format or how 'all' values expand, but for a simple two-parameter generator, the essential information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with enum descriptions for both parameters, providing a solid baseline. The description adds value by mapping animation names to concrete output types (fade-in, slide-up, scale-in, spring-pop, shimmer) and explaining that easing tokens are delivered as cubic-beziers with specific variants (decelerate/accelerate/standard/spring), which is not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Generate' and clearly identifies the resource: a motion system composed of easing tokens, duration tokens, and ready-to-paste keyframe animations. The enumerated target stacks (CSS, Framer Motion, SwiftUI) and animation names distinguish it from sibling tools like generate_design_tokens or generate_color_system.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: when a motion system, easing/duration tokens, or keyframe animations in supported stacks are needed. It does not explicitly exclude alternatives or name sibling tools, but the purpose is stated concretely enough that an agent can infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_type_scaleGenerate Type ScaleARead-onlyIdempotent
Generate a modular typographic scale from a base size and ratio: named steps (xs…6xl) with sizes, line-heights, letter-spacing, and optional fluid clamp() that scales display type down on small screens. Emits CSS custom properties and a Tailwind v4 @theme block. Deterministic real output. Pair with suggest_font_pairing and generate_design_tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | Base body size in px (default 16) | |
| fluid | No | Emit fluid clamp() for headings (default true) | |
| ratio | No | Modular ratio (default 1.25). Common: 1.2 minor-third, 1.25 major-third, 1.333 perfect-fourth, 1.5, 1.618 golden | |
| steps | No | Named steps above base (default 7 → up to 6xl) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint), so the description's job is to add behavior beyond that. It does: 'Deterministic real output' assures the agent the result is computed and genuine rather than a placeholder or sample, and the fluid clamp() behavior ('scales display type down on small screens') plus the emitted @theme block describe the output's runtime behavior. No contradiction with annotations — generating a scale as output is consistent with read-only, idempotent, non-destructive hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with zero filler: purpose and output contents first, then output format, then the behavioral guarantee and companion tools. Every sentence earns its place and the most decision-relevant information (what the tool generates) appears immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining return values — and it does, covering the named steps, sizes, line-heights, letter-spacing, CSS custom properties, and Tailwind @theme block. All 4 parameters are optional and fully documented in the schema, and the annotations disclose safety and idempotency. The only minor gap is that it doesn't explicitly state the output is returned as text/CSS for the agent to consume rather than written to files, but readOnlyHint largely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 with the schema doing the heavy lifting. The description adds marginal connective meaning — 'from a base size and ratio' maps to the base and ratio params, 'optional fluid clamp()' maps to the fluid boolean, and 'named steps (xs…6xl)' clarifies the steps param — but these largely restate what the schema already documents in adequate detail. No critical param semantics are added beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Generate a modular typographic scale from a base size and ratio') and enumerates the exact deliverables: named steps (xs…6xl), sizes, line-heights, letter-spacing, and optional fluid clamp(). It distinguishes itself from siblings like generate_color_system and generate_design_tokens by specifying its precise output surface (CSS custom properties and a Tailwind v4 @theme block), leaving no ambiguity about what this tool produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing line, 'Pair with suggest_font_pairing and generate_design_tokens,' gives concrete workflow context that positions this tool within a broader design-token pipeline. It implies when to use it (typography scale generation) via the focused subject matter, but it does not explicitly state when not to use it or which sibling (e.g., generate_design_tokens) would be the better choice for broader token needs. Clear context, no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_component_guidanceGet Component GuidanceARead-onlyIdempotent
Get expert guidance for designing one UI component or screen pattern (button, form, navigation, card, modal, hero, pricing page, onboarding, paywall, checkout, empty state, dashboard…). Returns the most relevant docs in full — specs, states, sizing, anti-patterns, and real-world patterns from top apps/sites. Use when designing a specific element; for copy-paste code use get_component_recipe, for annotated screenshots use get_design_examples.
| Name | Required | Description | Default |
|---|---|---|---|
| platform | No | Target platform ('mobile' | 'web' | 'macos') — strongly recommended so guidance matches the platform's conventions. | |
| component | Yes | Component or pattern name, e.g. 'primary button', 'signup form', 'bottom tab bar', 'hero section', 'paywall'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful behavioral detail about the output: it 'Returns the most relevant docs in full' and specifies the content categories such as specs, states, sizing, anti-patterns, and real-world patterns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then provides outcome details and routing guidance in a compact three-sentence structure. Every sentence contributes distinct information without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by explicitly describing what will be returned: full relevant docs covering specs, states, sizing, anti-patterns, and real-world patterns. It also covers when to use the tool and how to choose between its closest siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema. The description reinforces the component parameter with examples but does not add material meaning beyond the schema, especially for the platform parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get expert guidance for designing one UI component or screen pattern' and gives concrete examples. It also explicitly distinguishes itself from sibling tools by naming get_component_recipe and get_design_examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool: 'Use when designing a specific element.' It also gives direct alternative tool guidance: 'for copy-paste code use get_component_recipe, for annotated screenshots use get_design_examples.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_component_recipeGet Component RecipeARead-onlyIdempotent
Get production-ready, accessible reference CODE for a UI component in a chosen stack (react-tailwind, html-css, swiftui, compose) — not advice, actual copy-paste code with all states, ARIA/accessibility, keyboard support and correct motion, grounded in the SaglitzDesign specs. Use when you need to actually build a button, input, modal, toast, card, switch, tabs, empty-state, list-row, navigation, search, select, table, tooltip, form, pagination, skeleton, badge, or breadcrumb. Pair with get_component_guidance (the design rationale) and generate_design_tokens (the theme).
| Name | Required | Description | Default |
|---|---|---|---|
| stack | No | Target stack. Omit to get the spec + all available stacks. | |
| scales | No | Your ramps, keyed by step — {"neutral":{"50":"#…",…,"950":"#…"},"primary":{…},"danger":{…}} — exactly the `neutral`, `primary` and `danger` scales generate_color_system returns. Swapped step for step, which is what themes the dark-mode shades too; role tokens alone leave those behind. | |
| tokens | No | Your colours, so the code comes back in them instead of the house palette. Roles: primary, primaryHover, danger, dangerHover, background, textPrimary. Pass the values create_design_system or generate_color_system produced, e.g. {"primary":"#0F62FE","primaryHover":"#0043CE"}. Omit to get the recipe as written. | |
| component | Yes | Component name, e.g. 'button', 'input', 'modal', 'toast', 'card', 'switch', 'tabs', 'empty-state', 'list-row', 'navigation', 'search', 'select', 'table', 'tooltip', 'form', 'pagination', 'skeleton', 'badge', 'breadcrumb' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds valuable behavioral context: the output is production-ready code with all states, ARIA/accessibility, keyboard support, motion, and stack-specific variants. It does not detail edge cases or return structure, but this is a read-only lookup tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the first states the core value, the second gives concrete use cases, and the third points to related tools. The component list is long but directly informs when the tool applies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing what the code includes: states, accessibility, keyboard support, and motion. The optional stack behavior and scales/tokens semantics are already covered in the schema. Missing error/edge-case behavior is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains component, stack, scales, and tokens in detail. The description mainly reinforces that the stack is a choice and lists example components, adding little beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get production-ready, accessible reference CODE for a UI component.' It clearly differentiates from the sibling get_component_guidance by stating this is 'not advice' but actual copy-paste code, and by naming which components it supports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use when you need to actually build' a component, which gives clear situational guidance. It also names companion tools get_component_guidance and generate_design_tokens, but does not explicitly state when not to use this tool versus each sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_design_docGet Design DocARead-onlyIdempotent
Fetch one knowledge-base document in full by its id. Returns the whole document — title, metadata, prescriptive body, and cited sources. Ids come from list_design_knowledge or search_design_knowledge; if the id is unknown it suggests near matches.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Exact document id, e.g. 'buttons', 'material-3', 'accessibility', 'geo-tactics-checklist'. Get ids from list_design_knowledge or search results. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations: it specifies the full return contents (title, metadata, prescriptive body, cited sources) and the unknown-id fallback behavior of suggesting near matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and every sentence adds useful information. There is no filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one required parameter, full schema documentation, and annotations covering safety, the description is complete. It covers what is returned and how unknown ids are handled, so an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with the schema already explaining the id type, examples, and where to get ids. The description only reinforces 'by its id' and repeats the id-source guidance, adding minimal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') with a clear resource ('one knowledge-base document in full by its id'). It also distinguishes this tool from list/search siblings by positioning them as id sources and this tool as the retrieval action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent where valid ids come from (list_design_knowledge or search_design_knowledge), which is actionable usage guidance. It does not explicitly state when not to use it, but the single-document fetch purpose is clear enough to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_design_examplesGet Design ExamplesARead-onlyIdempotent
Fetch curated real-world examples of a design pattern from top apps and websites (paywalls, onboarding, auth, navigation, checkout, settings, empty states, heroes, pricing, features, social proof, signup, dashboards, footers). Returns, for each example, the app/site, what it does well, and a source link to view the screenshot. NOTE: this installation does not bundle the screenshot images (they are third-party assets, excluded from the published package), so the notes and links are returned WITHOUT inline images — open the links, or use your own browser tool, if you need to see them.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max examples to return (default 4; images are large) | |
| query | Yes | Pattern to see examples of, e.g. 'paywall', 'pricing section', 'dark hero', 'empty state' | |
| platform | No | 'mobile' for iOS and Android app screens, 'ios' or 'android' to pin one, 'web' for website examples |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag read-only/idempotent behavior, and the description adds significant value by disclosing that screenshot images are NOT bundled and links must be opened externally. This is a non-obvious behavioral trait that prevents false expectations about inline images.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler; the core purpose is front-loaded and the critical image caveat is clearly separated in a NOTE. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool, the description covers the return shape (app/site, what works well, source link) and the key limitation (no embedded images). With no output schema, this is sufficient for an agent to call it correctly, though it could mention default behavior or empty-result handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add much beyond the schema for parameters; it mainly lists example query values, which is helpful but not essential since query examples already appear in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Fetch curated real-world examples of a design pattern') and a clear resource (examples from top apps/sites), with a helpful list of pattern areas. It is distinct from sibling tools like list_design_knowledge or search_design_knowledge, though it does not explicitly name them as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implicit: an agent would call this when a user wants real-world examples of a pattern. However, it does not explicitly explain when to prefer this over related knowledge/design tools, nor state any exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_design_languageGet Design LanguageARead-onlyIdempotent
Fetch the full reference document for one modern design language or platform design system (Material 3, Apple HIG/Liquid Glass, iOS/Android/macOS, Apple Intelligence, visionOS, Fluent 2, 2026 web trends, design tokens). Returns the complete spec — rules, do/don't lists, numbers, and examples — for the chosen system. Use when you need the authoritative platform baseline before designing; for a specific component use get_component_guidance, and to plan a whole project use get_design_roadmap.
| Name | Required | Description | Default |
|---|---|---|---|
| language | Yes | Which reference to fetch. e.g. 'material-3' (Android/Material), 'apple-hig-liquid-glass' or 'ios-app-design' (iOS), 'macos-app-design', 'visionos-spatial-design' (Vision Pro), 'web-trends-2026', 'design-tokens-theming'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint, idempotentHint, and non-destructive behavior, so the description's additional job is to explain what happens when invoked. It does so by stating that the tool returns 'the complete spec' with rules, do/don't lists, numbers, and examples. This adds useful behavioral context beyond the annotations, though it does not disclose error or edge-case behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: purpose, return content, and usage alternatives are each front-loaded and directly useful. Every sentence earns its place, and the description is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, enum-constrained read-only tool, the description is complete: it states what is fetched, what the returned spec contains, and when to use it versus relevant siblings. No output schema is present, but the description adequately explains the return value's content. There is no significant missing context for correct tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the sole parameter 'language' already has a descriptive enum with concrete examples. The description lists many of the same enum values and calls the parameter a 'design language or platform design system,' but it adds little beyond the schema. Baseline 3 is appropriate because the schema carries the semantic weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch the full reference document for one modern design language or platform design system,' and then enumerates the exact systems covered. It clearly says what is returned (rules, do/don't lists, numbers, examples), distinguishing this tool from generic search or component-level tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is provided: use when you need the authoritative platform baseline before designing. It also names two direct alternatives with their conditions: get_component_guidance for a specific component and get_design_roadmap for whole-project planning. This leaves little ambiguity about when to choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_design_roadmapGet Design RoadmapARead-onlyIdempotent
The SaglitzDesign roadmap: a phased, expert design process for a given project type (website, landing page, iOS app, Android app, macOS app, SaaS web app). Each phase has a goal and the exact knowledge-base docs to consult. Use this FIRST when starting any design project, then fetch phase docs as you reach them.
| Name | Required | Description | Default |
|---|---|---|---|
| project_type | Yes | What is being designed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and non-destructive, so no contradiction exists. The description adds useful behavioral context by explaining that the roadmap is structured by phases, each with goals and associated docs, and that it is the intended entry point for design work.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two focused sentences with no filler. It front-loads what the roadmap is, then gives actionable usage guidance, making it efficient and easy for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema, and the description explains the contents of the roadmap and what the agent should do next. It doesn't describe the exact output format, but that omission is acceptable given the roadmap's clear purpose and the follow-up instruction to fetch phase docs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole parameter, project_type, already has an enum and a description. The description restates the project types, which adds minimal new semantic value beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as returning a phased, expert design roadmap for a specified project type. It names the exact project types and explains that each phase contains goals and knowledge-base docs, which distinguishes it from sibling tools like get_design_doc or audit tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use this FIRST when starting any design project' and instructs the agent to 'fetch phase docs as you reach them.' This gives clear timing guidance, though it does not explicitly name or exclude alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_design_tokensImport Design TokensARead-onlyIdempotent
Read an EXISTING design system and convert it: paste CSS custom properties (a Tailwind v4 @theme block, a shadcn :root block, plain CSS), a W3C DTCG token file, or a theme object as JSON, and get back the roles it names, the semantic roles it is missing, a WCAG contrast check on the pairs it defines, and the whole set re-emitted as CSS / Tailwind / SwiftUI / Compose / DTCG. The inverse of generate_design_tokens — use it to take a web theme to iOS or Android, to audit an inherited system, or to see what a third-party theme leaves undefined. Only NAMED tokens are read; a bare hex inside a rule carries no role and is never imported as one (use audit_design_system to count those). JavaScript configs are never evaluated.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Token set / brand name for the emitted artifacts (default 'Imported') | |
| format | No | Format to re-emit in (default 'all'). Use 'swiftui' or 'compose' to take a web theme to native. | |
| source | Yes | The token source: CSS custom properties, DTCG JSON, or a theme object as JSON. The format is detected automatically. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and non-destructive, and the description adds meaningful behavioral context beyond those: only NAMED tokens are imported, bare hex values are ignored, and JavaScript configs are never evaluated. It also discloses the output contents — named roles, missing semantic roles, WCAG contrast check, and re-emitted artifacts — so an agent understands what will happen when invoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact but information-dense; every sentence earns its place. It front-loads the core action and outputs, then covers use cases, exclusions, and limitations without tangents. The semicolon-separated list of input formats keeps the structure readable despite covering many formats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with no output schema, the description explains what the tool returns: named roles, missing semantic roles, WCAG contrast checks, and re-emitted code in multiple formats. It also covers accepted input formats, automatic detection, unsupported inputs, and alternative tools. For a tool of this complexity, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra meaning by explaining that the source format is detected automatically, clarifying that 'swiftui' or 'compose' is how to take a web theme to native, and framing what kinds of sources are acceptable. This enriches the schema definitions without merely repeating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read an EXISTING design system and convert it,' then enumerates accepted input formats and outputs. It further differentiates itself from siblings by calling itself the inverse of generate_design_tokens and distinguishing its scope from audit_design_system, so an agent can tell exactly what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is explicit: use it to port a web theme to iOS/Android, audit an inherited system, or discover undefined roles in a third-party theme. It also names alternatives for adjacent cases, e.g., bare hex values should go to audit_design_system, and states that JavaScript configs are never evaluated. This is strong when-to-use versus when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_freshnessKnowledge FreshnessARead-onlyIdempotent
Report how fresh each knowledge document is (age since last verification vs its category's staleness threshold). Use this to decide which docs need re-research; refresh workflow is documented in the repo's /refresh-knowledge command.
| Name | Required | Description | Default |
|---|---|---|---|
| only_stale | No | Return only docs past their staleness threshold (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description aligns with annotations (readOnlyHint=true, idempotentHint=true, destructiveHint=false) and adds useful context: it reports on freshness across documents, uses category-specific thresholds, and implicitly performs no mutation. The pointer to the refresh workflow further clarifies the boundary between reporting and acting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences deliver the core function and the decision context without filler. The freshness metric is front-loaded, and the workflow pointer is a valuable addition that does not bloat the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only reporting tool with one optional boolean and no output schema, the description adequately conveys what output is expected (a freshness report per document) and how the result should be used. Annotations cover the safety profile, and the refresh workflow pointer addresses the natural follow-up action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the only only_stale boolean parameter, so the schema fully documents its meaning and default. The description does not add parameter-level detail, but the baseline of 3 is appropriate given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and a well-defined resource ('knowledge documents') with a precise metric: age since last verification compared to each category's staleness threshold. It is clearly distinct from sibling tools, which all focus on design knowledge retrieval or audits rather than freshness reporting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the tool's intended use ('decide which docs need re-research') and directs the agent to the refresh workflow via /refresh-knowledge. It does not explicitly name alternatives or exclusion conditions, but no sibling tool appears to offer freshness reporting, making the context clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_design_knowledgeList Design KnowledgeARead-onlyIdempotent
List the knowledge-base index (design languages, UI components, UX, craft, books, process, marketing, SEO, GEO, patterns). Returns every document grouped by category — each with its id, title, platform, and tags. Use this first to discover what's available and get exact ids; then read one with get_design_doc, or search by need with search_design_knowledge.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Filter to one category, e.g. 'component', 'ux', 'marketing'. Omit for all. | |
| platform | No | Filter to one platform: 'mobile', 'web', or 'macos' (docs marked 'both' are always included). Omit for all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds behavioral context by stating that results are grouped by category and include id, title, platform, and tags, which is valuable beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the core purpose, and avoids redundancy with the schema. Every clause adds useful information: scope, return composition, and workflow guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by naming the exact fields returned (id, title, platform, tags) and the grouping behavior. Filters are fully documented in the input schema, and annotations cover the operational constraints, making this complete for a list/discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both category and platform already have clear descriptions and enums. The tool description does not add much parameter-level meaning beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List the knowledge-base index' and enumerates the content domains. It clearly distinguishes itself from siblings by explaining this is the discovery entry point, while read is get_design_doc and search is search_design_knowledge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance is given: 'Use this first to discover what's available and get exact ids; then read one with get_design_doc, or search by need with search_design_knowledge.' This tells an agent exactly when to call this tool versus its closest alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
measure_screenshotMeasure ScreenshotARead-onlyIdempotent
Measure a real screenshot from its actual pixels — the exact palette and how many distinct colours it really uses, true WCAG contrast ratios for the colour pairs on screen, whitespace/density, and structural detections (left-edge alignment, vertical rhythm, off-grid gaps) each carrying a confidence level. Returns a markdown measurement and, on request, a self-contained HTML report you can save, open and share. PNG only. Reads the local file you name; makes no network call. Use it before critiquing a UI so the review cites measured numbers instead of impressions; pair with fix_contrast for failing pairs and audit_design_system for the codebase behind the screen.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path to a .png screenshot. Absolute paths are strongly preferred — a relative path is resolved against the server's working directory, which is usually not your project folder. | |
| scale | No | Device pixel ratio of the screenshot (default 1). Pass 2 for a Retina/2× capture so lengths are reported in logical px instead of image px. | |
| format | No | 'markdown' (default) for the measurement text, 'html' for a self-contained report document to save and open, 'both' for each. | |
| max_colors | No | How many palette clusters to list (default 12). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already supply readOnlyHint=true, idempotentHint=true and destructiveHint=false; the description complements these by disclosing operational behavior: 'Reads the local file you name; makes no network call' (scope/privacy), 'PNG only' (input restriction), and the markdown/HTML output modes. The description is fully consistent with the annotations — no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: measurement capabilities, output formats, constraints/privacy, then workflow routing. The most decision-relevant information (what it measures, PNG-only, no network) is front-loaded, with zero tautology or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema, the description covers the essentials: measured dimensions (with confidence levels), returned artifacts (markdown, optional HTML), input constraints (PNG, local path), and workflow placement ('before critiquing a UI'). The return format is described narratively rather than structurally, and failure behavior for invalid paths is left unspecified — minor gaps given the strong schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter is already richly documented — path (absolute paths preferred, working-directory caveat), scale (Retina/2× pixel-ratio logic), format (markdown/html/both semantics), max_colors (cluster count, default 12). The description narratively reinforces the format behavior ('on request, a self-contained HTML report') but adds no genuinely new parameter meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource — 'Measure a real screenshot from its actual pixels' — and enumerates the exact outputs: palette, distinct colour count, WCAG contrast ratios, whitespace/density, and structural detections with confidence levels. No sibling tool measures screenshots; fix_contrast and audit_design_system are explicitly positioned as companions, not competitors, so the definition fully disambiguates the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to invoke it: 'Use it before critiquing a UI so the review cites measured numbers instead of impressions.' It also routes to the right companion tools — fix_contrast for failing pairs and audit_design_system for the codebase behind the screen — and states hard constraints (PNG only, local file, no network call) that shape when it applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_design_knowledgeSearch Design KnowledgeARead-onlyIdempotent
Search the whole knowledge base with a natural-language query — UI components, UX, accessibility, typography, color, motion, conversion, copywriting, SEO/GEO, platform design languages, craft standards, distilled design & marketing books, roadmaps, and real-world app/site patterns. Returns the top-matching documents, each with its single most relevant section excerpted and its id. Use for open-ended 'how should I…' questions; if you already know the id use get_design_doc, to browse everything use list_design_knowledge.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 5) | |
| query | Yes | What you need guidance on, e.g. 'primary button size mobile', 'pricing page layout', 'dark mode colors', 'llms.txt' | |
| category | No | Restrict to one category | |
| platform | No | Restrict to one platform |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior. The description adds useful behavioral detail beyond annotations: it searches the entire knowledge base, returns top-matching documents, and includes the single most relevant section excerpt plus document id. This is meaningful context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and resource, then covers what is returned and when to use it. The enumeration of knowledge domains is long but informative for an agent deciding whether this tool matches the task; it earns its place even if it could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 4 well-documented parameters and no output schema, the description provides the essential return shape (documents with excerpt and id) and usage context. It distinguishes the tool from the most relevant siblings, which is important given the large sibling list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that query should be a natural-language phrase and gives examples of open-ended questions, but it does not need to add much because the schema already documents limit, category, and platform clearly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search the whole knowledge base with a natural-language query' and enumerates the covered domains. It clearly differentiates from siblings by stating it returns top-matching documents with an excerpt and id, as opposed to fetching a known doc or listing everything.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the intended use case ('open-ended how should I… questions') and names alternatives: use get_design_doc when the id is already known, and list_design_knowledge to browse everything. This gives the agent actionable routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seo_geo_guideSeo Geo GuideARead-onlyIdempotent
SEO and GEO expertise for websites — classic SEO (technical, on-page, design-impact) and GEO, Generative Engine Optimization for AI answer engines (ChatGPT, Perplexity, Google AI Overviews, llms.txt, citations). Returns the full relevant guide docs, optionally narrowed to a topic. Use when planning or auditing a site's discoverability; pair with get_design_roadmap('website') for the full process.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | Yes | Which discipline: 'seo' (classic search), 'geo' (AI answer engines), or 'both'. | |
| topic | No | Optional narrower topic, e.g. 'core web vitals', 'llms.txt', 'structured data'. Omit to get the full guides. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the tool returns documentation and can be narrowed by topic, but it does not disclose any additional behavioral details such as output formatting, size limits, or whether external data is fetched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and organized: domain definition, return behavior, and usage guidance. The first sentence is somewhat long but contains useful disambiguation between SEO and GEO. Every sentence contributes, though the lead could be more action-forward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only reference tool with a required enum and one optional topic parameter, the description covers what the tool returns, when to use it, and how to combine it with a related tool. It does not explain what 'full relevant guide docs' looks like in practice, but with no output schema and low complexity, the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema. The description reinforces the optional narrowing behavior for topic and gives domain examples, but it does not add substantial meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens by naming the resource ('SEO and GEO expertise for websites') and specifies the verb/resource relationship: it returns full relevant guide docs, optionally narrowed by topic. This clearly distinguishes it from audit-style siblings like audit_seo_geo and design guidance tools by framing itself as a reference/knowledge retrieval tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: 'Use when planning or auditing a site's discoverability' and even suggests pairing with get_design_roadmap('website') for the full process. It does not explicitly list exclusions or alternative tools, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_font_pairingSuggest Font PairingARead-onlyIdempotent
Recommend production-ready font pairings for a brand/product from an intent or vibe (e.g. 'modern SaaS dashboard', 'luxury editorial', 'bold marketing landing', 'native iOS app', 'developer tool'). Returns matched heading + body (+ mono) with ready-to-paste CSS stacks, weights, source, the reason each pairing works, pairing rules, and a suggested type scale. Deterministic curated recommendations, not generic advice. Pair with generate_design_tokens to emit the fonts as tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many pairings to return (default 3) | |
| intent | Yes | The product/brand vibe or use case, e.g. 'trustworthy fintech dashboard', 'playful consumer app', 'minimal portfolio', 'AI developer product' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the read-only, idempotent, non-destructive profile. The description adds valuable behavioral context beyond annotations: deterministic curated recommendations rather than generic advice, and what the full result contains including CSS stacks, weights, source, reasoning, pairing rules, and type scale. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then details return value components, then adds the deterministic behavior note and a sibling integration tip. Every sentence earns its place with no filler or redundant restating of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description sufficiently enumerates what will be returned so an agent knows what to expect. It covers the input signal, output artifacts, behavioral guarantees, and the natural follow-on tool, making the definition self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both intent and limit already documented, so the schema carries most of the parameter meaning. The description reinforces the intent-as-vibe concept with examples but adds little new parameter-level detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: 'Recommend production-ready font pairings' from an intent/vibe, with concrete examples. The detailed output composition (heading + body + mono CSS stacks) makes it clearly distinct from related siblings like generate_type_scale or generate_design_tokens.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to call it: when the user describes a brand/product vibe and wants production font pairings. It also names a complementary next step (generate_design_tokens), though it does not explicitly state when not to use the tool or which alternatives to prefer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_icon_librarySuggest Icon LibraryARead-onlyIdempotent
Recommend the right icon library for a product from an intent/vibe/platform (e.g. 'minimal SaaS dashboard', 'friendly consumer app with personality', 'iOS app', 'Android Material app', 'dense admin panel'). Returns matched open-source (or platform-native) icon systems with license, install command, coverage, the reason each fits, usage rules, and universal icon best-practices. Deterministic curated guidance — icons are NOT bundled; install the chosen library in your own project. Pair with suggest_font_pairing and generate_color_system.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many libraries to return (default 3) | |
| intent | Yes | Product vibe / platform / use case, e.g. 'clean developer tool', 'premium fintech app', 'iOS native app', 'Material 3 Android app', 'data-dense dashboard' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful context beyond the annotations: 'Deterministic curated guidance' and the critical caveat 'icons are NOT bundled; install the chosen library in your own project'. The return-content list (license, install command, coverage) also clarifies what the agent can expect. Annotations already cover read-only/idempotent/non-destructive, so the description's additions are the right kind of value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with purpose, then return shape, then the key behavioral caveat, then composition guidance. Each sentence earns its place, though 'universal icon best-practices' is slightly verbose phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating return fields (license, install command, coverage, fit rationale, usage rules, best-practices). Combined with annotations covering the safety profile and a simple 2-param schema, almost everything an agent needs is present. Minor gaps: no mention of result ordering or no-match behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters have rich descriptions already. The tool description's intent examples ('minimal SaaS dashboard', 'iOS app') overlap heavily with the schema's own examples, so it adds little beyond it. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Recommend'), resource ('icon library'), and the input framing ('from an intent/vibe/platform') with concrete examples. The icon-specific focus clearly differentiates it from siblings like suggest_font_pairing and generate_color_system.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context via 'Pair with suggest_font_pairing and generate_color_system', explicitly naming the natural companion tools for a design-system flow. It lacks explicit when-not-to-use exclusions, but the intent/vibe input framing effectively implies the triggering scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
33 tool updates
v0.28.0- Added
audit_accessibility - Added
audit_android_ui - Added
audit_apple_ui - Added
audit_design_system - Added
audit_ethical_design - Added
audit_generic_design - Added
audit_performance - Added
audit_project - Added
audit_security - Added
audit_seo_geo - Added
audit_ux_copy - Added
compare_design_languages - Added
create_design_system - Added
design_lint - Added
fix_contrast - Added
generate_color_system - Added
generate_design_tokens - Added
generate_elevation_system - Added
generate_layout_system - Added
generate_motion - Added
generate_type_scale - Changed
get_component_guidance2 fields changed- changed
Input schema / properties / component / descriptionPrevious value: -"Component or pattern name, e.g. 'primary button', 'signup form', 'bottom tab bar', 'hero section', 'paywall'"New value: +"Component or pattern name, e.g. 'primary button', 'signup form', 'bottom tab bar', 'hero section', 'paywall'." - changed
Input schema / properties / platform / descriptionPrevious value: -"Target platform — strongly recommended"New value: +"Target platform ('mobile' | 'web' | 'macos') — strongly recommended so guidance matches the platform's conventions."
- Added
get_component_recipe - Changed
get_design_doc1 field changed- changed
Input schema / properties / id / descriptionPrevious value: -"Document id, e.g. 'buttons', 'material-3', 'geo-tactics-checklist'"New value: +"Exact document id, e.g. 'buttons', 'material-3', 'accessibility', 'geo-tactics-checklist'. Get ids from list_design_knowledge or search results."
- Changed
get_design_examples2 fields changed- changed
Input schema / properties / platform / descriptionPrevious value: -"'mobile' for iOS app screens, 'web' for website examples"New value: +"'mobile' for iOS and Android app screens, 'ios' or 'android' to pin one, 'web' for website examples" - changed
Input schema / properties / platform / enumPrevious value: -[ - "mobile", - "web" -]New value: +[ + "mobile", + "ios", + "android", + "web" +]
- Changed
get_design_language2 fields changed- changed
Input schema / properties / language / descriptionPrevious value: -"Which design language / platform reference to fetch"New value: +"Which reference to fetch. e.g. 'material-3' (Android/Material), 'apple-hig-liquid-glass' or 'ios-app-design' (iOS), 'macos-app-design', 'visionos-spatial-design' (Vision Pro), 'web-trends-2026', 'design-tokens-theming'." - changed
Input schema / properties / language / enumPrevious value: -[ - "material-3", - "apple-hig-liquid-glass", - "ios-app-design", - "android-app-design", - "macos-app-design", - "fluent-2", - "web-trends-2026", - "design-tokens-theming" -]New value: +[ + "material-3", + "apple-hig-liquid-glass", + "ios-app-design", + "android-app-design", + "macos-app-design", + "apple-intelligence-design", + "visionos-spatial-design", + "wwdc-design-principles", + "fluent-2", + "web-trends-2026", + "design-tokens-theming" +]
- Added
import_design_tokens - Changed
list_design_knowledge3 fields changed- changed
Input schema / properties / category / descriptionPrevious value: -"Filter by category"New value: +"Filter to one category, e.g. 'component', 'ux', 'marketing'. Omit for all." - changed
Input schema / properties / category / enumPrevious value: -[ - "design-language", - "component", - "ux", - "seo", - "geo", - "pattern", - "craft", - "book", - "process" -]New value: +[ + "design-language", + "component", + "ux", + "seo", + "geo", + "pattern", + "craft", + "book", + "process", + "marketing", + "security" +] - changed
Input schema / properties / platform / descriptionPrevious value: -"Filter by platform (docs marked 'both' always included)"New value: +"Filter to one platform: 'mobile', 'web', or 'macos' (docs marked 'both' are always included). Omit for all."
- Added
measure_screenshot - Changed
search_design_knowledge1 field changed- changed
Input schema / properties / category / enumPrevious value: -[ - "design-language", - "component", - "ux", - "seo", - "geo", - "pattern", - "craft", - "book", - "process" -]New value: +[ + "design-language", + "component", + "ux", + "seo", + "geo", + "pattern", + "craft", + "book", + "process", + "marketing", + "security" +]
- Changed
seo_geo_guide2 fields changed- changed
Input schema / properties / scope / descriptionPrevious value: -"Which discipline"New value: +"Which discipline: 'seo' (classic search), 'geo' (AI answer engines), or 'both'." - changed
Input schema / properties / topic / descriptionPrevious value: -"Optional narrower topic, e.g. 'core web vitals', 'llms.txt', 'structured data'"New value: +"Optional narrower topic, e.g. 'core web vitals', 'llms.txt', 'structured data'. Omit to get the full guides."
- Added
suggest_font_pairing - Added
suggest_icon_library
10 tool updates
v0.3.2- First observed
design_review_checklist - First observed
get_component_guidance - First observed
get_design_doc - First observed
get_design_examples - First observed
get_design_language - First observed
get_design_roadmap - First observed
knowledge_freshness - First observed
list_design_knowledge - First observed
search_design_knowledge - First observed
seo_geo_guide
TDQS
The set is organized into clear families (knowledge retrieval, generators, audits) and descriptions cross-reference each other heavily, which helps an agent pick between them. The main confusable pair is audit_design_system vs audit_project, which both report color/shadow/spacing sprawl, and get_design_doc vs get_design_language overlap slightly as document fetchers.
The vast majority follow a predictable verb_noun pattern across list_/search_/get_, generate_, audit_, suggest_, create_, compare_, measure_, and import_ families. A few outliers—design_review_checklist, seo_geo_guide, knowledge_freshness, and design_lint—break the pattern but are still readable.
36 tools is far above the 25+ threshold and creates a heavy selection burden for an agent, even though the scope is broad. Several audit_* and generate_* tools could plausibly be consolidated, with audit_project largely subsuming design_lint plus audit_design_system's consistency counting.
The surface covers the full design workflow: knowledge discovery and planning, component guidance and recipes, deterministic design-token/color/type/motion generation, import of existing systems, screenshot measurement, and a wide matrix of audits across accessibility, copy, ethical design, security, SEO, performance, Apple, and Android. There are no obvious dead ends or missing lifecycle stages for the stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
UI design from prompts, screenshots, and URLs for AI coding agents and theme tokens.
Growth marketing, SEO and GEO as agent tools: ranked moves, ship them, AI answer visibility.
Distribution memory and marketing context for coding agents, with SEO, content, media, and analytics
SEO & marketing toolkit for AI agents: GA4, Search Console, AdSense, GTM, PageSpeed, Trends.
Related MCP Servers
- FlicenseAqualityDmaintenanceProvides comprehensive design principles and best practices to help LLMs generate modern, accessible web pages through guidance on layouts, colors, and typography. It enables users to review design approaches and access expert recommendations for responsive design, component structure, and current industry trends.12163-
- AlicenseNot gradedqualityDmaintenanceGive your AI coding agent design taste. 104 curated design seeds with colors, fonts, spacing, and shadows. Query by vibe, brand, or style.43MIT
- FlicenseNot gradedqualityDmaintenanceExposes The Vibe Coder's Web Design Guide as tools for AI agents, enabling lookups of UI design patterns, CSS/JS snippets, and composition of optimized front-end prompts.-
- AlicenseNot gradedqualityAmaintenanceSTOP UI SLOP. Gives coding agents searchable evidence from 800,000+ real web and iOS screens, design contracts, and a hard UI finish gate.1MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/HalidSaglam/saglitzdesign-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server