melta-ds-mcp
Provides tools for enforcing a Tailwind CSS-based design system: retrieving design tokens, component contracts, and rules, and linting generated HTML/JSX class usage against prohibited Tailwind patterns via check_rule and check_html.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@melta-ds-mcpcheck this pricing card HTML for design rule violations"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
melta UI
AI 向けデザインガイドラインを、違反を止める実行可能な契約へ。
🇬🇧 English: README.en.md · Showcase: https://melta.tsubotax.com (ドキュメント・契約の正本はこのリポジトリ)
AI にガイドラインを読ませることはできる。守るかどうかは AI 任せになる。melta UI は、その「任せ」を機械に置き換える。生成の前(MCP で契約を参照させる)・直後(lint / hook が違反を突き返す)・マージ前(CI が止める)・その後(drift 検査がドキュメントと実装の腐りを検知し続ける)の 4 点で機械が関与する。読ませるだけでなく、守らせる。
境界: melta UI は完成済みの CSS コンポーネント集ではない。配るのは値(tokens)・規則(rules)・仕様(contracts)・検証器(lint / MCP)で、import して貼れば動く UI ライブラリではない。web の実装は HTML + Tailwind クラスの参照実装として同梱している。
誰のためのものか
向いている
AI コーディングエージェント(Claude Code / Cursor / Codex 等)で UI を生成・運用しているチーム
「ガイドラインは書いたのに守られない」を仕組みで潰したい個人・少人数チーム
web と React Native で 1 つのデザイン契約を共有したいプロダクト
向いていない
完成済みの Web コンポーネント集(React コンポーネントを install してすぐ使いたい)が欲しい場合
Tailwind / class ベース以外のスタイリング(CSS-in-JS の props 経由など)が主体で、スタイルがマークアップに現れない場合。静的 lint が効かない
Related MCP server: web-stylebook-mcp
Proof — 主張には検証経路をつける
107 禁止ルールのうち 49 / 107 を静的に自動検出。残りも「なぜ自動検出しないか」を
automationStatusで分類・可視化する(rules.json / 内訳は制約と正直な範囲)Playwright + axe-core 372 tests が CI 必須ゲート(.github/workflows/design-check.yml / 実行履歴)
5 種の代表的リセット CSS 環境で VRT 差分ピクセル 0。pixelmatch の literal 比較で機械検証(tests/reset-vrt.spec.ts、
npm run test:reset-vrt)npm 3 パッケージ + MCP Registry で配布(melta-contracts / melta-ds-mcp / melta-app、Registry ID
io.github.tsubotax/melta-ui)別リポジトリの React Native 実装が同じ契約を購読し、契約の破壊的変更は APP 側が契約バージョンを取り込んだ時点で consumer テストが検出する(melta-app / npm 公開版との互換は
npm run design:compatが publish 前に検査)外部プロジェクトで「AI が違反を書く → 即検出 → 自己修正」ループを実測(2026-08、非公開 RN アプリへ npm 経由で導入。melta-app README の成熟度・メンテナンス節)
drift 検査自身に負のテストがある(わざと壊して発火することを固定:tests/drift-heal.spec.ts)
配布物 — いま install できるもの
パッケージ | 役割 | 使い方 |
契約データ(tokens / rules / component contracts / recipes の JSON)。ビルド不要・フレームワーク非依存 |
| |
MCP サーバー + lint エンジン(このリポジトリ)。 |
| |
React Native 実装。消費者プロジェクト向け eslint plugin を同梱 |
|
melta-ds-mcp自体の bare import(import "melta-ds-mcp")は非サポート。entry は import しただけで stdio サーバーが起動する CLI なので、npx melta-ds-mcpか subpath 経由で使う。entry 規約・deep import 互換・パッケージ分割の予定は docs/distribution.md。
自分のデザインシステムで検査したい場合(BYO-DS):
melta-ds-mcpは起動時に読み込むアセット root を自分の DS bundle へ切り替えられる(melta のルールとは混在しない)。4 ファイルの最小構成から始める手順と限界は docs/distribution.md の BYO-DS 節。
前提条件・互換性
項目 | 値 |
Node | 22 以上(CI は 22 で検証) |
MCP クライアント | stdio MCP に対応したもの(Claude Code で検証。Cursor は同梱の |
スタイリング | Tailwind CSS の class ベース前提。静的 lint は class 属性 / HTML 属性 / DOM 構造を読む |
生成物の表示 | プロトタイプは Tailwind CDN + |
JSX / Vue | class 属性と HTML 属性の lint は効く。composition lint(ネスト構造・a11y DOM)は HTML のみ。JSX の変数経由 class・spread は静的には追えない |
ライセンス | MIT |
5 分クイックスタート
経路 A — npm(MCP サーバー、推奨)
clone せずに、契約参照と自己検証だけを既存プロジェクトへ足す経路。
claude mcp add melta-ui -- npx -y melta-ds-mcp
claude mcp list成功判定 — claude mcp list にこの行が出る:
melta-ui: npx -y melta-ds-mcp - ✔ Connected接続時に MCP instructions が渡るので、「melta は完成 CSS ライブラリではない」「先に melta://design-constitution を読む」「生成後は check_html で自己検証する」を利用側が毎回プロンプトに書く必要はない。あとは UI を指示するだけ:
ユーザー一覧のテーブルを作って
成功判定 — AI が生成 HTML を check_html に通し、この形の応答を得る(違反があれば修正して再検証する):
{
"passed": false,
"errorCount": 2,
"warnCount": 0,
"violations": [
{ "ruleId": "AI_NO_CARD_COLOR_BAR_TOP", "severity": "error", "token": "border-t-4",
"reason": "AI生成UIの典型パターン。装飾過剰で汎用性が低い",
"alternative": "border border-slate-200 のみでカードを構成" },
{ "ruleId": "COLOR_NO_BLUE_BG", "severity": "error", "token": "bg-blue-500",
"reason": "primaryで統一する", "alternative": "bg-primary-*" }
],
"coverage": { "automated": "...", "notAutomated": "..." }
}生成された HTML をブラウザで表示するには Tailwind と melta のトークン設定が要る。プロトタイプなら CDN でよい:
<script src="https://cdn.tailwindcss.com"></script>
<script>
// DESIGN.md「Quick Reference → HTML テンプレート」の tailwind.config をそのまま貼る。
// fontSize は 8 段すべて Tailwind デフォルトと異なる(本文 18px / 行間 2.0 が melta の核)。
</script>経路 B — clone(フルハーネス)
hook / CI / lint CLI まで含めた強制層が要る場合。npm install した消費者にはこの 3 層は届かない(制約と正直な範囲)。
git clone https://github.com/tsubotax/melta-ui.git
cd melta-ui && npm install
printf '<div class="text-black shadow-2xl">x</div>' > /tmp/melta-bad.html
npm run design:lint-generated -- /tmp/melta-bad.htmlnpm install で有効になるもの: .mcp.json(Claude Code へ MCP 自動接続)/ .cursor/mcp.json(Cursor 向けに同じ MCP サーバーの設定を同梱。有効化は Cursor 側の操作に従う。作業指示は AGENTS.md を読ませ、.cursor/rules/melta-ui.mdc は所在ポインタだけを置く)/ .claude/settings.json の PostToolUse hook / lint CLI。
成功判定 1 — 違反ファイルに lint CLI をかけると exit 1 で落ちる:
✗ [error] COLOR_NO_TEXT_BLACK: "text-black" → text-slate-900(純黒はコントラストが強すぎて長時間の利用で目が疲れる)
✗ [error] SPACE_NO_SHADOW_2XL: "shadow-2xl" → shadow-sm 〜 shadow-md(オーバーレイ: shadow-xl)(影が強すぎてノイズになる)
1 ファイル走査 / error 2 / warn 0
❌ FAILED成功判定 2 — Claude Code が .html / .tsx / .jsx / .vue を Write / Edit した直後、hook がこの JSON を返して修正ループに乗せる(warn のみなら additionalContext で助言注入):
{"decision":"block","reason":"melta UI 禁止パターン検出(error 2 / warn 0)。書き込まれたファイルを修正してください: ..."}仕組み — 契約・参照・検証・監視の 4 層
① 契約(SSOT) design/contracts/
tokens.json 101 デザイントークン
rules.json 107 禁止ルール(ID + severity + detector + alternative)
components/ 40 contract(web 28 / app 先行 12)
recipes/ プラットフォーム具象(web: 生成ミラー / app: RN styleRefs)
DESIGN.md / AGENTS.md AI が最初に読む憲法と作業ガイド
② 参照(生成の前) MCP サーバー(melta-ds-mcp)
必要な仕様・値・ルールだけをオンデマンドで渡す
③ 検証(生成の直後〜マージ前)
PostToolUse hook Write/Edit 直後に lint → error は block で自動修正
lint CLI / CI .github/workflows/design-check.yml
MCP check_html CI と同一ロジックの自己検証
④ 監視(その後) design:drift ドキュメント ↔ contracts の腐りを検知
design:compat npm 公開版との破壊的変更 × semver 検査
design:drift-heal drift を検出して derived のみ再生成(SSOT は human gate)MCP が公開するツール:
ツール | 説明 | 入力例 |
| トークン検索 |
|
| コンポーネント仕様取得(variants / sizes / stateSpecs / anatomy / a11y) |
|
| クラス文字列の禁止パターン検査(34パターン自動検出)。文脈依存は conditional 付き |
|
| 生成 HTML / JSX 全体を CI / hook と同一ロジックで lint |
|
| 107 禁止ルール参照(manual 含む全件、filter 対応) |
|
| 全文検索(最大 20 件 + truncated 通知) |
|
Resource は melta://design-constitution(DESIGN.md 全文)/ melta://tokens / melta://components / melta://components/{id} / melta://rules / melta://rules/auto-detectable。
web の実装対象は 28 コンポーネント + 13 ファウンデーション + 5 パターン。設計原則は Content First / WCAG 2.1 AA / Semantic Color / 3-Color Rule / 4px Grid / Minimal Elevation / No AI-ish Decoration の 7 つ(DESIGN.md)。
Web と APP — 1 つの契約が両方に降りる
同じ契約パッケージ(melta-contracts)を web(このリポジトリ / HTML + Tailwind)と APP(melta-app / React Native)の両実装が購読する。トークンを各実装にコピーして持つ経路は存在しない(二重化の物理防止)。
契約は規範と具象の 2 層。規範(components/*.contract.json)は variant の語彙・states・tokenRefs・a11y で、全プラットフォーム共通。分岐が正当な箇所(hover→pressed、elevation の表現差、タッチターゲット 44pt 等)は platformSemantics で意味論だけを宣言する。具象(recipes/)は web が契約の Tailwind からの導出ミラー(鮮度を CI が担保)、app が RN の styleRefs(色は 100% token 参照)を手書きする authoring source。
守らせる仕組みも双方向:
web 側 → 互換ゲート(
npm run design:compat): npm 公開版と HEAD の golden diff。token 削除・variant 削除・rule の意味変更を breaking 分類し、semver bump を機械強制するAPP 側 → consumer テスト: melta-app の CI が「契約 subset・token 実在・contractVersion 同期」を照合する。web 側が契約を壊すと APP のテストが赤くなる
melta-app は消費者プロジェクト向けの eslint plugin も npm で配っており、使う側のコードで生値の直書きが止まる。RN カタログの live showcase は https://app.melta.tsubotax.com。
制約と正直な範囲
49 / 107 の意味。「107 禁止ルールを強制する」とは言えない。静的に自動検出できるのは 49 件で、残りは検証経路を automationStatus で分類して可視化している(宣言だけのルールをゼロにするための棚卸し)。
経路 | 件数 | 内容 |
静的自動検証 | 49 / 107 | class マッチ 34(MCP |
interaction test | 3 |
|
静的検出 不能 | 3(うち error 3) |
|
LLM 審査候補 | 43(うち error 31) |
|
human-only | 9(うち error 9) | 人間レビューでのみ守る。 |
未分類 | 0(うち error 0) | 棚卸し未了(automationStatus 未宣言) |
この表は npm run design:coverage が contracts から生成し、鮮度を npm run design:drift が守る。数字は改善のたびに動く。各ルールの状態の SSOT は rules.json の automationStatus。
その他の制約:
clone 経路と npm 経路で届く層が違う。PostToolUse hook / CI / lint CLI は「このリポジトリを clone して使う」前提の層で、
npm installした消費者には届かない。npm 経路の強制層はmelta-ds-mcp/lint-core(class / html-attr lint のみ。composition lint は含まない)と MCP のcheck_html(composition 込み)の 2 つで、これを各プロジェクトのフック / CI に自前で組み込むclass ベースでないスタイリングは検査できない。スタイルがマークアップ(class / 属性)に現れないコードでは静的 lint が空振りする
JSX の composition lint は未対応。ネスト構造・a11y DOM の検査は HTML のみ。JSX は class / 属性 lint まで
BYO-DS はデータ差し替えのみ。自分の DS のトークン・ルールを JSON で書けば同じエンジンで検査できるが、component metadata の生成ツールは npm に未同梱で、detector は 6 種固定(
manualは自動検査しない宣言)。新しい検査ロジックはエンジン側の変更が要るcheck_html.passedは完成承認ではない。lint-clean draft であってブランド適合の判定ではなく、最終判断は人間に渡す
セキュリティ・データ境界
MCP サーバーも lint エンジンもローカルプロセスで完結する。生成コード・プロンプト・検査結果を外部へ送信する経路はなく、telemetry も持たない。ネットワークに出るのは npx によるパッケージ取得と、npm run design:compat / npm run check:pack が npm registry の公開バージョンを照会するときだけ。
成熟度とメンテナンス
個人メンテナンスの OSS(tsubotax)。SLA も専任チームもない。実運用の dogfood(web showcase / RN アプリ)で回している
0.x / 1.x の方針: 契約パッケージ
melta-contractsは 0.x で、破壊的変更は minor bump で入りうる。ただし破壊的変更の分類は人手ではなくnpm run design:compatの機械判定で、semver bump を強制する変更の通知経路は CHANGELOG.md。リリースごとに Added / Changed / Removed を残す
バグ・要望は GitHub Issues へ。手を動かすなら CONTRIBUTING.md(ローカル CI ミラー・生成物の扱い・PR 前チェックリスト)
脆弱性は issue ではなく SECURITY.md の非公開経路(GitHub private vulnerability reporting)へ。セキュリティ修正は各パッケージの最新 minor にのみ提供する
もっと知る
ドキュメント | 内容 |
デザイン憲法 + Quick Reference。これだけで基本 UI を生成できる | |
AI エージェント共通の作業ガイド(読み込みモード・タスク別ガイド・npm scripts) | |
SSOT 宣言と値競合時の優先順位 | |
loop / pipeline 自動化の統治原則(自動化 3 Level 分類・SSOT write-protect・Human Gate の Hard / Soft 2 層化・監査ログ)。現状 W2 drift repair が稼働 | |
ベンチマークのプロトコル(5 条件 × N トライアルで DS 準拠スコアの lift を測る)と既知の限界 | |
npm entry 規約・deep import 互換・BYO-DS(自分の DS を持ち込む)・パッケージ分割の予定 | |
AI-Ready 成熟度モデル(Lv0 None → Lv4 Verified)。任意のプロジェクトに当てられる | |
UI 変更 PR の Before/After を実キャプチャで比較する Claude Code plugin(別リポジトリ)。このリポジトリの | |
Google Labs design.md spec との対応表。melta の |
License
MIT License — LICENSE。同梱アイコンのライセンスは THIRD_PARTY_LICENSES.md を参照。
Acknowledgments: Charcoal Icons(pixiv Inc., Apache License 2.0)/ Lucide Icons(ISC License)/ Tailwind CSS
Available Tools
6 toolscheck_htmlA
Lint a full HTML/JSX source against melta-ui rules — the same checks as CI and the PostToolUse hook (class rules + html-attr rules + composition rules for HTML). Use this AFTER generating UI code to self-verify before presenting it. Response always includes coverage info (manual rules cannot be auto-checked).
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Full source code to lint (HTML / JSX / Vue template) | |
| sourceType | No | Source type. "html" (default) also runs composition lint (nested modal etc.); "jsx" runs class + attr lint only |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses CI/PostToolUse parity, the exact rule categories run, the html-vs-jsx behavioral split, and that coverage info is always returned. It omits failure modes or output detail beyond coverage, but for a pure lint operation the disclosed context is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all load-bearing: scope, when-to-use, and return-value caveat. The self-verification instruction is front-loaded and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by stating the response always includes coverage info and that manual rules cannot be auto-checked. Timing, scope, and return expectations are all present; only deeper output structure is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sourceType enum is already described in the schema. The description's mention of composition rules for HTML corroborates but does not extend the sourceType semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lint), resource (full HTML/JSX source), and the rule system it validates against (melta-ui rules). It further scopes what is checked (class rules, html-attr rules, composition), which separates it from siblings like check_rule and get_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this AFTER generating UI code to self-verify before presenting it" gives explicit timing and intent, which is real guidance. It does not name an alternative tool or state when not to use it (e.g., single-rule checks via check_rule), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_ruleB
Check class strings against this design system's prohibition rules. Returns violations with reasons and alternatives.
| Name | Required | Description | Default |
|---|---|---|---|
| classes | Yes | Space-separated class string to check (e.g. "text-black shadow-2xl bg-green-500") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the return payload ('violations with reasons and alternatives'), which is the most useful behavioral fact, but it never states that the call is non-mutating, requires no auth, or what an empty result looks like. 'Check' implies a read, but that is inference rather than disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the action is front-loaded ahead of the return-value note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only checker with no output schema, the description is nearly complete: it covers what is checked and what comes back. It could be fuller by noting the clean-result case or its relationship to check_html, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented in the schema with a concrete example. The description adds no extra meaning about the classes string (e.g. format constraints or whether multiple strings are accepted), so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (check) and resource (class strings against this design system's prohibition rules), which is enough to separate it from get_rules or get_token. It does not explicitly differentiate itself from the sibling check_html, so an agent must infer that one takes class strings and the other markup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus check_html, get_rules, or get_token, and no prerequisites such as running it before shipping markup. Usage is only implied by the phrase 'this design system's prohibition rules'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_componentA
Get component metadata including variants, sizes, accessibility requirements, HTML sample, per-state specs (stateSpecs: disabled/loading/open/empty etc. — each with delta class strings + aria changes), and anatomy parts (overlay/container/th etc. with element/roles/class). Prefer these structured fields over inferring states from the HTML sample.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Component ID (e.g. "button", "card", "table", "sidebar") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningfully disclose the return shape: per-state specs (disabled/loading/open/empty with delta class strings and aria changes) and anatomy parts with element/roles/class. This is real behavioral context beyond a bare "get," though it omits any statement of read-only safety, permissions, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the tool's purpose and contents, with no filler. The first sentence is dense with comma-separated field lists, but each clause maps to actual returned data and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must describe return values, and it does so thoroughly across variants, sizes, accessibility, HTML sample, stateSpecs, and anatomy. What remains missing is operational context (read-only nature, error cases on an unknown id), which is minor for a lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single id parameter, including example values, so the schema already does the work. The description adds no syntax, format, or lookup semantics for id beyond what is structured. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Get component metadata" states a specific verb and resource, and the description then enumerates exactly what that metadata includes (variants, sizes, accessibility requirements, HTML sample, stateSpecs, anatomy parts). The resource is unambiguous against siblings like get_token, get_rules, and check_html.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the resource: given a component id, fetch its metadata. The only explicit guidance — "Prefer these structured fields over inferring states from the HTML sample" — concerns how to interpret the output, not when to choose this tool over siblings such as check_html or get_token. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rulesA
Get melta-ui prohibition rules from rules.json (107 total). Use this to retrieve manual/contextual rules that check_rule cannot auto-detect. Supports filtering by category, severity, or detector.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Filter by category (e.g. "color", "spacing", "accessibility", "button", "modal") | |
| detector | No | Filter by detector type | |
| severity | No | Filter by severity |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the data source (rules.json) and corpus size (107), which is useful, but says nothing about whether all rules are returned by default, pagination/limits on a 107-item result, or permissions. Adequate but incomplete for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool returns and immediately followed by the routing condition. No sentence is redundant and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description should ideally describe the returned rule shape or the unfiltered default behavior. It covers the lookup purpose and filters adequately, but an agent still cannot predict what a call with no arguments yields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with two enum-constrained params and descriptive text for all three filters. The description merely restates that category/severity/detector filtering exists, adding no syntax, default, or combination semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Get) and resource (melta-ui prohibition rules from rules.json), gives the corpus size (107 total), and explicitly distinguishes the tool from its closest sibling check_rule. An agent can tell it apart from check_rule, search, and get_token without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to retrieve manual/contextual rules that check_rule cannot auto-detect" gives a clear selection condition and names the alternative tool. It stops short of stating when not to use it (e.g. when searching for arbitrary content, where `search` applies), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tokenA
Get a design token by dot-path. Returns the token object with its value and class mapping.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Dot-separated path to the token (e.g. "color.primary.600", "spacing.4", "radius.lg", "typography.fontSize.base") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the return shape ('token object with its value and class mapping'), which is genuine context. However, it says nothing about what happens on an invalid path, whether the lookup is exact-match or prefix-tolerant, or any read-only guarantee beyond the implicit verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and the key, followed by the return value. No filler or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with full schema coverage, the description supplies both the lookup semantics and a summary of the returned object, which is what an agent needs to invoke it. Only error/miss behavior is left unaddressed, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents 'path' with illustrative dot-path examples. The description's mention of dot-path adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('get a design token') plus the lookup mechanism ('by dot-path'), which is enough to distinguish it from siblings like get_component or check_rule. It stops short of explicitly contrasting itself with any named sibling, so it lands at clear-but-undifferentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (resolving a single token by path) but never states when to reach for this tool over search or get_rules. Usage is inferable from the resource name rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Search across tokens and components by keyword. Matches against names, values, class strings, and descriptions. Returns up to 20 results (truncated flag when more matched).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search keyword (e.g. "card", "primary", "shadow") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose two substantive traits: a hard cap of 20 results and a truncation flag when more matched. It stops short of explaining result ordering/relevance or that there is no pagination parameter to retrieve results beyond the cap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action, then matching scope, then result behavior. Every sentence adds information and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must cover behavior itself; it does so for result count and truncation. Gaps remain around relevance ordering and whether results are typed/labeled, and there is no hint of how to get past the 20-result cap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one parameter, so the schema already documents 'query' fully. The description adds what the keyword is matched against, which helps query construction, but no syntax, operators, or format rules beyond the schema's examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search across tokens and components by keyword') and enumerates the fields matched against (names, values, class strings, descriptions). This implicitly separates it from the get_* siblings, which fetch by identity, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: keyword search is what you reach for when you don't have an exact token/component name. The description never states when to prefer it over get_token or get_component, nor any exclusions, so the agent must infer the routing itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.8.0- First observed
check_html - First observed
check_rule - First observed
get_component - First observed
get_rules - First observed
get_token - First observed
search
TDQS
Scored across 6 tools
Most tools target distinct objects (get_component, get_token, search) with clear boundaries, but check_rule, check_html, and get_rules form a cluster around prohibition rules. The descriptions do distinguish them (apply to class strings vs. full HTML vs. retrieve raw rules), so an agent can reasonably select correctly.
Names mostly follow a verb_noun pattern (get_component, get_token, get_rules, check_rule, check_html), with only 'search' as a bare verb deviation. Still highly predictable and consistent in style.
Six tools is well-scoped for a design-system lookup and validation server. Each tool serves a distinct purpose (component metadata, tokens, search, three rule-checking variants) with no obvious filler.
The surface covers the core read-only lifecycle: discover (search), inspect (get_component/get_token), and validate (check_rule/check_html/get_rules). Only minor gaps exist, e.g. no explicit list-components or list-tokens endpoint, though search largely covers discovery.
Maintenance
Related MCP Connectors
Serves your design system and coding standards to coding agents, so they stop guessing.
Give your agent a real design system: tokens, measured WCAG contrast, and rules to follow.
Live React design-system APIs, patterns, and code validation so AI agents build real UI, not slop.
Design intelligence for coding agents: audits, design systems, and a taste profile agents consult.
Related MCP Servers
- AlicenseBqualityBmaintenanceEnables AI agents to access design system tokens and component contracts through MCP, reducing token usage and ensuring consistency.29MIT
- AlicenseAqualityBmaintenanceProvides deterministic, read-only design knowledge for AI coding agents to help them choose visual directions, plan UI states, and compose design tokens, all without network access.612 npm4MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to score live URLs against a 40-check design contract, validate DTCG tokens and Lottie animations, audit accessibility, and retrieve design-system contracts, catalogs, and review rubrics.35 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to query a workspace's design system before writing UI and validate generated code against the same system afterward, using configurable token and component sources.9 npmMIT