spec-score-mcp
Spec Score MCP
Claudeが仕様書から構築を始める前に、その仕様書をスコアリングします。
バランスの取れた仕様書は、バランスの取れたコードを生み出します。バランスの悪い仕様書は、創作的なフィクションを生み出します。
問題点
仕様書がある軸では詳細なのに、別の軸では曖昧な場合、Claudeは明確化を求めるのではなく、空白を埋めてしまいます。結果としてコンパイルは通り、テストもパスしますが、それは意図したものではありません。
このツールは、構築を始める前にその問題を検知します。仕様書を4つの軸でスコアリングし、最も弱い軸を指摘し、それを修正するための具体的なヒントを提供します。
Related MCP server: MCP Prompt Optimizer
4つの軸
軸 | 回答する問い |
完全性 (completeness) | Claudeは構築すべき範囲全体を理解できるか? |
明確性 (clarity) | この仕様書には唯一の解釈しかないか? |
制約 (constraints) | Claudeは「何を構築すべきでないか」を知っているか? |
具体性 (specificity) | 具体的なテスト可能な詳細があるか? |
各軸は0.0から1.0でスコアリングされます。バランススコアは、4つの軸がどれだけ均等にカバーされているかを測定します。
個々のスコアよりもバランスが重要です。 4つの軸すべてで0.50のスコア(バランス: 0.97)を持つ仕様書は、0.95 / 0.95 / 0.20 / 0.90(バランス: 0.58)のスコアを持つ仕様書よりも優れた出力を生成します。なぜでしょうか?その弱い軸である「制約(0.20)」こそが、Claudeが即興で補完してしまう部分だからです。何を構築するかは詳細に記述していても、何が範囲外かを書き忘れているため、Claudeは要求したすべてに加えて、要求していない機能まで構築してしまいます。
レーダーチャートでは、鋭いスパイクよりも均等なひし形の方が優れています。
判定結果
判定 | 意味 |
SHIP IT | 仕様書は準備完了 — Claudeは何を構築し、何を構築すべきでないかを理解している |
ALMOST | 開始前に1つの軸に小さな修正が必要 |
DRAFT | 複数の軸に改善が必要だが、構造はできている |
VAGUE | 整理はされているが、抽象的すぎて実行できない |
UNBOUNDED | 目標は明確だが境界がない — Claudeは過剰に構築してしまう |
OVER-CONSTRAINED | ルールは多いが、実際の目標が不明確 |
SKETCH | 出発点 — ほとんどの軸で詳細が必要 |
まだ「SHIP IT」ではありませんか?このツールは、どの軸が最も弱く、何を追加すべきかを教えてくれます。その軸を修正し、再スコアリングし、繰り返してください。ほとんどの仕様書は2〜3ラウンドで「SHIP IT」に到達します。
インストール
git clone https://github.com/openpoem/spec-score-mcp.git
cd spec-score-mcp && npm install && npm run build
claude mcp add spec-score -- node $(pwd)/dist/mcp.jsこれで、すべてのClaude Codeセッションで3つのツールが利用可能になります。
使用方法
スラッシュコマンド
このリポジトリをクローンして、組み込みのスラッシュコマンドを取得します:
/project:scan my-feature-spec.mdファイルを読み込んでスコアリングし、スコア、判定、ヒント、レーダーチャートを含む my-feature-spec.md.scored.md を書き出します。
/project:compare blueprint.md implementation.md両方のファイルをスコアリングし、並べて比較したレーダーチャートを含む compared.scored.md を書き出します。
直接的なツール使用
3つのMCPツールは、どのClaude Codeの会話でも機能します:
ツール | 機能 |
| 仕様書を4つの軸でスコアリングし、バランススコアと判定を返す |
| スコアからSVGレーダーチャートを生成する |
| 2つのスコアリング済み仕様書を並べて比較する |
Claudeに「Score this spec(この仕様書をスコアリングして)」、「Show me the radar chart(レーダーチャートを見せて)」、または「Compare these two specs(これら2つの仕様書を比較して)」と尋ねてください。
例:UNBOUNDEDからSHIP ITへ
このツール自身の仕様書をスコアリングした例です。4ラウンドにわたり、それぞれ最も弱い軸を修正しました:
ラウンド1:アイデア
仕様書スコアリングツールを構築する
UNBOUNDED 0.12 Tip: What does 'scoring' mean? What axes? What output?1つの軸(明確性 — 目標は明確)は高いですが、他はほぼゼロです。Claudeは何でも構築してしまうでしょう。Webアプリ?CLI?VS Code拡張機能?知る由もありません。
ラウンド2:コンテキストの追加
4つの軸(完全性、明確性、制約、具体性)で仕様書をスコアリングするMCPサーバーを構築する。各軸は0.0〜1.0。バランススコアと判定を返す。
ALMOST 0.67 Tip: What are the verdicts? What does the tool NOT do?これでClaudeは何を構築すべきかを知りました。しかし、制約はまだ弱いです。自動修正、CI統合、データベースなどを勝手に追加してしまうかもしれません。
ラウンド3:境界の追加
3つのツール:spec_score, spec_visualize, spec_compare。非目標:自動修正なし、CI統合なし、ストレージなし。
SHIP IT 0.84 Tip: Add testable criteria — what balance maps to which verdict?閾値を超えました。Claudeは今や、何を構築すべきか、そして何を構築すべきでないかを理解しています。具体性が依然として最も弱い軸です。
ラウンド4:テスト可能な詳細の追加
バランス = 1 - sqrt(分散)/平均。SHIP IT > 0.75, ALMOST > 0.60、およびパターンベースの判定。Node.js, MCP SDK, stdioトランスポート。
SHIP IT 0.95 Spec is ready for implementation.4ラウンド:0.12 → 0.67 → 0.84 → 0.95。各ラウンドで正確に1つずつ修正しました。
数学的な仕組み
Claudeが各軸をスコアリングする (0.0 - 1.0)
ベクトルの正規化:
v / ||v||バランス:
1 - sqrt(分散) / 平均判定: バランス閾値 + 軸のパターンマッチング
スコアリングの知能はアルゴリズムではなくClaudeから来ています。アルゴリズムはバランスを測定するだけです。
プロジェクト構造
src/
mcp.ts # MCP server (3 tools)
score.ts # Scoring engine
visualize.ts # SVG radar charts
.claude/
commands/
scan.md # /project:scan command
compare.md # /project:compare commandOpenPoem — spec-score-mcp
MITライセンス。
© 2026 OpenPoem. info@openpoem.org
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v2.0.2- First observed
spec_compare - First observed
spec_score - First observed
spec_visualize
TDQS
Scored across 3 tools
The tools have overlapping purposes that could cause confusion. spec_score and spec_visualize both score a spec on the same four axes and provide the same analysis, making them nearly redundant. Only spec_compare has a clearly distinct function by comparing two specs, but the other two tools are ambiguous in their differentiation.
The naming follows a consistent pattern with all tools using the prefix 'spec_' followed by a verb (compare, score, visualize). This makes the purpose of each tool predictable and readable, though the similarity in naming between spec_score and spec_visualize contributes to the disambiguation issue.
With 3 tools, the count is reasonable for a server focused on spec evaluation. It covers core functions like scoring, comparing, and visualizing specs, which aligns well with the server's purpose, though the overlap between spec_score and spec_visualize suggests the set could be streamlined without losing functionality.
The tool surface is mostly complete for spec evaluation, covering scoring, comparison, and visualization. However, there is a notable gap in tools for editing or updating specs based on the analysis, which could limit workflow coverage. The redundancy between spec_score and spec_visualize also indicates inefficiency rather than a functional gap.
Maintenance
Related MCP Connectors
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
Generate and validate a .specs/ bundle for your repo, then hand it to your AI coding agent
Commission a multi-model AI spec committee from your agent; get rubric-scored, build-ready specs.
Checks llms.txt, AI crawler access in robots.txt, and sitemap - with a 0-100 AI readiness score.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Spec-Driven Development toolkit that transforms LLMs into development agents by providing expert-crafted prompts for generating structured specifications and validating documents across the Requirements → Design → Tasks → Code workflow.1MIT
- AlicenseBqualityDmaintenanceAutomatically analyzes and optimizes AI prompts by calculating clarity scores, detecting risks, asking clarifying questions, and adding domain-specific requirements to improve AI interaction quality.1MIT
- AlicenseAqualityBmaintenanceVet ClawHub skills before installing them; detects prompt-injection, exfiltration, and other security issues, outputting a risk score with per-finding evidence.7MIT
- AlicenseAqualityFmaintenanceTurn rough requests into rigorously structured prompts for any coding agent. Quality-scored to ≥90/100 across 12 dimensions, calibrated on 1,000+ real coding cases.128 npm2MIT