spec-score-mcp
Spec Score MCP
在 Claude 根据规范构建之前对其进行评分。
平衡的规范产生平衡的代码。不平衡的规范产生创意小说。
问题所在
当你的规范在某些维度上很详细,但在其他维度上很模糊时,Claude 不会要求澄清——它会自行填补空白。结果虽然能编译通过,测试也能通过,但这并不是你的本意。
该工具在你开始构建之前就能发现这一点。它从 4 个维度对你的规范进行评分,告诉你哪个维度最薄弱,并提供具体的修复建议。
Related MCP server: MCP Prompt Optimizer
4 个维度
维度 | 回答的问题 |
完整性 | Claude 能否理解要构建内容的全部范围? |
清晰度 | 对此规范是否只有一种解读方式? |
约束 | Claude 是否知道什么“不”该构建? |
具体性 | 是否有具体、可测试的细节? |
每个维度的得分在 0.0 到 1.0 之间。平衡得分衡量 4 个维度的覆盖是否均匀。
平衡比单项得分更重要。 一个在 4 个维度上都得 0.50 分(平衡度:0.97)的规范,其产出效果要好于得分为 0.95 / 0.95 / 0.20 / 0.90(平衡度:0.58)的规范。为什么?那个薄弱的维度——约束(0.20)——正是 Claude 会进行即兴发挥的地方。你详细描述了要构建什么,但忘记说明什么不在范围内。因此,Claude 构建了你要求的所有内容,外加你没要求的特性。
在雷达图上:均匀的菱形胜过尖锐的突刺。
结论
结论 | 含义 |
SHIP IT | 规范已就绪——Claude 知道要构建什么以及不构建什么 |
ALMOST | 在开始之前,有一个维度需要小幅修正 |
DRAFT | 多个维度需要改进,但结构已经具备 |
VAGUE | 组织良好,但过于抽象,无法执行 |
UNBOUNDED | 目标明确但没有边界——Claude 会过度构建 |
OVER-CONSTRAINED | 规则很多,但不清楚实际目标是什么 |
SKETCH | 起点——大多数维度都需要细节 |
还没到 SHIP IT?该工具会告诉你哪个维度最薄弱以及需要添加什么。修复该维度,重新评分,重复此过程。大多数规范在 2-3 轮后即可达到 SHIP IT。
安装
git clone https://github.com/openpoem/spec-score-mcp.git
cd spec-score-mcp && npm install && npm run build
claude mcp add spec-score -- node $(pwd)/dist/mcp.js这 3 个工具现在可在每个 Claude Code 会话中使用。
使用方法
斜杠命令
克隆此仓库以获取内置的斜杠命令:
/project:scan my-feature-spec.md读取文件,对其评分,并写入一个包含分数、结论、提示和雷达图的 my-feature-spec.md.scored.md 文件。
/project:compare blueprint.md implementation.md对两个文件进行评分,并写入一个包含并排雷达图的 compared.scored.md 文件。
直接使用工具
这 3 个 MCP 工具可在任何 Claude Code 对话中使用:
工具 | 功能 |
| 从 4 个维度对规范评分,返回平衡得分和结论 |
| 根据分数生成 SVG 雷达图 |
| 对两个已评分的规范进行并排比较 |
询问 Claude:"Score this spec"(为该规范评分)、"Show me the radar chart"(向我展示雷达图)或 "Compare these two specs"(比较这两个规范)。
示例:从 UNBOUNDED 到 SHIP IT
该工具为自己的规范评分——四轮,每一轮都修复最薄弱的维度:
第 1 轮:想法
构建一个规范评分工具
UNBOUNDED 0.12 Tip: What does 'scoring' mean? What axes? What output?一个维度很高(清晰度——目标明确),其他几乎为零。Claude 会构建……任何东西。Web 应用?CLI?VS Code 扩展?无法得知。
第 2 轮:添加上下文
构建一个 MCP 服务器,从 4 个维度对规范进行评分:完整性、清晰度、约束、具体性。每个维度为 0.0-1.0。返回平衡得分和结论。
ALMOST 0.67 Tip: What are the verdicts? What does the tool NOT do?现在 Claude 知道要构建什么了。但约束仍然很薄弱——它可能会添加自动修复、CI 集成、数据库等。
第 3 轮:添加边界
三个工具:spec_score, spec_visualize, spec_compare。 非目标:无自动修复,无 CI 集成,无存储。
SHIP IT 0.84 Tip: Add testable criteria — what balance maps to which verdict?跨过了阈值。Claude 现在知道要构建什么以及“不”构建什么。具体性仍然是最薄弱的维度。
第 4 轮:添加可测试的细节
平衡度 = 1 - sqrt(方差)/平均值。SHIP IT > 0.75, ALMOST > 0.60, 加上基于模式的结论。Node.js, MCP SDK, stdio 传输。
SHIP IT 0.95 Spec is ready for implementation.四轮:0.12 → 0.67 → 0.84 → 0.95。每一轮都精确修复了一件事。
数学原理
Claude 对每个维度评分 (0.0 - 1.0)
归一化向量:
v / ||v||平衡度:
1 - sqrt(方差) / 平均值结论:平衡阈值 + 维度模式匹配
评分智能来自 Claude,而不是算法。算法仅用于衡量平衡度。
项目结构
src/
mcp.ts # MCP server (3 tools)
score.ts # Scoring engine
visualize.ts # SVG radar charts
.claude/
commands/
scan.md # /project:scan command
compare.md # /project:compare commandOpenPoem — spec-score-mcp
MIT 许可证。
© 2026 OpenPoem. info@openpoem.org
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v2.0.2- First observed
spec_compare - First observed
spec_score - First observed
spec_visualize
TDQS
Scored across 3 tools
The tools have overlapping purposes that could cause confusion. spec_score and spec_visualize both score a spec on the same four axes and provide the same analysis, making them nearly redundant. Only spec_compare has a clearly distinct function by comparing two specs, but the other two tools are ambiguous in their differentiation.
The naming follows a consistent pattern with all tools using the prefix 'spec_' followed by a verb (compare, score, visualize). This makes the purpose of each tool predictable and readable, though the similarity in naming between spec_score and spec_visualize contributes to the disambiguation issue.
With 3 tools, the count is reasonable for a server focused on spec evaluation. It covers core functions like scoring, comparing, and visualizing specs, which aligns well with the server's purpose, though the overlap between spec_score and spec_visualize suggests the set could be streamlined without losing functionality.
The tool surface is mostly complete for spec evaluation, covering scoring, comparison, and visualization. However, there is a notable gap in tools for editing or updating specs based on the analysis, which could limit workflow coverage. The redundancy between spec_score and spec_visualize also indicates inefficiency rather than a functional gap.
Maintenance
Related MCP Connectors
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
Generate and validate a .specs/ bundle for your repo, then hand it to your AI coding agent
Commission a multi-model AI spec committee from your agent; get rubric-scored, build-ready specs.
Checks llms.txt, AI crawler access in robots.txt, and sitemap - with a 0-100 AI readiness score.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Spec-Driven Development toolkit that transforms LLMs into development agents by providing expert-crafted prompts for generating structured specifications and validating documents across the Requirements → Design → Tasks → Code workflow.1MIT
- AlicenseBqualityDmaintenanceAutomatically analyzes and optimizes AI prompts by calculating clarity scores, detecting risks, asking clarifying questions, and adding domain-specific requirements to improve AI interaction quality.1MIT
- AlicenseAqualityBmaintenanceVet ClawHub skills before installing them; detects prompt-injection, exfiltration, and other security issues, outputting a risk score with per-finding evidence.7MIT
- AlicenseAqualityFmaintenanceTurn rough requests into rigorously structured prompts for any coding agent. Quality-scored to ≥90/100 across 12 dimensions, calibrated on 1,000+ real coding cases.128 npm2MIT