Skip to main content
Glama

Spec Score MCP

在 Claude 根据规范构建之前对其进行评分。

平衡的规范产生平衡的代码。不平衡的规范产生创意小说。

问题所在

当你的规范在某些维度上很详细,但在其他维度上很模糊时,Claude 不会要求澄清——它会自行填补空白。结果虽然能编译通过,测试也能通过,但这并不是你的本意。

该工具在你开始构建之前就能发现这一点。它从 4 个维度对你的规范进行评分,告诉你哪个维度最薄弱,并提供具体的修复建议。

Related MCP server: MCP Prompt Optimizer

4 个维度

维度

回答的问题

完整性

Claude 能否理解要构建内容的全部范围?

清晰度

对此规范是否只有一种解读方式?

约束

Claude 是否知道什么“不”该构建?

具体性

是否有具体、可测试的细节?

每个维度的得分在 0.0 到 1.0 之间。平衡得分衡量 4 个维度的覆盖是否均匀。

平衡比单项得分更重要。 一个在 4 个维度上都得 0.50 分(平衡度:0.97)的规范,其产出效果要好于得分为 0.95 / 0.95 / 0.20 / 0.90(平衡度:0.58)的规范。为什么?那个薄弱的维度——约束(0.20)——正是 Claude 会进行即兴发挥的地方。你详细描述了要构建什么,但忘记说明什么不在范围内。因此,Claude 构建了你要求的所有内容,外加你没要求的特性。

在雷达图上:均匀的菱形胜过尖锐的突刺。

结论

结论

含义

SHIP IT

规范已就绪——Claude 知道要构建什么以及不构建什么

ALMOST

在开始之前,有一个维度需要小幅修正

DRAFT

多个维度需要改进,但结构已经具备

VAGUE

组织良好,但过于抽象,无法执行

UNBOUNDED

目标明确但没有边界——Claude 会过度构建

OVER-CONSTRAINED

规则很多,但不清楚实际目标是什么

SKETCH

起点——大多数维度都需要细节

还没到 SHIP IT?该工具会告诉你哪个维度最薄弱以及需要添加什么。修复该维度,重新评分,重复此过程。大多数规范在 2-3 轮后即可达到 SHIP IT。

安装

git clone https://github.com/openpoem/spec-score-mcp.git
cd spec-score-mcp && npm install && npm run build
claude mcp add spec-score -- node $(pwd)/dist/mcp.js

这 3 个工具现在可在每个 Claude Code 会话中使用。

使用方法

斜杠命令

克隆此仓库以获取内置的斜杠命令:

/project:scan my-feature-spec.md

读取文件,对其评分,并写入一个包含分数、结论、提示和雷达图的 my-feature-spec.md.scored.md 文件。

/project:compare blueprint.md implementation.md

对两个文件进行评分,并写入一个包含并排雷达图的 compared.scored.md 文件。

直接使用工具

这 3 个 MCP 工具可在任何 Claude Code 对话中使用:

工具

功能

spec_score

从 4 个维度对规范评分,返回平衡得分和结论

spec_visualize

根据分数生成 SVG 雷达图

spec_compare

对两个已评分的规范进行并排比较

询问 Claude:"Score this spec"(为该规范评分)、"Show me the radar chart"(向我展示雷达图)或 "Compare these two specs"(比较这两个规范)。

示例:从 UNBOUNDED 到 SHIP IT

该工具为自己的规范评分——四轮,每一轮都修复最薄弱的维度:

第 1 轮:想法

构建一个规范评分工具

UNBOUNDED  0.12  Tip: What does 'scoring' mean? What axes? What output?

一个维度很高(清晰度——目标明确),其他几乎为零。Claude 会构建……任何东西。Web 应用?CLI?VS Code 扩展?无法得知。

第 2 轮:添加上下文

构建一个 MCP 服务器,从 4 个维度对规范进行评分:完整性、清晰度、约束、具体性。每个维度为 0.0-1.0。返回平衡得分和结论。

ALMOST  0.67  Tip: What are the verdicts? What does the tool NOT do?

现在 Claude 知道要构建什么了。但约束仍然很薄弱——它可能会添加自动修复、CI 集成、数据库等。

第 3 轮:添加边界

三个工具:spec_score, spec_visualize, spec_compare。 非目标:无自动修复,无 CI 集成,无存储。

SHIP IT  0.84  Tip: Add testable criteria — what balance maps to which verdict?

跨过了阈值。Claude 现在知道要构建什么以及“不”构建什么。具体性仍然是最薄弱的维度。

第 4 轮:添加可测试的细节

平衡度 = 1 - sqrt(方差)/平均值。SHIP IT > 0.75, ALMOST > 0.60, 加上基于模式的结论。Node.js, MCP SDK, stdio 传输。

SHIP IT  0.95  Spec is ready for implementation.

四轮:0.12 → 0.67 → 0.84 → 0.95。每一轮都精确修复了一件事。

数学原理

  1. Claude 对每个维度评分 (0.0 - 1.0)

  2. 归一化向量:v / ||v||

  3. 平衡度:1 - sqrt(方差) / 平均值

  4. 结论:平衡阈值 + 维度模式匹配

评分智能来自 Claude,而不是算法。算法仅用于衡量平衡度。

项目结构

src/
  mcp.ts        # MCP server (3 tools)
  score.ts      # Scoring engine
  visualize.ts  # SVG radar charts
.claude/
  commands/
    scan.md     # /project:scan command
    compare.md  # /project:compare command

OpenPoem — spec-score-mcp

MIT 许可证。

© 2026 OpenPoem. info@openpoem.org

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv2.0.2
    • First observedspec_compare
    • First observedspec_score
    • First observedspec_visualize

TDQS

A3.8/5.0

Scored across 3 tools

Disambiguation3/5

The tools have overlapping purposes that could cause confusion. spec_score and spec_visualize both score a spec on the same four axes and provide the same analysis, making them nearly redundant. Only spec_compare has a clearly distinct function by comparing two specs, but the other two tools are ambiguous in their differentiation.

Naming Consistency4/5

The naming follows a consistent pattern with all tools using the prefix 'spec_' followed by a verb (compare, score, visualize). This makes the purpose of each tool predictable and readable, though the similarity in naming between spec_score and spec_visualize contributes to the disambiguation issue.

Tool Count4/5

With 3 tools, the count is reasonable for a server focused on spec evaluation. It covers core functions like scoring, comparing, and visualizing specs, which aligns well with the server's purpose, though the overlap between spec_score and spec_visualize suggests the set could be streamlined without losing functionality.

Completeness4/5

The tool surface is mostly complete for spec evaluation, covering scoring, comparison, and visualization. However, there is a notable gap in tools for editing or updating specs based on the analysis, which could limit workflow coverage. The redundancy between spec_score and spec_visualize also indicates inefficiency rather than a functional gap.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A Spec-Driven Development toolkit that transforms LLMs into development agents by providing expert-crafted prompts for generating structured specifications and validating documents across the Requirements → Design → Tasks → Code workflow.
    1
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Automatically analyzes and optimizes AI prompts by calculating clarity scores, detecting risks, asking clarifying questions, and adding domain-specific requirements to improve AI interaction quality.
    1
    MIT