Skip to main content
Glama
Zoya-Ammar

AI Agent Release Assurance MCP

by Zoya-Ammar

AI Agent Release Assurance MCP

Version 0.1: 使用合成 QA 数据的发布智能基础。
AI 智能体评估能力计划在 Version 0.2 中提供。

一个可解释的 Model Context Protocol (MCP) 服务器,帮助 AI 客户端分析软件测试结果和缺陷,并生成基于证据的发布就绪建议。

本仓库中的所有发布、测试、缺陷和客户影响场景均为虚构。未使用任何雇主、客户、生产环境、个人或受监管数据。

为什么创建这个项目

发布决策通常需要分散在测试结果、缺陷记录和团队文档中的证据。

该服务器为 AI 客户端提供了一个小型只读接口,用于回答如下问题:

  • 某个特定发布是否应该上线?

  • 哪些失败的测试是潜在的发布阻塞项?

  • 未解决的缺陷风险集中在哪些地方?

  • 在定向回归测试中应优先执行哪些测试?

AI 不会凭空得出风险评分。服务器以确定性方式计算该评分,并返回底层证据、权重、阻塞项和建议的后续操作,供人工审查。

Related MCP server: QA Copilot AI

当前能力

类型

名称

用途

工具

assess_release_readiness

返回可解释的 GO、CONDITIONAL_GO 或 NO_GO 建议

工具

get_failed_tests

检索失败和阻塞的测试,支持可选的关键性筛选

工具

find_defect_hotspots

按严重性加权的未解决缺陷风险对组件进行排名

工具

recommend_regression_tests

创建有界、基于风险的回归计划

资源

qa://releases

列出可供分析的合成发布

提示词

release_go_no_go

引导基于证据的发布就绪审查

架构

flowchart TD
    A[AI host or MCP Inspector] -->|MCP request| B[Python MCP server]
    B --> C[QA service and risk rules]
    C --> D[(Synthetic SQLite data)]
    D --> C
    C -->|Structured evidence| B
    B -->|Tool result| A
    D -. optional migration .-> E[(Snowflake)]

SQLite 使 Version 0.1 可重现且无需凭据。可选的 snowflake/setup.sql 文件演示了一种可能的 Snowflake 原生 MCP 路径。

快速入门

环境要求

  • Python 3.10 或更高版本

  • uv

  • Node.js/npm,用于可视化 MCP Inspector

安装并运行

git clone https://github.com/Zoya-Ammar/ai-agent-release-assurance-mcp.git
cd ai-agent-release-assurance-mcp
uv sync --extra dev
uv run python -m banking_qa_mcp.seed
uv run mcp dev src/banking_qa_mcp/server.py

最后一条命令将启动 MCP Inspector。

打开 工具,选择 assess_release_readiness,并提供:

{
  "release_id": "REL-2026.08.1"
}

预期的主要结果:

{
  "recommendation": "NO_GO",
  "risk_score": 100,
  "test_pass_rate_percent": 62.5,
  "blockers": [
    "Open SEV1 defect",
    "Failed or blocked critical test",
    "Failed or blocked high-criticality test"
  ]
}

作为对比,REL-2026.08.2 返回 GO,风险评分为 7。

运行测试

运行完整的自动化测试套件:

uv run pytest -q

运行无依赖的核心验证:

uv run python scripts/smoke_test.py

Version 0.1 包含以下测试:

  • 高风险和较低风险发布建议

  • 测试结果筛选

  • 回归计划的限制和优先级排序

  • 无效的发布标识符

可解释的风险评分

评分上限为 100:

25 × failed or blocked critical tests
12 × failed or blocked high-criticality tests
35 × open SEV1 defects
18 × open SEV2 defects
 7 × open SEV3 defects
 2 × open SEV4 defects

未关闭的 SEV1 缺陷、失败或阻塞的关键测试,以及失败或阻塞的高关键性测试,也会被报告为明确的发布阻塞项。

这些权重是演示策略——而非通用的金融服务或软件质量标准。在生产环境中,阈值需要由相应的风险负责人批准,并经过版本控制、验证和定期审查。

安全注意事项

Version 0.1 在应用层刻意设计为只读。生产实现还应包括:

  • 身份验证和基于角色的授权

  • 最小权限的数据库角色和服务角色

  • 输入和输出验证

  • 工具调用和推荐建议的审计日志

  • 速率限制和可观测性

  • 机密管理和加密传输

  • 发布决策的人工审批

  • 对检索内容进行提示注入测试

可选的 Snowflake 示例包含一个用于沙盒演示的原生 SQL 执行工具。它应通过专用的只读角色加以限制,并在任何非演示用途之前进一步收紧权限。

Version 0.2 路线图

下一版本将把这个发布智能基础扩展为 AI 智能体保障系统。

计划中的能力包括:

  • 一个原创的 AI 智能体评估语料库

  • 事实锚定和引用验证

  • 提示注入抗性测试

  • 隐私和数据最小化检查

  • 可访问性和负向路径场景

  • 基线版与候选版的对比

  • 智能体版本之间的回归检测

  • 基于 Playwright 的 UI 和可访问性执行

  • 基于 Snowflake 的评估证据

  • 经人工审查的 AI 智能体发布建议

项目状态

本仓库是一个教育用途的作品集原型。它不是生产级银行系统、合规工具,也不是自主发布决策机构。

参考资料

许可证

本项目基于 MIT License 提供。

Available Tools

4 tools
assess_release_readinessB

Calculate an explainable GO, CONDITIONAL_GO, or NO_GO recommendation.

ParametersJSON Schema
NameRequiredDescriptionDefault
release_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the output is an explainable recommendation with three possible values, but it does not reveal how the recommendation is derived, whether it depends on external sources, or what 'explainable' means in practice. This is acceptable but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and outcome with no filler. It is appropriately sized for a one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only one parameter and no output schema, and the description names the three output categories, which covers the basic return shape. But it omits the criteria behind the recommendation, the source of the release ID, and any caveats, leaving the agent with an incomplete picture of how to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not elaborate on release_id beyond the schema's string type and title. Since the only parameter is central to the tool, the description should at least clarify what qualifies as a release_id and how it is used; it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear action ('Calculate') and a specific deliverable ('GO, CONDITIONAL_GO, or NO_GO recommendation'), which goes beyond the tool name. It is distinguishable from the sibling tools by its outcome-oriented purpose, though it does not explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied: this is the high-level readiness assessment tool, while siblings like get_failed_tests and find_defect_hotspots are lower-level diagnostic tools. However, the description never states when to use this tool versus its alternatives, so an agent must infer the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_defect_hotspotsB

Rank release components by the weighted risk of unresolved defects.

ParametersJSON Schema
NameRequiredDescriptionDefault
release_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It implies a read-only ranking operation and specifically scopes to unresolved defects, but it does not explain how 'weighted risk' is computed, whether historical data is considered, or what happens when no defects are found. Basic but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and object, then adds the precise qualifier. Every word earns its place with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, and the single parameter is simple. However, the description omits when to prefer this over sibling tools and does not clarify the meaning of 'components' or 'weighted risk.' It is minimally viable but leaves notable gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description should compensate, but it never explains what release_id means or how it relates to the ranking. The schema only shows it is a required string. The description uses 'release' in its wording, providing only a weak hint, not clear parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Rank'), a resource ('release components'), and a distinguishing criterion ('weighted risk of unresolved defects'). This clearly differentiates it from sibling tools like get_failed_tests or assess_release_readiness, which focus on different outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus the sibling tools. It does not mention alternatives, exclusions, or conditions under which another tool would be a better fit, leaving the agent to infer usage purely from the name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_failed_testsB

Return failed and blocked tests, optionally filtered by criticality.

ParametersJSON Schema
NameRequiredDescriptionDefault
release_idYes
criticalityNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the output (failed/blocked tests) but does not mention pagination, ordering, empty-result behavior, required release context, or consequences. Nothing contradicts annotations because none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word adds meaning, and the main result is stated before the optional filter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite a simple two-parameter shape and an output schema, the definition lacks enough context for confident invocation: no sibling differentiation, no release_id semantics, and no criticality value guidance. This is insufficient for a low-coverage schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only clarifies the optional criticality filter; it does not explain release_id or enumerate accepted criticality values, leaving a required parameter largely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Return failed and blocked tests'. This clearly distinguishes it from siblings like assess_release_readiness and recommend_regression_tests, which are analysis/recommendation tools rather than retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to prefer this tool over its siblings. The only usage hint is the optional criticality filter, which is more of a parameter option than a when-to-use instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_regression_testsC

Build a risk-based regression plan grounded in test and defect evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_testsNo
release_idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions that the plan is 'risk-based' and 'grounded in test and defect evidence,' but it does not disclose what the tool returns, how it uses release_id and max_tests, or whether it only reads data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It begins with the action and object and adds value by specifying risk-based and evidence-grounded characteristics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, no annotations, and no output schema, this one-line description is incomplete. It does not explain expected outputs, the role of max_tests, or selection criteria, leaving important context for correct invocation unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions release_id or max_tests. The phrase 'test and defect evidence' does not explain the required release parameter or the meaning of the max_tests default, so the agent gets no parameter help beyond field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action, 'Build a risk-based regression plan,' and a clear resource. It distinguishes itself from sibling tools by focusing on test recommendation and evidence grounding, though it does not explicitly name or contrast any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool versus assess_release_readiness, get_failed_tests, or find_defect_hotspots. There are no prerequisites or exclusions, so an agent must infer usage solely from the name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedassess_release_readiness
    • First observedfind_defect_hotspots
    • First observedget_failed_tests
    • First observedrecommend_regression_tests

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation4/5

Each tool produces a distinct output: a GO/NO-GO decision, a filtered list of test failures, a component risk ranking, and a regression test plan. find_defect_hotspots and recommend_regression_tests share an evidence base of defect/test risk, but their purposes are clearly separated by output type, so misselection is unlikely.

Naming Consistency5/5

All four tools follow a consistent verb_noun snake_case pattern (assess_release_readiness, get_failed_tests, find_defect_hotspots, recommend_regression_tests). The verb clearly signals the action (assess, get, find, recommend) and the noun signals the resource, making the pattern highly predictable.

Tool Count5/5

Four tools is on the lean side but well-scoped for release assurance: each tool fills a distinct role covering evidence gathering, risk analysis, planning, and final decision. There is no redundancy or bloat, and every tool earns its place in the pipeline.

Completeness4/5

The set forms a coherent end-to-end release readiness workflow: pull test failures, rank defect hotspots, build a regression plan from that evidence, and produce a final GO/NO-GO assessment. Minor gaps exist, such as no tool to drill into individual defect details or fetch component/change scope, but agents can work around these.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    B
    maintenance
    MCP server for AI-powered QA analysis. It enables analyzing test failures, identifying root causes, suggesting fixes, classifying defects, detecting flaky tests, and generating test cases and bug reports.
    10
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables evidence-first release readiness assessment by running or accepting build, API, browser, visual, performance, and security evidence, then returning SHIP, REVIEW, or HOLD recommendations with clustered regressions.
    2 npm
    MIT