AI Agent Release Assurance MCP
AI Agent Release Assurance MCP
Version 0.1: 합성 QA 데이터를 사용한 릴리스 인텔리전스 기반.
AI 에이전트 평가 기능은 Version 0.2에서 계획되어 있습니다.
이 저장소의 모든 릴리스, 테스트, 결함 및 고객 영향 시나리오는 가상의 것입니다. 고용주, 고객, 프로덕션, 개인 또는 규제 대상 데이터는 사용되지 않습니다.
이 프로젝트가 존재하는 이유
릴리스 결정에는 테스트 결과, 결함 기록, 팀 문서에 분산된 근거가 필요한 경우가 많습니다.
이 서버는 AI 클라이언트가 다음과 같은 질문에 답할 수 있도록 작은 읽기 전용 인터페이스를 제공합니다:
특정 릴리스가 출시되어야 할까요?
어떤 실패한 테스트가 잠재적 릴리스 차단 요인일까요?
미해결 결함 위험이 어디에 집중되어 있을까요?
대상 회귀 테스트 중 어떤 테스트에 우선순위를 두어야 할까요?
AI는 위험 점수를 지어내지 않습니다. 서버는 점수를 결정론적으로 계산하고, 인간 검토를 위해 기본 근거, 가중치, 차단 요인 및 권장 다음 작업을 반환합니다.
Related MCP server: QA Copilot AI
현재 기능
유형 | 이름 | 목적 |
도구 |
| 설명 가능한 |
도구 |
| 선택적 중요도 필터를 사용하여 실패 및 차단된 테스트를 검색합니다 |
도구 |
| 심각도 가중치가 적용된 미해결 결함 위험을 기준으로 구성 요소 순위를 매깁니다 |
도구 |
| 제한된 범위의 위험 기반 회귀 테스트 계획을 생성합니다 |
리소스 |
| 분석에 사용할 수 있는 합성 릴리스를 나열합니다 |
프롬프트 |
| 근거 기반 릴리스 출시 준비 검토를 안내합니다 |
아키텍처
flowchart TD
A[AI host or MCP Inspector] -->|MCP request| B[Python MCP server]
B --> C[QA service and risk rules]
C --> D[(Synthetic SQLite data)]
D --> C
C -->|Structured evidence| B
B -->|Tool result| A
D -. optional migration .-> E[(Snowflake)]SQLite는 Version 0.1을 재현 가능하게 유지하며 자격 증명이 필요 없습니다. 선택적 snowflake/setup.sql 파일은 가능한 Snowflake 네이티브 MCP 경로를 보여줍니다.
빠른 시작
요구 사항
Python 3.10 이상
시각적 MCP Inspector용 Node.js/npm
설치 및 실행
git clone https://github.com/Zoya-Ammar/ai-agent-release-assurance-mcp.git
cd ai-agent-release-assurance-mcp
uv sync --extra dev
uv run python -m banking_qa_mcp.seed
uv run mcp dev src/banking_qa_mcp/server.py마지막 명령이 MCP Inspector를 시작합니다.
도구를 열고 assess_release_readiness를 선택한 후 다음을 입력합니다:
{
"release_id": "REL-2026.08.1"
}예상되는 주요 결과:
{
"recommendation": "NO_GO",
"risk_score": 100,
"test_pass_rate_percent": 62.5,
"blockers": [
"Open SEV1 defect",
"Failed or blocked critical test",
"Failed or blocked high-criticality test"
]
}비교를 위해 REL-2026.08.2는 위험 점수 7로 GO를 반환합니다.
테스트 실행
전체 자동화 테스트 스위트를 실행합니다:
uv run pytest -q의존성 없는 핵심 검증을 실행합니다:
uv run python scripts/smoke_test.pyVersion 0.1에는 다음에 대한 테스트가 포함됩니다:
고위험 및 저위험 릴리스 권고
테스트 결과 필터링
회귀 테스트 계획 제한 및 우선순위 지정
잘못된 릴리스 식별자
설명 가능한 위험 점수 산정
점수는 100으로 상한이 적용됩니다:
25 × failed or blocked critical tests
12 × failed or blocked high-criticality tests
35 × open SEV1 defects
18 × open SEV2 defects
7 × open SEV3 defects
2 × open SEV4 defects오픈 상태의 SEV1 결함, 실패하거나 차단된 중요(critical) 테스트, 또는 실패하거나 차단된 고중요도(high-criticality) 테스트도 명시적 릴리스 차단 요인으로 보고됩니다.
이러한 가중치는 데모 정책일 뿐, 보편적인 금융 서비스 또는 소프트웨어 품질 표준이 아닙니다. 프로덕션에서는 임계값에 대해 적절한 위험 소유자의 승인, 버전 관리, 검증 및 정기 검토가 필요합니다.
보안 고려 사항
Version 0.1은 애플리케이션 계층에서 의도적으로 읽기 전용입니다. 프로덕션 구현에는 다음도 포함되어야 합니다:
인증 및 역할 기반 권한 부여
최소 권한 데이터베이스 및 서비스 역할
입력 및 출력 검증
도구 호출 및 권고에 대한 감사 로그
속도 제한 및 관측 가능성
비밀 관리 및 암호화된 전송
릴리스 결정에 대한 인간 승인
검색된 콘텐츠에 대한 프롬프트 인젝션 테스트
선택적 Snowflake 예제에는 샌드박스 데모 목적의 네이티브 SQL 실행 도구가 포함되어 있습니다. 전용 읽기 전용 역할을 통해 제한되어야 하며, 데모 이외의 용도로 사용하기 전에 범위를 더욱 좁혀야 합니다.
Version 0.2 로드맵
다음 버전은 이 릴리스 인텔리전스 기반을 AI 에이전트 보증 시스템으로 확장할 것입니다.
계획된 기능은 다음과 같습니다:
독창적인 AI 에이전트 평가 말뭉치
근거(grounding) 및 인용 검증
프롬프트 인젝션 저항성 테스트
개인정보 보호 및 데이터 최소화 검사
접근성 및 부정 경로 시나리오
기준선 대비 후보 비교
에이전트 버전 간 회귀 탐지
Playwright 기반 UI 및 접근성 실행
Snowflake 기반 평가 증거
인간 검토를 거친 AI 에이전트 릴리스 권고
프로젝트 상태
이 저장소는 교육용 포트폴리오 프로토타입입니다. 이는 프로덕션 뱅킹 시스템, 컴플라이언스 도구 또는 자율 릴리스 권한이 아닙니다.
참고 자료
라이선스
이 프로젝트는 MIT License에 따라 이용할 수 있습니다.
Available Tools
4 toolsassess_release_readinessB
Calculate an explainable GO, CONDITIONAL_GO, or NO_GO recommendation.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the output is an explainable recommendation with three possible values, but it does not reveal how the recommendation is derived, whether it depends on external sources, or what 'explainable' means in practice. This is acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and outcome with no filler. It is appropriately sized for a one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one parameter and no output schema, and the description names the three output categories, which covers the basic return shape. But it omits the criteria behind the recommendation, the source of the release ID, and any caveats, leaving the agent with an incomplete picture of how to invoke and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not elaborate on release_id beyond the schema's string type and title. Since the only parameter is central to the tool, the description should at least clarify what qualifies as a release_id and how it is used; it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear action ('Calculate') and a specific deliverable ('GO, CONDITIONAL_GO, or NO_GO recommendation'), which goes beyond the tool name. It is distinguishable from the sibling tools by its outcome-oriented purpose, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: this is the high-level readiness assessment tool, while siblings like get_failed_tests and find_defect_hotspots are lower-level diagnostic tools. However, the description never states when to use this tool versus its alternatives, so an agent must infer the boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_defect_hotspotsB
Rank release components by the weighted risk of unresolved defects.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It implies a read-only ranking operation and specifically scopes to unresolved defects, but it does not explain how 'weighted risk' is computed, whether historical data is considered, or what happens when no defects are found. Basic but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and object, then adds the precise qualifier. Every word earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, and the single parameter is simple. However, the description omits when to prefer this over sibling tools and does not clarify the meaning of 'components' or 'weighted risk.' It is minimally viable but leaves notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description should compensate, but it never explains what release_id means or how it relates to the ranking. The schema only shows it is a required string. The description uses 'release' in its wording, providing only a weak hint, not clear parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Rank'), a resource ('release components'), and a distinguishing criterion ('weighted risk of unresolved defects'). This clearly differentiates it from sibling tools like get_failed_tests or assess_release_readiness, which focus on different outputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus the sibling tools. It does not mention alternatives, exclusions, or conditions under which another tool would be a better fit, leaving the agent to infer usage purely from the name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_failed_testsB
Return failed and blocked tests, optionally filtered by criticality.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes | ||
| criticality | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the output (failed/blocked tests) but does not mention pagination, ordering, empty-result behavior, required release context, or consequences. Nothing contradicts annotations because none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word adds meaning, and the main result is stated before the optional filter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite a simple two-parameter shape and an output schema, the definition lacks enough context for confident invocation: no sibling differentiation, no release_id semantics, and no criticality value guidance. This is insufficient for a low-coverage schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only clarifies the optional criticality filter; it does not explain release_id or enumerate accepted criticality values, leaving a required parameter largely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Return failed and blocked tests'. This clearly distinguishes it from siblings like assess_release_readiness and recommend_regression_tests, which are analysis/recommendation tools rather than retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to prefer this tool over its siblings. The only usage hint is the optional criticality filter, which is more of a parameter option than a when-to-use instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_regression_testsC
Build a risk-based regression plan grounded in test and defect evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| max_tests | No | ||
| release_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It mentions that the plan is 'risk-based' and 'grounded in test and defect evidence,' but it does not disclose what the tool returns, how it uses release_id and max_tests, or whether it only reads data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It begins with the action and object and adds value by specifying risk-based and evidence-grounded characteristics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, no annotations, and no output schema, this one-line description is incomplete. It does not explain expected outputs, the role of max_tests, or selection criteria, leaving important context for correct invocation unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions release_id or max_tests. The phrase 'test and defect evidence' does not explain the required release parameter or the meaning of the max_tests default, so the agent gets no parameter help beyond field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action, 'Build a risk-based regression plan,' and a clear resource. It distinguishes itself from sibling tools by focusing on test recommendation and evidence grounding, though it does not explicitly name or contrast any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus assess_release_readiness, get_failed_tests, or find_defect_hotspots. There are no prerequisites or exclusions, so an agent must infer usage solely from the name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
assess_release_readiness - First observed
find_defect_hotspots - First observed
get_failed_tests - First observed
recommend_regression_tests
TDQS
Scored across 4 tools
Each tool produces a distinct output: a GO/NO-GO decision, a filtered list of test failures, a component risk ranking, and a regression test plan. find_defect_hotspots and recommend_regression_tests share an evidence base of defect/test risk, but their purposes are clearly separated by output type, so misselection is unlikely.
All four tools follow a consistent verb_noun snake_case pattern (assess_release_readiness, get_failed_tests, find_defect_hotspots, recommend_regression_tests). The verb clearly signals the action (assess, get, find, recommend) and the noun signals the resource, making the pattern highly predictable.
Four tools is on the lean side but well-scoped for release assurance: each tool fills a distinct role covering evidence gathering, risk analysis, planning, and final decision. There is no redundancy or bloat, and every tool earns its place in the pipeline.
The set forms a coherent end-to-end release readiness workflow: pull test failures, rank defect hotspots, build a regression plan from that evidence, and produce a final GO/NO-GO assessment. Minor gaps exist, such as no tool to drill into individual defect details or fetch component/change scope, but agents can work around these.
Maintenance
Related MCP Connectors
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
QA platform for agents: coverage signals, in-repo test plans, verified tests and release governance.
Diagnose AI workflows for failure, security, and handoff risks — RED/AMBER/GREEN per node.
Read-only, deterministic AI triage and readiness tools implementing Sophon's published rubrics.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables intelligent analysis of regression test failures and automatic discovery of solutions in JIRA. Analyzes test logs using AI-driven algorithms and matches errors with relevant JIRA issues through natural language interactions.-
- FlicenseBqualityBmaintenanceMCP server for AI-powered QA analysis. It enables analyzing test failures, identifying root causes, suggesting fixes, classifying defects, detecting flaky tests, and generating test cases and bug reports.10-
- AlicenseNot gradedqualityCmaintenanceEnables evidence-first release readiness assessment by running or accepting build, API, browser, visual, performance, and security evidence, then returning SHIP, REVIEW, or HOLD recommendations with clustered regressions.2 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables evaluating AI applications, inspecting reliability evidence, and gating releases from development and CI workflows.32Apache 2.0