AI Agent Release Assurance MCP
AI Agent Release Assurance MCP
Version 0.1: Grundlage der Release-Intelligence unter Verwendung synthetischer QA-Daten.
KI-Agenten-Evaluierungsfunktionen sind für Version 0.2 geplant.
Ein erklärbarer Model Context Protocol (MCP)-Server, der KI-Clients dabei hilft, Software-Testergebnisse und Fehler zu analysieren und evidenzbasierte Empfehlungen zur Release-Bereitschaft zu erstellen.
Alle Releases, Tests, Fehler und Szenarien mit Kundenauswirkung in diesem Repository sind fiktiv. Es werden keine Arbeitgeber-, Kunden-, Produktions-, personenbezogenen oder regulierten Daten verwendet.
Warum dieses Projekt existiert
Release-Entscheidungen erfordern häufig Belege, die über Testergebnisse, Fehlerdatensätze und Teamdokumentation verteilt sind.
Dieser Server bietet einem KI-Client eine kleine, schreibgeschützte Schnittstelle für Fragen wie:
Sollte ein bestimmtes Release ausgeliefert werden?
Welche fehlgeschlagenen Tests sind potenzielle Release-Blocker?
Wo konzentriert sich das ungelöste Fehlerrisiko?
Welche Tests sollten bei gezielten Regressionstests priorisiert werden?
Die KI erfindet den Risikowert nicht. Der Server berechnet ihn deterministisch und gibt die zugrunde liegenden Belege, Gewichtungen, Blocker und empfohlenen nächsten Schritte zur menschlichen Prüfung zurück.
Related MCP server: QA Copilot AI
Aktuelle Funktionen
Typ | Name | Zweck |
Tool |
| Gibt eine erklärbare Empfehlung ( |
Tool |
| Ruft fehlgeschlagene und blockierte Tests mit optionalem Kritikalitätsfilter ab |
Tool |
| Reiht Komponenten nach schwerwiegend gewichtetem, ungelöstem Fehlerrisiko |
Tool |
| Erstellt einen begrenzten, risikobasierten Regressionsplan |
Resource |
| Listet die für die Analyse verfügbaren synthetischen Releases auf |
Prompt |
| Führt durch eine evidenzbasierte Release-Bereitschaftsprüfung |
Architektur
flowchart TD
A[AI host or MCP Inspector] -->|MCP request| B[Python MCP server]
B --> C[QA service and risk rules]
C --> D[(Synthetic SQLite data)]
D --> C
C -->|Structured evidence| B
B -->|Tool result| A
D -. optional migration .-> E[(Snowflake)]SQLite hält Version 0.1 reproduzierbar und ohne Anmeldedaten. Die optionale Datei snowflake/setup.sql demonstriert einen möglichen Snowflake-nativen MCP-Pfad.
Schnellstart
Voraussetzungen
Python 3.10 oder neuer
Node.js/npm für den visuellen MCP Inspector
Installieren und ausführen
git clone https://github.com/Zoya-Ammar/ai-agent-release-assurance-mcp.git
cd ai-agent-release-assurance-mcp
uv sync --extra dev
uv run python -m banking_qa_mcp.seed
uv run mcp dev src/banking_qa_mcp/server.pyDer letzte Befehl startet den MCP Inspector.
Öffnen Sie Tools, wählen Sie assess_release_readiness, und geben Sie Folgendes ein:
{
"release_id": "REL-2026.08.1"
}Erwartetes Hauptergebnis:
{
"recommendation": "NO_GO",
"risk_score": 100,
"test_pass_rate_percent": 62.5,
"blockers": [
"Open SEV1 defect",
"Failed or blocked critical test",
"Failed or blocked high-criticality test"
]
}Zum Vergleich: REL-2026.08.2 liefert GO mit einem Risikowert von 100.
Tests ausführen
Führen Sie die vollständige automatisierte Testsuite aus:
uv run pytest -qFühren Sie die abhängigkeitsfreie Kernverifikation aus:
uv run python scripts/smoke_test.pyVersion 0.1 enthält Tests für:
Release-Empfehlungen mit hohem und niedrigerem Risiko
Filterung von Testergebnissen
Grenzen und Priorisierung von Regressionsplänen
Ungültige Release-Kennungen
Erklärbare Risikobewertung
Der Risikowert ist auf 100 begrenzt:
25 × failed or blocked critical tests
12 × failed or blocked high-criticality tests
35 × open SEV1 defects
18 × open SEV2 defects
7 × open SEV3 defects
2 × open SEV4 defectsEin offener SEV1-Fehler, ein fehlgeschlagener oder blockierter kritischer Test oder ein fehlgeschlagener oder blockierter Test mit hoher Kritikalität wird ebenfalls als expliziter Release-Blocker gemeldet.
Diese Gewichtungen sind eine Demonstrationsrichtlinie – kein allgemeingültiger Standard für Finanzdienstleistungen oder Softwarequalität. In der Produktion würden Schwellenwerte eine Genehmigung, Versionskontrolle, Validierung und regelmäßige Überprüfung durch die zuständigen Risikoverantwortlichen erfordern.
Sicherheitshinweise
Version 0.1 ist in der Anwendungsschicht bewusst schreibgeschützt. Eine Produktionsimplementierung sollte außerdem Folgendes enthalten:
Authentifizierung und rollenbasierte Autorisierung
Datenbank- und Diensterollen mit geringsten Rechten
Eingabe- und Ausgabevalidierung
Audit-Logs für Tool-Aufrufe und Empfehlungen
Rate-Limiting und Beobachtbarkeit
Geheimnisverwaltung und verschlüsselter Transport
Freigabe durch Menschen für Release-Entscheidungen
Prompt-Injection-Tests für abgerufene Inhalte
Das optionale Snowflake-Beispiel enthält ein natives Tool zur SQL-Ausführung für Sandbox-Demonstrationszwecke. Es sollte über eine dedizierte schreibgeschützte Rolle eingeschränkt und vor jeder Nicht-Demo-Nutzung weiter eingegrenzt werden.
Roadmap für Version 0.2
Die nächste Version erweitert diese Release-Intelligence-Grundlage zu einem KI-Agenten-Assurance-System.
Geplante Funktionen umfassen:
Einen originären KI-Agenten-Evaluierungskorpus
Validierung von Grounding und Zitierungen
Tests zur Prompt-Injection-Resistenz
Prüfungen zu Datenschutz und Datenminimierung
Szenarien für Barrierefreiheit und Negativpfade
Vergleiche zwischen Baseline und Kandidat
Erkennung von Regressionen zwischen Agentenversionen
Playwright-basierte UI- und Barrierefreiheitsausführung
Snowflake-gestützte Evaluierungsnachweise
Von Menschen geprüfte KI-Agenten-Release-Empfehlungen
Projektstatus
Dieses Repository ist ein pädagogischer Portfolio-Prototyp. Es ist kein Produktionsbanksystem, kein Compliance-Tool und keine autonome Freigabeinstanz für Releases.
Referenzen
Lizenz
Dieses Projekt ist unter der MIT-Lizenz verfügbar.
Available Tools
4 toolsassess_release_readinessB
Calculate an explainable GO, CONDITIONAL_GO, or NO_GO recommendation.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the output is an explainable recommendation with three possible values, but it does not reveal how the recommendation is derived, whether it depends on external sources, or what 'explainable' means in practice. This is acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and outcome with no filler. It is appropriately sized for a one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one parameter and no output schema, and the description names the three output categories, which covers the basic return shape. But it omits the criteria behind the recommendation, the source of the release ID, and any caveats, leaving the agent with an incomplete picture of how to invoke and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not elaborate on release_id beyond the schema's string type and title. Since the only parameter is central to the tool, the description should at least clarify what qualifies as a release_id and how it is used; it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear action ('Calculate') and a specific deliverable ('GO, CONDITIONAL_GO, or NO_GO recommendation'), which goes beyond the tool name. It is distinguishable from the sibling tools by its outcome-oriented purpose, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: this is the high-level readiness assessment tool, while siblings like get_failed_tests and find_defect_hotspots are lower-level diagnostic tools. However, the description never states when to use this tool versus its alternatives, so an agent must infer the boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_defect_hotspotsB
Rank release components by the weighted risk of unresolved defects.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It implies a read-only ranking operation and specifically scopes to unresolved defects, but it does not explain how 'weighted risk' is computed, whether historical data is considered, or what happens when no defects are found. Basic but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and object, then adds the precise qualifier. Every word earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, and the single parameter is simple. However, the description omits when to prefer this over sibling tools and does not clarify the meaning of 'components' or 'weighted risk.' It is minimally viable but leaves notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description should compensate, but it never explains what release_id means or how it relates to the ranking. The schema only shows it is a required string. The description uses 'release' in its wording, providing only a weak hint, not clear parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Rank'), a resource ('release components'), and a distinguishing criterion ('weighted risk of unresolved defects'). This clearly differentiates it from sibling tools like get_failed_tests or assess_release_readiness, which focus on different outputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus the sibling tools. It does not mention alternatives, exclusions, or conditions under which another tool would be a better fit, leaving the agent to infer usage purely from the name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_failed_testsB
Return failed and blocked tests, optionally filtered by criticality.
| Name | Required | Description | Default |
|---|---|---|---|
| release_id | Yes | ||
| criticality | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the output (failed/blocked tests) but does not mention pagination, ordering, empty-result behavior, required release context, or consequences. Nothing contradicts annotations because none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word adds meaning, and the main result is stated before the optional filter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite a simple two-parameter shape and an output schema, the definition lacks enough context for confident invocation: no sibling differentiation, no release_id semantics, and no criticality value guidance. This is insufficient for a low-coverage schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only clarifies the optional criticality filter; it does not explain release_id or enumerate accepted criticality values, leaving a required parameter largely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Return failed and blocked tests'. This clearly distinguishes it from siblings like assess_release_readiness and recommend_regression_tests, which are analysis/recommendation tools rather than retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to prefer this tool over its siblings. The only usage hint is the optional criticality filter, which is more of a parameter option than a when-to-use instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_regression_testsC
Build a risk-based regression plan grounded in test and defect evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| max_tests | No | ||
| release_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It mentions that the plan is 'risk-based' and 'grounded in test and defect evidence,' but it does not disclose what the tool returns, how it uses release_id and max_tests, or whether it only reads data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It begins with the action and object and adds value by specifying risk-based and evidence-grounded characteristics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, no annotations, and no output schema, this one-line description is incomplete. It does not explain expected outputs, the role of max_tests, or selection criteria, leaving important context for correct invocation unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions release_id or max_tests. The phrase 'test and defect evidence' does not explain the required release parameter or the meaning of the max_tests default, so the agent gets no parameter help beyond field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action, 'Build a risk-based regression plan,' and a clear resource. It distinguishes itself from sibling tools by focusing on test recommendation and evidence grounding, though it does not explicitly name or contrast any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus assess_release_readiness, get_failed_tests, or find_defect_hotspots. There are no prerequisites or exclusions, so an agent must infer usage solely from the name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
assess_release_readiness - First observed
find_defect_hotspots - First observed
get_failed_tests - First observed
recommend_regression_tests
TDQS
Scored across 4 tools
Each tool produces a distinct output: a GO/NO-GO decision, a filtered list of test failures, a component risk ranking, and a regression test plan. find_defect_hotspots and recommend_regression_tests share an evidence base of defect/test risk, but their purposes are clearly separated by output type, so misselection is unlikely.
All four tools follow a consistent verb_noun snake_case pattern (assess_release_readiness, get_failed_tests, find_defect_hotspots, recommend_regression_tests). The verb clearly signals the action (assess, get, find, recommend) and the noun signals the resource, making the pattern highly predictable.
Four tools is on the lean side but well-scoped for release assurance: each tool fills a distinct role covering evidence gathering, risk analysis, planning, and final decision. There is no redundancy or bloat, and every tool earns its place in the pipeline.
The set forms a coherent end-to-end release readiness workflow: pull test failures, rank defect hotspots, build a regression plan from that evidence, and produce a final GO/NO-GO assessment. Minor gaps exist, such as no tool to drill into individual defect details or fetch component/change scope, but agents can work around these.
Maintenance
Related MCP Connectors
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
QA platform for agents: coverage signals, in-repo test plans, verified tests and release governance.
Diagnose AI workflows for failure, security, and handoff risks — RED/AMBER/GREEN per node.
Read-only, deterministic AI triage and readiness tools implementing Sophon's published rubrics.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables intelligent analysis of regression test failures and automatic discovery of solutions in JIRA. Analyzes test logs using AI-driven algorithms and matches errors with relevant JIRA issues through natural language interactions.-
- FlicenseBqualityBmaintenanceMCP server for AI-powered QA analysis. It enables analyzing test failures, identifying root causes, suggesting fixes, classifying defects, detecting flaky tests, and generating test cases and bug reports.10-
- AlicenseNot gradedqualityCmaintenanceEnables evidence-first release readiness assessment by running or accepting build, API, browser, visual, performance, and security evidence, then returning SHIP, REVIEW, or HOLD recommendations with clustered regressions.2 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables evaluating AI applications, inspecting reliability evidence, and gating releases from development and CI workflows.32Apache 2.0