Skip to main content
Glama

The-QA-Skill

The AI-native QA operating system for coding agents.

Turns Claude Code, Cursor, Copilot, Windsurf, Codex, Cline, Gemini CLI, and Zed into an elite, evidence-driven QA organization — deterministic engines first, AI only where semantics demand it.

Sponsored by xShredo.dev CI License: MIT Node

qa impact — smallest high-confidence test set for this PR · qa triage — 87 failures → real defects vs flakes vs env · qa heal — repair without weakening intent, with confidence


Sponsored by xShredo.dev — opening the link supports this project.

Why

Coding agents are great at writing tests and terrible at doing QA. They jump from "user asked for tests" straight to Playwright code, with no risk model, no selection discipline, no triage, no evidence, and no way to say "I verified this" honestly. The-QA-Skill is the missing operating layer: 28 agent-readable skills (Layer A) backed by a deterministic execution platform (Layer B) with a documented risk algorithm, a 12-category failure classifier, confidence-tiered self-healing, evidence bundles, a quality graph, and a benchmark corpus that measures all of it — 13 broken apps with expected diagnoses, graded honestly.

Related MCP server: rush

The lifecycle (every QA task, no exceptions)

DISCOVER → MODEL → PLAN → GENERATE → VALIDATE → EXECUTE →
OBSERVE → TRIAGE → HEAL/FIX → VERIFY → MEASURE → LEARN

The LifecycleTracker mechanically refuses phase-skipping: you cannot EXECUTE without PLAN, cannot HEAL without TRIAGE, cannot MEASURE without VERIFY. "User asked for tests → write Playwright code" is now an exception, not a behavior.

Killer commands

Command

What it actually does

qa impact --range HEAD~1..HEAD

Parses the diff, classifies areas, builds the import graph, selects the smallest high-confidence test set — every selected test carries explicit reasons; CSS-only changes never trigger API suites

qa risk

8-factor explainable risk engine (0–100): businessCriticality ×0.20, changeSurface ×0.15, userImpact ×0.15, defectHistory ×0.12, integrationDepth ×0.12, codeComplexity ×0.10, securitySensitivity ×0.08, dataSensitivity ×0.08 + documented tier floors (payment/migration ≥ 0.9 → critical) — always prints + payment logic changed style reasons

qa triage --evidence failures.json

Classifies failures into 12 categories with confidence, signals, contradicting signals, root-cause hypothesis, and recommended action; clusters 87 failures into signatures (primary vs cascade)

qa heal --evidence … --apply --confirm-risk

HIGH tier (pure selector swap, observed in DOM, assertions untouched) auto-applies with .pre-heal.bak backup; MEDIUM proposes; LOW never touches. Never weakens assertions, never raises timeouts, never deletes tests

qa release

PASS / PASS_WITH_WARNINGS / BLOCKED / FAIL / UNKNOWN — a subset passing is never PASS; empty evidence is never PASS

qa doctor

12 health checks (node, git, playwright, browsers, ports, env, secrets hygiene…) → score 0–100 with fixes

Plus: init · discover · plan · generate · review · test · flake · coverage · report · explain — all with --json --quiet --verbose --dry-run where applicable, machine envelope {schemaVersion, command, ok, data, label}.

Architecture

Layer A — skills the agent reads          Layer B — the execution platform
┌────────────────────────────┐            ┌─────────────────────────────────┐
│ skills/ (28 SKILL.md)      │            │ packages/core     engines       │
│  qa-enterprise (entry)     │──calls──▶  │ packages/agents   8 + orchestr. │
│  qa-orchestrator (govern)  │            │ packages/runners   pw/vitest/py  │
│  qa-failure-triage         │            │ packages/healing  tiered heals  │
│  qa-llm-testing (flagship) │            │ packages/graph    traceability  │
│  qa-agent-evaluation       │            │ packages/reporting 4 audiences  │
│  …24 more                  │            │ packages/data      factories    │
└────────────────────────────┘            │ packages/reasoning multi-model  │
                                          │ packages/mcp-server 11 MCP tools│
                                          │ packages/cli        16 commands │
                                          └─────────────────────────────────┘
  • Deterministic first: parsers → analyzers → matchers → only then optional LLMs via the provider-agnostic ReasoningProvider (OpenAI / Anthropic / Gemini / local behind one interface; the default provider is deterministic and always configured).

  • Explainability everywhere: every conclusion carries {result, confidence, evidence[], assumptions[], fallbackUsed} and a verification label — NOT_VERIFIED / NOT_RUN / INFERRED / OBSERVED / CONFIRMED.

  • Safe automation: every action classified READ_ONLY / LOW_RISK_WRITE / HIGH_RISK; HIGH_RISK requires explicit --confirm-risk.

  • No AI magic: "classified FLAKE (0.88) because: failed then passed on retry, no relevant change, historically intermittent" — never "AI thinks it's probably flaky".

The quality graph

Requirement → Feature → Code → API → UI → Test → Execution → Evidence → Defect
Commit → Changed file → Symbol → Affected feature → Risk → Relevant tests

(packages/graph) — traceability queries: traceRequirement, affectedTests, GraphViz export, JSON persistence. See docs/quality-graph.md.

MCP server (agent-native)

npm run mcp starts a dependency-free stdio MCP server exposing 11 tools: discover_project, analyze_risk, list_relevant_tests, generate_tests, run_tests, get_failure_evidence, triage_failure, propose_test_heal, analyze_flake, generate_quality_report, evaluate_release. Planning tools never write; run_tests defaults to dry-run; healing tools propose, never apply. (docs/architecture.md)

Benchmark (the honest moat)

13 realistic broken apps — auth-bug, flaky-test, selector-change, api-regression, db-regression, a11y-regression, payment-regression, visual-regression, race-condition, missing-coverage, weak-assertion, false-positive, false-negative — each with materializable git history, recorded failure evidence, and machine-checkable expected diagnoses.

npm run benchmark        # → benchmarks/agentic-qa/results.md

Latest measured run on this repo (deterministic engines, no LLM): triage 18/18 · quality 9/9 · coverage 1/1 · risk 1/1 expectation checks green — reproducible byte-for-byte. Numbers come only from the harness on this repo's fixtures; we publish misses, never invented ones. Methodology: docs/benchmark.md.

Quick start

# in your project repo
npx the-qa-skill init          # or clone this repo and npm install && npm run build
qa doctor                      # 12 checks + fixes
qa discover                    # stack, tests, config inventory
qa impact --range HEAD~1..HEAD # smallest high-confidence test set for your change
qa test --policy pr            # execute with evidence bundles
qa triage                      # classify what failed, with proof
qa release                     # honest gate verdict

For agents: read skills/qa-enterprise/SKILL.md first — it routes everything else. For MCP clients, register packages/mcp-server (stdio).

Documentation

Architecture · Agent model · Quality model · Risk engine · Test selection · Failure triage · Self-healing · Quality graph · Governance · CI integration · Security · Benchmark · Gap analysis · Migration map

The 15 golden rules (embedded in the orchestrator)

  1. Never weaken an assertion to make a test pass. (mechanically enforced in healing)

  2. Never hide a regression behind retries. (enforced in triage)

  3. Never use arbitrary sleeps as first fix. (penalized in quality scoring)

  4. Never claim verification without execution evidence. (enforced by labels + gate)

  5. Never generate massive redundant E2E suites. (pyramid demotion in selection)

  6. Never destroy test isolation for speed.

  7. Never auto-delete tests without strong evidence. (deletion always requires human approval)

  8. Never expose secrets in test artifacts. (scrubbed before write)

  9. Never perform destructive production actions without authorization. (--confirm-risk)

  10. Always distinguish product defects from test defects. (triage categories)

  11. Prefer the smallest test that catches the defect. (selection ranking)

  12. Preserve reproducibility. (evidence metadata: commit, seed, env)

  13. Preserve evidence. (append-only bundles + .pre-heal.bak)

  14. Explain important decisions. (ExplainableConclusion everywhere)

  15. Optimize for signal, not test count.

Project status

v0.1.0 — full Layer A + Layer B build-out, 480+ tests green, benchmark corpus measured. Semantic versioning; see docs/migration-map.md for capability parity guarantees and CHANGELOG.md for history.

Contributing

Issues and PRs welcome — see CONTRIBUTING.md. Every new engine must ship with tests, docs, and a fixture expectation; every AI-adjacent feature must expose its evidence, confidence, and fallback.

License

MIT © The-QA-Skill contributors

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides MCP tools that give LLM agents a full QA engineer workflow: scanning projects, generating deterministic test suites, executing them across browser/API/mobile, diagnosing failures, and proposing fixes that require human approval.
    -
  • A
    license
    C
    quality
    A
    maintenance
    Agentic code-quality CLI and stdio MCP server that runs real Python and JS/TS quality engines (Ruff, pytest, ESLint, Prettier, Vitest, etc.) for linting, testing, and reviewing code.
    79
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Provides coding agents with visibility into test health through tools for flaky test detection, test quality linting, and LLM evaluation harness, enabling them to triage failures, review test quality, and check prompt changes for regressions.
    9
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI clients to assess infrastructure-as-code deployment risk, generate incident response runbooks, review CI/CD delivery readiness, and calculate SLO error budgets through structured advisory tools. Runs locally over stdio without cloud credentials or external API calls, returning guidance for platform work, pull request review, and production readiness checks.
    12
    2,512 npm
    MIT