TestCraft
Exports generated BDD scenarios and step definitions as Gherkin .feature files for use with Cucumber/Behave.
Uses local Ollama models as a provider to generate test design packages from acceptance criteria and descriptions.
Uses OpenAI-compatible model providers to generate test design packages from acceptance criteria and descriptions.
Exports generated tests as parametrized Python test modules using pytest.mark.parametrize.
Exports generated test suites as standard .robot files for Robot Framework.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@TestCraftGenerate pairwise test cases for login form fields"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TestCraft Engine
TestCraft Engine is a Python 3.9+ library for deterministic constrained test-design calculations and traceability diagnostics. The base installation contains only Pydantic. LangGraph, model providers, and dotenv loading are optional.
What the library guarantees
The pairwise and N-way APIs validate their complete input before generating cases. A report is complete only when every feasible target has a generated full assignment. A target that cannot be extended to any assignment satisfying all constraints is reported in infeasible_targets; it is not silently counted as covered.
The legacy functions remain available:
from testcraft.combinatorics import generate_pairwise_combinations
combinations = generate_pairwise_combinations(
{"os": ["Windows", "Linux"], "browser": ["Chrome", "Firefox"]},
constraints=[{"os": "Windows", "browser": "Firefox"}],
)Use the structured API when diagnostics are needed:
from testcraft.combinatorics import generate_pairwise_report
report = generate_pairwise_report(
{"os": ["Windows", "Linux"], "browser": ["Chrome", "Firefox"]},
constraints=[{"os": "Windows", "browser": "Firefox"}],
)
print(report.is_complete)
print(report.coverage_percentage)
print(report.infeasible_targets)
print(report.to_dict())generate_nway_report accepts n_way=3 and higher. If the requested arity is larger than the number of parameters, the effective arity is the number of parameters and the report says so in diagnostics. Exact enumeration and a minimum-size proof are attempted only for bounded small spaces. optimal is true only when that proof completed; otherwise it is None and the strategy is reported. A greedy fallback is still checked for complete feasible coverage, but is not described as optimal.
Constraints are prohibited conjunctions. A constraint may mention one parameter or a high-arity subset. For a lower-arity target, at least one complete feasible continuation is required before the target is considered feasible.
Related MCP server: mcp-testing-tools
Other combinatorics
analyze_boundary_values accepts [min, max] ranges or mappings with min, max, optional step, and explicit inclusive_min and inclusive_max flags. Values must be finite real numbers; booleans, strings, non-finite values, invalid ranges, and invalid steps are rejected. The report separates valid and invalid points according to the declared step and endpoint semantics.
generate_decision_table validates condition and action schemas. Multiple actions or more than three conditions require explicit rules so that mutually exclusive business actions are not invented. A small synthetic table is marked is_complete=false; the CLI requires --allow-unresolved before persisting or returning success for it. allow_unresolved=True returns a clearly marked condition-only diagnostic table; it is not a completed business decision table.
Traceability and pipeline
calculate_traceability_matrix returns the historical AC-to-TC mapping and coverage statistics. The statistics also contain tc_to_ac, duplicate diagnostics, unknown references, and an explicit coverage_ok result. calculate_bidirectional_traceability returns both directions directly.
The pipeline has two execution backends with the same node functions:
LangGraph when the
aiextra is installed and the backend imports successfully.A dependency-free sequential runner otherwise.
Both backends use the same coverage gate. A package is complete only when link coverage passes and no validation or generation errors remain. Link coverage means that every criterion ID is referenced by at least one test case. Semantic coverage is not inferred from an ID reference and is reported as not_assessed unless an external process supplies that assessment.
Without a model, the pipeline creates a deterministic draft test case for each criterion. This provides a measurable link-coverage draft, not a semantic quality guarantee. A requested model is not silently replaced by a fallback: partial JSON, unknown references, duplicate IDs, invalid fields, or missing criteria produce diagnostics and an incomplete package.
Each BDD scenario has covers_ac links. Feature serialization strips control characters, normalizes tags, and keeps scenario steps on single Gherkin lines. The generated feature also emits safe @ac-... tags when possible.
Installation
pip install .
pip install .[dev]Optional extras are separate:
pip install .[ai]
pip install .[openai]
pip install .[ollama]
pip install .[env]
pip install ".[ai,openai,env]"The ai extra installs LangGraph and core LangChain support, openai installs the OpenAI-compatible provider used by OpenAI, DeepSeek, and custom endpoints, and ollama installs local model support. The dev extra does not install model or LangGraph packages, so the core test suite is intended to run without them.
Command line
testcraft --version
testcraft pairwise parameters.json --constraints constraints.json --output pairwise.json --json
testcraft bva ranges.json --output boundaries.json --json
testcraft decision-table table.json --output table-result.json --json
testcraft design criteria.json --title "Example Suite" --output package.json --gherkin feature.feature --jsondecision-table accepts an object containing conditions, actions, and optional rules. Constraints passed to pairwise must exist; a missing constraints file is an error. Existing output files are protected unless --force is supplied. Output and Gherkin paths must be different. Writes use a same-directory temporary file followed by an atomic replace. JSON input is read as UTF-8 with optional BOM support and is size-limited.
The design command can use an explicitly requested model:
testcraft --env-file .env design criteria.json --llm --provider openai --output package.jsonEnvironment files are not loaded at import time. --env-file is explicit. Provider keys are isolated: an OpenAI key is not reused for DeepSeek or a custom OpenAI-compatible endpoint. DeepSeek uses DEEPSEEK_API_KEY, custom uses CUSTOM_API_KEY, and Ollama uses its host configuration. Unknown providers raise ValueError. Timeout, token, retry, and Ollama-specific options are available through the Python factory.
When a model is used, acceptance criteria and descriptions are sent to the selected external provider. The CLI enables best-effort redaction of common API keys, bearer tokens, password assignments, and private-key blocks before building the architect prompt. This redaction is not a complete data-loss-prevention system; review sensitive data before enabling an external provider. No API key is placed in package provenance. When available, provenance records provider, model, backend, and whether redaction was requested.
Exit codes and diagnostics
0: command completed and any required coverage gate passed.1: invalid input, JSON, model configuration, missing file, or I/O error.2: command-line usage error from argparse.3: design package is incomplete or its link-coverage gate failed, or an unresolved decision table was rejected.4: output-path collision or overwrite-protection failure.
Use --json to receive machine-readable results and errors. The design command always returns a non-zero incomplete status by default and has no success bypass for a failed link-coverage gate.
Pluggable Combinatorial Backends & Zero-Trust Audit
TestCraft supports pluggable generation backends:
BuiltinPairwiseBackend: Fast, pure-Python deterministic solver (zero dependencies).CoverTableBackend: Adapter for CoverTable AETG algorithm (pip install testcraft-engine[covertable]).PictBackend: Adapter for Microsoft PICT CLI (pict.exe/pict).
Regardless of backend, all generated suites are validated by an independent Zero-Trust auditor:
from testcraft.combinatorics import verify_pairwise_coverage
audit = verify_pairwise_coverage(parameters, combinations, n_way=2, constraints=constraints)
print(audit.is_valid) # True if 100% feasible tuples covered & 0 constraints violated
print(audit.coverage_percentage) # Exact mathematical coverageBDD & Test Execution Exporters
TestCraft is an orchestration layer, exporting directly to test runners:
Behave / Cucumber:
package.to_behave()(Gherkin.feature) andpackage.to_behave_bundle()(includes Python step definitions).Robot Framework:
package.to_robot()(standard.robottest suites).Pytest:
package.to_pytest()(parametrized test module using@pytest.mark.parametrize).
Model Context Protocol (MCP)
TestCraft provides native MCP tools for AI agents:
generate_pairwise: Combinatorial N-way / pairwise test generator.analyze_bva: Valid and invalid boundary value analyzer.build_traceability: Bidirectional matrix and gate calculator.design_test_package: End-to-end requirement-to-test-suite pipeline.
Run as an MCP tool provider:
python -m testcraft.mcp.serverDevelopment
ruff check src tests examples
mypy
pytest -p no:cacheprovider
pytest --cov=testcraft --cov-report=term-missing -p no:cacheproviderThe project keeps a src layout and supports Python 3.9 and newer. The included CI workflow runs core tests on Linux and Windows across Python 3.9 and newer, and performs a package build/install smoke test where the environment permits it.
License
MIT. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
100+ MCP tools for AI agents: content metadata, trade intelligence, business-expertise analysis.
Official MCP server for Qase — manage test cases, runs, suites, defects via AI tools.
AI workflow/MCP implementation package planner.
One MCP tool for verified AI-agent outcomes with success-only charging.
Related MCP Servers
- FlicenseAqualityDmaintenanceAn intelligent MCP toolset for software testers that monitors code changes, analyzes test impact, recommends tests, and assesses risk, supporting multiple AI coding frameworks via dual transport modes.6-
- AlicenseAqualityDmaintenanceProvides testing and quality assurance tools for AI agents via MCP, enabling generation of test cases, mock data, API mocks, coverage analysis, and assertions.540 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables generation of test cases, edge cases, and test matrices for software testing, integrated with MCP protocol and EU AI Act compliance.3 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables traceable requirement discovery, technical alignment, and ISO-aligned process checking through deterministic MCP tools and resources, without requiring an embedded LLM.1Apache 2.0