ai-delivery
by Danzxz
README.md
# AI Delivery
**A personal design for AI-native software delivery in organizations:** let agents move work forward, while people keep authority over intent, risk, and release.
The missing layer is shared delivery memory. A coding agent can produce a patch, but an organization needs every agent and person to work from the same approved intent, hand work across roles without losing context, and show evidence tied to the exact work and source revision under review. A green check is useful only when the team can tell what it checked, against which revision, in which environment, and under whose authority.
AI Delivery is an original, small reference implementation of that idea. It was informed by public work on AI-driven development lifecycles, agent workflows, MCP, and reusable agent skills. It does not implement or claim conformance with any one organization's development method. The AI-DLC references below are inspiration and background, not dependencies.
## What exists today
This repository contains a runnable Python package, an offline demonstration, and an MCP server using one shared service layer. The implemented lifecycle includes:
- versioned work items with compare-and-swap updates (`expected_revision`);
- role-tagged handoffs tied to a work revision;
- append-only evidence tied to a work revision, actual Git SHA-1/SHA-256, and named environment;
- a QA gate with mandatory `implementation` and `qa` evidence, latest-result semantics, work-status checks, and stale-evidence rejection;
- a separate release-owner attestation that requires the QA gate to pass for approval;
- local JSON persistence with atomic replacement and POSIX file locking on macOS and Linux; Windows is not supported by this file adapter.
A `decided_by` value is an audit claim supplied by the caller. This local reference server does not authenticate users or enforce organizational roles. It has no deployment integration, so recording an approval never deploys software. The JSON adapter is suitable for a local demonstration; it is not a network service, distributed store, or complete enterprise control plane.
## Run it
Python 3.11 or newer is required.
```bash
python -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[dev]'
ai-delivery demo --store /tmp/ai-delivery-demo.json
pytest
```
The demo creates a fictional reading-list export item, shows a closed gate before evidence exists, records implementation and QA evidence for one source revision in `qa`, and records a separate release-owner attestation. The demo makes no network calls and performs no deployment.
## Connect an MCP client
Install the package in an environment available to the MCP host, then launch the stdio server:
```bash
AI_DELIVERY_STORE=/path/to/your/delivery.json ai-delivery-mcp
```
A client configuration commonly looks like this; exact configuration keys depend on the client:
```json
{
"mcpServers": {
"ai-delivery": {
"command": "python",
"args": ["-m", "ai_delivery.mcp_server"],
"env": {
"AI_DELIVERY_STORE": "/path/to/your/delivery.json"
}
}
}
}
```
The server currently exposes eight tools: `list_work_items`, `get_context`, `create_work_item`, `update_work_item`, `handoff_work`, `record_evidence`, `evaluate_qa_gate`, and `record_release_decision`. Every update, handoff, evidence append, and release decision names the expected work revision. A stale write is rejected, so the caller must reread context before proceeding. The server exposes tools only; MCP resources and prompts are possible future interfaces, not implemented features.
For a real organization, put the MCP endpoint behind authenticated identity, authorization, audit, and network controls. Give each role the narrow operations it needs: product can draft intent, engineering can create implementation revisions, QA can record test evidence, and an authorized human release owner can attest to a release decision. Those boundaries are design goals here; this demo accepts caller-supplied role labels and does not enforce them.
## A delivery memory with three layers
AI Delivery separates durable authority from agent convenience. This is a design proposal for an organization-scale implementation; only versioned work items, handoffs, evidence, and the gate are present in this repository today.
| Layer | What belongs there | Authority and freshness |
| --- | --- | --- |
| **Canonical delivery record** | Approved specs, ADRs, work-item revisions, handoffs, evidence manifests, gate results, and human decisions | Authoritative. Each record has an immutable revision or identifier, actor, time, parent references, and content/source digest. A change creates a new revision; it does not silently rewrite history. |
| **Working agent memory** | Current assignment, relevant context excerpts, decisions still in progress, and next actions | A disposable, task-scoped cache. It carries source references and the revision/hash it was read from. Before a write or handoff, the agent refreshes if that source has advanced. It cannot override the canonical record. |
| **Derived search index** | Embeddings, keyword indexes, summaries, and cross-project discovery metadata | Rebuildable projection. Every indexed chunk points to its source record and exact revision/hash. Results are hints; the reader fetches and validates the canonical source before acting. |
In the proposed Git-backed mode, reviewed specifications, ADRs, and work-item manifests live in version control. Evidence can be a small signed or immutable manifest in the repository, with large reports and logs in an approved artifact store addressed by digest. A portable transactional store can serve the same record contract where Git is not the right persistence layer. Git commits make changes reviewable; a storage adapter provides queries and concurrent writes. Neither an index nor an agent's notes becomes the source of truth.
Freshness is explicit. A handoff names the work revision. Evidence names both the work revision and source commit digest, plus the environment and outcome. A later revision does not inherit earlier evidence. A search result or working-memory note is accepted only if its source revision still matches the canonical record. Retention is policy-driven: preserve decisions and evidence for the required audit period, expire disposable agent memory quickly, and rebuild or delete derived indexes according to their source and retention policy.
Credentials, tokens, private keys, and personal data do not belong in Git specs, agent memory, evidence summaries, or search indexes. A production adapter should use a secrets manager or short-lived identity exchange and store only a reference to the secret or credential event where audit requires one. Access to records and tools should follow least privilege, with environment and project boundaries enforced by the service rather than by skill instructions alone.
## Gates and role skills
The proposed lifecycle uses five review points. Each has a small `SKILL.md` example under [`skills/gates`](skills/gates/README.md), following the public Agent Skills file format. Skills tell an agent how to prepare work and evidence; a trusted service and an authorized person must enforce decisions. A skill cannot grant access or prove that a human approved something.
| Gate | Agent prepares | Required exit evidence | Status here |
| --- | --- | --- | --- |
| **Inception** | Intent, users, constraints, success measures, assumptions, and a small set of work items | Human-confirmed scope and acceptance criteria | Proposed skill; no inception gate in the package |
| **Design / ready for development** | Design notes, ADR proposals, dependencies, risks, and implementation slices | Reviewed design where needed; accepted ready-for-dev revision | Proposed skill; current package can version the spec but has no design approval policy |
| **Implementation / review** | Code changes, test results, source commit, and review response | Implementation evidence tied to exact work revision and source hash; review disposition | Package can record implementation evidence; it does not run tests or perform code review |
| **QA** | Test plan, environment, scenario results, and failures | Passing QA evidence for the same work revision and source hash | Implemented as a simple required-evidence gate |
| **Release / post-release** | Release notes, rollout and rollback plan, monitoring signals, and post-release observations | Human release decision; later, deployment and post-release evidence | Package records a separate release attestation; it has no deployment or post-release gate |
The hard-coded minimum gate policy requires the most recent relevant result for both `implementation` and `qa` to be `passed`, the current work status to be `implemented` or `qa_ready`, and all evidence to match the current revision, source hash, and QA environment. A later failure blocks the gate until a newer passing result is recorded. Callers may add evidence kinds; they cannot remove the two minimum kinds. The local service does not establish who is allowed to set a work status or write a result.
## Walkthrough
1. Product records a concise outcome and acceptance criteria as revision 1.
2. The work is handed to engineering with a role-tagged note that points to revision 1.
3. Engineering creates revision 2, records a source commit hash, and moves it to a testable status.
4. Engineering and QA record results against revision 2, the same source hash, and `qa`.
5. The service evaluates the gate. Missing, stale, wrong-environment, or failed latest evidence keeps it closed.
6. A release owner reviews the evidence and records a decision for the target environment. This record is separate from QA and does not deploy.
7. In a future operations stage, deployment and post-release observations would be recorded against the released source revision and environment, then fed into future work as linked evidence and new requirements.
If a new work revision is created, evidence from the old revision remains available for audit but does not satisfy the current gate. If two writers update from the same revision, the second write sees the new revision and fails with a conflict rather than replacing the first writer's work.
## Architecture
```mermaid
flowchart LR
Product[Product agent / human] --> Canonical[Versioned delivery records]
Engineering[Engineering agent] --> Canonical
QA[QA agent] --> Canonical
Owner[Human release owner] --> Canonical
Canonical --> Service[Policy and revision service]
Service --> MCP[MCP tools]
MCP --> Agents[Compatible agent clients]
Canonical --> Index[Derived search index]
Index -->|source refs and revision checks| Service
Service --> Gate[QA evidence gate]
Gate -->|passing QA evidence| HumanDecision[Separate release decision]
HumanDecision -->|future integration| Operations[Deployment and post-release evidence]
```
The current implementation is deliberately smaller: one Python service, a local JSON adapter, an offline CLI, and an MCP stdio adapter. It does not contain the Git canonical backend, transactional database, search index, authenticated role policy, skill runner, CI/deployment connector, or post-release observation pipeline shown as future design.
## Design principles and public references
This project synthesizes a few useful public ideas rather than claiming a new standard:
- AWS's [AI-Driven Development Life Cycle overview](https://aws.amazon.com/blogs/devops/ai-driven-development-life-cycle/) describes Inception, Construction, and Operations, with AI planning and execution under human oversight. The [AWS open-source adaptive workflow](https://github.com/awslabs/aidlc-workflows) is a separate, evolving implementation of that methodology.
- Anthropic's [Building Effective Agents](https://www.anthropic.com/engineering/building-effective-agents) recommends simple composable patterns and describes where fixed workflows, gates, and more autonomous agents fit. AI Delivery applies that spirit to explicit revision and evidence contracts.
- The [MCP specification](https://modelcontextprotocol.io/specification/2025-11-25) defines a protocol for exposing tools, resources, and prompts. AI Delivery currently uses MCP tools as the portable action interface; transport and authorization remain separate concerns.
- The [Agent Skills specification](https://agentskills.io/specification) defines a portable folder format centered on `SKILL.md`, with optional scripts, references, and assets. The example gate skills here are workflow instructions, not policy enforcement.
## Roadmap
The code-backed MVP is the revision, handoff, evidence, QA gate, release-attestation, CLI, MCP, and local-store flow. The next design steps are a Git-backed canonical adapter, a proper multi-user transactional store, authenticated role authorization, policy configuration owned by the service, review and post-release evidence schemas, and integration tests across concurrent writers. Search indexing and deployment should follow only when their source provenance, access policy, and recovery behavior can be made explicit.
## License
MIT. See [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues