Skip to main content
Glama
mypoorbrain

delivery-intelligence-mcp

by mypoorbrain
README.md
# Delivery Intelligence MCP Workbench

**A governed AI/MCP delivery-intelligence showcase over a fully synthetic enterprise programme.**

It helps a delivery lead answer: what changed, what is blocked, which dependency matters next, what a change request will affect, and which claims are supported by evidence. The core value works without an API key.

![Generated preview of the Delivery Intelligence MCP Workbench](docs/workbench-preview.svg)

## 60-Second Review Path

1. Open the generated workbench: [workbench/index.html](workbench/index.html)
2. Inspect the architecture: [docs/architecture.md](docs/architecture.md)
3. Review the engine/tool boundary: [delivery_intelligence/engine.py](delivery_intelligence/engine.py) and [delivery_intelligence/mcp_server.py](delivery_intelligence/mcp_server.py)
4. Check the evaluation harness: [docs/evaluation.md](docs/evaluation.md)
5. Run the project:

```bash
python -m pip install -e .
python -m delivery_intelligence build
python -m delivery_intelligence validate
python -m unittest discover -s tests
```

## What This Proves

| Capability | Public evidence |
| --- | --- |
| AI/MCP tool design | Eight focused MCP tools over a coherent programme model, with structured outputs and validated arguments. |
| Delivery/programme thinking | Milestones, RAID, decisions, change requests, readiness gates, dependencies and steering context pack. |
| Hallucination controls | Unsupported claims return `insufficient_evidence` instead of becoming facts. |
| Evidence traceability | Outputs separate source facts, deterministic derivations and recommendations with evidence references. |
| Evaluation discipline | Deterministic checks cover tool contracts, evidence coverage, refusal behavior, recursive change impact, snapshot deltas and repeatability. |
| Cost/latency awareness | Tool outputs include latency and estimated-cost telemetry; deterministic core has zero model cost. |

## Architecture

```mermaid
flowchart LR
    A[Synthetic programme fixture] --> B[Deterministic delivery engine]
    B --> C[Evidence ledger]
    B --> D[Tool registry]
    D --> E[Optional MCP server]
    D --> F[Evaluation harness]
    D --> G[Generated workbench]
    H[Future data adapter] -. documented boundary .-> A
    I[Optional narrative adapter] -. recommendations only .-> D
```

The public showcase stands alone. It does not import another portfolio repo and does not expose private operational, job-search, email, salary, eligibility, credential or account data.

## Tool Surface

| Tool | Purpose |
| --- | --- |
| `get_program_health` | Explainable health score with formula, limitations and evidence. |
| `list_priority_risks` | Deterministic RAID prioritisation by severity, probability, impact and overdue status. |
| `trace_dependency_impact` | Downstream dependency traversal from a milestone. |
| `assess_change_request` | Schedule, cost, scope and readiness impact for a change request. |
| `get_changes_since_last_review` | Snapshot delta across health, risks, decisions, milestones, readiness gates, change requests and evidence keys. |
| `get_blocked_decisions` | Blocked decisions, blockers and missing evidence. |
| `build_steering_context_pack` | Board-ready context pack with facts, derivations and recommendations separated. |
| `get_evidence_for_claim` | Evidence lookup or refusal when support is insufficient. |

## Facts vs Derivations vs Recommendations

| Layer | Meaning | Example |
| --- | --- | --- |
| Source facts | Synthetic fixture records: milestones, RAID, dependencies, decisions, gates and snapshots. | Payment freeze forecast moved to day 92. |
| Deterministic derivations | Python-calculated scores, risk ranks, impact paths and evidence coverage. | Payment freeze slippage propagates to pilot launch. |
| Recommendations | Policy suggestions generated from traceable facts and derivations. | Split or defer a change request unless sponsor accepts schedule impact. |
| Optional AI narrative | Future adapter boundary only; not required for tests or demo. | A model may rewrite a steering summary, but cannot create unsupported facts. |

## MCP Usage

The core package has no runtime dependencies. To run the real MCP server using the official Python SDK:

```bash
python -m pip install -e ".[mcp]"
delivery-intelligence-mcp
```

The MCP extra is optional because the deterministic engine, workbench and evaluation harness should remain reviewable without service credentials or model access.

## Live Demo Readiness

The repo includes a root [index.html](index.html) entrypoint for GitHub Pages. Pages is not currently enabled on the public repository. To make the live route work, enable GitHub Pages from `main` / root in repository settings; the expected URL will be `https://mypoorbrain.github.io/delivery-intelligence-mcp-workbench/`.

## Evaluation Results

Run:

```bash
python -m delivery_intelligence eval
```

The harness checks:

- tool-call contract validity;
- evidence coverage for health and steering outputs;
- unsupported-claim refusal;
- change/dependency impact correctness against the known fixture;
- recursive change-impact propagation;
- snapshot delta detection;
- repeatability across deterministic runs;
- latency and estimated-cost telemetry presence.

## Safety Boundaries

- Synthetic programme only.
- No external actions.
- No broad file-system, network or shell tools.
- No vector database or orchestration framework.
- No API key required.
- No private programme, job-search, resume, recruiter, account, salary, legal, eligibility or credential data.
- Future adapters must preserve the same evidence boundary.

## Repository Layout

```text
delivery_intelligence/
  fixtures.py       synthetic programme records
  engine.py         health, risk, dependency, change and evidence reasoning
  tools.py          validated tool registry and telemetry wrapper
  mcp_server.py     optional real MCP server using the Python SDK
  evaluation.py     deterministic evaluation harness
  workbench.py      generated HTML/SVG workbench artifacts
tests/              engine, tool, evaluation and visual-output checks
docs/               architecture, tool, evaluation and portfolio notes
workbench/          generated self-contained workbench
```

## License

MIT.