Skip to main content
Glama
sergio4848

p2p-finance

by sergio4848
README.md
# p2p-mcp-agent

An MCP server built on top of a mock procurement (Coupa-like) API and a mock payment provider, plus
a LangGraph agent that processes purchase-to-pay invoices through it end to end: fetch, three-way
match, classify the exception, request a human approval or hold, and schedule payment only once a
human has approved it. Every step that touches money is authenticated, logged, idempotent and
recorded in a tamper-evident audit log.

This is a portfolio/demo project: a solid, tested example of owning an MCP layer over finance
systems, not a production payment platform. See [Limitations](#limitations).

## Why this exists

A client project needs MCP servers built on top of Finance systems such as Coupa and payment
platforms, with reliability, security, versioning and documentation owned end to end, and logging,
audit trails and human-in-the-loop steps in every workflow that touches a financial transaction.
This repo is a self-contained demonstration of that shape of problem: a mock system standing in for
Coupa, a mock payment rail, an MCP layer over both reached only through their REST APIs (never
their databases directly, see [ADR 0001](docs/adr/0001-mcp-over-rest-not-direct-db.md)), and an
agent workflow that cannot pay an invoice without a human's sign-off, enforced in two independent
places.

## Architecture

```
                         ┌─────────────────────┐
                         │  LangGraph agent     │
                         │  (p2p_finance.agent) │
                         └──────────┬───────────┘
                                    │ in-process (thin adapter, ADR 0001)
                                    ▼
                         ┌─────────────────────┐        any MCP client
                         │  MCP server          │◄────── (desktop app,
                         │  "p2p-finance"        │        mcp dev, custom)
                         │  (p2p_finance.mcp_    │
                         │   server)             │
                         └──────────┬───────────┘
                                    │ httpx (REST, API key)
                                    ▼
                    ┌───────────────────────────────┐
                    │  Procurement API (mock Coupa)  │
                    │  suppliers · POs · receipts ·  │
                    │  invoices · matching · audit   │
                    └───────┬───────────────┬────────┘
                             │               │
                    read/pay/approve   httpx (Idempotency-Key)
                    scoped API keys           │
                             │                 ▼
                    ┌────────┴──────┐  ┌───────────────────┐
                    │  Human via     │  │  Payment provider  │
                    │  `p2p approve` │  │  (mock, idempotent) │
                    └────────────────┘  └───────────────────┘
```

| Concern | Where | Notes |
|---|---|---|
| Mock procurement API | `src/p2p_finance/procurement/` | FastAPI + SQLAlchemy 2 + SQLite; suppliers, POs, receipts, invoices, matching, approvals, audit |
| Mock payment provider | `src/p2p_finance/payments/` | FastAPI + SQLite; idempotent payment creation |
| MCP layer | `src/p2p_finance/mcp_server/` | Official Python MCP SDK; tools/resources over `httpx`, correlation-id logging, versioning rule |
| Agent workflow | `src/p2p_finance/agent/` | `langgraph.StateGraph`; `RuleBackend` (deterministic) and `OpenAIBackend` (structured output) |
| Evaluation | `src/p2p_finance/evals/` | 20 labelled scenarios, decision + exception-type accuracy, confusion table |
| Tests | `tests/` | 44 tests: matching, auth/scopes, idempotency, approvals, audit tamper detection, MCP protocol contract, agent end to end |
| Delivery | `Dockerfile`, `docker-compose.yml`, `.github/workflows/ci.yml` | Non-root image, three services, CI runs lint + tests + eval + a compose smoke test |
| Decisions | `docs/adr/` | Four ADRs |

## Quickstart

```bash
python -m venv .venv && . .venv/bin/activate      # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
cp .env.example .env                                # then edit the keys for anything beyond local use

# Terminal 1: mock payment provider
uvicorn p2p_finance.payments.app:app --port 8002
# Terminal 2: mock procurement API (Coupa-like)
uvicorn p2p_finance.procurement.app:app --port 8001
# Seed ~15 purchase orders/invoices covering a clean match and every exception type
p2p seed

# Terminal 3: MCP server, stdio transport (for an MCP client) or streamable-http
python -m p2p_finance.mcp_server.server
MCP_TRANSPORT=streamable-http python -m p2p_finance.mcp_server.server   # http://localhost:8000
```

Or the whole stack with Docker:

```bash
docker compose up --build
```

### Running the agent against a seeded invoice

```bash
python -c "
import asyncio
from p2p_finance.agent.adapter import P2PTools
from p2p_finance.agent.backends import RuleBackend
from p2p_finance.agent.graph import build_graph, initial_state
from p2p_finance.mcp_server.clients import procurement_client
from p2p_finance.config import get_settings

async def main():
    tools = P2PTools(procurement_client(get_settings()))
    graph = build_graph(tools, RuleBackend())
    result = await graph.ainvoke(initial_state(1))
    print(result['decision'], result['outcome'])

asyncio.run(main())
"
```

### Resolving an approval as a human

```bash
p2p approve 3 --as-user "finance-manager@example.com"
p2p reject 4 --as-user "finance-manager@example.com"
```

The `p2p approve`/`p2p reject` CLI authenticates with the *approver* API key, which the agent's own
key does not have -- see [ADR 0002](docs/adr/0002-approval-scope-separation.md).

### Verifying the audit log

```bash
p2p verify-audit
# OK: 17 audit events, chain intact
```

## MCP tools and resources

| Name | Type | What it does |
|---|---|---|
| `get_invoice` | tool | Fetch one invoice: lines, PO number, status, latest approval if any |
| `list_open_invoices` | tool | List invoices, optionally filtered by status |
| `three_way_match` | tool | Match an invoice against its PO and goods receipt; returns `matched`, `exception_type`, per-line variances |
| `request_approval` | tool | Create a pending human approval for an invoice |
| `get_approval_status` | tool | Check whether an approval is pending, approved or rejected |
| `schedule_payment` | tool | Schedule payment; refuses without an approved approval and a passing match; forwards an idempotency key |
| `p2p://policy` | resource | Current matching tolerances and the approval threshold, as JSON |
| `p2p://audit/recent` | resource | The most recent audit log entries, as JSON |

Every tool call is logged as a single structured line with a correlation id, e.g.
`cid=1f9c2ab4e7d1 tool=schedule_payment version=1 status=ok ms=12.40`
(`p2p_finance/mcp_server/server.py`). Tool versioning is documented in
[ADR 0004](docs/adr/0004-tool-schema-versioning.md).

### Connecting a client

Any MCP client can talk to the server over stdio or streamable-http. As one example, add this to your client's
`mcpServers` configuration (adjust the paths):

```json
{
  "mcpServers": {
    "p2p-finance": {
      "command": "C:/path/to/p2p-mcp-agent/.venv/Scripts/python.exe",
      "args": ["-m", "p2p_finance.mcp_server.server"],
      "env": {
        "P2P_PROCUREMENT_BASE_URL": "http://localhost:8001",
        "P2P_AGENT_API_KEY": "dev-agent-key"
      }
    }
  }
}
```

## Human in the loop and audit

- Every invoice that is not both cleanly matched *and* already carrying a human-approved approval
  needs one before payment. `request_approval` creates a **pending** approval; only a caller with
  the `approve` scope (the human-held key, via `p2p approve`/`p2p reject`) can resolve it. The
  agent's own key has `read` and `pay`, never `approve` -- see
  [ADR 0002](docs/adr/0002-approval-scope-separation.md).
- `schedule_payment` is refused, both by the agent's own code guard
  (`p2p_finance.agent.graph.make_schedule_payment_step`) and independently by the procurement
  service itself, unless an approval with `status == "approved"` exists for that exact invoice and
  its latest three-way match passed. Defense in depth: either guard failing on its own still stops
  the payment.
- Every write (`match_run`, `approval_requested`, `approval_resolved`, `payment_scheduled`) appends
  one row to a hash-chained, append-only audit log in the same transaction as the change it
  describes. `p2p verify-audit` recomputes the chain from the SQLite file directly and reports the
  first tampered event id, if any -- see
  [ADR 0003](docs/adr/0003-idempotency-and-hash-chained-audit.md).
- The payment provider requires an `Idempotency-Key` header and replays the stored response for a
  repeated key with the same payload, or rejects a repeated key with a different payload, so a
  retried request is never charged twice.

## Tests and evaluation

```bash
ruff check .    # clean
pytest          # 44 tests, no network, ~4 s
p2p-eval --backend rules      # writes eval/results/rules-<timestamp>.json and eval/results/report.md
p2p-eval --backend openai     # same dataset and graph against a real OpenAI-compatible model
```

Each of the 44 tests runs against an isolated in-memory procurement + payment stack over
`httpx.ASGITransport` (`p2p_finance.harness`) -- no running server, no network, and every request
still goes through real FastAPI routing, Pydantic validation and API-key auth. Coverage includes
matching/tolerance logic, auth scope separation, payment idempotency, approval workflow, audit
hash-chain tamper detection (including a direct SQLite edit bypassing the API), an MCP protocol
contract test driven through a real `mcp.client.ClientSession` over an in-memory transport, and
agent end-to-end runs for auto-pay, needs-approval, hold and duplicate-invoice scenarios.

Rules backend on the bundled 20-scenario dataset (2026-09-25, real numbers from `p2p-eval`, not
estimated):

| Metric | Value |
|---|---|
| Decision accuracy (auto_pay / needs_approval / hold) | 1.000 |
| Exception-type accuracy | 1.000 |
| Decisions | 3 auto_pay, 9 needs_approval, 8 hold |
| Latency p50 / p95 | 16.7 ms / 34.0 ms |

The decision and exception-type accuracy are 1.000 because both are deterministic business rules
(`p2p_finance.procurement.matching`), not a model call -- the point of the harness here is proving
the rule engine matches its own specification on boundary cases (a price variance at exactly the
2% tolerance line, a quantity variance at exactly the 20% severe line, a missing receipt that
overrides a pre-existing approval, a duplicate invoice number for the same supplier). The one step
that genuinely varies by backend, the exception explanation, is covered by
`p2p_finance.agent.backends.OpenAIBackend` and can be run against the same dataset with
`p2p-eval --backend openai` for comparison; that run is not included here because it needs a live
API key.

## Operations

- `GET /healthz` on both the procurement and payment services.
- Structured, single-line logs: FastAPI request logs on both services, and per-tool-call
  correlation-id logs on the MCP server.
- `Dockerfile`: multi-stage, non-root (`uid 10001`), the same image runs any of the three processes
  depending on the container's `command:`.
- `docker-compose.yml`: procurement + payment + MCP server (streamable-http transport on `:8000`),
  wired together with per-service API keys.
- `.github/workflows/ci.yml`: ruff, pytest, `p2p-eval` (uploaded as an artifact) on Python 3.11 and
  3.12, then a Docker build and a compose-based smoke test.

## Repository layout

```
src/p2p_finance/
  procurement/   mock Coupa-like API: models, matching, auth/scopes, audit, seed data, FastAPI app
  payments/      mock payment provider: idempotent payment creation
  mcp_server/    MCP tools/resources over httpx, correlation-id logging, tool versioning
  agent/         LangGraph workflow, RuleBackend/OpenAIBackend, thin MCP-tool adapter
  evals/         20-scenario dataset and the eval harness
  cli.py         `p2p seed` / `p2p approve` / `p2p reject` / `p2p verify-audit`
  harness.py     in-process test/eval stack (ASGITransport, no network)
tests/           44 tests
docs/adr/        4 architecture decision records
```

## Limitations

- The procurement and payment services are **mocks built for this repo**. They do not call Coupa's
  or any real payment provider's API, and nothing here has been tested against one. Connecting a
  real system means replacing `p2p_finance.mcp_server.clients`/`tools` targets and, most likely,
  the auth scheme -- the REST-over-HTTP boundary in ADR 0001 is what makes that a contained change.
- Matching supports one line per PO/invoice in the seed data and the eval dataset; the data model
  (`POLine`/`InvoiceLine`) supports multiple lines, but multi-line variance aggregation is
  summarised (worst-line variance), not itemised per line in the agent's decision.
- The audit hash chain detects tampering; it does not prevent someone with direct database write
  access from deleting rows outright. A production system would also ship these events to an
  external, append-only sink.
- No retry/backoff logic on the `httpx` calls between services; a transient failure surfaces as a
  tool error rather than being retried automatically.
- API keys are two environment variables per service, not a real identity provider. The scope model
  they enforce (ADR 0002) is the part meant to generalise, not the storage.

## License

MIT