p2p-finance
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@p2p-financeProcess pending invoices and list any that need human approval."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
p2p-mcp-agent
An MCP server built on top of a mock procurement (Coupa-like) API and a mock payment provider, plus a LangGraph agent that processes purchase-to-pay invoices through it end to end: fetch, three-way match, classify the exception, request a human approval or hold, and schedule payment only once a human has approved it. Every step that touches money is authenticated, logged, idempotent and recorded in a tamper-evident audit log.
This is a portfolio/demo project: a solid, tested example of owning an MCP layer over finance systems, not a production payment platform. See Limitations.
Why this exists
A client project needs MCP servers built on top of Finance systems such as Coupa and payment platforms, with reliability, security, versioning and documentation owned end to end, and logging, audit trails and human-in-the-loop steps in every workflow that touches a financial transaction. This repo is a self-contained demonstration of that shape of problem: a mock system standing in for Coupa, a mock payment rail, an MCP layer over both reached only through their REST APIs (never their databases directly, see ADR 0001), and an agent workflow that cannot pay an invoice without a human's sign-off, enforced in two independent places.
Related MCP server: Payables MCP
Architecture
┌─────────────────────┐
│ LangGraph agent │
│ (p2p_finance.agent) │
└──────────┬───────────┘
│ in-process (thin adapter, ADR 0001)
▼
┌─────────────────────┐ any MCP client
│ MCP server │◄────── (desktop app,
│ "p2p-finance" │ mcp dev, custom)
│ (p2p_finance.mcp_ │
│ server) │
└──────────┬───────────┘
│ httpx (REST, API key)
▼
┌───────────────────────────────┐
│ Procurement API (mock Coupa) │
│ suppliers · POs · receipts · │
│ invoices · matching · audit │
└───────┬───────────────┬────────┘
│ │
read/pay/approve httpx (Idempotency-Key)
scoped API keys │
│ ▼
┌────────┴──────┐ ┌───────────────────┐
│ Human via │ │ Payment provider │
│ `p2p approve` │ │ (mock, idempotent) │
└────────────────┘ └───────────────────┘Concern | Where | Notes |
Mock procurement API |
| FastAPI + SQLAlchemy 2 + SQLite; suppliers, POs, receipts, invoices, matching, approvals, audit |
Mock payment provider |
| FastAPI + SQLite; idempotent payment creation |
MCP layer |
| Official Python MCP SDK; tools/resources over |
Agent workflow |
|
|
Evaluation |
| 20 labelled scenarios, decision + exception-type accuracy, confusion table |
Tests |
| 44 tests: matching, auth/scopes, idempotency, approvals, audit tamper detection, MCP protocol contract, agent end to end |
Delivery |
| Non-root image, three services, CI runs lint + tests + eval + a compose smoke test |
Decisions |
| Four ADRs |
Quickstart
python -m venv .venv && . .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
cp .env.example .env # then edit the keys for anything beyond local use
# Terminal 1: mock payment provider
uvicorn p2p_finance.payments.app:app --port 8002
# Terminal 2: mock procurement API (Coupa-like)
uvicorn p2p_finance.procurement.app:app --port 8001
# Seed ~15 purchase orders/invoices covering a clean match and every exception type
p2p seed
# Terminal 3: MCP server, stdio transport (for an MCP client) or streamable-http
python -m p2p_finance.mcp_server.server
MCP_TRANSPORT=streamable-http python -m p2p_finance.mcp_server.server # http://localhost:8000Or the whole stack with Docker:
docker compose up --buildRunning the agent against a seeded invoice
python -c "
import asyncio
from p2p_finance.agent.adapter import P2PTools
from p2p_finance.agent.backends import RuleBackend
from p2p_finance.agent.graph import build_graph, initial_state
from p2p_finance.mcp_server.clients import procurement_client
from p2p_finance.config import get_settings
async def main():
tools = P2PTools(procurement_client(get_settings()))
graph = build_graph(tools, RuleBackend())
result = await graph.ainvoke(initial_state(1))
print(result['decision'], result['outcome'])
asyncio.run(main())
"Resolving an approval as a human
p2p approve 3 --as-user "finance-manager@example.com"
p2p reject 4 --as-user "finance-manager@example.com"The p2p approve/p2p reject CLI authenticates with the approver API key, which the agent's own
key does not have -- see ADR 0002.
Verifying the audit log
p2p verify-audit
# OK: 17 audit events, chain intactMCP tools and resources
Name | Type | What it does |
| tool | Fetch one invoice: lines, PO number, status, latest approval if any |
| tool | List invoices, optionally filtered by status |
| tool | Match an invoice against its PO and goods receipt; returns |
| tool | Create a pending human approval for an invoice |
| tool | Check whether an approval is pending, approved or rejected |
| tool | Schedule payment; refuses without an approved approval and a passing match; forwards an idempotency key |
| resource | Current matching tolerances and the approval threshold, as JSON |
| resource | The most recent audit log entries, as JSON |
Every tool call is logged as a single structured line with a correlation id, e.g.
cid=1f9c2ab4e7d1 tool=schedule_payment version=1 status=ok ms=12.40
(p2p_finance/mcp_server/server.py). Tool versioning is documented in
ADR 0004.
Connecting a client
Any MCP client can talk to the server over stdio or streamable-http. As one example, add this to your client's
mcpServers configuration (adjust the paths):
{
"mcpServers": {
"p2p-finance": {
"command": "C:/path/to/p2p-mcp-agent/.venv/Scripts/python.exe",
"args": ["-m", "p2p_finance.mcp_server.server"],
"env": {
"P2P_PROCUREMENT_BASE_URL": "http://localhost:8001",
"P2P_AGENT_API_KEY": "dev-agent-key"
}
}
}
}Human in the loop and audit
Every invoice that is not both cleanly matched and already carrying a human-approved approval needs one before payment.
request_approvalcreates a pending approval; only a caller with theapprovescope (the human-held key, viap2p approve/p2p reject) can resolve it. The agent's own key hasreadandpay, neverapprove-- see ADR 0002.schedule_paymentis refused, both by the agent's own code guard (p2p_finance.agent.graph.make_schedule_payment_step) and independently by the procurement service itself, unless an approval withstatus == "approved"exists for that exact invoice and its latest three-way match passed. Defense in depth: either guard failing on its own still stops the payment.Every write (
match_run,approval_requested,approval_resolved,payment_scheduled) appends one row to a hash-chained, append-only audit log in the same transaction as the change it describes.p2p verify-auditrecomputes the chain from the SQLite file directly and reports the first tampered event id, if any -- see ADR 0003.The payment provider requires an
Idempotency-Keyheader and replays the stored response for a repeated key with the same payload, or rejects a repeated key with a different payload, so a retried request is never charged twice.
Tests and evaluation
ruff check . # clean
pytest # 44 tests, no network, ~4 s
p2p-eval --backend rules # writes eval/results/rules-<timestamp>.json and eval/results/report.md
p2p-eval --backend openai # same dataset and graph against a real OpenAI-compatible modelEach of the 44 tests runs against an isolated in-memory procurement + payment stack over
httpx.ASGITransport (p2p_finance.harness) -- no running server, no network, and every request
still goes through real FastAPI routing, Pydantic validation and API-key auth. Coverage includes
matching/tolerance logic, auth scope separation, payment idempotency, approval workflow, audit
hash-chain tamper detection (including a direct SQLite edit bypassing the API), an MCP protocol
contract test driven through a real mcp.client.ClientSession over an in-memory transport, and
agent end-to-end runs for auto-pay, needs-approval, hold and duplicate-invoice scenarios.
Rules backend on the bundled 20-scenario dataset (2026-09-25, real numbers from p2p-eval, not
estimated):
Metric | Value |
Decision accuracy (auto_pay / needs_approval / hold) | 1.000 |
Exception-type accuracy | 1.000 |
Decisions | 3 auto_pay, 9 needs_approval, 8 hold |
Latency p50 / p95 | 16.7 ms / 34.0 ms |
The decision and exception-type accuracy are 1.000 because both are deterministic business rules
(p2p_finance.procurement.matching), not a model call -- the point of the harness here is proving
the rule engine matches its own specification on boundary cases (a price variance at exactly the
2% tolerance line, a quantity variance at exactly the 20% severe line, a missing receipt that
overrides a pre-existing approval, a duplicate invoice number for the same supplier). The one step
that genuinely varies by backend, the exception explanation, is covered by
p2p_finance.agent.backends.OpenAIBackend and can be run against the same dataset with
p2p-eval --backend openai for comparison; that run is not included here because it needs a live
API key.
Operations
GET /healthzon both the procurement and payment services.Structured, single-line logs: FastAPI request logs on both services, and per-tool-call correlation-id logs on the MCP server.
Dockerfile: multi-stage, non-root (uid 10001), the same image runs any of the three processes depending on the container'scommand:.docker-compose.yml: procurement + payment + MCP server (streamable-http transport on:8000), wired together with per-service API keys..github/workflows/ci.yml: ruff, pytest,p2p-eval(uploaded as an artifact) on Python 3.11 and 3.12, then a Docker build and a compose-based smoke test.
Repository layout
src/p2p_finance/
procurement/ mock Coupa-like API: models, matching, auth/scopes, audit, seed data, FastAPI app
payments/ mock payment provider: idempotent payment creation
mcp_server/ MCP tools/resources over httpx, correlation-id logging, tool versioning
agent/ LangGraph workflow, RuleBackend/OpenAIBackend, thin MCP-tool adapter
evals/ 20-scenario dataset and the eval harness
cli.py `p2p seed` / `p2p approve` / `p2p reject` / `p2p verify-audit`
harness.py in-process test/eval stack (ASGITransport, no network)
tests/ 44 tests
docs/adr/ 4 architecture decision recordsLimitations
The procurement and payment services are mocks built for this repo. They do not call Coupa's or any real payment provider's API, and nothing here has been tested against one. Connecting a real system means replacing
p2p_finance.mcp_server.clients/toolstargets and, most likely, the auth scheme -- the REST-over-HTTP boundary in ADR 0001 is what makes that a contained change.Matching supports one line per PO/invoice in the seed data and the eval dataset; the data model (
POLine/InvoiceLine) supports multiple lines, but multi-line variance aggregation is summarised (worst-line variance), not itemised per line in the agent's decision.The audit hash chain detects tampering; it does not prevent someone with direct database write access from deleting rows outright. A production system would also ship these events to an external, append-only sink.
No retry/backoff logic on the
httpxcalls between services; a transient failure surfaces as a tool error rather than being retried automatically.API keys are two environment variables per service, not a real identity provider. The scope model they enforce (ADR 0002) is the part meant to generalise, not the storage.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Autonomous commerce for AI agents: discover, quote, order, pay, verify.
Agentic workflow budget approvals with usage receipts.
Hosted AI agents and workflows with app OAuth, human approval gates, and a run ledger.
Agentic rails for complex workflows with receipts, fees, and MCP tool access.
Related MCP Servers
FlicenseNot gradedqualityDmaintenanceEnables AI agents to request purchase approval from humans, receive scoped virtual cards, complete checkout, and report receipts for audit.-- FlicenseAqualityBmaintenanceEnables an AI agent to handle accounts payable tasks against a mock ERP, including reading and writing bills and vendors, checking duplicates, matching invoices, recommending approvals, and queuing payment releases, with configurable profiles that limit available tools.11-
- FlicenseNot gradedqualityCmaintenanceEnables accounts payable teams to extract invoice data from PDFs and images, detect duplicates, normalize vendor names, calculate payment terms, and validate invoice completeness. Supports local extraction for text PDFs and optional vision providers for scanned documents.-
- AlicenseNot gradedqualityBmaintenanceEnables agents to process supplier invoices through a staged workflow with separate read and commit servers, so reads and reversible drafting are separated from irreversible posting. Enforces server-side approval gates, typed errors, and audit logging for safe back-office automation.40 PyPIMIT