Innkeeper
by edycutjong
README.md
<div align="center">
<img src="docs/icon.svg" alt="Innkeeper" width="144" height="144" />
<h1>๐๏ธ Innkeeper</h1>
<p><em>The night audit finally sleeps โ an autopilot night auditor that reconciles three ledgers, adjudicates every mismatch with cited evidence, and signs the books.</em></p>
<img src="docs/readme-hero.svg" alt="Innkeeper โ the night audit finally sleeps" width="100%" />
<br/><br/>
[](https://innkeeper.edycu.dev)
[](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/health)
[](https://youtu.be/LuK4dX-PRd0)
[](https://devpost.com/software/innkeeper)
[](https://innkeeper.edycu.dev/pitch/)
[](DEMO.md)
[](https://qwencloud-hackathon.devpost.com/)
<br/>







[](LICENSE)
[](https://github.com/edycutjong/innkeeper/actions/workflows/ci.yml)
</div>
> An autopilot night auditor for small hotels. Every night at 2 AM it reconciles
> the property-management ledger against the card processor's settlements and the
> OTA's PDF statement, **adjudicates every mismatch with cited evidence and
> competing hypotheses**, auto-clears the confident ones inside an expected-loss
> policy gate, and queues only the material discrepancies for the owner's coffee โ
> then signs the night's books cryptographically.
*A husband-and-wife inn runs 14 rooms; every night one of them stays up past
midnight squinting at three statements that disagree by $6.67 โ because the last
time they let it slide, it was $2,300 by month-end.*
> Innkeeper is a **CLI + MCP servers**, not a hosted web app โ but the audit
> engine is **deployed live on Alibaba Function Compute**
> ([`/health` ยท `/verify` ยท `/run`](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/health),
> offline, zero keys โ see [โ๏ธ Deployed](#-deployed-on-alibaba-function-compute)).
> Everything above also runs locally, offline, with zero API keys; see
> [`DEMO.md`](./DEMO.md) for the exact copy-paste script.
---
> ### โ ๏ธ MOCK SYSTEMS โ read this first
>
> The three source systems (PMS, card processor, OTA) are **mocks**, shipped in
> this repo as three **MCP-compatible servers** (`mcp/`) that read from a
> committed, deterministically **seeded** 14-room month. This is the honest path
> to realism: the data is a coherent month with planted, ground-truth-labeled
> discrepancy archetypes and **real reportlab-rendered OTA PDFs**.
>
> **Which transport runs, precisely:** the offline demo and every test run on
> **`FakeQwen`** โ a deterministic stand-in that computes each verdict from the
> mismatch's own arithmetic and memos (the same signals `qwen3.7-max` reasons
> over) and parses the committed statement sidecar; **it makes no VL call and
> never reads the ground-truth labels.** The real **`qwen3-vl-plus`** two-pass
> read of the 8-pt PDF and **`qwen3.7-max`** adjudication run only under
> **`--live`** with a `DASHSCOPE_API_KEY` โ a real, key-gated path you can run
> yourself (see the **Run the real Qwen path (`--live`)** section below);
> the reportlab PDFs exist so that path has a genuine document to read.
>
> **No live PMS/processor/OTA is contacted, and Innkeeper is decision support
> with signed evidence, not an accounting system of record.** The graded core
> runs **offline with zero API keys.**
---
## ๐ Quickstart (offline, no keys)
```bash
python3.12 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
innkeeper seed --nights 30 # deterministic month + rendered OTA statement PDFs
innkeeper run --night 2026-07-04 # fetch โ extract โ match โ adjudicate โ gate โ signed close
innkeeper replay --night 2026-07-04 # re-derive byte-identical, zero keys (invariant I4)
innkeeper verify-chain # recompute every root, check every signature
innkeeper bench # 277/281 auto-cleared ยท 0 false clears ยท accuracy 0.9964
pytest -q # 404 passed
```
The full demo script is in [`DEMO.md`](./DEMO.md).
## โ๏ธ Deployed on Alibaba Function Compute
Innkeeper is **deployed live on Alibaba Function Compute** (FC 3.0, managed
`python3.10` runtime โ no container, no ACR) at
**<https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run>**. Every endpoint runs
**offline on the committed `fixtures/` + `ledger/`** โ zero keys, no network, no
vision model โ so a judge can verify the signed books straight from a browser:
| Endpoint | What it does |
|---|---|
| [`GET /health`](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/health) | liveness |
| [`GET /verify`](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/verify) | re-verifies the 30 signed closes **in the cloud** โ roots, Ed25519 signatures, evidence sha256 bindings, a byte-identical replay of 2026-07-04, and a one-byte tamper caught (I2 / I3 / I4) |
| [`GET /run?night=2026-07-04`](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/run?night=2026-07-04) | reproduces one deterministic FakeQwen night audit โ 39 txns, 12 mismatches, 11 auto-cleared, 1 queued, Merkle root `d175694cโฆ` |
The live `qwen3-vl-plus` / `qwen3.7-max` path is **wired and verified with a real
DashScope smoke call**, and stays key-gated behind `--live` (next section); the
deployed endpoints are the **offline-deterministic engine**, byte-for-byte
replayable. Full request/response transcript:
[`docs/proof/DEPLOY_PROOF.md`](./docs/proof/DEPLOY_PROOF.md). The `02:00` timer
trigger is configured in [`infra/fc/s.yaml`](./infra/fc/s.yaml); a captured
console recording of the cron firing on its own is the one remaining artifact.
## ๐ Run the real Qwen path (`--live`)
The offline demo above proves the whole pipeline with **zero keys**. To exercise
the actual Qwen Cloud models on the real reportlab-rendered statement โ the one
step this project can't fake โ add your key and the `live` extra:
```bash
pip install -e ".[live]" # openai + pypdfium2 rasterizer
export DASHSCOPE_API_KEY=sk-โฆ # dashscope.console.aliyun.com/apiKey
innkeeper run --night 2026-07-04 --live # SAME pipeline, real models
```
Under `--live` the extractor is `LiveQwen`: it **rasterizes the committed 8-pt
OTA PDF and calls `qwen3-vl-plus` twice** (temperatures 0.0 / 0.4) โ agreement
becomes confidence, disagreement escalates (I5) โ and the mismatch residue is
adjudicated by **`qwen3.7-max` + thinking** returning typed JSON the gate
computes over. The evidence citations are bound to the on-disk sha256 hashes
either way, so a live verdict can never cite a hash that doesn't resolve. This
is the single key-gated path; nothing else needs a network.
## ๐๏ธ The pipeline
```mermaid
flowchart LR
TMR["Function Compute<br/>(deployed ยท 02:00 timer configured)"] --> RUN["audit run"]
RUN -->|"MCP tools"| M["3 mock servers โ PMS ยท processor ยท OTA-PDF"]
M --> EX["extract: qwen3-vl-plus two-pass + bbox (I5)"]
RUN --> MT["deterministic 3-tier matcher<br/>clears ~77% โ zero model calls"]
EX --> MT
MT -->|"mismatch residue"| ADJ["adjudicate: qwen3.7-max + thinking โ typed Verdict"]
ADJ <--> T{{"Qwen transport"}}
T --> FQ["FakeQwen โ default, no key"]
T --> LQ["LiveQwen โ qwen3.7-max ยท qwen3-vl-plus"]
ADJ --> G["E-loss gate: conf โฅ .85 โง โค $50 โง โ true_error"]
G --> CL["signed Merkle night-close โ prev_root chain (I1โI4)"]
```
<sub>**As built** = the three source systems are in-repo **mocks** (disclosed above); everything is green on FakeQwen, keyless. Live `qwen3-vl-plus` / `qwen3.7-max` sit behind `DASHSCOPE_API_KEY`; the app is **deployed live on Function Compute** (managed python3.10 โ `/health` ยท `/verify` ยท `/run`, offline) and the 02:00 timer trigger is configured in `infra/fc/s.yaml`. Plain-text view below.</sub>
```
fetch (3ร MCP) โ extract (qwen3-vl-plus, two-pass + bbox) โ deterministic match
โ adjudicate the residue (qwen3.7-max + thinking) โ E[loss] policy gate
โ signed Merkle night-close โ chain
```
A deterministic three-tier matcher (`ref_exact โ fuzzy-ref โ amount+date`) clears
~77% of transactions with **zero model calls**. Only the mismatch residue reaches
the language model. Every verdict is a typed [`Verdict`](./src/innkeeper_audit/schemas.py):
a classification, **evidence citations from โฅ2 systems** (each bound by sha256),
competing hypotheses, and a confidence โ so the gate is *math over typed fields*,
not vibes.
**The policy gate is expected-loss math, not a button:**
```
auto_clear โบ confidence โฅ 0.85 โง materiality โค $50 โง class โ true_error
generalised: E[loss] = amount ร (1 โ confidence) โค ฯ
```
Two hard constraints can only ever *queue*: a `true_error` classification (I1)
and a two-pass extraction disagreement (I5).
## ๐ Invariants (tested, not promised)
| | Invariant | Test |
|---|---|---|
| **I1** | **Zero false auto-clears on the planted true errors across all 30 nights** โ the load-bearing one | `test_invariants.py::test_I1_*` |
| I2 | Every verdict cites โฅ2 systems with resolvable sha256 hashes | `test_I2_*` (parametrised ร 30 nights) |
| I3 | The signed close chain verifies; a one-byte tamper (verdict / evidence / root) fails | `test_I3_*` |
| I4 | `replay --night N` reproduces identical verdicts, root, and signature | `test_I4_*` (ร 30 nights) |
| I5 | A two-pass extraction disagreement always escalates, never averages | `test_I5_*` |
## ๐ Benchmark
Seeded month, 30 nights, 1199 transactions, FakeQwen (deterministic, offline).
Regenerate with `python scripts/bench.py` โ [`docs/BENCH.md`](./docs/BENCH.md).
| metric | value | target |
|---|---|---|
| classification accuracy | **0.9964** | โฅ 0.92 |
| HITL action accuracy | 1.0000 | โ |
| **false auto-clears on true errors** | **0** | **0 (invariant)** |
| auto-cleared / mismatches | 277/281 (98.58%) | โ |
| queue precision / recall | 1.00 / 1.00 | โ |
| residue fraction (LLM-touched) | 23.44% | small |
| runtime (30 nights, offline) | 0.16s | < 5 min/night |
| modelled cost / night | $0.0327 | ~$0.15 |
The single class disagreement is the page-broken row **correctly declining to
classify** (`unknown`) while still queueing โ the right call under I5. Across the
whole **ฯ-sweep** the false-clear column stays **0**: the true-error and
escalation constraints hard-queue at every threshold, so the risk knob trades
automation against review load without ever touching the safety floor.
## ๐งฉ Why only Qwen Cloud
| Qwen surface | What it does here | Without it |
|---|---|---|
| **`qwen3-vl-plus`** | reads the 8-pt OTA statement table, **bbox-cited** per figure | a document-AI vendor bill + losing the evidence-highlight |
| two-pass VL agreement | agreement = confidence, disagreement = escalation (I5) | single-pass misreads corrupt money decisions silently |
| **`qwen3.7-max` + thinking** | ranks *competing hypotheses over evidence* | plus-tier collapses to the first plausible story |
| structured output | the E[loss] gate has typed fields to compute over | free-text verdicts can't be gated, benched, or replayed |
| function calling / **3ร MCP** | typed, logged tool calls = the replayable decision log | bespoke connectors; forfeits the rubric's named example |
Remove Qwen Cloud and a 14-room inn needs a document-AI vendor, a frontier LLM,
and a tool-orchestration framework โ three bills it will never pay, and no
single-vendor trail to sign into the books.
## ๐๏ธ Layout
```
src/innkeeper_audit/ amounts ยท config ยท schemas ยท crypto ยท matcher ยท policy
pipeline ยท report ยท verify ยท benchmark ยท cli
qwen/ fake (offline) ยท live (DashScope) ยท base
mcp/ server (JSON-RPC) ยท tools ยท pms/processor/ota
mcp/ pms_server.py ยท processor_server.py ยท ota_server.py (stdio)
scripts/ bench.py ยท verify_offline.py ยท check_submission_readiness.py
infra/fc/ s.yaml ยท wsgi.py (deployed HTTP handler) ยท audit_handler.py ยท PROOF.md
tests/ 404 tests
docs/ BENCH.md ยท friction-log.md ยท proof/DEPLOY_PROOF.md
```
## โ
Testing & CI
**5-stage pipeline:** Quality โ Security โ Build โ Offline Proof โ Deploy Gate.
```bash
# โโ Code Quality โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
ruff check . # lint
mypy src # type check (advisory on qwen/live.py)
pytest --cov=innkeeper_audit --cov-report=term # 404 tests, 100% coverage
# โโ Offline Proof (the invariants, live) โโโโ
innkeeper seed --nights 30
innkeeper run --night 2026-07-04
innkeeper replay --night 2026-07-04 # I4: byte-identical re-derivation
innkeeper verify-chain # I3: every root + signature checked
python scripts/verify_offline.py # socket-guarded tamper proof
innkeeper bench # I1: 0 false auto-clears, always
# โโ Security โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
pip-audit # dependency vulnerability scan
```
| Layer | Tool | Status |
|---|---|---|
| Code Quality | ruff + mypy | โ
|
| Unit Testing | pytest, 404 tests, 100% coverage | โ
|
| Invariant Testing | I1โI5, parametrised ร 30 nights | โ
|
| Security (SAST) | CodeQL (`python`) | โ
|
| Security (SCA) | Dependabot (`pip`) + `pip-audit` | โ
|
| Secret Scanning | TruffleHog | โ
|
| Build Verification | sdist/wheel + CLI entrypoint smoke test | โ
|
CI runs on every push/PR to `main` โ see [`.github/workflows/ci.yml`](./.github/workflows/ci.yml).
## ๐ Status โ honest
- โ
**Offline core is complete and green**: seed, matcher, adjudication, gate,
crypto, MCP, CLI, `bench`, `verify_offline`, **404 passing tests**.
- ๐ **Live Qwen path is wired but behind `DASHSCOPE_API_KEY`** (`--live`):
`qwen3-vl-plus` two-pass extraction and `qwen3.7-max` adjudication via the
DashScope OpenAI-compatible endpoint. Not exercised in the offline suite by
design; the deterministic `FakeQwen` computes verdicts from the same signals.
- โ๏ธ **Deployed live on Alibaba Function Compute** (managed `python3.10`) โ
[`/health` ยท `/verify` ยท `/run`](https://innkeeper-temfmzpqug.ap-southeast-1.fcapp.run/health)
run **offline** on the committed ledger (`/verify` re-verifies the signed-close
chain in the cloud; `/run` reproduces a night audit). Transcript in
[`docs/proof/DEPLOY_PROOF.md`](./docs/proof/DEPLOY_PROOF.md). The `02:00` timer
trigger is configured ([`infra/fc/s.yaml`](./infra/fc/s.yaml)); a captured
console recording of the cron firing on its own is the one remaining step.
- โ **`report.html` is optional** (single-page render via `innkeeper report
--night N --html`); the primary report is Markdown. No Next.js UI shipped โ
the CLI + signed ledger are the product surface.
Built solo in a five-project sprint. The night auditor finally sleeps.
## ๐ค Contributing & Security
Issue/PR templates, Code of Conduct, and a security policy live under
[`.github/`](./.github/) โ see [`CONTRIBUTING.md`](./.github/CONTRIBUTING.md)
and [`SECURITY.md`](./.github/SECURITY.md).
## ๐ License
[MIT](LICENSE) ยฉ 2026 Edy Cu
## ๐ท๏ธ Versioning
This project uses [Semantic Versioning](https://semver.org) with **fully automated** version
management driven by [Conventional Commits](https://www.conventionalcommits.org) โ the version is
never edited by hand.
| Commit type | Bump | Example |
|---|---|---|
| `fix: โฆ` | patch | 1.0.0 โ 1.0.1 |
| `feat: โฆ` | minor | 1.0.0 โ 1.1.0 |
| `feat!: โฆ` or `BREAKING CHANGE:` footer | major | 1.0.0 โ 2.0.0 |
[python-semantic-release](https://python-semantic-release.readthedocs.io) keeps the version in sync
across `pyproject.toml` and `src/innkeeper_audit/__init__.py`.
- **In CI/CD:** Stage 6 of the pipeline (`.github/workflows/ci.yml`) runs on every push to `main`,
computes the next version from the commits since the last tag, then commits + tags it automatically.
- **Locally:**
```bash
pip install -e ".[release]"
semantic-release version # compute + apply the next version and tag
```
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues