Skip to main content
Glama
kt2191886

vendor-crosswalk

by kt2191886
README.md
# vendor-crosswalk-mcp

A governed vendor-identity crosswalk. Three finance systems hold a record for
the same vendor under three different IDs. This package proposes which records
belong together, shows its evidence, keeps a person on every link, and refuses
to touch tax IDs or bank details. It ships as a Python library, a one-command
demo, and an MCP server.

It is the public, synthetic-data version of the flagship workflow in an
eight-week Finance AI pilot I proposed for a $17M+ cultural institution. The
proposed pilot would reconcile Sage Intacct, Airbase, and Airtable. This one
reconciles `ledger`, `spend`, and `ops`, with vendors that do not exist.

```
pip install -e ".[mcp]"
vendor-crosswalk demo
```

## The problem, in one paragraph

Bills and card charges sync between the ledger and the spend platform. Vendor
records do not. The name moves; the address, billing, and tax details do not.
Finance creates the vendor in both systems, links them by hand, and keeps a
third operational record with the department, owner, and contract dates. Each
system uses its own ID. Matching them depends on whoever has been there longest.
When records drift apart, nobody notices until a payment fails or a 1099 is
wrong. That is a control problem, not a data-entry problem.

## What it does

```
records (3 systems)  ->  propose()  ->  Proposals with evidence + tier
                                  ->  Crosswalk.from_proposals()  ->  rows waiting for a person
                                  ->  build_review_packet(row)    ->  the one thing the reviewer reads
                                  ->  Crosswalk.decide(...)       ->  the ONLY write, and it is to the crosswalk
                                  ->  evaluate()                  ->  did "high" mean high?
```

**Evidence, not vibes.** A proposal carries the kinds of evidence it found:
`name_exact`, `name_close`, `address`, `domain`, `zip`. Tiers are a contract
with the reviewer:

| tier | rule | who acts |
|---|---|---|
| high | exact normalized name plus address or email domain, or close name plus both; no competing candidate, no stale marker, no entity-suffix change | still a person, but the packet is short |
| medium | one strong signal, or two with a competitor within the ambiguity margin | a person, with the discrepancies laid out |
| low | a single weak signal | shown so nothing is silently dropped |

**A field-authority matrix instead of a fourth vendor master.** `authority.py`
says which system is trusted for which field. Legal name comes from the tax
document (status only). Accounting defaults come from the ledger. Department
and owner come from ops. Addresses are three different things, not one field to
overwrite. It is data, so a finance reviewer can read and approve it.

**A redaction boundary with a tripwire.** `redact.scrub` strips restricted
fields; `redact.assert_clean` raises if one reaches an output. The matcher never
reads a TIN. Matching on tax IDs would be both a data-boundary violation and a
false comfort: the crosswalk exists to reconcile the records people actually
look at.

**The only write is a human decision.** `Crosswalk.decide(index, reviewer,
approve)` is the single path to `linked`. It writes to the crosswalk table
(native IDs, status, evidence, reviewer, date). It never writes to a source
system. There is no tool that does.

## The eval

The pilot's gate: **zero incorrect high-confidence matches.** Not "high
precision." Zero. Everything else is reported so the number can be argued with.

Default benchmark (`--seed 7`, 40 vendors, 120 records, 119 true cross-system
pairs). The generator includes near-duplicate businesses, vendors missing from
one system, stale duplicate ledger records, entity reorganizations (LLC to
Inc.), address-type drift, and missing or expired tax documents.

| tier | proposed | correct | incorrect | precision | coverage of true pairs |
|---|---|---|---|---|---|
| high | 98 | 98 | 0 | 1.000 | 0.824 |
| medium | 21 | 21 | 0 | 1.000 | 0.176 |
| low | 34 | 0 | 34 | 0.000 | 0.000 |

False negatives: 0. Proposals demoted for a competing candidate: 12. The gate
holds across seeds 1, 2, 3, 7, 11, 42, 99, 123, 500, 2026 (`tests/` checks
seven of them). High-tier coverage ranges 0.70 to 0.84 across seeds; the rest
lands in medium, where a person looks. That is the design, not a shortfall.

What the eval does not measure, because synthetic data cannot: reviewer change
rate (how often a person materially edits the packet), active minutes per
completed vendor against a baseline, and cost per packet. Those are Week 1
baseline measurements in the real pilot, not claims in a README.

Run it yourself:

```
vendor-crosswalk eval --seed 2026
vendor-crosswalk packet 3
python -m unittest discover -s tests
```

## The MCP server

```
python -m vendor_crosswalk.server
```

Tools: `load_synthetic_records`, `find_vendor_matches`, `check_crosswalk`,
`build_review_packet_tool`, `run_eval`. All read only against the sources. The
model can propose, check, and package. It cannot link, and it cannot write to a
source system, because the server registers no such tool: the human decision
(`record_reviewer_decision`, which calls `Crosswalk.decide`) exists in the module
for the CLI and a future authenticated review UI, and is deliberately kept out
of the MCP tool registry. `tests/test_mcp_boundary.py` fails if that changes.
Authority is not something an agent should be able to ask for.

Demo limitation: `find_vendor_matches` rebuilds the in-memory crosswalk, so
repeated matching discards earlier decisions. A real deployment would persist
decisions in a separate store; this synthetic demo does not.

Claude Desktop config:

```json
{ "mcpServers": { "vendor-crosswalk": { "command": "python", "args": ["-m", "vendor_crosswalk.server"] } } }
```

## Design decisions

- **Tiers over scores.** A score of 0.83 invites arguing with the number. A
  tier invites checking the evidence. The score is still there for sorting.
- **Ambiguity demotes both candidates.** If a record's best and second-best
  matches in a target system are within 0.10, both are demoted and each names
  the other. Nothing is quietly picked.
- **Stale markers and suffix changes cap the tier.** "(do not use)" and "OLD"
  are things people leave in systems on purpose. An LLC that became an Inc. is
  probably the same vendor. Probably is a person's call.
- **Standard library only.** The matcher, crosswalk, packet, and eval have no
  dependencies. The MCP server is an optional extra.
- **The generator is honest.** It is built to include the cases that break
  naive matching. If you make it easier, the eval stops meaning anything.

## Files

```
src/vendor_crosswalk/
  authority.py   field-authority matrix + RESTRICTED set (data, not code)
  redact.py      scrub + tripwire
  normalize.py   names, addresses, domains, token similarity
  match.py       propose(): evidence, tiers, ambiguity, notes
  crosswalk.py   the linking table; decide() is the only write
  packet.py      build_review_packet(): checklist, discrepancies, change lists
  evaluate.py    the gate, precision and coverage by tier, false negatives
  synth.py       synthetic benchmark with hidden ground truth
  server.py      MCP server (FastMCP, stdio)
  cli.py         demo / eval / packet
tests/           unittest; the gate is a test
HUMANS.md        who is on the other end, and what they get back
STOP.md          when this must be paused, and who decides
```

## Who made this

Kai Chieh Tu. Dramaturg by training, finance operator by trade, builds with
Claude Code. I wrote the pilot this comes from because our vendor setup was a
control problem dressed as a typing problem. MIT license. No real vendor,
person, or institution appears in this repository.