yarrowdalegraph
by J-X0
README.md
# yarrowdalegraph
Graph-based fraud propagation scoring for Elmsworth Diagnostics (Project
Yarrowdale). It runs as an MCP server that exposes one tool, `score_graph`, over
JSON-RPC 2.0 on stdio.
## What it does
Clinical-diagnostics fraud is rarely visible in a single entity. A colluding
provider, the billing accounts it controls, the devices those accounts log in
from, and downstream referrers form a connected structure. Yarrowdale models
that structure as a graph and diffuses fraud risk across it.
The pipeline is:
1. Build an entity graph: nodes are providers, patients, accounts, devices and
claims; edges are observed relationships with a strength weight. Analysts can
attach a `seed` in `[0,1]` to any node (a confirmed-fraud entity is `1`).
2. Obtain a per-node **prior** risk from a graph-ML model via the
`RiskPriorProvider` interface.
3. **Propagate** those priors across the graph with a personalised-PageRank-style
diffusion until convergence, producing a final risk score per node.
### The hard constraint: deterministic fallback
Scoring must keep working when the model is unavailable. That is not handled by
a `try/catch` bolted on at the edge — it drives the architecture:
- The model is reached only through `RiskPriorProvider`
(`src/providers/base.ts`). Its sole job is to turn a graph into prior scores.
- `FraudScorer` (`src/scorer.ts`) consults the provider behind three guards.
Any of them routes to the deterministic path:
- `available()` returns false (or throws),
- `priors()` throws,
- `priors()` returns a map that fails validation (wrong size / out of range).
- The fallback (`src/fallback.ts`) computes priors from graph structure alone —
analyst seeds plus weighted neighbour evidence. No network, no model, fully
deterministic.
- Both paths then run the **same** propagation algorithm, so a fallback score
differs from a model score only in its priors, never in the method. That keeps
results explainable to a reviewer.
- Every degradation decision is recorded through the `audit` callback, and the
result carries `priorSource` (`'model'` | `'fallback'`) and `fallbackReason`.
The test suite runs entirely offline against `StubProvider`
(`src/providers/stub.ts`), which also has flags (`down`, `throwOnPriors`,
`returnIncomplete`) to exercise each degradation branch.
## Layout
```
src/
types.ts domain model (nodes, edges, params, results)
graph.ts validation + row-normalised transition matrix
propagation.ts core diffusion algorithm
fallback.ts deterministic structural priors
scorer.ts orchestration + degradation + audit
config.ts environment-driven configuration
logger.ts structured JSON logging to stderr
validation.ts untrusted-input parsing + resource limits
server.ts MCP JSON-RPC request handling
main.ts stdio entry point
index.ts public exports
providers/
base.ts RiskPriorProvider interface
stub.ts deterministic offline provider
forced-fallback.ts operator pin to the offline path
tests/ node:test suites
```
## Architecture
Request flow for a `score_graph` call:
```
stdin (JSON-RPC line)
-> main.ts framing, ordering, fatal-error guard
-> server.ts JSON-RPC dispatch, error codes, timing, audit->log
-> validation.ts untrusted JSON -> FraudGraph, resource limits
-> scorer.ts model-or-fallback decision (the degradation point)
-> providers/ model priors, or ...
-> fallback.ts deterministic structural priors
-> propagation.ts diffusion shared by both paths
-> stdout (JSON-RPC line) logs -> stderr
```
The decisions that had a real alternative are recorded in
[`docs/adr/`](docs/adr/): the provider-interface fallback boundary, the shared
propagation path, undirected iterate-don't-invert diffusion, the zero-dependency
MCP transport, and fail-fast configuration.
## Requirements
Node 22 or newer. Node runs the TypeScript sources directly, so there is no
build step and no `dist/`.
## Install and test
```sh
npm ci
npm test
```
`npm test` runs the built-in `node --test` runner over `tests/**/*.test.ts`.
Optional type check (requires devDependencies installed):
```sh
npm run typecheck
```
## Running the server
```sh
npm start
```
This starts the MCP server on stdio. It reads one JSON-RPC message per line on
stdin and writes one response per line on stdout; all logs go to stderr so they
never corrupt the protocol stream. A minimal session:
```sh
printf '%s\n%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"score_graph","arguments":{"graph":{"nodes":[{"id":"p1","kind":"provider","seed":1},{"id":"acc","kind":"account"}],"edges":[{"source":"p1","target":"acc","weight":2}]}}}}' \
| npm start --silent
```
The `score_graph` result is JSON text containing `priorSource`
(`model` | `fallback`), `fallbackReason`, `converged`, `iterations`,
`elapsedMs`, and a `scores` map.
### Configuration
All configuration is read from the environment at startup. A malformed value is
a fatal startup error (exit code 78), never silently ignored.
| Variable | Default | Meaning |
| --- | --- | --- |
| `YARROWDALE_ALPHA` | `0.85` | Diffusion mixing, strictly in `(0,1)` |
| `YARROWDALE_TOLERANCE` | `1e-9` | Convergence threshold |
| `YARROWDALE_MAX_ITERATIONS` | `200` | Iteration cap |
| `YARROWDALE_MAX_NODES` | `50000` | Reject graphs with more nodes |
| `YARROWDALE_MAX_EDGES` | `500000` | Reject graphs with more edges |
| `YARROWDALE_FORCE_FALLBACK` | `false` | Pin scoring to the offline path |
| `YARROWDALE_LOG_LEVEL` | `info` | `debug` \| `info` \| `warn` \| `error` |
`YARROWDALE_FORCE_FALLBACK=true` is the operator switch for a model incident: it
routes every request through the deterministic fallback without a redeploy.
## Library usage
```ts
import { FraudScorer, StubProvider } from './src/index.ts';
const scorer = new FraudScorer(new StubProvider(), {
audit: (e) => console.error(`[audit] ${e.event}: ${e.detail}`),
});
const result = await scorer.score({
nodes: [
{ id: 'p1', kind: 'provider', seed: 1 }, // confirmed fraud
{ id: 'acc', kind: 'account' },
{ id: 'dev', kind: 'device' },
],
edges: [
{ source: 'p1', target: 'acc', weight: 2, relation: 'billing' },
{ source: 'acc', target: 'dev', weight: 1, relation: 'login' },
],
});
console.log(result.priorSource); // 'model' or 'fallback'
console.log(result.scores.get('acc'));
```
## Algorithm notes
Propagation solves the fixed point
```
s = alpha * W s + (1 - alpha) * p
```
where `W` is the row-normalised undirected adjacency and `p` the prior vector.
For `alpha` in `(0,1)` the iteration is a contraction and converges to the
unique solution `s = (1 - alpha)(I - alpha W)^{-1} p`. Edges are treated as
undirected because collusion evidence is symmetric. We iterate rather than
invert: the graphs are sparse and a few hundred sweeps are cheaper than an
`O(n^3)` solve. Default `alpha = 0.85`, tolerance `1e-9`, cap `200` iterations.
## Known limitations
- The learned `RiskPriorProvider` (graph neural network) is not yet shipped;
only the interface and the deterministic stub are. `main.ts` wires in the stub
and will select the model provider from config once it exists. See the TODO in
`src/providers/base.ts`.
- The stdio transport handles one message per line and serialises request
handling. That is enough for a single MCP client; JSON-RPC batch arrays and
concurrent in-flight scoring are not implemented.
- Scoring is synchronous and in-memory. A graph near the configured node/edge
limits is scored in one event-loop-blocking pass; there is no chunking or
backpressure. The limits exist to keep that pass bounded, not to stream.
- Propagation treats every relationship type the same once weighted. Distinct
edge semantics (a shared device vs. a referral) are captured only through the
caller-supplied `weight`, not through per-relation propagation rules.
---
*Elmsworth Diagnostics is an illustrative client; this repository is a self-directed reference implementation built to work end to end.*
TDQS
A3.9/5.0
Scored across 1 tool
Disambiguation5/5
With only one tool, there is no potential for overlap or confusion. The tool's purpose is clearly distinct by default.
Naming Consistency5/5
score_graph is a clear, snake_case verb-noun name that accurately describes the operation. It follows a consistent, predictable convention.
Tool Count3/5
A single tool is thin for a server whose name suggests a graph-oriented domain, though it may be intentionally focused. The one tool is substantive, but users may expect supporting operations.
Completeness4/5
The tool covers the core fraud-scoring workflow with propagation, seeds, priors, and a fallback. Missing auxiliary operations (e.g., graph management or seed configuration) are likely out of scope.
Maintenance
ActivityInactive
ResponsivenessNo issues