privacy
by J-X0
README.md
# privacy
k-anonymity enforcement with utility measurement, built to run entirely inside
the boundary that holds the regulated data. There is no runtime network access:
records are loaded, generalized, and released in-process, and a test
(`tests/airgap.test.ts`) fails the build if any source file imports a network or
child-process module.
## What it does
Given a table of records and a description of which columns are quasi-identifiers
(age, ZIP, gender — the fields that, combined, re-identify a person), it produces
a release where every combination of quasi-identifier values occurs at least `k`
times. It does this by full-domain generalization (coarsening values along a
taxonomy or into numeric bins) plus suppression of the records that cannot be
protected any other way. It then reports what that cost in utility.
## Architecture
The core is a pure library with no I/O:
- `generalization.ts` turns a raw quasi-identifier value into a coarser label at
a chosen level (numeric bins or a categorical taxonomy).
- `kanonymity.ts` runs a greedy full-domain search over those levels, suppressing
the residual small groups, and provides an independent `isKAnonymous` audit.
- `utility.ts` scores what enforcement cost (suppression, information loss,
discernibility, class size).
Around that core sits the server: `tools.ts` validates untrusted JSON and applies
resource limits, `mcp/server.ts` dispatches JSON-RPC, and `index.ts` frames it
over stdio. Any model-assisted step goes through the `PrivacyAdvisor` interface,
whose only in-tree implementation is a deterministic local stub — so the whole
system runs offline and, by construction, never sends data off the host. The
decisions behind this shape are recorded in [`docs/adr/`](docs/adr/).
## Requirements
Node 22.6 or newer (the code is TypeScript executed directly via Node's type
stripping; there is no build step).
## Install and test
```sh
npm ci
npm test
```
`npm test` runs the suite with `node --test`; no API key or network is needed.
## Usage
```ts
import { enforce, isKAnonymous } from './src/kanonymity.ts';
import { NumericGeneralizer, CategoricalGeneralizer } from './src/generalization.ts';
import { utilityReport } from './src/utility.ts';
const records = [
{ age: 21, zip: 'A1', diagnosis: 'flu' },
{ age: 22, zip: 'A2', diagnosis: 'cold' },
// ...
];
const quasiIdentifiers = [
new NumericGeneralizer('age', 20, 60, 5),
new CategoricalGeneralizer('zip', {
A1: ['A1', 'A', '*'],
A2: ['A2', 'A', '*'],
}),
];
const result = enforce(records, { k: 2, quasiIdentifiers, maxSuppression: 0.05 });
result.satisfied; // true if k met within the suppression ceiling
result.released; // generalized records safe to publish
result.suppressedCount; // records dropped to protect the rest
isKAnonymous(result.released, ['age', 'zip'], 2); // independent audit check
utilityReport(result, quasiIdentifiers); // suppression, info loss, discernibility
```
### Field roles without touching the data
`StubAdvisor` classifies columns into roles (identifier, quasi-identifier,
sensitive, other) from their **names only** — never their values — so schema
review can happen without exposing cells. It is the deterministic, offline
implementation of the `PrivacyAdvisor` interface in `src/providers/base.ts`; a
networked model could replace it, but only behind that same interface and only
by explicit choice.
```ts
import { StubAdvisor } from './src/providers/stub.ts';
await new StubAdvisor().suggestRoles(['patient_id', 'age', 'zip', 'diagnosis']);
```
## How enforcement chooses levels
The search is greedy. Starting from raw values it repeatedly raises the
generalization level of whichever quasi-identifier most reduces suppression,
breaking ties toward lower information loss, until suppression falls within the
configured ceiling or every field is fully generalized. Greedy full-domain
recoding is fast and produces auditable releases (one level per column), but it
is not guaranteed optimal — a lattice search such as Incognito would find the
minimal generalization at higher cost. That tradeoff is deferred; see the TODO
in `src/kanonymity.ts`.
## Running as an MCP server
The entry point is a Model Context Protocol server speaking JSON-RPC 2.0 over
stdio. It has no runtime dependencies and never opens a socket, so it runs where
the data lives.
```sh
node src/index.ts
```
It exposes three tools:
- `classify_fields` — assign privacy roles to column names (names only)
- `enforce_kanonymity` — generalize/suppress records to meet `k`, with a utility report
- `check_kanonymity` — audit an existing release
Example exchange (one JSON object per line on stdin, responses on stdout):
```
{"jsonrpc":"2.0","id":1,"method":"initialize"}
{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"enforce_kanonymity","arguments":{"k":2,"records":[{"age":21,"city":"London"},{"age":22,"city":"London"}],"quasiIdentifiers":[{"field":"age","kind":"numeric","min":20,"max":60,"step":5},{"field":"city","kind":"categorical","taxonomy":{"London":["London","*"]}}]}}}
```
Bad input (malformed JSON, unknown tool, oversized request) is reported without
taking the stream down: caller mistakes come back as JSON-RPC errors or
`isError` tool results, and only genuinely unexpected faults return
`-32603`. Logs are structured JSON on stderr — stdout carries the protocol only.
### Configuration
Set via environment variables; all are optional.
| Variable | Default | Meaning |
| --- | --- | --- |
| `PRIVACY_LOG_LEVEL` | `info` | `debug`, `info`, `warn`, or `error` |
| `PRIVACY_MAX_RECORDS` | `1000000` | reject requests with more rows |
| `PRIVACY_MAX_QUASI_IDENTIFIERS` | `32` | cap quasi-identifier columns |
| `PRIVACY_MAX_TAXONOMY_ENTRIES` | `100000` | cap categorical taxonomy size |
The limits exist so one oversized request cannot exhaust memory on the host that
holds the regulated data. Invalid values fail fast at startup.
## Known limitations
- k-anonymity alone does not defend against attribute disclosure: if every
member of an equivalence class shares the same sensitive value, k-anonymity
holds yet the value leaks. l-diversity or t-closeness would address this and
are not implemented.
- The generalization search is greedy, not optimal (see ADR 0002). It can pick a
higher level than a full lattice search would.
- Generalizers require explicit bounds/taxonomies; they are not inferred from the
data. This is deliberate — inferring bins from values would couple the metadata
step to the data it must stay clear of — but it means the caller supplies them.
- Enforcement holds the dataset in memory. `PRIVACY_MAX_RECORDS` bounds this;
there is no streaming or on-disk path for datasets larger than RAM.
- The `StubAdvisor` classifies by English field-name substrings; non-English or
opaque schemas fall through to `other`.
## Layout
- `src/kanonymity.ts` — enforcement and the independent `isKAnonymous` check
- `src/generalization.ts` — numeric and categorical generalizers
- `src/utility.ts` — suppression rate, information loss, discernibility
- `src/providers/` — the advisor interface and its offline stub
- `src/mcp/server.ts` — JSON-RPC 2.0 message dispatch
- `src/tools.ts` — input validation, resource limits, tool wiring
- `src/config.ts`, `src/logger.ts` — configuration and structured logging
- `src/index.ts` — stdio entry point (`runServer`)
- `tests/` — behaviour tests plus the air-gap guard
This server cannot be deployed
Maintenance
ActivityInactive
ResponsivenessNo issues