retest-mcp
by vtsouval
README.md
# ReTest
**“Not detected” is not the end of the story.**
ReTest uses **TabPFN-3.5** to prioritize sensitive PFAS retesting after older
PFOA/PFOS tests reported nondetects. It pairs targeted discovery with a randomized
audit, so the decision-maker can see both what was found and how uncertain the
untested remainder remains.
[](docs/DEMO.md)
Built for the [TabPFN-3.5 Hackathon](https://platform.priorlabs.ai/hackathon-3.5).
Runnable dashboard · real model results · local inference · six MCP tools · public EPA data.
## Run the demo in two commands
Requires Python 3.11+ and [uv](https://docs.astral.sh/uv/getting-started/installation/).
```sh
uv sync --locked
uv run --no-sync retest serve
```
Open **http://127.0.0.1:8765**. The default demo uses bundled predictions from
actual flagship TabPFN-3.5 executions; it needs **no API key and no GPU**.
Choose a budget, freeze a plan, reveal historical lab outcomes, inspect the audit,
compare every baseline, and export the selection as CSV. [Two-minute walkthrough](docs/DEMO.md)
## The hard problem
The older EPA UCMR3 reporting limits were **20 ng/L for PFOA and 40 ng/L for PFOS**.
UCMR5 uses **4 ng/L** for each. An old nondetect therefore cannot establish absence
at the later reporting limit.
We join the public monitoring cycles into **4,265 matched water systems**, each
with only nondetecting pre-2016 PFOA/PFOS records. From **149 historical features**,
TabPFN predicts whether either compound is detected in the first complete later
paired sample. The system preserves measured values, nondetect bounds, and missing
assays. Later records never become predictor features.
This is a retrospective screening-priority experiment, not live water monitoring.
The benchmark systems already have later test results; the shortlist is historical.
[Who would use it, and what is still unvalidated?](docs/USE_CASE.md)
A changed detection can reflect assay sensitivity, changed sources or treatment,
sampling differences, or actual change. One sample cannot clear an entire system.
[EPA source and dictionaries](https://www.epa.gov/dwucmr/occurrence-data-unregulated-contaminant-monitoring-rule)
## Actual results, including where baselines win
Flagship **TabPFN-3.5**, four ensemble members, one NVIDIA RTX 3090 Ti. All models
use the same training rows and permitted feature information. Baseline settings
were chosen on development data; the final geographic groups stayed untouched.
| Evaluation | Test systems / detections | TabPFN AP / AUROC | Found with 20% budget | Random expectation |
| --- | ---: | ---: | ---: | ---: |
| Unseen state groups | 357 / 38 | 0.350 / 0.804 | **22 / 38** using 72 priorities | 7.7 |
| Later collection year, secondary | 551 / 48 | 0.431 / 0.876 | **34 / 48** using 111 priorities | 9.7 |
That is **2.87×** and **3.52×** the random expected yield. These figures are for
targeting only; reserving an audit changes the selection. Histogram boosting wins
geographic AP (0.386) and finds 24 detections at the same budget. In the temporal
track, TabPFN has the best AUROC and ties ExtraTrees for detections at 20%; boosting
has slightly higher AP (0.440). State-bootstrap intervals do **not** establish
TabPFN superiority over the strongest baselines.
The full [report](artifacts/demo/report.json) includes logistic regression, an
inventory-only control, histogram boosting, ExtraTrees, CatBoost, TabPFN, timings,
paired state-bootstrap intervals, and every budget curve. Per-system predictions
are included, with hashes and model versions. The temporal track is a secondary
collection-year analysis, not a historical publication-delay backtest.
The [validation record](docs/VALIDATION.md) describes the executed checks.
## The audit is independent of model confidence
1. Select the targeted portion by TabPFN score, with outcomes hidden.
2. Draw the audit portion uniformly without replacement from the remaining pool.
3. Reveal selected historical lab results.
4. Invert exact hypergeometric tails to bound positive endpoints still untested.
This gives a conservative 95% **fixed-plan sampling interval** without requiring
calibrated model probabilities or independent sites. It requires complete random
audit outcomes and a fixed pool; it is not valid for cherry-picked plans or
repeated optional stopping. Zero audit detections do not certify zero remaining.
The default geographic replay deliberately retains the prespecified seed 35:
its interval **misses** the known remainder (17–138 estimated range, 16 actual).
The UI flags this. A 95% procedure can miss; changing the seed to hide that would
be misleading. Exact design-coverage checks and exhaustive small-population tests
verify the procedure. [Audit design evidence](artifacts/demo/audit_design_check.json)
## Execute the models yourself
```sh
uv sync --locked --all-extras
cp .env.example .env
# Add TABPFN_TOKEN locally; accept the TabPFN-3.5 license in your Prior Labs account.
uv run --no-sync retest prepare
CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync retest benchmark
```
Download and prepare uses the official EPA ZIP files (roughly 22 MB compressed).
Outputs go to `data/epa/` and `outputs/benchmark/`, both ignored by Git. The
configuration is [configs/benchmark.json](configs/benchmark.json). A GPU is
recommended; set its visible index or UUID to avoid other workloads. The dashboard
also has a **Re-run TabPFN locally** button once data and model dependencies exist.
To regenerate the bundled dashboard artifacts and twelve model explanations:
```sh
CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync python scripts/export_demo.py
```
License/account setup: [Prior Labs licenses](https://platform.priorlabs.ai/account/licenses)
and [API keys](https://platform.priorlabs.ai/account/api-keys). Credentials remain
in `.env`; do not commit them. No paid inference endpoint is used by the local path.
Raw EPA URLs can be revised upstream; compare archive hashes in the manifest with
the bundled report before treating a rerun as the identical dataset snapshot.
## Agent tools and explanations
```sh
uv sync --locked --extra mcp
uv run --no-sync retest-mcp
```
An LLM can inspect evidence, request real local inference, rank candidates, freeze
a plan, replay it, and retrieve model explanations. Predictions and counts always
come from tools. The application does not require an LLM and does not imitate one.
[MCP setup and tool contracts](docs/AGENT_GUIDE.md)
Twelve leading geographic candidates have executed grouped-reference perturbation
explanations. They are model sensitivities, **not SHAP values or causal effects**.
Both recorded and explainer-query scores are retained because query batching can
introduce small numerical differences. Unavailable explanations are labeled.
## Reproducibility and scope
```sh
uv sync --locked --extra dev --extra mcp
uv run --no-sync pytest -q
uv run --no-sync ruff check src tests scripts
```
Tests cover censoring, future-outcome leakage, first-sample selection, geographic
feature separation, exact audit coverage, API/export behavior, and an actual MCP
protocol exchange. CI runs without model credentials or GPUs.
With the server running, optional browser checks also refresh the screenshots:
```sh
uv run --no-sync playwright install chromium
uv run --no-sync python scripts/check_browser.py
```
- [Frozen evaluation protocol](docs/PROTOCOL.md)
- [Closest prior work and contribution boundaries](docs/RELATED_WORK.md)
- [Submission description](docs/SUBMISSION.md)
The matched cohort overrepresents larger public systems. Results do not validate
private wells, unmonitored systems, causal pollution sources, regulatory compliance,
or health outcomes. The tool supports research about screening priorities; it
does not replace laboratory testing or justify delaying required monitoring.
**License:** source code is Apache-2.0, as required by the hackathon. EPA data and
TabPFN code, weights, and outputs retain their applicable terms; model weights
are not bundled. TabPFN-3.5 has separate non-commercial/non-production restrictions.
See [NOTICE](NOTICE) and the [model license](https://huggingface.co/Prior-Labs/tabpfn_3_5/blob/main/LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues