Skip to main content
Glama

ReTest

“Not detected” is not the end of the story.

ReTest uses TabPFN-3.5 to prioritize sensitive PFAS retesting after older PFOA/PFOS tests reported nondetects. It pairs targeted discovery with a randomized audit, so the decision-maker can see both what was found and how uncertain the untested remainder remains.

ReTest dashboard

Built for the TabPFN-3.5 Hackathon. Runnable dashboard · real model results · local inference · six MCP tools · public EPA data.

Run the demo in two commands

Requires Python 3.11+ and uv.

uv sync --locked
uv run --no-sync retest serve

Open http://127.0.0.1:8765. The default demo uses bundled predictions from actual flagship TabPFN-3.5 executions; it needs no API key and no GPU. Choose a budget, freeze a plan, reveal historical lab outcomes, inspect the audit, compare every baseline, and export the selection as CSV. Two-minute walkthrough

Related MCP server: WEATHGARDS

The hard problem

The older EPA UCMR3 reporting limits were 20 ng/L for PFOA and 40 ng/L for PFOS. UCMR5 uses 4 ng/L for each. An old nondetect therefore cannot establish absence at the later reporting limit.

We join the public monitoring cycles into 4,265 matched water systems, each with only nondetecting pre-2016 PFOA/PFOS records. From 149 historical features, TabPFN predicts whether either compound is detected in the first complete later paired sample. The system preserves measured values, nondetect bounds, and missing assays. Later records never become predictor features.

This is a retrospective screening-priority experiment, not live water monitoring. The benchmark systems already have later test results; the shortlist is historical. Who would use it, and what is still unvalidated? A changed detection can reflect assay sensitivity, changed sources or treatment, sampling differences, or actual change. One sample cannot clear an entire system. EPA source and dictionaries

Actual results, including where baselines win

Flagship TabPFN-3.5, four ensemble members, one NVIDIA RTX 3090 Ti. All models use the same training rows and permitted feature information. Baseline settings were chosen on development data; the final geographic groups stayed untouched.

Evaluation

Test systems / detections

TabPFN AP / AUROC

Found with 20% budget

Random expectation

Unseen state groups

357 / 38

0.350 / 0.804

22 / 38 using 72 priorities

7.7

Later collection year, secondary

551 / 48

0.431 / 0.876

34 / 48 using 111 priorities

9.7

That is 2.87× and 3.52× the random expected yield. These figures are for targeting only; reserving an audit changes the selection. Histogram boosting wins geographic AP (0.386) and finds 24 detections at the same budget. In the temporal track, TabPFN has the best AUROC and ties ExtraTrees for detections at 20%; boosting has slightly higher AP (0.440). State-bootstrap intervals do not establish TabPFN superiority over the strongest baselines.

The full report includes logistic regression, an inventory-only control, histogram boosting, ExtraTrees, CatBoost, TabPFN, timings, paired state-bootstrap intervals, and every budget curve. Per-system predictions are included, with hashes and model versions. The temporal track is a secondary collection-year analysis, not a historical publication-delay backtest.

The validation record describes the executed checks.

The audit is independent of model confidence

  1. Select the targeted portion by TabPFN score, with outcomes hidden.

  2. Draw the audit portion uniformly without replacement from the remaining pool.

  3. Reveal selected historical lab results.

  4. Invert exact hypergeometric tails to bound positive endpoints still untested.

This gives a conservative 95% fixed-plan sampling interval without requiring calibrated model probabilities or independent sites. It requires complete random audit outcomes and a fixed pool; it is not valid for cherry-picked plans or repeated optional stopping. Zero audit detections do not certify zero remaining.

The default geographic replay deliberately retains the prespecified seed 35: its interval misses the known remainder (17–138 estimated range, 16 actual). The UI flags this. A 95% procedure can miss; changing the seed to hide that would be misleading. Exact design-coverage checks and exhaustive small-population tests verify the procedure. Audit design evidence

Execute the models yourself

uv sync --locked --all-extras
cp .env.example .env
# Add TABPFN_TOKEN locally; accept the TabPFN-3.5 license in your Prior Labs account.
uv run --no-sync retest prepare
CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync retest benchmark

Download and prepare uses the official EPA ZIP files (roughly 22 MB compressed). Outputs go to data/epa/ and outputs/benchmark/, both ignored by Git. The configuration is configs/benchmark.json. A GPU is recommended; set its visible index or UUID to avoid other workloads. The dashboard also has a Re-run TabPFN locally button once data and model dependencies exist.

To regenerate the bundled dashboard artifacts and twelve model explanations:

CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync python scripts/export_demo.py

License/account setup: Prior Labs licenses and API keys. Credentials remain in .env; do not commit them. No paid inference endpoint is used by the local path. Raw EPA URLs can be revised upstream; compare archive hashes in the manifest with the bundled report before treating a rerun as the identical dataset snapshot.

Agent tools and explanations

uv sync --locked --extra mcp
uv run --no-sync retest-mcp

An LLM can inspect evidence, request real local inference, rank candidates, freeze a plan, replay it, and retrieve model explanations. Predictions and counts always come from tools. The application does not require an LLM and does not imitate one. MCP setup and tool contracts

Twelve leading geographic candidates have executed grouped-reference perturbation explanations. They are model sensitivities, not SHAP values or causal effects. Both recorded and explainer-query scores are retained because query batching can introduce small numerical differences. Unavailable explanations are labeled.

Reproducibility and scope

uv sync --locked --extra dev --extra mcp
uv run --no-sync pytest -q
uv run --no-sync ruff check src tests scripts

Tests cover censoring, future-outcome leakage, first-sample selection, geographic feature separation, exact audit coverage, API/export behavior, and an actual MCP protocol exchange. CI runs without model credentials or GPUs.

With the server running, optional browser checks also refresh the screenshots:

uv run --no-sync playwright install chromium
uv run --no-sync python scripts/check_browser.py

The matched cohort overrepresents larger public systems. Results do not validate private wells, unmonitored systems, causal pollution sources, regulatory compliance, or health outcomes. The tool supports research about screening priorities; it does not replace laboratory testing or justify delaying required monitoring.

License: source code is Apache-2.0, as required by the hackathon. EPA data and TabPFN code, weights, and outputs retain their applicable terms; model weights are not bundled. TabPFN-3.5 has separate non-commercial/non-production restrictions. See NOTICE and the model license.

Related MCP Connectors

Related MCP Servers