retest-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@retest-mcprank our water systems by TabPFN score and freeze a 20% plan"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ReTest
“Not detected” is not the end of the story.
ReTest uses TabPFN-3.5 to prioritize sensitive PFAS retesting after older PFOA/PFOS tests reported nondetects. It pairs targeted discovery with a randomized audit, so the decision-maker can see both what was found and how uncertain the untested remainder remains.

Built for the TabPFN-3.5 Hackathon. Runnable dashboard · real model results · local inference · six MCP tools · public EPA data.
Run the demo in two commands
Requires Python 3.11+ and uv.
uv sync --locked
uv run --no-sync retest serveOpen http://127.0.0.1:8765. The default demo uses bundled predictions from actual flagship TabPFN-3.5 executions; it needs no API key and no GPU. Choose a budget, freeze a plan, reveal historical lab outcomes, inspect the audit, compare every baseline, and export the selection as CSV. Two-minute walkthrough
Related MCP server: WEATHGARDS
The hard problem
The older EPA UCMR3 reporting limits were 20 ng/L for PFOA and 40 ng/L for PFOS. UCMR5 uses 4 ng/L for each. An old nondetect therefore cannot establish absence at the later reporting limit.
We join the public monitoring cycles into 4,265 matched water systems, each with only nondetecting pre-2016 PFOA/PFOS records. From 149 historical features, TabPFN predicts whether either compound is detected in the first complete later paired sample. The system preserves measured values, nondetect bounds, and missing assays. Later records never become predictor features.
This is a retrospective screening-priority experiment, not live water monitoring. The benchmark systems already have later test results; the shortlist is historical. Who would use it, and what is still unvalidated? A changed detection can reflect assay sensitivity, changed sources or treatment, sampling differences, or actual change. One sample cannot clear an entire system. EPA source and dictionaries
Actual results, including where baselines win
Flagship TabPFN-3.5, four ensemble members, one NVIDIA RTX 3090 Ti. All models use the same training rows and permitted feature information. Baseline settings were chosen on development data; the final geographic groups stayed untouched.
Evaluation | Test systems / detections | TabPFN AP / AUROC | Found with 20% budget | Random expectation |
Unseen state groups | 357 / 38 | 0.350 / 0.804 | 22 / 38 using 72 priorities | 7.7 |
Later collection year, secondary | 551 / 48 | 0.431 / 0.876 | 34 / 48 using 111 priorities | 9.7 |
That is 2.87× and 3.52× the random expected yield. These figures are for targeting only; reserving an audit changes the selection. Histogram boosting wins geographic AP (0.386) and finds 24 detections at the same budget. In the temporal track, TabPFN has the best AUROC and ties ExtraTrees for detections at 20%; boosting has slightly higher AP (0.440). State-bootstrap intervals do not establish TabPFN superiority over the strongest baselines.
The full report includes logistic regression, an inventory-only control, histogram boosting, ExtraTrees, CatBoost, TabPFN, timings, paired state-bootstrap intervals, and every budget curve. Per-system predictions are included, with hashes and model versions. The temporal track is a secondary collection-year analysis, not a historical publication-delay backtest.
The validation record describes the executed checks.
The audit is independent of model confidence
Select the targeted portion by TabPFN score, with outcomes hidden.
Draw the audit portion uniformly without replacement from the remaining pool.
Reveal selected historical lab results.
Invert exact hypergeometric tails to bound positive endpoints still untested.
This gives a conservative 95% fixed-plan sampling interval without requiring calibrated model probabilities or independent sites. It requires complete random audit outcomes and a fixed pool; it is not valid for cherry-picked plans or repeated optional stopping. Zero audit detections do not certify zero remaining.
The default geographic replay deliberately retains the prespecified seed 35: its interval misses the known remainder (17–138 estimated range, 16 actual). The UI flags this. A 95% procedure can miss; changing the seed to hide that would be misleading. Exact design-coverage checks and exhaustive small-population tests verify the procedure. Audit design evidence
Execute the models yourself
uv sync --locked --all-extras
cp .env.example .env
# Add TABPFN_TOKEN locally; accept the TabPFN-3.5 license in your Prior Labs account.
uv run --no-sync retest prepare
CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync retest benchmarkDownload and prepare uses the official EPA ZIP files (roughly 22 MB compressed).
Outputs go to data/epa/ and outputs/benchmark/, both ignored by Git. The
configuration is configs/benchmark.json. A GPU is
recommended; set its visible index or UUID to avoid other workloads. The dashboard
also has a Re-run TabPFN locally button once data and model dependencies exist.
To regenerate the bundled dashboard artifacts and twelve model explanations:
CUDA_VISIBLE_DEVICES=0 OMP_NUM_THREADS=6 uv run --no-sync python scripts/export_demo.pyLicense/account setup: Prior Labs licenses
and API keys. Credentials remain
in .env; do not commit them. No paid inference endpoint is used by the local path.
Raw EPA URLs can be revised upstream; compare archive hashes in the manifest with
the bundled report before treating a rerun as the identical dataset snapshot.
Agent tools and explanations
uv sync --locked --extra mcp
uv run --no-sync retest-mcpAn LLM can inspect evidence, request real local inference, rank candidates, freeze a plan, replay it, and retrieve model explanations. Predictions and counts always come from tools. The application does not require an LLM and does not imitate one. MCP setup and tool contracts
Twelve leading geographic candidates have executed grouped-reference perturbation explanations. They are model sensitivities, not SHAP values or causal effects. Both recorded and explainer-query scores are retained because query batching can introduce small numerical differences. Unavailable explanations are labeled.
Reproducibility and scope
uv sync --locked --extra dev --extra mcp
uv run --no-sync pytest -q
uv run --no-sync ruff check src tests scriptsTests cover censoring, future-outcome leakage, first-sample selection, geographic feature separation, exact audit coverage, API/export behavior, and an actual MCP protocol exchange. CI runs without model credentials or GPUs.
With the server running, optional browser checks also refresh the screenshots:
uv run --no-sync playwright install chromium
uv run --no-sync python scripts/check_browser.pyThe matched cohort overrepresents larger public systems. Results do not validate private wells, unmonitored systems, causal pollution sources, regulatory compliance, or health outcomes. The tool supports research about screening priorities; it does not replace laboratory testing or justify delaying required monitoring.
License: source code is Apache-2.0, as required by the hackathon. EPA data and TabPFN code, weights, and outputs retain their applicable terms; model weights are not bundled. TabPFN-3.5 has separate non-commercial/non-production restrictions. See NOTICE and the model license.
This server cannot be deployed
Maintenance
Related MCP Connectors
Goal and task planning MCP for Codex and AI agents, with evidence-backed completion.
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
LLM Orchestration Agent (Mcp)
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceConnects local LLMs to external tools (calculator, knowledge base) via MCP protocol, enabling automatic tool detection and execution to enhance query responses.MIT
- FlicenseNot gradedqualityDmaintenanceExposes MCP tools that enable remote LLMs to query local Docker containers, OS processes, and system services in real time.-
- AlicenseNot gradedqualityDmaintenanceEnables multi-model leader-worker agent orchestration, workflow execution, and deterministic validation via structured MCP tools.7 npmApache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to run project-specific checks, replays, simulations, and queries as MCP tools, providing ground-truth feedback on config edits instead of guessing.1MIT