Skip to main content
Glama

mri-preread

A local brain MRI viewer and experimental AI review assistant for neurologists and radiologists. Import a scan, inspect its sequences in a browser, and connect your chosen model through MCP. The model runtime can be local or institution-approved hosted infrastructure. An optional MedGemma runner produces a research pre-read with exact images, prompts and responses available for inspection; its current bundled backend is documented below.

The intended workflow is scan → image preparation → visualization → AI draft → clinician review. The current MedGemma configuration is experimental: it flagged every sampled plane in our initial pilot. It does not provide a validated diagnosis or automatically prioritize the clinical reading queue.

Not a medical device. Not a diagnosis. The heterogeneity maps and the LLM pre-read point at places to look. They miss findings and they flag normal anatomy and artifacts. A qualified radiologist must review every scan.

Hospitals, clinics and public source

Visit the website and interactive public MRI demo or clone the repository.

Start with the hospital integration guide for pipeline commands, MCP configuration, image-delivery integration and current gaps. The hardware guide explains model-host choices, the bundled backend and the limits of the tested laptop demonstration.

The product website includes an image-first interactive MRI header, a real public MRI viewer and actual feature screenshots derived from CC0 OpenNeuro data. Private scans are excluded. To preview:

git clone https://github.com/vallverdu/mri-preread.git
cd mri-preread
python3 -m http.server 8080 --bind 127.0.0.1 --directory site

Open http://127.0.0.1:8080/. The demo has the full viewer, measurement, clinician note and mask tools. Example drawings are not clinical labels or AI predictions. See dataset credits and reproduction and third-party notices. Follow public-release.md before publishing source, and CONTRIBUTING.md to contribute. The MIT license covers the application; model weights have separate terms. The website photographs use their separate Unsplash license.

Anatomical scope: the current preparation, brain masking, hemispheric asymmetry and candidate rules are brain-specific. Other organs can have different geometry and normal asymmetry. This is not a validated general-purpose MRI anomaly detector. The browser rendering components could be adapted, but other anatomy requires a separately implemented and evaluated pipeline.

Related MCP server: MedVision MCP

Guide

What it does

  • Extract: reads a DICOM export (CD, USB, PACS download), groups images by series, and writes NIfTI volumes with patient geometry. It can also write PNGs. It guesses which series is T1, FLAIR, T2, SWI, DWI and ADC. import builds a study from NIfTI files instead. Any subset of the six sequences works.

  • Analyze:

    • Removes the skull (morphological).

    • Rigidly registers sequences to the selected reference (T1 when available), using mutual information and an automatic fallback to scanner coordinates when a registration drifts.

    • Corrects the intensity bias field (N4).

    • Computes two heterogeneity maps:

      • Outliers: a Gaussian-mixture model of this brain's own 6-channel intensities (T1, T2, FLAIR, SWI, DWI, ADC). Voxels whose combination fits none of its tissue classes score high.

      • Asymmetry: the brain is registered to its mirror image, and regions that differ from the other hemisphere, beyond the typical left/right difference, score high.

    • Clusters both maps into hotspots and labels them with simple rules, for example "DWI distortion artifact", "vein-like on SWI" or "fluid-like".

  • Viewer: brain_viewer.html is self-contained and needs no server or libraries:

    • WebGL2 ray-marching with Surface, Volume and MIP modes.

    • Starts with the whole volume. Slice navigation or box sliders define cuts; cut faces show the selected sequence.

    • Linked axial, coronal and sagittal slices.

    • Heat-map overlays, visible in 3D and through the surface.

    • A ruler and a circular ROI with statistics on every sequence, exportable as CSV.

    • Clinician notes with 2D/3D point locations, manual voxel masks, brush/eraser, undo/redo and portable JSON.

    • The AI pre-read list, with numbered markers and inspectable input/response traces.

  • LLM pre-read (MCP): 16 tools let an LLM:

    • see slices, as crops and montages with coordinates

    • compare a region with the mirror side

    • review the algorithm's candidates

    • write prioritised "area to review" annotations

    The annotations show up live in the viewer at http://127.0.0.1:8765/. The LLM can move your crosshair and read where you are looking. The server's instructions forbid diagnoses and "normal" verdicts.

Privacy

Everything runs locally. The code never uploads anything, with one exception: when you connect an LLM, the rendered images go to that LLM's provider. The optional local MedGemma runner below keeps inference on your Mac; only its separate model-download step needs internet access.

  • A study folder, and the brain_viewer.html built from it, contain the full scan. A surface rendering of a 3D T1 shows the face. Treat them like the original DICOM.

  • study.json retains description, date, scanner, age and sex, but excludes patient name, ID and birth date. This is not a guarantee of anonymization: descriptions and dates can still identify someone. Worklist input snapshots retain the original DICOM files and their metadata.

  • .gitignore excludes data/, DICOM, NIfTI, .npz, viewers and annotations, so a study inside the repo (for example data/<study>/) stays out of git. Keep it that way, and don't put patient screenshots in issues or PRs.

Install

Run the commands below in a terminal from the repository root. Paths beginning with data/ or models/ are relative to that directory. Replace /path/to/... with your own input locations; quote paths containing spaces. Keep source exports and generated studies in separate folders.

The viewer and preprocessing require Python 3.10 or newer. The local MedGemma runner additionally requires an Apple Silicon Mac and Metal GPU access; this repository does not currently provide a CUDA, Windows, or CPU inference backend. Local inference has been exercised on an M5 MacBook Air with 16 GB unified memory. A separate GPU is not needed for that tested configuration.

After obtaining a checkout of this repository:

cd /path/to/mri-preread
python3 -m venv .venv
.venv/bin/python -m pip install -e .

For compressed DICOM exports, install the optional pixel decoders:

.venv/bin/python -m pip install -e ".[dicom-compressed]"

To enable local MedGemma, install its runtime, download the weights once, and check inference using a synthetic image. The download uses the internet; the smoke test and MRI review use local weights.

.venv/bin/python -m pip install -e ".[medgemma]"
.venv/bin/mri-preread medgemma download
.venv/bin/mri-preread medgemma smoke-test

The default model is mlx-community/medgemma-1.5-4b-it-4bit, saved in models/medgemma-1.5-4b-it-4bit (about 3.4 GB of weights). The smoke test asks the model to describe a red square; it verifies that the runtime works, not that the model can interpret MRI correctly. The download records its repository revision in download-provenance.json. Review runs reuse those local weights; they do not download a new model. If you choose another storage location, pass the same --model-dir /path/to/model to download, smoke-test, review, and worklist.

From a scan to a viewable AI report

This walkthrough uses data/my-study as the generated study directory. There is no browser upload form: provide a local DICOM folder or NIfTI files through the CLI, or use the automatic watcher. You do not need lesion annotations or an expert report to run inference. You do need a reference to measure whether its findings are correct.

flowchart TD
    A[DICOM export folder] -->|extract| C[Study folder and sequence assignments]
    B[NIfTI volumes] -->|import| C
    C --> D[Check roles.json against series.json]
    D -->|analyze| E[Prepared images and statistical candidate maps]
    E -->|build| F[Browser MRI viewer]
    E -->|medgemma review| G[Local sampled-image review]
    G --> H[HTML trace, JSON record, exact input PNGs]
    F --> I[Clinician inspects original images and AI draft]
    H --> I

1. Provide one scan

Choose one of these input routes.

DICOM export. Use an unzipped export folder containing one study. Nested series folders are allowed. The extractor reads classic DICOM images with patient orientation and position; do not combine exports from different patients or studies in a manual extraction. The automatic watcher additionally rejects mixed studies and unsupported multi-frame images. This is not a general-purpose enhanced-MRI importer.

.venv/bin/mri-preread extract "/path/to/dicom-export" data/my-study

This creates NIfTI volumes, series.json, study.json, and an initial roles.json. Add --png if you also want exported 8-bit slice images; these PNGs are optional and are not the default MedGemma inputs. The source export remains in its original location.

NIfTI volumes. If the scan is already in .nii or .nii.gz format, import its sequences directly:

.venv/bin/mri-preread import data/my-study \
  --flair "/path/to/flair.nii.gz" \
  --dwi "/path/to/dwi.nii.gz" \
  --adc "/path/to/adc.nii.gz"

The supported roles are t1, t2, flair, swi, dwi, and adc. Supply only the sequences you actually have; all files must belong to the same study and carry valid spatial geometry. An anatomical volume is not a replacement for a missing ADC map. The import command creates the same study layout as extraction.

2. Verify which sequence is which

Open data/my-study/series.json and data/my-study/roles.json in your editor. DICOM roles are guesses from series descriptions; verify them before processing. NIfTI roles come from the flags you supplied.

For example, if series.json lists 001_FLAIR, 002_DWI, and 003_ADC, use:

{
  "series": {
    "t1": null,
    "t2": null,
    "flair": "001_FLAIR",
    "swi": null,
    "dwi": "002_DWI",
    "adc": "003_ADC"
  },
  "register": {}
}

Use the actual names in your own series.json, without the .nii.gz suffix. Missing sequences should be null or omitted. Any nonempty subset works technically, but available sequences constrain what can be assessed. Check that diffusion images and ADC are not swapped, and that localizers or unintended acquisitions have not been selected. For 4D inputs, the pipeline selects the DWI volume with the lowest mean intensity as a high-b heuristic, and the first volume for other roles; this does not verify the acquisition's b-values. Prefer a checked 3D volume when supplying multi-volume diffusion data.

Optional settings in the same JSON object:

Setting

Effect

"reference": "flair"

Selects the reference grid. Otherwise the first available of T1, FLAIR, T2, SWI, DWI, ADC is used.

"register": {"t2": false}

Uses scanner coordinates for T2 instead of rigid registration.

"register": {"t2": true}

Forces retention of registration even if it exceeds the usual drift limit; inspect the alignment carefully.

"skull_strip": "t1"

Forces the pipeline's skull stripping when automatic detection would keep already-stripped input. Default: "auto".

ADC units are normalized automatically to ×10⁻⁶ mm²/s. These heuristics, role guesses, and registration still need visual checking. Editing roles.json does not update existing analysis: rerun the following steps after a correction, and generate a new AI report.

3. Prepare the images

.venv/bin/mri-preread analyze data/my-study

This is image processing, not a MedGemma call. It constructs the reference grid, performs skull stripping and bias correction, aligns sequences, normalizes ADC units, and computes statistical outlier and left/right asymmetry maps. It saves prepared volumes, metadata, and algorithmic candidates under data/my-study/analysis/. Processing time depends on image size and registration; allow a few minutes.

Review the terminal messages and analysis/meta.json for available sequences and alignment results. When registration is rejected, the pipeline can fall back to scanner coordinates; completion alone does not establish correct alignment. Statistical candidate maps also do not establish a lesion diagnosis.

4. Build and inspect the viewer

.venv/bin/mri-preread build data/my-study

Open data/my-study/brain_viewer.html in a browser by double-clicking it. The HTML contains the scan and works without a server. Use a browser with WebGL2 support. For optional local HTTP access, run this in a second terminal and leave it running:

.venv/bin/python -m http.server 8798 --bind 127.0.0.1 --directory data/my-study

Then visit the study viewer. This static server exposes files in that study directory to local clients; stop it with Ctrl+C when finished.

Use Four-up or Slices, move through the brain with the slice scroll controls, and switch sequences at the same crosshair position. Check image coverage, orientation, brain boundaries, and cross-sequence alignment before reviewing AI output. The 3D rendering is a navigation aid; inspect the source slice contrast as well. The controls table below covers navigation and measurement.

The all command combines extraction, analysis, and viewer building, but bypasses the manual pause for checking roles. Use the staged commands above for a new export protocol. all does not run MedGemma.

5. Run the local MedGemma pre-read

After model setup and image preparation, run:

.venv/bin/mri-preread medgemma review data/my-study \
  --step-mm 10 --max-tokens 128 \
  --output data/my-study/analysis/medgemma-review-01.json

Use a new output name for every run; the command refuses to overwrite an existing trace or image folder. Omitting --output generates a timestamped name automatically. The output directory must already exist. The terminal prints the HTML report path and progress such as Slice 3/14.

The runner starts its own read-only MCP session, obtains aligned images, and sends them to MedGemma locally. You do not need to configure a separate chat client or start serve. By default it samples axial planes roughly every 10 mm and sends all available sequences together for each plane. This is sampled review, not an exhaustive 3D read. See the input workflow for image preparation and decoding details.

Useful options:

Option

Purpose

--sequences flair dwi adc

Restrict inputs to these sequences; all requested roles must exist. Omit to use all available roles.

--step-mm 5

Sample more densely than the default 10 mm; increases work and still does not establish detection accuracy. Allowed range: 1–30 mm.

--max-tokens 128

Limit generated tokens per plane. CLI default is 256; this walkthrough and the watcher use 128. Truncation is recorded.

--max-slices 2

Short development check only. It limits the beginning of the sampling plan and must not be presented as a complete review.

--model-dir /path/to/model

Use weights downloaded to a different local directory.

6. Open the report and inspect the evidence

For the explicit output name above, open:

data/my-study/analysis/medgemma-review-01.html

With the optional static server from step 4 still running, visit the AI input/output report. You can open it during inference; it refreshes every 15 seconds. Each plane shows the exact sequence images, model decision, explanation, timing, and expandable prompt/raw response and geometry details. The saved response is the model's generated explanation, not access to internal reasoning.

Check the report status and reviewed/planned slice counts first. complete means the requested run finished; it does not mean the model is correct or the whole scan was inspected. A development limit can produce a completed but deliberately partial run.

Model label

Meaning within this research workflow

yes

The model flagged this sampled plane for review. It is not a confirmed lesion.

no

The model did not identify a focal finding in these supplied images. It does not clear the scan.

uncertain

The model expressed uncertainty. Human inspection is still required.

invalid

The response could not be parsed into the expected decision and observation.

Current quality result: the tested setup returned yes for every plane in the pilot, sometimes while its explanation said no abnormality was visible. Treat its statements as unverified research output. Without an independent expert reference, a private-scan run has no measured accuracy score.

There are three different outputs: statistical maps, point annotations written by an MCP client, and MedGemma's sampled-plane report. A successful medgemma review now rebuilds the viewer with a dedicated MedGemma review group. Select an entry to navigate to that axial plane, see its explanation and exact input images, and highlight the plane in the other views. Click an input thumbnail to enlarge it; Esc closes the image. These reports do not supply lesion coordinates: the viewer draws a plane indicator, not a lesion contour or invented point marker. Existing MCP annotations are preserved separately.

Ordinary build automatically checks the newest analysis/medgemma-review-*.json. Incomplete or mismatched reports are not attached. To attach a report saved elsewhere, including a report from a separate pilot directory, use:

.venv/bin/mri-preread build data/my-study \
  --medgemma-report /path/to/completed-report.json

The builder compares reference geometry and regenerates every supplied plane's input PNGs to verify their SHA-256 hashes against the report. Explicitly attaching a mismatched report fails the build. The verified images and model text are embedded in the standalone viewer. This adds build time and file size; it does not rerun inference or validate the medical statements. Reload an already-open viewer to see the update.

Files to keep together

data/my-study/
  study.json                     retained study metadata
  series.json                    extracted/imported series inventory
  roles.json                     sequence assignments and processing settings
  nifti/                         source volumes for this study
  png/                           optional extraction PNGs (--png)
  analysis/
    meta.json                    sequence and alignment metadata
    viewer_data.npz              prepared viewer data
    findings.json                statistical candidates, not Gemma findings
    annotations.json             optional MCP-client annotations
    medgemma-review-01.json       complete machine-readable trace
    medgemma-review-01.html       human-readable trace
    medgemma-review-01.assets/    exact model input PNGs
    ...                          prepared volumes, maps and transforms
  brain_viewer.html              self-contained MRI viewer

Keep the report HTML, JSON, and .assets/ directory together with their relative paths intact. The report contains image hashes, model provenance, sampling information, prompts, responses, and decoding settings. The viewer and reports contain medical data; keep them under ignored data/ and out of commits.

Viewer

Clinician notes and segmentation

Open Clinician annotations → New annotation to create a doctor's annotation. Enter a title, clinical note and optional author; text saves when you leave the field or click Save note. Use Place point to locate the note on a 2D slice or the visible 3D surface. Clicking the annotation in the list brings all slice views to its location. Clinician markers use D1, D2, etc., and remain separate from MCP annotations and MedGemma's unverified flags.

Choose Paint mask and drag to segment a region, or Erase mask to correct it. The brush radius is measured in millimetres and respects the reference image's voxel spacing. In a 2D view, painting affects one reference slice; scrolling while a clinician tool is selected advances one reference slice. In 3D, painting uses a spherical brush centred on the visible anatomical surface or cut face, so it can include deeper voxels. Use the cut controls to reach an internal region and inspect the mask in all three slice views. Navigate returns to normal 3D rotation; shift/right-drag pans while annotating.

The segmentation is a voxel mask shared by all views. The selected annotation appears orange and other clinician annotations cyan. In 3D, masks are shown through anatomy and follow the current cut limits; they are manual delineations, not model predictions. Undo / Redo covers notes, points, brush strokes, imports and deletion, retaining up to 30 actions within a voxel-history memory budget. Esc cancels the current stroke. Erasing one annotation preserves overlapping masks belonging to other annotations.

Edits save in IndexedDB in the current browser, keyed to the reference scan's content and physical geometry. They are not sent to the MCP server, S3, or other website visitors, and they are not embedded automatically in a rebuilt HTML file. Browser profiles, origins, and devices have separate workspaces; clearing browser data removes local drafts. The status line reports whether saving succeeded. If another tab changes the same scan, the viewer prevents an overwrite and asks you to export and reload.

Export notes + masks (JSON) creates a portable copy containing text, author, timestamps, voxel locations, image affine, and losslessly encoded mask runs (x + nx * (y + ny * z)). Import JSON adds annotations to the matching scan; mismatched geometry, malformed masks and duplicate annotation IDs are rejected without replacing current work. This is the viewer's annotation format, not DICOM SEG or a NIfTI export. Export before switching computers or sharing a review. The workspace supports up to 64 annotations and two million painted voxels across their masks. A browser save failure leaves editing available but requires export to keep the work.

Navigation and display

Layout

toolbar: 3D + slices, Four-up, Slices. Drag the divider to widen the slice column; double-click it to reset.

Controls

Click a left-panel group heading to expand/collapse it. Drag the panel's right edge to resize; double-click to reset. The panel has a styled scrollbar.

Slice order

Drag the ⠿ AXIAL / CORONAL / SAGITTAL handles onto another slice view. Focus a handle and use arrow keys for keyboard reordering. Order, panel widths and collapsed groups persist in this browser.

3D cuts

The initial volume is whole. With Cut follows slice navigation, click/drag or scroll a 2D view to cut at that view's plane. Keep chooses which side remains. Moving a cut slider switches to a manual box cut; Reset cut restores the whole volume.

MedGemma

Expand MedGemma review, choose a reviewed plane, and inspect its explanation, input images and raw response. Purple dashed indicators mark the selected plane, not a segmented abnormality.

Focus

⤢ on a pane, double-click it, or keys 1 3D, 2 axial, 3 coronal, 4 sagittal: maximize. Esc restores.

Full screen

F or ⛶ Full screen. P or ☰ Panel hides the control panel.

Links

brain_viewer.html#layout=quad, #max=ax&panel=0 open the viewer straight into a layout.

3D

drag to rotate, right/shift-drag to pan, scroll to zoom

Slices

click/drag to move the crosshair, scroll to step through slices

Measure

Ruler: drag in a slice, or click two points on the 3D surface. ROI: drag a circle. Export CSV.

Publish a standalone viewer

After building and checking a study, copy brain_viewer.html to your static website's project folder as index.html. The volume data, annotations and any attached MedGemma review are embedded in that file; the source DICOM directory and Python server are not required. The hosted page displays the saved review; it does not run MedGemma inference. Live MCP annotation polling is enabled only on loopback hosts.

For S3 hosting, set the HTML object's Content-Type to text/html; charset=utf-8. If uploading a gzip compressed copy, also set Content-Encoding: gzip, while keeping the object key index.html. Add a cover image and link to the page from your site's project catalog. When replacing a cached catalog or viewer, refresh its CloudFront cache entry. Verify the public URL loads and that the downloaded, decompressed HTML matches your local build.

Publishing this file also publishes its embedded scan and model output: visitors can download them. Choose the study and attached report with that visibility in mind.

LLM pre-read (MCP)

Copy .mcp.json.example to .mcp.json and fill in the absolute paths (for Claude Code in this folder). For another MCP client, add the same command / args entry to its config. Then:

  1. Start the client and approve the mri-preread server.

  2. Open http://127.0.0.1:8765/ (set BRAIN_VIEWER_PORT to change the port).

  3. Ask, for example: "Use the mri-preread tools to do an early pre-read and annotate the areas a radiologist should look at."

  4. Afterwards, rebuild_standalone_viewer (or mri-preread build) embeds the list in brain_viewer.html.

Tools: get_study_overview, list_candidate_regions, view_region, view_overview, region_stats, add_annotation, update_annotation, remove_annotation, list_annotations, clear_annotations, set_summary, focus_viewer, get_viewer_state, rebuild_standalone_viewer, get_slice_plan, view_slice_images.

Local MedGemma (Apple Silicon, experimental)

Start with MedGemma 1.5 4B, using the MLX community 4-bit conversion. It runs on the Mac's built-in GPU using MLX-VLM. The weights occupy about 3.4 GB. An M5 MacBook Air with 16 GB unified memory has run both the synthetic test and an MRI/MCP trial successfully; a separate GPU is not needed for this starting point.

Why this model: Google's MedGemma 1.5 model card explicitly adds CT/MRI volume support. Its internal MRI classification benchmark reports 64.7% macro accuracy for 1.5 4B versus 57.4% for the older multimodal 27B. This does not establish accuracy on our studies, our rendered montages, or quantized weights. The 27B text variant cannot interpret images. Google also notes that MedGemma has not been optimized or evaluated for multi-turn applications.

Follow installation and the walkthrough for runnable commands. The following describes what the default runner actually gives the model.

review requires an analyzed study and uses a fixed, read-only MCP workflow. It saves a separate research report without changing viewer annotations. The default slices workflow:

  • Samples the brain-mask extent every 10 mm (--step-mm, 1–30 mm). Selection never uses lesion masks, algorithmic candidates or previous annotations. Sampling can miss lesions between planes.

  • Reads the float reference/registered NIfTIs through view_slice_images, avoiding the viewer's prior 8-bit quantization. Applies a fixed window per volume (0–99.5th nonzero percentile; ADC 0–2400 ×10⁻⁶ mm²/s).

  • Produces a separate 896×896 image for each sequence at the same reference plane, preserving physical aspect ratio with black padding. No mosaics, crosshairs, annotations or reference-mask overlays.

  • Sends all available sequences together, in a documented order. --sequences flair dwi adc restricts the input. This is sampled axial review, not a full-volume or adjacent-slice interpretation.

  • Requests a review decision and short visible evidence. The first output token is constrained to yes, no or uncertain, followed by a newline; the explanation is unconstrained. This makes labels parseable, not calibrated or trustworthy. The raw response, prompt, assistant prefix and decoding settings are recorded. Contradictions between the decision and explanation remain possible.

Each run writes analysis/medgemma-review-<timestamp>.json, a matching HTML report, and an .assets/ directory containing the exact input PNGs with SHA-256 hashes. Open the HTML to see every image, decision, prompt and raw response. It refreshes every 15 seconds while a run is active. --output /path/to/new.json selects a destination; existing reports/assets are refused. Preserve the HTML, JSON and image directory together. These are patient data and belong outside git. Graceful failures/interruption retain completed responses with a failure status; a killed process can leave in_progress.

--max-tokens bounds each response; a truncated explanation is marked explicitly. --max-slices is for development smoke tests, not full evaluation. Use --model-dir /absolute/path outside the repository. The runtime requires Metal access; run in a local Terminal if a sandbox cannot access the GPU.

The earlier montage workflow remains available with --workflow montage --max-candidates 1. Its images can contain old annotation markers, so it should not be used for a blinded evaluation. The new slice path removes this source of leakage. Neither workflow gives the model autonomous control of tools.

Inspectable pilot against expert masks

This is a separate evaluation workflow, not a prerequisite for analyzing your own scan. It requires the prepared and analyzed benchmark studies in data/benchmark/ plus data/benchmark/_key/key.json and the referenced expert masks. Those medical datasets are not included in a fresh checkout; the pilot command does not download or prepare them. The existing benchmark preparation scripts are described in Evaluation. Use --root /path/to/prepared-benchmark for another benchmark location. An optional private study must also have completed analyze first.

Choose a new output directory for inference:

.venv/bin/python eval/medgemma_pilot.py --out data/medgemma-pilot-new \
  --private-study data/my-study --step-mm 10 --max-tokens 128
# Optional local browser access (the reports also work as local files):
.venv/bin/python -m http.server 8799 --bind 127.0.0.1 --directory data/medgemma-pilot-new

The default pilot uses four ISLES mask cases (case_02–case_05) and two controls (case_10, case_15). These were selected by source before inference; development case_01 is excluded. Model calls receive images and a fixed question, not source labels, old findings or ground truth. After inference, predictions and input image hashes are frozen in FROZEN.json. Only then does scoring load _key/key.json, align expert masks using the saved DWI transform, and create orange reference overlays for human inspection. Previous benchmark scores, frozen annotations and viewers are preserved.

Open index.html for metrics and links to every input/output trace; open private.html for the optional private study. The private scan has no accuracy score because no expert reference mask is supplied. --score-only verifies hashes and regenerates evaluation artifacts without rerunning inference.

Scores measure flags on sampled planes containing stroke-mask pixels, plus flags on planes without labelled stroke. A positive anywhere on a lesion-containing plane counts, even if Gemma describes the wrong location. They are not comparable with the previous point-within-10-mm benchmark below. Controls have no masks and are reported separately as assumed negative. Uncertain/invalid responses are counted explicitly, never silently treated as correct negatives. Slices from one patient are correlated; this small pilot is not clinical validation or evidence of improved accuracy over the montage input.

Initial pilot result (2026-09-27): this configuration failed to discriminate. It selected yes on all 91 sampled planes: 18/18 planes containing expert-mask lesions, 44/44 planes without labelled stroke in the mask cases, and 29/29 control planes. Therefore the apparent 100% lesion-plane flag rate is not useful detection. The constrained label sometimes contradicted the generated explanation. This result applies to this quantized model, prompt and decoding setup; it does not establish MedGemma's performance under other workflows. The remaining benchmark cases have not been used in this pilot.

Development checks on M5/16 GB: three-sequence inputs take roughly 4–5 seconds per plane and 4.7 GB peak MLX allocation (not total system memory). Better inputs have not solved model reliability: report-style responses can ignore instructions, and constrained decisions can over-flag or contradict their explanation. Do not treat a flag as a confirmed finding, or a negative response as reassurance about the whole scan.

Your application code remains MIT licensed. MedGemma weights are governed separately by the Health AI Developer Foundations terms; the community conversion does not make them MIT licensed. Weights and caches are excluded from git.

Automatic local review worklist

Use this route to process successive DICOM exports without manually running each pipeline command. Complete the local model installation and smoke test first. The watcher uses the same analyze, build, and medgemma review pipeline as the walkthrough, with 10 mm sampling and a 128-token response limit. It does not require a separately running MCP client.

Start the watcher

From the repository root, run the following and leave the terminal open:

.venv/bin/mri-preread worklist --inbox data/incoming-dicom --workspace data/worklist

Open the local worklist. Put each study in its own immediate subfolder of data/incoming-dicom. The exporter must create an empty .ready file last, after every DICOM file has been written and closed. Without that marker the study stays waiting; a quiet folder is not considered complete. This contract needs to be implemented by the export integration.

Deliver a completed study

For example, arrange the inbox like this (the series directories can have other names):

data/incoming-dicom/
  study-001/
    series-a/                      original DICOM files
    series-b/                      original DICOM files
    roles.json                     optional checked sequence mapping
    .ready                         create only after the export is complete
  study-002/                       still exporting; no .ready yet

To prepare a destination for an exporter:

mkdir -p data/incoming-dicom/study-001

Export or copy all DICOM files into that directory. Check that the export has finished and all files are closed; then create the marker as a separate final action:

touch data/incoming-dicom/study-001/.ready

Do not create the marker at the beginning of a transfer. If correcting an export, remove its marker first, finish the correction, then recreate the marker. Keep the inbox and workspace separate and non-nested. Folder names are shown in the worklist; use neutral study labels if possible.

Follow processing and review the draft

The worker checks that the folder contains one MR study, snapshots its inputs, then runs extraction, image preparation, viewer generation, and local MedGemma review. It supports classic single-frame DICOM with patient geometry; enhanced multi-frame MRI is rejected. An optional roles.json in the export folder overrides sequence guesses, using the extracted series names (001_FLAIR, for example). Otherwise the draft explicitly warns that sequence assignments need human verification.

The worklist links to the MRI viewer and the exact AI images, prompts, and outputs. It records durable job states, processing failures, human priority, and review completion in SQLite. Identical exports are deduplicated by content and DICOM instance IDs; changed images or sequence mappings create a new job. Interrupted jobs become visible failures on restart and can be retried. Rejected exports must be corrected; intake checks them again automatically. Processing logs and input snapshots live under the workspace. Keep this workspace inside ignored data/, since snapshots retain the original DICOM metadata.

One local worker processes studies in receipt order. Intake is checked between jobs; the browser remains available during processing. Human-assigned Expedite entries appear first in the displayed worklist. AI output never changes priority: the current MedGemma setup failed the preliminary benchmark. Every draft remains unverified and every study still requires clinical review. This is a single-operator, loopback-only research prototype, without PACS integration, authenticated reviewer identities, or clinical alerting. The export marker does not establish acquisition completeness or image quality.

The browser states distinguish workflow progress from clinical findings:

Display/state

What to do

Waiting for export completion

Finish the export, then create .ready. No model run has started.

Queued

A completed export has been accepted and is waiting for the worker.

Processing

Check the displayed stage: extraction, image preparation, viewer building, or local AI review.

Draft ready

Open MRI viewer and AI input/output trace, inspect limitations, then perform human review.

Needs technical review / failed

Read the error and job log. Correct rejected exports, or use Retry processing after fixing a processing failure.

Reviewed

Someone used Mark reviewed. This records a workflow action, not validation of the AI findings.

Human priority offers Unassigned, Routine, and Expedite. It affects the displayed order only; processing remains in receipt order. No AI prediction sets this value.

Job artifacts are stored at data/worklist/jobs/<job-id>/: processing.log contains subprocess output, input-<timestamp>/ contains a source snapshot, and study/ contains the viewer and analysis reports. data/worklist/worklist.sqlite3 stores job states and action events. Use the ID shown in the browser to find the matching log. Preserve the workspace to retain history; avoid editing its generated study in place. Correct source sequence mappings in the export and submit a completed version instead.

Stop the watcher with Ctrl+C and restart with the same command and workspace. Finished drafts remain available, while jobs interrupted during processing become visible failures that can be retried. Only one worker may own a workspace. This command is a foreground process, not an installed background service.

Use --model-dir, --port, or --poll-seconds to change defaults; --once processes the completed exports currently discovered and exits. See the clinician-assistant design and validation plan for the path toward evidence-linked findings and validated queue assistance.

Troubleshooting

Symptom

Check / next action

.venv/bin/mri-preread does not exist

Run installation from the repository root and confirm the virtual environment was created there.

No DICOM images found

Unzip the export; point to the directory containing the actual image files. The extractor expects classic DICOM with a DICM preamble, pixel data, and patient geometry. PDFs/screenshots are not volumetric inputs.

Compressed pixel decoding fails

Install ".[dicom-compressed]" in the same environment, then retry extraction. Unsupported transfer syntaxes may require another export format.

Wrong or missing sequences

Compare roles.json with series.json, correct the mappings, rerun analyze and build, and create a new MedGemma report.

Misaligned sequences or unexpected orientation

Inspect the slices and alignment metadata. A registration fallback is not proof of correct alignment; resolve the input/registration problem before interpreting cross-sequence output.

Study has no viewer data

Run analyze on the generated study directory, not the raw DICOM folder.

Local weights missing

Run medgemma download; use the same --model-dir for download and review.

Optional runtime missing

Install ".[medgemma]" on an Apple Silicon Mac. Installing the extra on another platform does not add an alternative backend.

Metal/GPU unavailable in a sandbox

Run the command in a local Mac terminal with GPU access. The current runner does not fall back to CPU.

Memory pressure during inference

Close other large applications and avoid concurrent model runs. Fewer valid input sequences may reduce the workload but also reduce the evidence available.

Output already exists

Choose a fresh --output name or omit the option for a timestamped report. The old run is intentionally preserved.

Report images missing

Keep the HTML, JSON, and matching .assets/ directory together; serve their common parent directory.

Report says failed_or_interrupted or stays in_progress after the process exits

It is incomplete. Inspect the terminal/job log and start a new run with a new output name. Completed responses may remain for inspection.

All planes say yes, or label and explanation disagree

This occurred in the recorded pilot. It is a model/workflow reliability failure, not evidence that every plane contains disease.

Worklist never receives a folder

It must be an immediate child of the configured inbox and contain a .ready file after export completion. Intake is checked between processing jobs.

Export rejected

Check for multiple Study UIDs, duplicate SOP instances, symlinks, non-MR images, or unsupported geometry/multi-frame DICOM. Correct the export and recreate .ready; intake checks it again.

Port already in use

Stop the conflicting server or choose another port: worklist --port 8801, or http.server 8797. MCP viewer ports use BRAIN_VIEWER_PORT.

For command-specific flags, run .venv/bin/mri-preread --help, .venv/bin/mri-preread medgemma review --help, or .venv/bin/mri-preread worklist --help.

Limitations

  • The 3D view is best with a thin-slice 3D T1 (or 3D FLAIR) as the reference.

  • Thick-slice sequences (often 4-5 mm FLAIR, T2, DWI) are resampled to about 2 mm, so small lesions blur.

  • Skull stripping is threshold and morphology based, not a trained model: it keeps the brainstem and upper cord, and can fail on unusual contrast. Already skull-stripped data is detected and used as is. SynthStrip or HD-BET are more robust.

  • Registration is rigid only. DWI inherits the ADC transform.

  • The hotspot labels are simple rules, not validated classifiers. Nothing here has been clinically validated.

  • The heterogeneity maps miss large lesions: a lesion covering much of a hemisphere becomes one of the mixture model's "normal" tissue classes, and the asymmetry map removes slow left/right trends (see Evaluation).

  • DICOM extraction is tested on one GE 3T study. Other vendors' series naming may need a hand-written roles.json.

Evaluation (blinded benchmark, 40 cases)

Data.

  • 30 cases drawn at random (seed 20260928) from the ISLES 2022 stroke training set (CC BY 4.0). Each has FLAIR, DWI, ADC and an expert lesion mask.

  • 10 healthy volunteers from OpenNeuro ds007401 (IDEAS II, CC0). Their ADC and b1000 DWI were computed from the multi-shell diffusion data. They were skull-stripped with this pipeline's T1 mask so that they look like ISLES data.

  • Five volunteers were excluded and replaced by the next ones in the random draw: four have no DWI files on OpenNeuro, and one failed the quality gates.

Blinding.

  • Every case got a neutral id and only FLAIR, DWI and ADC.

  • 8 independent LLM readers (Claude, 5 cases each) worked only through the MCP tools (eval/mcpcall.py). They were told that some scans may be normal.

  • An audit of all 503 tool calls found no access to the key, the masks or case files.

  • The annotations were frozen (SHA-256) before scoring with eval/score_benchmark.py. A point counts as a hit within 10 mm of a lesion.

  • To reproduce, run eval/prepare_benchmark.py, then eval/analyze_all.py, then the readers, then eval/score_benchmark.py.

LLM pre-read (high/medium items)

algorithm "review" candidates

stroke cases whose lesion was pointed at

24 / 29 (83%, 95% CI 66-92%)

20 / 29 (69%)

items that lie on a lesion (stroke cases)

53 / 62 (85%, CI 75-92%)

51 / 249 (21%)

healthy controls with any item

0 / 10 (CI 0-28%)

10 / 10 (11 candidates each)

lesion components ≥ 10 ml / 1-10 ml

13 / 13 · 10 / 16

lesion components 0.1-1 ml / < 0.1 ml

18 / 70 · 8 / 127

  • One ISLES case has an empty mask (no labelled lesion). It is left out of the stroke rows. The reader flagged two FLAIR white-matter areas in it.

  • Four of the five missed stroke cases had lesions of 1.2 ml or less. The fifth (7.2 ml) was missed while the reader flagged FLAIR white-matter change elsewhere.

  • Small lesions are the weak spot: most of the 226 lesion components are tiny satellite foci under 1 ml. On 4-6 mm slices those often span one or two slices.

  • The heterogeneity maps alone flag something in every scan, healthy ones included, and only about 1 in 5 of their candidates lies on a lesion. The value comes from the reader checking them visually, and from the systematic screening that also finds lesions the maps miss (for example very large ones).

What this does not show.

  • Blinding is imperfect: the controls come from one other site and scanner, so a reader could partly tell them apart.

  • There are only 10 controls (a 0/10 false-alarm rate still allows up to 28%), and all abnormal cases are acute strokes, which DWI shows well. Other diseases, subtle findings and contrast-enhanced studies are untested.

  • The readers are the same model family that built the tool, and there is no radiologist comparison.

  • Controls have no labels, so an incidental real finding would have counted as a false alarm.

  • This is a feasibility check, not a clinical validation.

  • An earlier 3-case pilot on ISLES found bugs that were fixed before this run: ADC units, the registration drift check, oblique voxel spacing, and skull stripping of defaced or loosely masked T1s.

None of the building blocks is new. What this project adds is the combination: a dependency-free multi-sequence viewer, classical anomaly candidates, and an MCP loop in which the LLM has to verify each candidate visually and writes review areas live into the viewer the human is using.

License

MIT © 2026 Jordi Vallverdu. The software is provided "as is", without warranty of any kind. It is not intended for clinical use.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to query and analyze medical imaging metadata from DICOM servers, including patient information, studies, series, and instances, as well as extract text from encapsulated PDF documents.
    101
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides AI-powered medical image analysis tools for LLM agents, enabling tasks such as X-ray classification, interactive segmentation, and visual question answering. It supports multi-step diagnostic reasoning and clinical workflows through a suite of specialized medical AI models.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to control a local CBCT viewer for dental and maxillofacial scans, with navigation verbs like open scan, set window, navigate slices, and snapshot, without executing code or returning interpretations.
    7
    AGPL 3.0