A consensus engine for chest imaging — built to be right on scans it has never seen.
Fifteen independent, published models read every chest X-ray and vote. We study how to make that ensemble generalise across scanners, sites and populations — through foundation-encoder selection, careful fine-tuning, honest calibration and a discipline of proving (and rejecting) hypotheses. MRI, CT and other X-rays are joining this vertical next.
Many validated readers, one calibrated consensus.
Each reader is an independent model with open weights and a measured operating point. They disagree usefully — a finding is surfaced only when readers agree; lone flags are demoted for review, not hidden.
Canonicalize
De-identify & normalize away vendor LUT / windowing shift.
Gate
Abstain on non-frontal, wrong-modality or unreadable inputs.
Read
Every reader scores independently, with localisation.
Consensus
Weighted agreement + per-site calibration + OOD advisory.
Review
Calibrated draft for radiologist review & export.
Every finding, localized and attributed on the image.
The ensemble doesn't just say what — it says where, and who saw it. Attention boxes and trained segmentation masks are drawn back onto the scan; a finding is marked consensus only when readers agree.
Reading the annotation
Each box is a claim a reader can defend — and the colour tells you which kind of evidence produced it.
Localization is normalized back to the original scan and colour-coded by finding category; attention boxes are tightened to the hot core so they mark the finding, not a third of the film.
Fine-tuning & encoder studies, scored the honest way.
Every number below is from our own experiments, scored leave-one-site-out. Specific model and dataset names are withheld; methods are described. Values are research findings, not regulatory claims.
Foundation-encoder bake-off
The model that wins at home can lose away
Calibration: making confidence mean something
Diverse normals raise safe rule-out yield
Which findings are trustworthy — and where they aren't.
The hardest lesson in the whole programme: reliability is finding-specific. Some findings learn cleanly everywhere; some are intrinsically hard; and some collapse only because their report-mined labels are noise, not because the disease is. This matrix is why we build clean-anchored specialists for exactly the red cells.
cohort A
cohort B
cohort C
labels
Qualitative summary of our cross-cohort validation; cohorts anonymised. Effusion, cardiomegaly and pneumothorax are genuinely reliable; consolidation and pneumonia are mostly label problems (fixed by adding a clean anchor); atelectasis is near its intrinsic ceiling even with perfect labels.
What we've proven — and, just as usefully, rejected.
A research group is only as honest as the ideas it's willing to kill. These are hypotheses we tested on held-out data and let the numbers decide.
The real failure mode is the site, not the model.
In-domain finding-detection AUC ~0.90 collapses to ~0.70 on an unseen scanner — a shift driven by label definitions and acquisition, not the disease.
Add one clean-labelled anchor and noisy sources stop hurting.
Training a reader on report-mined labels alone is the failure; adding a radiologist-labelled anchor to the mix beats noisy-only by up to +0.43 AUC on a clean test.
Calibration fixes confidence without touching discrimination.
Per-pathology recalibration cut calibration error ~75× while AUC stayed identical — the slope is clamped so recalibration can never re-rank.
A self-supervised medical encoder generalises best.
Across three radiologist cohorts, a self-supervised foundation encoder beat a generic one by several AUC points and out-travelled a heavier supervised model externally.
A heavier, supervised-pretrained encoder generalises better.
It won in-domain (+0.033 AUC) but lost on a truly external site (−0.014). The apparent edge was pretraining overlap, not transfer.
Single-site full fine-tuning is the way to adapt.
Full fine-tuning gains a little in-domain but loses out-of-distribution — it distorts pretrained features to the source site.
More models in the ensemble always help; so does flip-TTA.
A larger cross-domain ensemble lowered in-domain AUC, and horizontal-flip test-time augmentation regressed laterality-dependent findings.
Conditioning cardiomegaly on view (AP) fixes false positives.
The assumed view-conditioning fix didn't hold on our data; the real driver was a flat cardiothoracic-ratio threshold over-calling ~44% of screening films.
MRI, CT and other X-rays — same method, new modalities.
The generalization gap isn't unique to chest radiographs. We're extending the ensemble to published, open-weight models for other modalities, carrying the same discipline: adapt lightly, validate leave-one-site-out, calibrate per site, keep only what survives.
CT
Volumetric chest CT — nodule characterisation and risk models, evaluated for cross-scanner robustness before anything is adopted.
MRI
Sequence-aware models where acquisition shift is even larger than in X-ray — a demanding test of our normalization-first approach.
Other radiographs
Musculoskeletal and abdominal projections, reusing the consensus, calibration and abstention machinery already validated on chest.
Have imaging data from a scanner we've never seen?
That's exactly the test our method is built for — and the best kind of collaboration.