A detailed account of what the system does, how broadly it has been validated, and how well it performs — for technical and clinical due diligence. Model architectures, training datasets, and methods are deliberately omitted here; those live in a separate confidential engineering record shared under agreement.
reader panel, each with a named operating point
countries · 15 public datasets · 2 labelling regimes
never seen in training — incl. a fresh multi-site benchmark
~6 s/scan · no GPU required · DICOM in/out
An assistive, non-autonomous second reader for adult PA/AP chest radiographs, in two roles: (1) worklist prioritisation via an advisory normal/abnormal score; (2) finding-level decision support, each finding returning present/flagged with a stated operating point, to be confirmed by a qualified radiologist. Not a diagnosis; not validated for paediatric, lateral-only, or portable-ICU edge cases without site validation.
Fifteen findings plus an advisory triage gate, combined by a reliability-weighted multi-reader consensus that fails open (a crashed reader never reads as "normal"). Each finding ships a measured operating point — the shipped point targets ≈90% specificity; a high-sensitivity point is available per finding. These are strong rule-outs (high NPV).
| Finding | AUC | Sens | Spec | NPV |
|---|---|---|---|---|
| Pleural effusion | 0.91 | 0.76 | 0.90 | 0.95 |
| Pneumothorax | 0.90 | 0.76 | 0.90 | 0.99 |
| Cardiomegaly | 0.85 | 0.68 | 0.90 | 0.95 |
| Atelectasis | 0.84 | 0.55 | 0.90 | 0.98 |
| Normal / abnormal triage (advisory) | 0.90 | 0.66 | 0.90 | 0.74 |
| Finding | AUC | Op point (sens/spec) |
|---|---|---|
| Diaphragmatic hernia | 0.92 | 0.85 / 0.90 |
| Pulmonary fibrosis | 0.76 | 0.39 / 0.90 |
| Consolidation | 0.75 | 0.43 / 0.90 |
| Calcification | 0.75 | 0.43 / 0.90 |
| Pleural thickening | 0.74 | 0.31 / 0.90 |
| Finding | AUC | Note |
|---|---|---|
| Tuberculosis screen | 0.72 | honest cross-site figure across two cohorts |
| Aortic enlargement | 0.97 ⚠ | single-source — optimistic; a second cohort is required |
| Mass | 0.72 | modest cross-site; advisory weight |
| Emphysema | 0.71 | modest cross-site; advisory weight |
| Nodule / mass (with localisation) | 0.87 (in-domain) | real bounding boxes from a trained detector |
Every headline number is measured leave-site-out — thresholds chosen on some cohorts, performance reported on a held-out one — across 15 public datasets spanning 6 countries and both labelling regimes (expert-annotated and report-derived). Beyond that, the deployed readers were stress-tested unchanged on three cohorts never seen in training: a US community hospital, an independent ~6,000-study multi-site benchmark, and an East-Asian cohort. Discrimination (AUC) transfers; the operating point is re-calibrated per deployment site.
| Finding | Leave-site-out | US community hospital | Independent benchmark |
|---|---|---|---|
| Pleural effusion | 0.90 | 0.95 | 0.92 |
| Cardiomegaly | 0.90 | 0.93 | 0.88 |
| Pneumothorax | 0.88 | 0.95 | 0.87 |
| Consolidation | 0.82 | 0.96 | 0.82 |
| Atelectasis | 0.82 | 0.84 | 0.82 |
The featured panel holds across every external cohort — the strongest evidence in this document that performance survives a change of hospital and scanner.
~6 s/scan, no GPU needed — deployable on-prem or at the edge
standards import + structured export (SEG / SR); PACS pilot-ready
scanner-normalised input; new-scanner calibration advisory
The system is the product of a disciplined, measure-before-you-build programme: 15 hypotheses tested — several rejected on the data, five foundation encoders benchmarked head-to-head, and a reproducible method for separating genuine "hard finding" from "noisy label." The result is a panel where we can state, per finding, exactly how far it has been proven and where it stops. (The methods and datasets behind these results are in the confidential engineering record.)