Chikitsa

Chest-radiograph AI — capability datasheet

A detailed account of what the system does, how broadly it has been validated, and how well it performs — for technical and clinical due diligence. Model architectures, training datasets, and methods are deliberately omitted here; those live in a separate confidential engineering record shared under agreement.

Retrospective validation only · not cleared for clinical use · assistive decision support, confirmed by a qualified radiologist.

Findings

15

reader panel, each with a named operating point

Validation breadth

6

countries · 15 public datasets · 2 labelling regimes

External cohorts

3

never seen in training — incl. a fresh multi-site benchmark

Runs on

CPU

~6 s/scan · no GPU required · DICOM in/out

Intended use

An assistive, non-autonomous second reader for adult PA/AP chest radiographs, in two roles: (1) worklist prioritisation via an advisory normal/abnormal score; (2) finding-level decision support, each finding returning present/flagged with a stated operating point, to be confirmed by a qualified radiologist. Not a diagnosis; not validated for paediatric, lateral-only, or portable-ICU edge cases without site validation.

What it detects — the reader panel

Fifteen findings plus an advisory triage gate, combined by a reliability-weighted multi-reader consensus that fails open (a crashed reader never reads as "normal"). Each finding ships a measured operating point — the shipped point targets ≈90% specificity; a high-sensitivity point is available per finding. These are strong rule-outs (high NPV).

Tier 1 — featured (cross-site validated, strong)

FindingAUCSensSpecNPV
Pleural effusion0.910.760.900.95
Pneumothorax0.900.760.900.99
Cardiomegaly0.850.680.900.95
Atelectasis0.840.550.900.98
Normal / abnormal triage (advisory)0.900.660.900.74

Tier 2 — supporting (cross-site, moderate)

FindingAUCOp point (sens/spec)
Diaphragmatic hernia0.920.85 / 0.90
Pulmonary fibrosis0.760.39 / 0.90
Consolidation0.750.43 / 0.90
Calcification0.750.43 / 0.90
Pleural thickening0.740.31 / 0.90

Tier 3 — investigational (weaker / advisory only)

FindingAUCNote
Tuberculosis screen0.72honest cross-site figure across two cohorts
Aortic enlargement0.97 ⚠single-source — optimistic; a second cohort is required
Mass0.72modest cross-site; advisory weight
Emphysema0.71modest cross-site; advisory weight
Nodule / mass (with localisation)0.87 (in-domain)real bounding boxes from a trained detector

How broadly it generalises

Every headline number is measured leave-site-out — thresholds chosen on some cohorts, performance reported on a held-out one — across 15 public datasets spanning 6 countries and both labelling regimes (expert-annotated and report-derived). Beyond that, the deployed readers were stress-tested unchanged on three cohorts never seen in training: a US community hospital, an independent ~6,000-study multi-site benchmark, and an East-Asian cohort. Discrimination (AUC) transfers; the operating point is re-calibrated per deployment site.

FindingLeave-site-outUS community hospitalIndependent benchmark
Pleural effusion0.900.950.92
Cardiomegaly0.900.930.88
Pneumothorax0.880.950.87
Consolidation0.820.960.82
Atelectasis0.820.840.82

The featured panel holds across every external cohort — the strongest evidence in this document that performance survives a change of hospital and scanner.

How the results are delivered

Deployment characteristics

Compute

CPU-first

~6 s/scan, no GPU needed — deployable on-prem or at the edge

Interoperability

DICOM

standards import + structured export (SEG / SR); PACS pilot-ready

Robustness

fail-open

scanner-normalised input; new-scanner calibration advisory

Depth of the underlying research

The system is the product of a disciplined, measure-before-you-build programme: 15 hypotheses tested — several rejected on the data, five foundation encoders benchmarked head-to-head, and a reproducible method for separating genuine "hard finding" from "noisy label." The result is a panel where we can state, per finding, exactly how far it has been proven and where it stops. (The methods and datasets behind these results are in the confidential engineering record.)

Honest limitations

What a partner funds next

← Summary← Overview