Reading pathogens straight from the sample — and asking if AI generalises there too.
Metagenomic next-generation sequencing (mNGS) sequences everything in a clinical sample at once, so it can name a pathogen without being told what to look for. We're building an analysis pipeline around it — and testing whether the same generalisation gap, label-noise and calibration problems we mapped in imaging reappear when the signal is a genome instead of a pixel.
From clinical sample to pathogen & resistance call.
An unbiased, assembly-aware pipeline: deplete the human host, classify what's left against reference genomes, and read antimicrobial-resistance gene content — with the same de-identification and reproducibility discipline as our imaging work.
The questions we're setting out to answer.
This vertical is early — so these are stated as falsifiable hypotheses, each with the leave-one-run-out validation that will prove or reject it. That's the point of the ledger: publish the question before the answer.
The generalization gap has a genomic twin.
A pathogen classifier trained on one sequencing platform, depth or centre will under-detect on another — a batch/depth shift analogous to acquisition shift in imaging.
Assembly-free profiles catch low-abundance pathogens earlier.
K-mer and embedding representations of raw reads flag rare organisms below the depth where de-novo assembly succeeds.
AMR gene content predicts resistance phenotype before culture.
Resistance-gene signal in mNGS reads anticipates the antibiogram, shortening time-to-appropriate-therapy.
Unbiased sequencing finds what a targeted panel misses.
Host-depletion + untargeted sequencing detects co-infections and off-panel organisms that syndromic panels cannot, without prior hypothesis.
Cheap signals, fused well, may beat expensive ones alone.
Routine blood counts, inflammatory markers and molecular assays each carry partial information. The open question is whether fusing them — and fusing them with imaging — produces triage and stratification that no single modality reaches, at a cost that works in real clinics.
Imaging + a compact marker panel triages better than either alone.
Fusing the chest-imaging reader outputs with routine blood and inflammatory markers improves early triage over the best single modality.
Marker trajectories stratify treatment response early.
Longitudinal biomarker trends separate responders from non-responders sooner than endpoint tests — for example in tuberculosis therapy.
Widely-available markers can approximate costly assays for screening.
For triage — not diagnosis — inexpensive, ubiquitous markers may stand in for expensive assays with acceptable loss.
The same calibration discipline transfers across omics.
Per-site calibration and abstention — proven in imaging — make marker and genomic scores equally trustworthy across populations.
These are research hypotheses under active investigation, not validated capabilities or clinical claims. Under our proven-capability rule, nothing is adopted until it measures well on data it has never seen.
Work in sequencing, microbiology or biomarkers?
These questions need paired data and clinical ground truth — the kind of collaboration that turns a hypothesis into a result.