Brilliant at home.
Blind on the next scanner.
Across medical imaging, pathogen genomics and health markers, models ace the data they were trained on and quietly fail on the next scanner, sequencer or population. Chikitsa studies why — and proves what actually generalises.
One method, wherever medical AI meets the real world.
Deepest in chest imaging today; extending the same generalisation-first method into genomics and biomarkers.
Medical imaging
A 15-reader consensus engine for chest X-ray — detection, localisation, calibration, cross-site robustness. MRI & CT next.
Genomics & pathogens
Metagenomic sequencing (mNGS) for unbiased pathogen and AMR detection — asking if the gap has a genomic twin.
Health markers
Can a compact panel of routine markers, fused with imaging, triage better than any single modality alone?
The model works. The data doesn't travel.
A model can learn the hospital from the pixels — a shortcut that inflates in-house scores and collapses on deployment. Four ways it goes wrong:
It aced its home data.
In-domain finding-detection reaches ~0.90 AUC — and everyone celebrates.
It fails on the next scanner.
On an unseen site the same model drops to ~0.70. Nothing about the disease changed.
The labels don't agree.
The collapse is usually label definitions and acquisition shift — not pixels, not biology.
The confidence can't be trusted.
An uncalibrated "80%" means nothing across sites — so no one can act on it.
Same model, measured three ways.
Chest-imaging finding detection
Robustness is engineered, not hoped for.
Discrimination comes from balanced multi-source training. Confidence comes from per-site calibration. Safety comes from consensus, out-of-distribution flags, and knowing when to abstain.
Every claim is scored leave-one-site-out — because within-site is the number that misleads.
We only ship what we can measure.
No capability enters the pipeline without measured accuracy on labelled data, published validation, or deterministic correctness.
The gap we close
Cross-site AUC recovery — the lever the whole programme is built around.
Calibration error cut
Per-pathology recalibration, identical AUC — confidence you can trust.
Rule-out yield
Diverse normals nearly doubled safe auto-clear at 98% sensitivity.
3 countries, 2 label regimes
Radiologist- and report-labelled, validated leave-one-site-out.
Figures are research findings from internal validation, not regulatory or performance claims. Outputs are AI-generated drafts for expert review, not a cleared diagnostic device.
Two ways to score a model. Only one holds up.
The number you report decides whether the model survives contact with the real world.
The number that flatters
The number that holds
Hypothesis → Adapt → Validate → keep only what survives.
Every capability, in every vertical, has to earn its place on data it has never seen.
State a falsifiable claim
A specific, testable question with a metric that can prove it wrong.
Fine-tune proven models
Frozen features, LoRA or light probes on open-weight models — not fragile full fine-tunes.
Leave-one-site-out
Scored on held-out sites and populations, never the training data. Label-noise audited.
Only survivors ship
Confirmed hypotheses enter the pipeline; rejected ones are documented.
Bring a hypothesis and data. Get a validated answer.
The reproducible harness we run every experiment through — de-identification, cross-site validation and calibrated outputs, without building any of it.
Ingest
De-identify & canonicalize — remove site shortcuts.
Gate
Abstain on out-of-scope or unreadable inputs.
Model
Adapted, open-weight models score independently.
Validate
Consensus + calibration + leave-one-site-out.
Review
Calibrated draft for expert review & export.
Start a research conversation.
Bring a new scanner or sequencer, a hard case, a cohort, or a research question. We collaborate with clinicians, academic labs, screening programmes and research engineers. We read every message and reply personally.
Chikitsa is a research collaborative. We do not sell a product; we study how to make healthcare AI reliable across the real world.
Let's study healthcare AI that survives contact with the real world.
Whether your signal is an image, a genome, or a biomarker — the generalization gap only closes together.