Healthcare AI that holds up in the real world.
Across medical imaging, pathogen genomics and health markers, the same thing goes wrong: a model that looks brilliant on its home data quietly fails on the next scanner, sequencer, site or population. Chikitsa is a research collaborative that studies why — and proves what actually generalises. We pick a hypothesis, adapt proven models, validate leave-one-site-out, and keep only what survives.
One method, applied wherever medical AI meets the real world.
We are deepest in chest imaging today, and extending the same generalisation-first, hypothesis-driven method into genomics and biomarkers.
Medical imaging
A 15-reader consensus engine for chest X-ray — finding detection, localisation, calibration and cross-site robustness. MRI, CT and other X-rays are joining the vertical next.
- Multi-reader consensus & abstention
- Leave-one-site-out validation
- MRI / CT / other X-rays — coming
Genomics & pathogens
Metagenomic next-generation sequencing (mNGS) for unbiased pathogen detection and antimicrobial-resistance signal — asking whether the imaging generalisation gap has a genomic twin.
- Assembly-free pathogen profiling
- AMR-gene & resistance-phenotype signal
- Platform / batch-shift robustness
Health markers & multi-omics
Can a compact panel of routine blood, inflammatory and molecular markers — fused with imaging — triage and stratify patients better than any single modality alone?
- Multi-modal fusion hypotheses
- Treatment-response stratification
- Low-cost, deployable panels
Hypothesis → Adapt → Validate → keep only what survives.
Every capability, in every vertical, has to earn its place. We don't ship what merely sounds good — we ship what measures well on data it has never seen.
State a falsifiable claim
A specific, testable question — "does a heavier encoder generalise better?" — with a metric that can prove it wrong.
Fine-tune proven models
Frozen features, LoRA or light probes on published, open-weight models — not fragile single-site full fine-tunes.
Leave-one-site-out
Scored on held-out sites, cohorts and populations — never the data it trained on. Label-noise audited.
Only survivors ship
Confirmed hypotheses enter the pipeline; rejected ones are documented so we — and you — don't repeat them.
This loop is why our results include as many rejected hypotheses as proven ones. See the imaging hypothesis ledger →
The generalization gap is medical AI's real failure mode.
A model can learn to recognise the instrument or site from the data itself — a shortcut that inflates in-house scores and collapses on deployment. The cause is usually label definitions and acquisition shift, not the biology. Closing that gap is the point of our work, whether the signal is a pixel or a read.
Same model, measured three ways
Robustness is engineered, not hoped for.
We treat cross-machine reliability as three separate, measurable problems — and study each on its own terms:
Discrimination comes from balanced multi-source training. Confidence comes from per-site calibration. Safety comes from consensus, out-of-distribution flags, and knowing when to abstain.
Every claim is scored leave-one-site-out, never within-site — because within-site is the number that misleads.
Bring a hypothesis and data. Get a validated answer.
Our infrastructure isn't a demo — it's the reproducible harness we run every experiment through. Collaborators plug their data in and get de-identification, cross-site validation and calibrated, human-in-the-loop outputs, without building any of it themselves.
Ingest
De-identify & canonicalize — remove instrument / site shortcuts.
Gate
Abstain on out-of-scope or unreadable inputs.
Model
Adapted, open-weight models score independently.
Validate
Consensus + calibration + leave-one-site-out scoring.
Review
Calibrated draft for expert review & export.
We only study what we can measure.
A hard rule across every vertical: nothing enters the pipeline without measured accuracy on labelled data, published validation, or deterministic correctness. Selected findings from our own imaging experiments:
The gap we study
Cross-site AUC recovery is the lever the whole programme is built around.
Calibration error cut
Per-pathology recalibration, with identical AUC — confidence you can trust.
Rule-out yield
Geographically-diverse normals nearly doubled safe auto-clear at 98% sensitivity.
3 countries, 2 label regimes
Radiologist- and report-labelled, validated leave-one-site-out.
Figures are research findings from internal validation, not regulatory or performance claims. Chikitsa is a non-commercial research project; outputs are AI-generated drafts for expert review, not a cleared diagnostic device.
Start a research conversation.
Bring a new scanner or sequencer, a hard case, a cohort, or a research question. We collaborate with clinicians, academic labs, screening programmes and research engineers. We read every message and reply personally.
Chikitsa is a research collaborative. We do not sell a product; we study how to make healthcare AI reliable across the real world.
Let's study healthcare AI that survives contact with the real world.
Whether your signal is an image, a genome, or a biomarker — the generalization gap only closes together.