Research note: in our cross-site study, leading CXR models drop ~0.90 → ~0.70 AUC on unseen scanners — the gap our whole method is built to close. See the finding →
A healthcare-AI research collaborative

Healthcare AI that holds up in the real world.

Across medical imaging, pathogen genomics and health markers, the same thing goes wrong: a model that looks brilliant on its home data quietly fails on the next scanner, sequencer, site or population. Chikitsa is a research collaborative that studies why — and proves what actually generalises. We pick a hypothesis, adapt proven models, validate leave-one-site-out, and keep only what survives.

3 research verticals — imaging, genomics, health markers
15+ independently-validated models in one consensus engine
Open CPU-first research stack · reproducible · human-in-the-loop
How we work

Hypothesis → Adapt → Validate → keep only what survives.

Every capability, in every vertical, has to earn its place. We don't ship what merely sounds good — we ship what measures well on data it has never seen.

01 · HYPOTHESIS

State a falsifiable claim

A specific, testable question — "does a heavier encoder generalise better?" — with a metric that can prove it wrong.

02 · ADAPT

Fine-tune proven models

Frozen features, LoRA or light probes on published, open-weight models — not fragile single-site full fine-tunes.

03 · VALIDATE

Leave-one-site-out

Scored on held-out sites, cohorts and populations — never the data it trained on. Label-noise audited.

04 · KEEP / REJECT

Only survivors ship

Confirmed hypotheses enter the pipeline; rejected ones are documented so we — and you — don't repeat them.

This loop is why our results include as many rejected hypotheses as proven ones. See the imaging hypothesis ledger →

The thread that unifies every vertical

The generalization gap is medical AI's real failure mode.

A model can learn to recognise the instrument or site from the data itself — a shortcut that inflates in-house scores and collapses on deployment. The cause is usually label definitions and acquisition shift, not the biology. Closing that gap is the point of our work, whether the signal is a pixel or a read.

Same model, measured three ways

Chest-imaging finding detection · AUC
In-domaintested on its training site
~0.90
Cross-siteheld-out unseen scanner
~0.70
Our stackmulti-source + calibrated
recovered
Illustrative of our leave-one-site-out experiments across five independent cohorts in three countries. Specific model and dataset names withheld; methods are described on the imaging page.
Our thesis

Robustness is engineered, not hoped for.

We treat cross-machine reliability as three separate, measurable problems — and study each on its own terms:

Discrimination comes from balanced multi-source training. Confidence comes from per-site calibration. Safety comes from consensus, out-of-distribution flags, and knowing when to abstain.

Every claim is scored leave-one-site-out, never within-site — because within-site is the number that misleads.

A working research pipeline

Bring a hypothesis and data. Get a validated answer.

Our infrastructure isn't a demo — it's the reproducible harness we run every experiment through. Collaborators plug their data in and get de-identification, cross-site validation and calibrated, human-in-the-loop outputs, without building any of it themselves.

01

Ingest

De-identify & canonicalize — remove instrument / site shortcuts.

02

Gate

Abstain on out-of-scope or unreadable inputs.

03

Model

Adapted, open-weight models score independently.

04

Validate

Consensus + calibration + leave-one-site-out scoring.

05

Review

Calibrated draft for expert review & export.

See the full pipeline & how it helps research →
Evidence, not adjectives

We only study what we can measure.

A hard rule across every vertical: nothing enters the pipeline without measured accuracy on labelled data, published validation, or deterministic correctness. Selected findings from our own imaging experiments:

0.70→0.90

The gap we study

Cross-site AUC recovery is the lever the whole programme is built around.

~75×

Calibration error cut

Per-pathology recalibration, with identical AUC — confidence you can trust.

+20 pts

Rule-out yield

Geographically-diverse normals nearly doubled safe auto-clear at 98% sensitivity.

5 cohorts

3 countries, 2 label regimes

Radiologist- and report-labelled, validated leave-one-site-out.

Figures are research findings from internal validation, not regulatory or performance claims. Chikitsa is a non-commercial research project; outputs are AI-generated drafts for expert review, not a cleared diagnostic device.

Collaborate with us

Start a research conversation.

Bring a new scanner or sequencer, a hard case, a cohort, or a research question. We collaborate with clinicians, academic labs, screening programmes and research engineers. We read every message and reply personally.

Chikitsa is a research collaborative. We do not sell a product; we study how to make healthcare AI reliable across the real world.

Let's study healthcare AI that survives contact with the real world.

Whether your signal is an image, a genome, or a biomarker — the generalization gap only closes together.