The generalization gap

Brilliant at home.
Blind on the next scanner.

Across medical imaging, pathogen genomics and health markers, models ace the data they were trained on and quietly fail on the next scanner, sequencer or population. Chikitsa studies why — and proves what actually generalises.

3research verticals
15+validated readers · one consensus engine
5cohorts · 3 countries · leave-one-site-out
02 / 04 01HYPOTHESIS 02ADAPT 03VALIDATE 04KEEP·REJECT
Problem

The model works. The data doesn't travel.

A model can learn the hospital from the pixels — a shortcut that inflates in-house scores and collapses on deployment. Four ways it goes wrong:

01

It aced its home data.

In-domain finding-detection reaches ~0.90 AUC — and everyone celebrates.

02

It fails on the next scanner.

On an unseen site the same model drops to ~0.70. Nothing about the disease changed.

03

The labels don't agree.

The collapse is usually label definitions and acquisition shift — not pixels, not biology.

04

The confidence can't be trusted.

An uncalibrated "80%" means nothing across sites — so no one can act on it.

The thread that unifies every vertical

Same model, measured three ways.

Chest-imaging finding detection

AUC · illustrative of leave-one-site-out
In-domaintrained site
~0.90
Cross-siteunseen scanner
~0.70
Our stackmulti-source + calibrated
recovered
Five independent cohorts across three countries. Dataset & model names withheld.
Our thesis

Robustness is engineered, not hoped for.

Discrimination comes from balanced multi-source training. Confidence comes from per-site calibration. Safety comes from consensus, out-of-distribution flags, and knowing when to abstain.

Every claim is scored leave-one-site-out — because within-site is the number that misleads.

Evidence, not adjectives

We only ship what we can measure.

No capability enters the pipeline without measured accuracy on labelled data, published validation, or deterministic correctness.

0.70→0.90

The gap we close

Cross-site AUC recovery — the lever the whole programme is built around.

~75×

Calibration error cut

Per-pathology recalibration, identical AUC — confidence you can trust.

+20 pts

Rule-out yield

Diverse normals nearly doubled safe auto-clear at 98% sensitivity.

5 cohorts

3 countries, 2 label regimes

Radiologist- and report-labelled, validated leave-one-site-out.

Figures are research findings from internal validation, not regulatory or performance claims. Outputs are AI-generated drafts for expert review, not a cleared diagnostic device.

The method, in one picture

Two ways to score a model. Only one holds up.

The number you report decides whether the model survives contact with the real world.

Trained & tested on one site

The number that flatters

ATrain on Site A
ATest on Site A 0.90
BDeploy to Site B 0.70
Looks brilliant. Breaks on deployment.
Validated leave-one-site-out

The number that holds

Train on four sites
5Test on the held-out fifth → the real score
Calibrate per site, then deploy
The score you can actually trust.
How we work

Hypothesis → Adapt → Validate → keep only what survives.

Every capability, in every vertical, has to earn its place on data it has never seen.

01 · HYPOTHESIS

State a falsifiable claim

A specific, testable question with a metric that can prove it wrong.

02 · ADAPT

Fine-tune proven models

Frozen features, LoRA or light probes on open-weight models — not fragile full fine-tunes.

03 · VALIDATE

Leave-one-site-out

Scored on held-out sites and populations, never the training data. Label-noise audited.

04 · KEEP / REJECT

Only survivors ship

Confirmed hypotheses enter the pipeline; rejected ones are documented.

A working research pipeline

Bring a hypothesis and data. Get a validated answer.

The reproducible harness we run every experiment through — de-identification, cross-site validation and calibrated outputs, without building any of it.

01

Ingest

De-identify & canonicalize — remove site shortcuts.

02

Gate

Abstain on out-of-scope or unreadable inputs.

03

Model

Adapted, open-weight models score independently.

04

Validate

Consensus + calibration + leave-one-site-out.

05

Review

Calibrated draft for expert review & export.

Try the live demo ↗See the full pipeline →
Collaborate with us

Start a research conversation.

Bring a new scanner or sequencer, a hard case, a cohort, or a research question. We collaborate with clinicians, academic labs, screening programmes and research engineers. We read every message and reply personally.

Chikitsa is a research collaborative. We do not sell a product; we study how to make healthcare AI reliable across the real world.

Let's study healthcare AI that survives contact with the real world.

Whether your signal is an image, a genome, or a biomarker — the generalization gap only closes together.