Dhup Bhukdee / Experimental open

The case-mix illusion

Change the sample

Area under the ROC curve (AUC) can look excellent in a pooled sample and much less impressive within each setting. Change who enters the sample. The discrimination inside each setting stays the same.

Synthetic population mathematics. No patient data, uploads or clinical validation.

The static referral example is shown. Interactive controls are preparing; reload if they remain unavailable.

Pooled AUC0.946All case-control pairs
Within setting A0.760A cases vs A controls
Within setting B0.760B cases vs B controls
The pooled AUC is 0.185 higher than the AUC inside either setting.

The pooled score distributions

Each curve has total area 1
CasesControls

In the referral example, the case density is mostly shifted right of the control density. Enable JavaScript to draw and change these distributions.

The curve shapes change because the mixture changes. The component distributions stay fixed while you adjust the first two sliders.

The ROC changes, too

PooledWithin either setting

The referral mixture has pooled AUC 0.946; the within-setting curve has AUC 0.760. Enable JavaScript to draw the receiver operating characteristic (ROC) curves.

AUC asks a pairwise question

Draw one case and one control independently. How often does the case have the higher score?

Pooling includes pairs from different settings. If cases mostly come from B and controls mostly come from A, the setting offset helps rank those pairs.

The pooled AUC is valid for the specified mixture. It does not establish discrimination inside a setting or in a different population.

Four kinds of comparison

Each card shows the probability that the case outscores the control. The bar shows how much of the pooled pair population that comparison occupies.

A case / A control
0.760
9.0% of all pairs
A case / B control
0.079
1.0% of all pairs
B case / A control
0.998
81.0% of all pairs
B case / B control
0.760
9.0% of all pairs
Show the weighted calculation
AUC weighted contributions by setting pair
Case / controlPair weightAUC contribution
A / A9.0%0.0684
A / B1.0%0.0008
B / A81.0%0.8081
B / B9.0%0.0684
Total100.0%0.9457

Contributions use unrounded values. Displayed numbers may not add exactly.

A population belongs beside an AUC.

For this toy, a change in referral or sampling patterns can raise or lower pooled AUC without changing either setting's conditional distributions. Try Reversed mix to see pooled discrimination fall below chance even though both within-setting AUCs remain above chance.

Changing disease prevalence alone, while holding the case and control score distributions fixed, does not change this population AUC. The sliders change the composition within cases and controls. That is a different operation.

Model, formula, and limits

Let d be disease status and g be setting, each coded 0 or 1. The score is S = d + δg + ε, with independent ε ~ N(0,1). The score is an arbitrary measurement, not a calibrated risk.

AUC = Σᵢ Σⱼ wᵢ⁺ wⱼ⁻ Φ((1 + δ(i − j)) / √2)

Here w⁺ = [1 − p, p] is the case mix; w⁻ = [1 − q, q] is the control mix. Φ is the standard normal CDF. Each within-setting AUC is Φ(1/√2) ≈ 0.76025. There are no ties in this continuous model.

Values are stipulated population quantities, not estimates from sampled patients. At a slider endpoint, a setting may contain no sampled cases or controls; its displayed AUC still describes the stipulated model and would not be estimable from that sample. No sample sizes, uncertainty intervals, threshold utility, or clinical validation are implied.

The setting offset is deliberately generic. It could represent a measurement shift or another setting-linked score feature. This example illustrates a mechanism; it does not attribute a real study's performance to that mechanism.

Reading behind the experiment

Hanley & McNeil, Radiology, 1982 established the pairwise probability interpretation used here.

Ransohoff & Feinstein, NEJM, 1978 examined spectrum and bias in diagnostic-test evaluation. The distributions and numbers on this page are our synthetic illustration, not their study results.

Experiment updated 2026-10-01. Part of dhupbh.com. IBM Plex, SIL Open Font License. Slider values stay in this browser and are not recorded by page analytics.

Change the sample