The case-mix illusion
Area under the ROC curve (AUC) can look excellent in a pooled sample and much less impressive within each setting. Change who enters the sample. The discrimination inside each setting stays the same.
Synthetic population mathematics. No patient data, uploads or clinical validation.
The static referral example is shown. Interactive controls are preparing; reload if they remain unavailable.
02 / Compare discrimination
The pooled score distributions
Each curve has total area 1In the referral example, the case density is mostly shifted right of the control density. Enable JavaScript to draw and change these distributions.
The curve shapes change because the mixture changes. The component distributions stay fixed while you adjust the first two sliders.
The ROC changes, too
The referral mixture has pooled AUC 0.946; the within-setting curve has AUC 0.760. Enable JavaScript to draw the receiver operating characteristic (ROC) curves.
AUC asks a pairwise question
Draw one case and one control independently. How often does the case have the higher score?
Pooling includes pairs from different settings. If cases mostly come from B and controls mostly come from A, the setting offset helps rank those pairs.
The pooled AUC is valid for the specified mixture. It does not establish discrimination inside a setting or in a different population.
Four kinds of comparison
Each card shows the probability that the case outscores the control. The bar shows how much of the pooled pair population that comparison occupies.
Show the weighted calculation
| Case / control | Pair weight | AUC contribution |
|---|---|---|
| A / A | 9.0% | 0.0684 |
| A / B | 1.0% | 0.0008 |
| B / A | 81.0% | 0.8081 |
| B / B | 9.0% | 0.0684 |
| Total | 100.0% | 0.9457 |
Contributions use unrounded values. Displayed numbers may not add exactly.
A population belongs beside an AUC.
For this toy, a change in referral or sampling patterns can raise or lower pooled AUC without changing either setting's conditional distributions. Try Reversed mix to see pooled discrimination fall below chance even though both within-setting AUCs remain above chance.
Changing disease prevalence alone, while holding the case and control score distributions fixed, does not change this population AUC. The sliders change the composition within cases and controls. That is a different operation.
Model, formula, and limits
Let d be disease status and g be setting, each coded 0 or 1. The score is S = d + δg + ε, with independent ε ~ N(0,1). The score is an arbitrary measurement, not a calibrated risk.
Here w⁺ = [1 − p, p] is the case mix; w⁻ = [1 − q, q] is the control mix. Φ is the standard normal CDF. Each within-setting AUC is Φ(1/√2) ≈ 0.76025. There are no ties in this continuous model.
Values are stipulated population quantities, not estimates from sampled patients. At a slider endpoint, a setting may contain no sampled cases or controls; its displayed AUC still describes the stipulated model and would not be estimable from that sample. No sample sizes, uncertainty intervals, threshold utility, or clinical validation are implied.
The setting offset is deliberately generic. It could represent a measurement shift or another setting-linked score feature. This example illustrates a mechanism; it does not attribute a real study's performance to that mechanism.
Reading behind the experiment
Hanley & McNeil, Radiology, 1982 established the pairwise probability interpretation used here.
Ransohoff & Feinstein, NEJM, 1978 examined spectrum and bias in diagnostic-test evaluation. The distributions and numbers on this page are our synthetic illustration, not their study results.
Experiment updated 2026-10-01. Part of dhupbh.com. IBM Plex, SIL Open Font License. Slider values stay in this browser and are not recorded by page analytics.