When Every Subgroup Agrees but the Total Reverses the Result
- 0 views
- Last updated
- Statistics
A concrete treatment study reveals Simpson's paradox: treatment succeeds more often among both mild and severe cases, yet appears much worse after the groups are pooled. Actual counts, weighted averages, and a geometric mixture diagram expose how unequal group sizes and sharply different baseline risks create the reversal. The lecture concludes with practical guidance on when pooled rates answer a useful question and when subgroup adjustment is essential.
Suppose two hospitals are comparing a treatment with ordinary care. The treatment will perform better among mild cases, and better among severe cases. Then we will pool everyone, and the apparent winner will reverse. That reversal is Simpson's paradox. Here is the question in its sharpest form. Can every clinically relevant subgroup favor the treatment while the total favors no treatment? Yes. The important task is not merely to be surprised. It is to see exactly how the reversal is made. We will use one thousand treated patients and one thousand untreated patients. The entries are successes over total cases, followed by the corresponding success rate. Start with mild cases. Among the treated patients, ninety of one hundred succeed. That is a ninety percent success rate. Without treatment, seven hundred twenty of nine hundred mild cases succeed. That is eighty percent. Within mild cases, treatment is ahead by ten percentage points. Now take severe cases. Among treated patients, two hundred seventy of nine hundred succeed, giving thirty percent. Without treatment, twenty of one hundred severe cases succeed, giving twenty percent. Treatment is again ahead by ten percentage points. Read the mild row once more. Ninety percent with treatment is ten points above eighty percent without treatment. Read the severe row. Thirty percent with treatment is ten points above twenty percent without treatment. There is no subgroup here where ordinary care has the higher success rate. Before revealing the total, make a prediction. Treatment wins by the same margin in both rows. It is natural to expect treatment to win when the rows are added. Now add the actual counts. Treatment has three hundred sixty successes among one thousand patients, for thirty-six percent. No treatment has seven hundred forty successes among one thousand, for seventy-four percent. The pooled comparison therefore says the opposite. Thirty-six minus seventy-four is negative thirty-eight percentage points. Within mild cases, treatment wins by ten points. Within severe cases, treatment wins by ten points. After pooling, treatment loses by thirty-eight points. Every fraction is correct, and the reversal is not an arithmetic error. What changed was the comparison itself. The subgroup rows compare patients with similar severity. The total compares two populations with radically different mixtures of mild and severe cases. That mixture is where the paradox lives.
The pooled rates become understandable as soon as we draw who entered each comparison. Blue means mild cases and red means severe cases. Each strip contains one thousand patients. The treatment strip is almost entirely severe cases. Nine hundred are severe, and only one hundred are mild. The no-treatment strip has exactly the opposite composition. Only one hundred cases are severe, while nine hundred are mild. These labels name the two populations. Treatment is being judged mostly on severe patients. No treatment is being judged mostly on mild patients. Severity matters enormously even before considering treatment. Without treatment, mild cases succeed eighty percent of the time, while severe cases succeed only twenty percent. That is a sixty-point baseline-risk gap. It is much larger than the treatment's ten-point advantage inside either severity group. A pooled rate is a weighted average. For treatment, mild cases have weight one tenth, while severe cases have weight nine tenths. Those weights place ninety percent beside a small group and thirty percent beside a large group. The weighted average is thirty-six percent. For no treatment, the weights switch. Mild cases now receive weight nine tenths, and severe cases receive only one tenth. That weighted average is seventy-four percent. It is high because the no-treatment population contains many more of the patients who had a high baseline chance of success. Pooling has not averaged the same ingredients in the same proportions. It has compared one average dominated by severe cases with another average dominated by mild cases. We can separate the reversal into two forces. First comes the treatment effect visible within either group: a positive ten percentage points. Second comes the composition effect. The treatment population has eighty percentage points fewer mild cases. Mild and severe baseline success rates are sixty points apart. Multiplying those differences gives negative forty-eight points. The unfavorable case mix costs forty-eight points in the pooled comparison. Combine a positive ten-point treatment effect with a negative forty-eight point composition effect. The result is negative thirty-eight points, exactly the aggregate reversal we observed. The total answers a real question: what fraction succeeded in each observed population? It does not answer the like-for-like question: what would happen if treatment and no treatment were applied to populations with the same severity mix?
Now turn the arithmetic into geometry. The horizontal axis records the fraction of cases that are mild. The vertical axis records the resulting success rate. At the far left, the population is entirely severe. Treatment succeeds at thirty percent, while no treatment succeeds at twenty percent. At the far right, the population is entirely mild. Treatment succeeds at ninety percent, while no treatment succeeds at eighty percent. Every mixture between those endpoints is a weighted average. Moving right gives more weight to mild cases, so both overall success rates rise. Most importantly, the red treatment line stays ten percentage points above the green control line at every common mixture. Start with a population that is ten percent mild. At this same mixture, treatment is thirty-six percent and no treatment is twenty-six percent. Sweep the common mixture toward ninety percent mild. The two points move together, and the yellow gap remains ten points from one end to the other. The same-mix formula confirms the picture. Both lines gain sixty points for a full change in mild share, so those matching terms cancel and leave a ten-point treatment advantage. If both groups had the same case mix, every vertical comparison would favor treatment. That is the geometric form of the two subgroup results. But the study did not observe both groups at one horizontal position. Remove the common-mixture comparison and mark the two populations that were actually observed. The treated population sits here, at ten percent mild and thirty-six percent successful. The untreated population sits far to the right, at ninety percent mild and seventy-four percent successful. Those two points are not vertically aligned. The control point benefits from a much easier case mix, so it can sit higher even though its entire green line lies below the red line. The observed comparison reads diagonally across the picture: left-hand red against right-hand green. The like-for-like comparison reads vertically: red against green at one shared mild-case fraction. That is the reversal made visible. Weighted averages are points between subgroup endpoints. Different weights select different horizontal positions, and a large enough horizontal shift can overwhelm the vertical treatment advantage.
The practical question is not whether totals are forbidden. Totals are useful summaries. The question is whether the pooled total answers the decision you believe it answers. Begin with severity. In our example, severity strongly predicts the outcome. Mild cases have a much higher baseline chance of success than severe cases. Severity also predicts which comparison group a patient enters. The treatment group received mostly severe cases, while the untreated group received mostly mild cases. Treatment choice may itself affect the outcome. But severity points toward both treatment choice and outcome, creating a backdoor route that mixes treatment effects with case selection. This is why the first question matters. Does the subgroup variable predict the outcome even without treatment? If it does, changing the subgroup mix can change an aggregate rate. Second, are treatment and comparison cases distributed differently across those subgroups? A risk factor cannot create this particular distortion if both groups carry the same mix. Third, which population is the decision about? A hospital may need the expected result for its actual patient mix. A causal comparison may instead need both options evaluated on one common mix. Pooling is informative when treatment assignment is balanced across the important risk groups. Then each overall rate combines comparable ingredients in comparable proportions. Pooling can also be appropriate when the target mixture is deliberately specified. If a health system expects seventy percent mild and thirty percent severe cases, it can standardize both options to that same mix. In that setting, the pooled answer is not pretending to be universal. It is an answer for a named population, using weights chosen to represent that population. Pooling becomes dangerous when baseline risks differ greatly, treatment exposure is uneven across the risk groups, and the crude total is read as though it were a like-for-like treatment effect. Our example satisfies all three warning conditions. Severity changes baseline success by sixty points. The mild share changes from ten percent to ninety percent. The pooled rates then compare unlike populations. A responsible analysis would report the raw total and the severity-specific results, then use stratification, standardization, regression, matching, or a suitable study design to compare like with like. One final caution. Finding a Simpson reversal does not automatically prove that the subgroup result is causal. Severity may be measured imperfectly, other confounders may remain, and the chosen subgroups may themselves be consequences of earlier decisions. The reversal is a diagnostic signal. It tells us that aggregation and composition matter, and that we should investigate how patients entered each group before drawing a treatment conclusion. So ask what the total literally measures. In this study, thirty-six and seventy-four percent describe outcomes in two observed populations. They are valid summaries of those populations. Then ask what comparison the decision requires. If the goal is a treatment effect, compare treatment and no treatment at the same severity mix. If the goal is operational forecasting, use the mix expected in the population where the decision will be applied. The lesson of Simpson's paradox is not simply to distrust averages. It is to identify their weights, understand what creates those weights, and make sure both sides of a comparison answer the same question for the same target population.
Loading discussion…