{"version":1,"lectureId":"01M14V08QADQC9R72387G10J4E","attempt":0,"publication":{"slug":"simpsons-paradox","title":"When Every Subgroup Agrees but the Total Reverses the Result","subject":"statistics","summary":"A concrete treatment study reveals Simpson's paradox: treatment succeeds more often among both mild and severe cases, yet appears much worse after the groups are pooled. Actual counts, weighted averages, and a geometric mixture diagram expose how unequal group sizes and sharply different baseline risks create the reversal. The lecture concludes with practical guidance on when pooled rates answer a useful question and when subgroup adjustment is essential.","metaDescription":"A concrete treatment example reveals Simpson's paradox, shows its unequal-weight geometry, and explains when pooled rates mislead.","transcript":"Suppose two hospitals are comparing a treatment with ordinary care. The treatment will perform better among mild cases, and better among severe cases. Then we will pool everyone, and the apparent winner will reverse. That reversal is Simpson's paradox. Here is the question in its sharpest form. Can every clinically relevant subgroup favor the treatment while the total favors no treatment? Yes. The important task is not merely to be surprised. It is to see exactly how the reversal is made. We will use one thousand treated patients and one thousand untreated patients. The entries are successes over total cases, followed by the corresponding success rate. Start with mild cases. Among the treated patients, ninety of one hundred succeed. That is a ninety percent success rate. Without treatment, seven hundred twenty of nine hundred mild cases succeed. That is eighty percent. Within mild cases, treatment is ahead by ten percentage points. Now take severe cases. Among treated patients, two hundred seventy of nine hundred succeed, giving thirty percent. Without treatment, twenty of one hundred severe cases succeed, giving twenty percent. Treatment is again ahead by ten percentage points. Read the mild row once more. Ninety percent with treatment is ten points above eighty percent without treatment. Read the severe row. Thirty percent with treatment is ten points above twenty percent without treatment. There is no subgroup here where ordinary care has the higher success rate. Before revealing the total, make a prediction. Treatment wins by the same margin in both rows. It is natural to expect treatment to win when the rows are added. Now add the actual counts. Treatment has three hundred sixty successes among one thousand patients, for thirty-six percent. No treatment has seven hundred forty successes among one thousand, for seventy-four percent. The pooled comparison therefore says the opposite. Thirty-six minus seventy-four is negative thirty-eight percentage points. Within mild cases, treatment wins by ten points. Within severe cases, treatment wins by ten points. After pooling, treatment loses by thirty-eight points. Every fraction is correct, and the reversal is not an arithmetic error. What changed was the comparison itself. The subgroup rows compare patients with similar severity. The total compares two populations with radically different mixtures of mild and severe cases. That mixture is where the paradox lives. The pooled rates become understandable as soon as we draw who entered each comparison. Blue means mild cases and red means severe cases. Each strip contains one thousand patients. The treatment strip is almost entirely severe cases. Nine hundred are severe, and only one hundred are mild. The no-treatment strip has exactly the opposite composition. Only one hundred cases are severe, while nine hundred are mild. These labels name the two populations. Treatment is being judged mostly on severe patients. No treatment is being judged mostly on mild patients. Severity matters enormously even before considering treatment. Without treatment, mild cases succeed eighty percent of the time, while severe cases succeed only twenty percent. That is a sixty-point baseline-risk gap. It is much larger than the treatment's ten-point advantage inside either severity group. A pooled rate is a weighted average. For treatment, mild cases have weight one tenth, while severe cases have weight nine tenths. Those weights place ninety percent beside a small group and thirty percent beside a large group. The weighted average is thirty-six percent. For no treatment, the weights switch. Mild cases now receive weight nine tenths, and severe cases receive only one tenth. That weighted average is seventy-four percent. It is high because the no-treatment population contains many more of the patients who had a high baseline chance of success. Pooling has not averaged the same ingredients in the same proportions. It has compared one average dominated by severe cases with another average dominated by mild cases. We can separate the reversal into two forces. First comes the treatment effect visible within either group: a positive ten percentage points. Second comes the composition effect. The treatment population has eighty percentage points fewer mild cases. Mild and severe baseline success rates are sixty points apart. Multiplying those differences gives negative forty-eight points. The unfavorable case mix costs forty-eight points in the pooled comparison. Combine a positive ten-point treatment effect with a negative forty-eight point composition effect. The result is negative thirty-eight points, exactly the aggregate reversal we observed. The total answers a real question: what fraction succeeded in each observed population? It does not answer the like-for-like question: what would happen if treatment and no treatment were applied to populations with the same severity mix? Now turn the arithmetic into geometry. The horizontal axis records the fraction of cases that are mild. The vertical axis records the resulting success rate. At the far left, the population is entirely severe. Treatment succeeds at thirty percent, while no treatment succeeds at twenty percent. At the far right, the population is entirely mild. Treatment succeeds at ninety percent, while no treatment succeeds at eighty percent. Every mixture between those endpoints is a weighted average. Moving right gives more weight to mild cases, so both overall success rates rise. Most importantly, the red treatment line stays ten percentage points above the green control line at every common mixture. Start with a population that is ten percent mild. At this same mixture, treatment is thirty-six percent and no treatment is twenty-six percent. Sweep the common mixture toward ninety percent mild. The two points move together, and the yellow gap remains ten points from one end to the other. The same-mix formula confirms the picture. Both lines gain sixty points for a full change in mild share, so those matching terms cancel and leave a ten-point treatment advantage. If both groups had the same case mix, every vertical comparison would favor treatment. That is the geometric form of the two subgroup results. But the study did not observe both groups at one horizontal position. Remove the common-mixture comparison and mark the two populations that were actually observed. The treated population sits here, at ten percent mild and thirty-six percent successful. The untreated population sits far to the right, at ninety percent mild and seventy-four percent successful. Those two points are not vertically aligned. The control point benefits from a much easier case mix, so it can sit higher even though its entire green line lies below the red line. The observed comparison reads diagonally across the picture: left-hand red against right-hand green. The like-for-like comparison reads vertically: red against green at one shared mild-case fraction. That is the reversal made visible. Weighted averages are points between subgroup endpoints. Different weights select different horizontal positions, and a large enough horizontal shift can overwhelm the vertical treatment advantage. The practical question is not whether totals are forbidden. Totals are useful summaries. The question is whether the pooled total answers the decision you believe it answers. Begin with severity. In our example, severity strongly predicts the outcome. Mild cases have a much higher baseline chance of success than severe cases. Severity also predicts which comparison group a patient enters. The treatment group received mostly severe cases, while the untreated group received mostly mild cases. Treatment choice may itself affect the outcome. But severity points toward both treatment choice and outcome, creating a backdoor route that mixes treatment effects with case selection. This is why the first question matters. Does the subgroup variable predict the outcome even without treatment? If it does, changing the subgroup mix can change an aggregate rate. Second, are treatment and comparison cases distributed differently across those subgroups? A risk factor cannot create this particular distortion if both groups carry the same mix. Third, which population is the decision about? A hospital may need the expected result for its actual patient mix. A causal comparison may instead need both options evaluated on one common mix. Pooling is informative when treatment assignment is balanced across the important risk groups. Then each overall rate combines comparable ingredients in comparable proportions. Pooling can also be appropriate when the target mixture is deliberately specified. If a health system expects seventy percent mild and thirty percent severe cases, it can standardize both options to that same mix. In that setting, the pooled answer is not pretending to be universal. It is an answer for a named population, using weights chosen to represent that population. Pooling becomes dangerous when baseline risks differ greatly, treatment exposure is uneven across the risk groups, and the crude total is read as though it were a like-for-like treatment effect. Our example satisfies all three warning conditions. Severity changes baseline success by sixty points. The mild share changes from ten percent to ninety percent. The pooled rates then compare unlike populations. A responsible analysis would report the raw total and the severity-specific results, then use stratification, standardization, regression, matching, or a suitable study design to compare like with like. One final caution. Finding a Simpson reversal does not automatically prove that the subgroup result is causal. Severity may be measured imperfectly, other confounders may remain, and the chosen subgroups may themselves be consequences of earlier decisions. The reversal is a diagnostic signal. It tells us that aggregation and composition matter, and that we should investigate how patients entered each group before drawing a treatment conclusion. So ask what the total literally measures. In this study, thirty-six and seventy-four percent describe outcomes in two observed populations. They are valid summaries of those populations. Then ask what comparison the decision requires. If the goal is a treatment effect, compare treatment and no treatment at the same severity mix. If the goal is operational forecasting, use the mix expected in the population where the decision will be applied. The lesson of Simpson's paradox is not simply to distrust averages. It is to identify their weights, understand what creates those weights, and make sure both sides of a comparison answer the same question for the same target population.","watch":{"version":1,"scenes":[{"title":"The Reversal in the Counts","start":0,"end":167.9937083333333,"objects":{"card":"a Title that says \"Applied Statistics for Professionals — When Every Subgroup Agrees but the Total Reverses the Result\"","counts":"a Table [text] that says \"Cases Treatment No treatment Mild $90/100=90%$ $720/900=80%$ Severe $270/900=30%$ $20/100=20%$ All cases $360/1000=36%$ $740/1000=74%$\" (rows=(('Cases', 'Treatment', 'No treatment'), ('Mild', '$90/100=90%$…, header=True)","heading":"a Heading that says \"One Study, Three Comparisons\"","mild_difference":"a Math [text] that says \"$90% - 80% = 10 thin upright(\"percentage points\")$\"","overall_difference":"a Math [text] that says \"$36% - 74% = -38 thin upright(\"percentage points\")$\"","question":"a Panel that says \"Can a treatment have a higher success rate for mild cases and for severe cases, yet have a lower success rate when all cases are pooled?\"","severe_difference":"a Math [text] that says \"$30% - 20% = 10 thin upright(\"percentage points\")$\"","verdict":"a Math [text] that says \"$upright(\"within each group\"): thin T > C quad upright(\"after pooling\"): thin T < C$\""},"beats":[{"start":0,"say":"Suppose two hospitals are comparing a treatment with ordinary care. The treatment will perform better among mild cases, and better among severe cases. Then we will pool everyone, and the apparent winner will reverse. That reversal is Simpson's paradox.","live":[],"does":[[0,"card is shown on the screen, written out."],[1.5,"card: enter:write-left-to-right."],[14.663,"card is hidden from the screen — left the board."]]},{"start":15.863,"say":"Here is the question in its sharpest form. Can every clinically relevant subgroup favor the treatment while the total favors no treatment? Yes. The important task is not merely to be surprised. It is to see exactly how the reversal is made.","live":null,"does":[[15.863,"question is shown on the screen, written out."],[31.6295,"question is hidden from the screen — left the board."]]},{"start":32.2295,"say":"We will use one thousand treated patients and one thousand untreated patients. The entries are successes over total cases, followed by the corresponding success rate.","live":null,"does":[[32.2295,"heading is shown on the screen, written out."],[37.650999999999996,"counts is shown on the screen, written out."]]},{"start":43.301500000000004,"say":"Start with mild cases. Among the treated patients, ninety of one hundred succeed. That is a ninety percent success rate.","live":["heading"],"does":[[44.16100000000001,"counts is shown on the screen, written out."]]},{"start":52.62050000000001,"say":"Without treatment, seven hundred twenty of nine hundred mild cases succeed. That is eighty percent. Within mild cases, treatment is ahead by ten percentage points.","live":null,"does":[[59.528000000000006,"counts (the \"row=2\" part) is emphasized."],[63.81250000000001,"counts (the \"row=2\" part) is no longer emphasized."]]},{"start":64.41250000000001,"say":"Now take severe cases. Among treated patients, two hundred seventy of nine hundred succeed, giving thirty percent.","live":null,"does":[[65.31800000000001,"counts is shown on the screen, written out."]]},{"start":72.93050000000001,"say":"Without treatment, twenty of one hundred severe cases succeed, giving twenty percent. Treatment is again ahead by ten percentage points.","live":null,"does":[[73.488,"counts (the \"row=3\" part) is emphasized."],[81.38300000000001,"counts (the \"row=3\" part) is no longer emphasized."]]},{"start":81.983,"say":"Read the mild row once more. Ninety percent with treatment is ten points above eighty percent without treatment.","live":null,"does":[[82.70299999999999,"counts (the \"row=2\" part) is emphasized."],[88.84450000000001,"counts (the \"row=2\" part) is no longer emphasized."]]},{"start":89.4445,"say":"Read the severe row. Thirty percent with treatment is ten points above twenty percent without treatment. There is no subgroup here where ordinary care has the higher success rate.","live":null,"does":[[90.211,"counts (the \"row=3\" part) is emphasized."],[100.21900000000001,"counts (the \"row=3\" part) is no longer emphasized."]]},{"start":100.819,"say":"Before revealing the total, make a prediction. Treatment wins by the same margin in both rows. It is natural to expect treatment to win when the rows are added.","live":null,"does":[]},{"start":112.1115,"say":"Now add the actual counts. Treatment has three hundred sixty successes among one thousand patients, for thirty-six percent. No treatment has seven hundred forty successes among one thousand, for seventy-four percent.","live":null,"does":[[112.646,"counts is shown on the screen, written out."]]},{"start":126.98,"say":"The pooled comparison therefore says the opposite. Thirty-six minus seventy-four is negative thirty-eight percentage points.","live":null,"does":[[130.498,"counts (the \"row=4\" part) is emphasized."],[134.7475,"counts is hidden from the screen — left the board."],[134.7475,"counts (the \"row=4\" part) is no longer emphasized."]]},{"start":135.3475,"say":"Within mild cases, treatment wins by ten points. Within severe cases, treatment wins by ten points. After pooling, treatment loses by thirty-eight points. Every fraction is correct, and the reversal is not an arithmetic error.","live":null,"does":[[135.3475,"mild_difference is shown on the screen, written out."],[139.43399999999997,"severe_difference is shown on the screen, written out."],[143.265,"overall_difference is shown on the screen, written out."],[146.632,"verdict is shown on the screen, written out."]]},{"start":151.67849999999999,"say":"What changed was the comparison itself. The subgroup rows compare patients with similar severity. The total compares two populations with radically different mixtures of mild and severe cases. That mixture is where the paradox lives.","live":["heading","mild_difference","severe_difference","overall_difference","verdict"],"does":[[164.52999999999997,"A box is drawn around overall_difference."],[166.95204166666664,"heading is hidden from the screen — left the board."],[166.95204166666664,"mild_difference is hidden from the screen — left the board."],[166.95204166666664,"overall_difference is hidden from the screen — left the board."],[166.95204166666664,"severe_difference is hidden from the screen — left the board."],[166.95204166666664,"verdict is hidden from the screen — left the board."]]}]},{"title":"Why Pooling Flips the Result","start":167.9937083333333,"end":339.2772083333333,"objects":{"baseline":"a Math [text] that says \"$upright(\"mild baseline\") = 80% quad upright(\"severe baseline\") = 20%$\"","combined_effect":"a Math [text] that says \"$10% - 48% = -38%$\"","composition_effect":"a Math [text] that says \"$upright(\"composition effect\") = (10% - 90%)(60%) = -48%$\"","control_label":"a Math [text] that says \"$upright(\"no treatment\")$\" drawn in mix","control_mild":"a Polygon [blue] drawn in mix (vertices=((100.0, 0.0), (1000.0, 0.0), (1000.0, 1.0), (100.0, 1.0)), fill_opacity=0.6)","control_mild_label":"a Math [text] that says \"$upright(\"900 mild\")$\" drawn in mix","control_severe":"a Polygon [red] drawn in mix (vertices=((0.0, 0.0), (100.0, 0.0), (100.0, 1.0), (0.0, 1.0)), fill_opacity=0.48)","control_severe_label":"a Math [text] that says \"$upright(\"100 severe\")$\" drawn in mix","control_weighted":"a Math [text] that says \"$C = 0.9(80%) + 0.1(20%)$\"","gap":"a Math [text] that says \"$80% - 20% = 60 thin upright(\"percentage points\")$\"","heading_mix":"a Heading that says \"The Two Groups Have Opposite Case Mixes\"","heading_split":"a Heading that says \"Treatment Effect Versus Composition Effect\"","heading_weights":"a Heading that says \"Each Total Is a Different Weighted Average\"","mix":"a Figure (x_range=(0.0, 1000.0), y_range=(-0.2, 3.2), aspect=(5.0, 2.0))","treated_label":"a Math [text] that says \"$upright(\"treatment\")$\" drawn in mix","treated_mild":"a Polygon [blue] drawn in mix (vertices=((900.0, 2.0), (1000.0, 2.0), (1000.0, 3.0), (900.0, 3.0)), fill_opacity=0.6)","treated_mild_label":"a Math [text] that says \"$upright(\"100 mild\")$\" drawn in mix","treated_severe":"a Polygon [red] drawn in mix (vertices=((0.0, 2.0), (900.0, 2.0), (900.0, 3.0), (0.0, 3.0)), fill_opacity=0.48)","treated_severe_label":"a Math [text] that says \"$upright(\"900 severe\")$\" drawn in mix","treatment_weighted":"a Math [text] that says \"$T = 0.1(90%) + 0.9(30%)$\"","within_effect":"a Math [text] that says \"$upright(\"within-group advantage\") = 10%$\""},"beats":[{"start":167.9937083333333,"say":"The pooled rates become understandable as soon as we draw who entered each comparison. Blue means mild cases and red means severe cases. Each strip contains one thousand patients.","live":[],"does":[[167.9937083333333,"heading_mix is shown on the screen, written out."],[167.9937083333333,"mix is shown on the screen, written out."]]},{"start":180.8072083333333,"say":"The treatment strip is almost entirely severe cases. Nine hundred are severe, and only one hundred are mild.","live":["mix","heading_mix"],"does":[[183.1527083333333,"treated_severe is shown on the screen, written out."],[184.8017083333333,"treated_severe_label is shown on the screen, written out."],[186.6707083333333,"treated_mild_label is shown on the screen, written out."],[187.3447083333333,"treated_mild is shown on the screen, written out."]]},{"start":188.7687083333333,"say":"The no-treatment strip has exactly the opposite composition. Only one hundred cases are severe, while nine hundred are mild.","live":["mix","heading_mix","treated_severe","treated_mild","treated_severe_label","treated_mild_label"],"does":[[193.3317083333333,"control_severe_label is shown on the screen, written out."],[194.32970833333331,"control_severe is shown on the screen, written out."],[195.3517083333333,"control_mild_label is shown on the screen, written out."],[196.0247083333333,"control_mild is shown on the screen, written out."]]},{"start":197.4952083333333,"say":"These labels name the two populations. Treatment is being judged mostly on severe patients. No treatment is being judged mostly on mild patients.","live":["mix","heading_mix","treated_severe","treated_mild","treated_severe_label","treated_mild_label","control_severe","control_mild","control_severe_label","control_mild_label"],"does":[[200.8167083333333,"treated_label is shown on the screen, written out."],[204.1247083333333,"control_label is shown on the screen, written out."]]},{"start":207.9177083333333,"say":"Severity matters enormously even before considering treatment. Without treatment, mild cases succeed eighty percent of the time, while severe cases succeed only twenty percent.","live":["mix","heading_mix","treated_severe","treated_mild","treated_severe_label","treated_mild_label","control_severe","control_mild","control_severe_label","control_mild_label","treated_label","control_label"],"does":[[208.2197083333333,"baseline is shown on the screen, written out."]]},{"start":219.5357083333333,"say":"That is a sixty-point baseline-risk gap. It is much larger than the treatment's ten-point advantage inside either severity group.","live":["mix","baseline","heading_mix","treated_severe","treated_mild","treated_severe_label","treated_mild_label","control_severe","control_mild","control_severe_label","control_mild_label","treated_label","control_label"],"does":[[219.5357083333333,"gap is shown on the screen, written out."],[227.5232083333333,"baseline moves to a new place on the board."],[227.5232083333333,"gap moves to a new place on the board."],[227.5232083333333,"heading_mix is hidden from the screen — left the board."],[227.5232083333333,"mix is hidden from the screen — left the board."],[227.5232083333333,"treated_severe is hidden from the screen — mix left the board."],[227.5232083333333,"treated_mild is hidden from the screen — mix left the board."],[227.5232083333333,"treated_severe_label is hidden from the screen — mix left the board."],[227.5232083333333,"treated_mild_label is hidden from the screen — mix left the board."],[227.5232083333333,"control_severe is hidden from the screen — mix left the board."],[227.5232083333333,"control_mild is hidden from the screen — mix left the board."],[227.5232083333333,"control_severe_label is hidden from the screen — mix left the board."],[227.5232083333333,"control_mild_label is hidden from the screen — mix left the board."],[227.5232083333333,"treated_label is hidden from the screen — mix left the board."],[227.5232083333333,"control_label is hidden from the screen — mix left the board."]]},{"start":228.1232083333333,"say":"A pooled rate is a weighted average. For treatment, mild cases have weight one tenth, while severe cases have weight nine tenths.","live":["baseline","gap"],"does":[[228.1232083333333,"heading_weights is shown on the screen, written out."],[229.5627083333333,"treatment_weighted is shown on the screen, written out."],[233.4177083333333,"treatment_weighted (the \"0.1\" part) is emphasized."],[236.0067083333333,"treatment_weighted (the \"0.1\" part) is no longer emphasized."],[236.0067083333333,"treatment_weighted (the \"0.9\" part) is emphasized."]]},{"start":237.81420833333328,"say":"Those weights place ninety percent beside a small group and thirty percent beside a large group. The weighted average is thirty-six percent.","live":["baseline","gap","treatment_weighted","heading_weights"],"does":[[237.81420833333328,"treatment_weighted (the \"0.9\" part) is no longer emphasized."],[244.55970833333328,"treatment_weighted becomes \"$T = 0.1(90%) + 0.9(30%) = 36%$\"."]]},{"start":246.6222083333333,"say":"For no treatment, the weights switch. Mild cases now receive weight nine tenths, and severe cases receive only one tenth.","live":null,"does":[[248.4337083333333,"control_weighted is shown on the screen, written out."],[251.5217083333333,"control_weighted (the \"0.9\" part) is emphasized."],[254.45870833333328,"control_weighted (the \"0.1\" part) is emphasized."],[254.45870833333328,"control_weighted (the \"0.9\" part) is no longer emphasized."]]},{"start":256.25470833333327,"say":"That weighted average is seventy-four percent. It is high because the no-treatment population contains many more of the patients who had a high baseline chance of success.","live":["baseline","gap","treatment_weighted","control_weighted","heading_weights"],"does":[[256.25470833333327,"control_weighted (the \"0.1\" part) is no longer emphasized."],[257.8807083333333,"control_weighted becomes \"$C = 0.9(80%) + 0.1(20%) = 74%$\"."]]},{"start":267.2337083333333,"say":"Pooling has not averaged the same ingredients in the same proportions. It has compared one average dominated by severe cases with another average dominated by mild cases.","live":null,"does":[[277.8107083333333,"baseline is hidden from the screen — left the board."],[277.8107083333333,"control_weighted is hidden from the screen — left the board."],[277.8107083333333,"gap is hidden from the screen — left the board."],[277.8107083333333,"heading_weights is hidden from the screen — left the board."],[277.8107083333333,"treatment_weighted is hidden from the screen — left the board."]]},{"start":279.0107083333333,"say":"We can separate the reversal into two forces. First comes the treatment effect visible within either group: a positive ten percentage points.","live":[],"does":[[279.0107083333333,"heading_split is shown on the screen, written out."],[286.9407083333333,"within_effect is shown on the screen, written out."]]},{"start":289.0497083333333,"say":"Second comes the composition effect. The treatment population has eighty percentage points fewer mild cases. Mild and severe baseline success rates are sixty points apart.","live":["within_effect","heading_split"],"does":[[290.2457083333333,"composition_effect is shown on the screen, written out."],[293.4497083333333,"composition_effect (the \"10% - 90%\" part) is emphasized."],[299.6837083333333,"composition_effect (the \"10% - 90%\" part) is no longer emphasized."],[299.6837083333333,"composition_effect (the \"60%\" part) is emphasized."],[301.1817083333333,"composition_effect (the \"60%\" part) is no longer emphasized."]]},{"start":301.7817083333333,"say":"Multiplying those differences gives negative forty-eight points. The unfavorable case mix costs forty-eight points in the pooled comparison.","live":["within_effect","composition_effect","heading_split"],"does":[[304.2547083333333,"composition_effect (the \"-48%\" part) is indicated — a transient flash."]]},{"start":310.9382083333333,"say":"Combine a positive ten-point treatment effect with a negative forty-eight point composition effect. The result is negative thirty-eight points, exactly the aggregate reversal we observed.","live":null,"does":[[311.2867083333333,"combined_effect is shown on the screen, written out."],[311.8437083333333,"combined_effect (the \"10%\" part) is emphasized."],[314.3637083333333,"combined_effect (the \"- 48%\" part) is emphasized."],[314.3637083333333,"combined_effect (the \"10%\" part) is no longer emphasized."],[318.8327083333333,"combined_effect (the \"- 48%\" part) is no longer emphasized."],[318.8327083333333,"combined_effect (the \"-38%\" part) is emphasized."],[323.0012083333333,"combined_effect (the \"-38%\" part) is no longer emphasized."]]},{"start":323.6012083333333,"say":"The total answers a real question: what fraction succeeded in each observed population? It does not answer the like-for-like question: what would happen if treatment and no treatment were applied to populations with the same severity mix?","live":["within_effect","composition_effect","combined_effect","heading_split"],"does":[[331.33370833333333,"A box is drawn around combined_effect."],[338.2355416666667,"combined_effect is hidden from the screen — left the board."],[338.2355416666667,"composition_effect is hidden from the screen — left the board."],[338.2355416666667,"heading_split is hidden from the screen — left the board."],[338.2355416666667,"within_effect is hidden from the screen — left the board."]]}]},{"title":"The Geometry of Unequal Weights","start":339.2772083333333,"end":498.0758125,"objects":{"actual_control":"a PlotPoint [green] labelled \"(90%, 74%)\" drawn in axes (target='control_line', x=0.9)","actual_treatment":"a PlotPoint [red] labelled \"(10%, 36%)\" drawn in axes (target='treatment_line', x=0.1)","axes":"an Axes (y_range=(0.1, 1.0), x_ticks_every=0.1, y_ticks_every=0.1)","common_control":"a PlotPoint [green] drawn in axes (target='control_line', x=<VariableNumber common_mix = 0.9>)","common_gap":"a Line [yellow] labelled \"10%\" drawn in axes (start=(<VariableNumber common_mix = 0.9>, (0.2 + (0.6 * common_mix))), end=(<VariableNumber common_mix = 0.9>, (0.3 + (0.6 * common_mix))))","common_mix":"a VariableNumber (initial_value=0.1)","common_treatment":"a PlotPoint [red] drawn in axes (target='treatment_line', x=<VariableNumber common_mix = 0.9>)","control_line":"a FunctionPlot [green] labelled \"upright(\"no treatment\")\" drawn in axes (function=<function>, x_range=(0.0, 1.0))","heading":"a Heading that says \"Success Rate as the Case Mix Changes\"","lesson":"a Math [text] that says \"$upright(\"same mix\") arrow.r upright(\"like-for-like comparison\")$\"","line":"a Line [yellow] drawn in axes (start=(0.1, 0.36), end=(0.0, 0.36), dashed=True)","line_2":"a Line [yellow] drawn in axes (start=(0.1, 0.36), end=(0.1, 0.0), dashed=True)","line_3":"a Line [yellow] drawn in axes (start=(0.1, 0.36), end=(0.0, 0.36), dashed=True)","line_4":"a Line [yellow] drawn in axes (start=(0.1, 0.36), end=(0.1, 0.0), dashed=True)","line_5":"a Line [yellow] drawn in axes (start=(0.9, 0.74), end=(0.0, 0.74), dashed=True)","line_6":"a Line [yellow] drawn in axes (start=(0.9, 0.74), end=(0.9, 0.0), dashed=True)","mild_control":"a PlotPoint [green] labelled \"80%\" drawn in axes (target='control_line', x=1.0)","mild_treatment":"a PlotPoint [red] labelled \"90%\" drawn in axes (target='treatment_line', x=1.0)","observed_mix":"a Math [text] that says \"$T(10%) = 36% quad C(90%) = 74%$\"","point":"a Point [yellow] drawn in axes (location=(0.1, 0.36))","point_2":"a Point [yellow] drawn in axes (location=(0.1, 0.36))","point_3":"a Point [yellow] drawn in axes (location=(0.9, 0.74))","same_mix":"a Math [text] that says \"$T(w) - C(w) = (30% + 60%w) - (20% + 60%w) = 10%$\"","severe_control":"a PlotPoint [green] labelled \"20%\" drawn in axes (target='control_line')","severe_treatment":"a PlotPoint [red] labelled \"30%\" drawn in axes (target='treatment_line')","treatment_line":"a FunctionPlot [red] labelled \"upright(\"treatment\")\" drawn in axes (function=<function>, x_range=(0.0, 1.0))"},"beats":[{"start":339.2772083333333,"say":"Now turn the arithmetic into geometry. The horizontal axis records the fraction of cases that are mild. The vertical axis records the resulting success rate.","live":[],"does":[[339.2772083333333,"heading is shown on the screen, written out."],[339.2772083333333,"axes is shown on the screen, written out."]]},{"start":350.4192083333333,"say":"At the far left, the population is entirely severe. Treatment succeeds at thirty percent, while no treatment succeeds at twenty percent.","live":["axes","heading"],"does":[[354.9242083333333,"treatment_line is shown on the screen, drawn."],[355.99220833333334,"severe_treatment is shown on the screen, written out."],[357.3392083333333,"control_line is shown on the screen, drawn."],[358.62720833333333,"severe_control is shown on the screen, written out."]]},{"start":360.4467083333333,"say":"At the far right, the population is entirely mild. Treatment succeeds at ninety percent, while no treatment succeeds at eighty percent.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control"],"does":[[365.6132083333333,"mild_treatment is shown on the screen, written out."],[368.1792083333333,"mild_control is shown on the screen, written out."]]},{"start":370.2647083333333,"say":"Every mixture between those endpoints is a weighted average. Moving right gives more weight to mild cases, so both overall success rates rise.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control"],"does":[]},{"start":380.4542083333333,"say":"Most importantly, the red treatment line stays ten percentage points above the green control line at every common mixture.","live":null,"does":[[382.0792083333333,"common_treatment is shown on the screen, written out."],[383.2752083333333,"common_gap is shown on the screen, written out."],[384.9932083333333,"common_control is shown on the screen, written out."]]},{"start":388.6007083333333,"say":"Start with a population that is ten percent mild. At this same mixture, treatment is thirty-six percent and no treatment is twenty-six percent.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control","common_control","common_treatment","common_gap"],"does":[[389.0072083333333,"point is shown on the screen, grown."],[390.6672083333333,"line is shown on the screen, drawn."],[394.6952083333333,"line_2 is shown on the screen, drawn."],[396.6952083333333,"point is hidden from the screen."],[396.6952083333333,"line is hidden from the screen."],[396.6952083333333,"line_2 is hidden from the screen."]]},{"start":398.9992083333333,"say":"Sweep the common mixture toward ninety percent mild. The two points move together, and the yellow gap remains ten points from one end to the other.","live":null,"does":[[399.4052083333333,"common_control is redrawn as the numbers it depends on change."],[399.4052083333333,"common_treatment is redrawn as the numbers it depends on change."],[399.4052083333333,"common_gap is redrawn as the numbers it depends on change."],[399.4052083333333,"common_mix ticks to 0.9."]]},{"start":408.8177083333333,"say":"The same-mix formula confirms the picture. Both lines gain sixty points for a full change in mild share, so those matching terms cancel and leave a ten-point treatment advantage.","live":null,"does":[[410.00220833333333,"same_mix is shown on the screen, written out."],[413.4032083333333,"same_mix (the \"60%w\" part) is emphasized."],[416.7702083333333,"same_mix (the \"60%w\" part) is no longer emphasized."],[416.7702083333333,"same_mix (the \"60%w#2\" part) is emphasized."],[418.77920833333326,"same_mix (the \"10%\" part) is emphasized."],[418.77920833333326,"same_mix (the \"60%w#2\" part) is no longer emphasized."],[420.9617083333333,"same_mix (the \"10%\" part) is no longer emphasized."]]},{"start":421.5617083333333,"say":"If both groups had the same case mix, every vertical comparison would favor treatment. That is the geometric form of the two subgroup results.","live":["axes","same_mix","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control","common_control","common_treatment","common_gap"],"does":[[424.8822083333333,"common_gap is indicated — a transient flash."],[430.89620833333333,"axes moves to a new place on the board."],[430.89620833333333,"same_mix is hidden from the screen — left the board."]]},{"start":431.4962083333333,"say":"But the study did not observe both groups at one horizontal position. Remove the common-mixture comparison and mark the two populations that were actually observed.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control","common_control","common_treatment","common_gap"],"does":[[436.4072083333333,"common_gap is hidden from the screen."],[436.4072083333333,"common_control is hidden from the screen."],[436.4072083333333,"common_treatment is hidden from the screen."]]},{"start":442.06920833333334,"say":"The treated population sits here, at ten percent mild and thirty-six percent successful.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control"],"does":[[443.82220833333326,"actual_treatment is shown on the screen, written out."],[443.82220833333326,"point_2 is shown on the screen, grown."],[444.67020833333333,"line_3 is shown on the screen, drawn."],[446.14420833333327,"line_4 is shown on the screen, drawn."],[448.14420833333327,"point_2 is hidden from the screen."],[448.14420833333327,"line_3 is hidden from the screen."],[448.14420833333327,"line_4 is hidden from the screen."]]},{"start":448.7297083333333,"say":"The untreated population sits far to the right, at ninety percent mild and seventy-four percent successful.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control","actual_treatment"],"does":[[451.0862083333333,"actual_control is shown on the screen, written out."],[451.0862083333333,"point_3 is shown on the screen, grown."],[452.0042083333333,"line_5 is shown on the screen, drawn."],[453.3502083333333,"line_6 is shown on the screen, drawn."],[455.3502083333333,"point_3 is hidden from the screen."],[455.3502083333333,"line_5 is hidden from the screen."],[455.3502083333333,"line_6 is hidden from the screen."]]},{"start":455.9592083333333,"say":"Those two points are not vertically aligned. The control point benefits from a much easier case mix, so it can sit higher even though its entire green line lies below the red line.","live":["axes","heading","treatment_line","control_line","severe_treatment","severe_control","mild_treatment","mild_control","actual_treatment","actual_control"],"does":[[459.2102083333333,"actual_control is indicated — a transient flash."],[466.3612083333333,"actual_treatment is indicated — a transient flash."]]},{"start":467.8902083333333,"say":"The observed comparison reads diagonally across the picture: left-hand red against right-hand green. The like-for-like comparison reads vertically: red against green at one shared mild-case fraction.","live":null,"does":[[468.4242083333333,"observed_mix is shown on the screen, written out."],[472.0932083333333,"observed_mix (the \"T(10%)\" part) is emphasized."],[473.2542083333333,"observed_mix (the \"C(90%)\" part) is emphasized."],[473.2542083333333,"observed_mix (the \"T(10%)\" part) is no longer emphasized."],[481.7637083333333,"axes moves to a new place on the board."],[481.7637083333333,"observed_mix is hidden from the screen — left the board."],[481.7637083333333,"observed_mix (the \"C(90%)\" part) is no longer emphasized."]]},{"start":482.3637083333333,"say":"That is the reversal made visible. Weighted averages are points between subgroup endpoints. Different weights select different horizontal positions, and a large enough horizontal shift can overwhelm the vertical treatment advantage.","live":null,"does":[[484.2912083333333,"lesson is shown on the screen, written out."],[495.1582083333333,"A box is drawn around lesson."],[497.0341458333333,"axes is hidden from the screen — left the board."],[497.0341458333333,"treatment_line is hidden from the screen — axes left the board."],[497.0341458333333,"control_line is hidden from the screen — axes left the board."],[497.0341458333333,"severe_treatment is hidden from the screen — axes left the board."],[497.0341458333333,"severe_control is hidden from the screen — axes left the board."],[497.0341458333333,"mild_treatment is hidden from the screen — axes left the board."],[497.0341458333333,"mild_control is hidden from the screen — axes left the board."],[497.0341458333333,"actual_treatment is hidden from the screen — axes left the board."],[497.0341458333333,"actual_control is hidden from the screen — axes left the board."],[497.0341458333333,"heading is hidden from the screen — left the board."],[497.0341458333333,"lesson is hidden from the screen — left the board."]]}]},{"title":"When Pooling Answers the Question","start":498.0758125,"end":739.0531874999999,"objects":{"appropriate":"a Block [text] that says \"Assignment is balanced across important risk groups. The target population mix is explicitly specified. The pooled rate is described as a result for that mixture.\"","appropriate_label":"a Tex [text] that says \"Pooling is informative\" (underline=True)","diagram":"a Figure (x_range=(0.0, 6.0), y_range=(0.0, 4.0), aspect=(3.0, 2.0))","final_rule":"a Math [text] that says \"$upright(\"same question\") + upright(\"same target mix\") arrow.r upright(\"meaningful comparison\")$\"","heading_cause":"a Heading that says \"Why Severity Must Stay in the Analysis\"","heading_decision":"a Heading that says \"Pooling Can Answer Different Questions\"","heading_questions":"a Heading that says \"Three Questions Before You Pool\"","misleading":"a Block [text] that says \"Baseline outcome rates differ greatly across groups. Treatment exposure is very uneven across those groups. The crude total is interpreted as a like-for-like treatment effect.\"","misleading_label":"a Tex [text] that says \"Pooling can mislead\" (underline=True)","outcome":"a Math [green] that says \"$upright(\"outcome\")$\" drawn in diagram","q1":"a Text [text] that says \"Does the subgroup variable strongly predict the outcome even without the treatment?\"","q2":"a Text [text] that says \"Are treatment and comparison cases distributed differently across those subgroups?\"","q3":"a Text [text] that says \"Which population mix does the decision actually concern?\"","severity":"a Math [red] that says \"$upright(\"severity\")$\" drawn in diagram","severity_to_outcome":"an Arrow [red] drawn in diagram (start=(1.65, 1.9), end=(4.35, 1.9))","severity_to_treatment":"an Arrow [red] drawn in diagram (start=(1.6, 2.25), end=(2.45, 2.85))","treatment":"a Math [blue] that says \"$upright(\"treatment choice\")$\" drawn in diagram","treatment_to_outcome":"an Arrow [blue] drawn in diagram (start=(3.55, 2.9), end=(4.65, 2.25))","warning":"a Panel that says \"A Simpson reversal is a signal to inspect subgroup composition and the data-generating process. By itself, it does not prove which comparison is causal.\""},"beats":[{"start":498.0758125,"say":"The practical question is not whether totals are forbidden. Totals are useful summaries. The question is whether the pooled total answers the decision you believe it answers.","live":[],"does":[[498.0758125,"heading_questions is shown on the screen, written out."],[498.0758125,"diagram is shown on the screen, written out."]]},{"start":509.3683125,"say":"Begin with severity. In our example, severity strongly predicts the outcome. Mild cases have a much higher baseline chance of success than severe cases.","live":["diagram","heading_questions"],"does":[[510.4718125,"severity is shown on the screen, written out."],[513.9198125,"severity_to_outcome is shown on the screen, written out."],[514.4068125,"outcome is shown on the screen, written out."]]},{"start":521.2878125,"say":"Severity also predicts which comparison group a patient enters. The treatment group received mostly severe cases, while the untreated group received mostly mild cases.","live":["diagram","heading_questions","severity","outcome","severity_to_outcome"],"does":[[522.5298124999999,"severity_to_treatment is shown on the screen, written out."],[523.1568125,"treatment is shown on the screen, written out."]]},{"start":532.2793125,"say":"Treatment choice may itself affect the outcome. But severity points toward both treatment choice and outcome, creating a backdoor route that mixes treatment effects with case selection.","live":["diagram","heading_questions","severity","outcome","severity_to_outcome","treatment","severity_to_treatment"],"does":[[533.9738125,"treatment_to_outcome is shown on the screen, written out."],[539.2688125,"severity_to_treatment is indicated — a transient flash."],[539.7678125,"severity_to_outcome is indicated — a transient flash."]]},{"start":543.5493125,"say":"This is why the first question matters. Does the subgroup variable predict the outcome even without treatment? If it does, changing the subgroup mix can change an aggregate rate.","live":["diagram","heading_questions","severity","outcome","severity_to_outcome","treatment","severity_to_treatment","treatment_to_outcome"],"does":[[544.5828124999999,"diagram moves to a new place on the board."],[544.5828124999999,"q1 is shown on the screen, written out."]]},{"start":555.5738125,"say":"Second, are treatment and comparison cases distributed differently across those subgroups? A risk factor cannot create this particular distortion if both groups carry the same mix.","live":["q1","diagram","heading_questions","severity","outcome","severity_to_outcome","treatment","severity_to_treatment","treatment_to_outcome"],"does":[[556.0498125,"q2 is shown on the screen, written out."]]},{"start":568.4913125,"say":"Third, which population is the decision about? A hospital may need the expected result for its actual patient mix. A causal comparison may instead need both options evaluated on one common mix.","live":["q1","q2","diagram","heading_questions","severity","outcome","severity_to_outcome","treatment","severity_to_treatment","treatment_to_outcome"],"does":[[569.0488124999999,"q3 is shown on the screen, written out."],[583.1318125,"diagram is hidden from the screen — left the board."],[583.1318125,"severity is hidden from the screen — diagram left the board."],[583.1318125,"outcome is hidden from the screen — diagram left the board."],[583.1318125,"severity_to_outcome is hidden from the screen — diagram left the board."],[583.1318125,"treatment is hidden from the screen — diagram left the board."],[583.1318125,"severity_to_treatment is hidden from the screen — diagram left the board."],[583.1318125,"treatment_to_outcome is hidden from the screen — diagram left the board."],[583.1318125,"heading_questions is hidden from the screen — left the board."],[583.1318125,"q1 is hidden from the screen — left the board."],[583.1318125,"q2 is hidden from the screen — left the board."],[583.1318125,"q3 is hidden from the screen — left the board."]]},{"start":583.7318124999999,"say":"Pooling is informative when treatment assignment is balanced across the important risk groups. Then each overall rate combines comparable ingredients in comparable proportions.","live":[],"does":[[583.7318124999999,"heading_decision is shown on the screen, written out."],[584.2078125,"appropriate_label is shown on the screen, written out."],[586.7388125,"appropriate is shown on the screen, written out."]]},{"start":595.3263125,"say":"Pooling can also be appropriate when the target mixture is deliberately specified. If a health system expects seventy percent mild and thirty percent severe cases, it can standardize both options to that same mix.","live":["appropriate_label","appropriate","heading_decision"],"does":[[597.6828125,"appropriate (the \"target population mix\" part) is emphasized."],[608.9328125,"appropriate (the \"target population mix\" part) is no longer emphasized."]]},{"start":609.5328125,"say":"In that setting, the pooled answer is not pretending to be universal. It is an answer for a named population, using weights chosen to represent that population.","live":null,"does":[[615.3028125,"appropriate (the \"that mixture\" part) is emphasized."],[620.0048125,"appropriate (the \"that mixture\" part) is no longer emphasized."]]},{"start":620.6048125,"say":"Pooling becomes dangerous when baseline risks differ greatly, treatment exposure is uneven across the risk groups, and the crude total is read as though it were a like-for-like treatment effect.","live":null,"does":[[621.9518125,"misleading_label is shown on the screen, written out."],[622.7068125,"misleading is shown on the screen, written out."]]},{"start":633.0708125,"say":"Our example satisfies all three warning conditions. Severity changes baseline success by sixty points. The mild share changes from ten percent to ninety percent. The pooled rates then compare unlike populations.","live":["appropriate_label","appropriate","misleading_label","misleading","heading_decision"],"does":[[637.2968125,"misleading (the \"Baseline outcome rates\" part) is emphasized."],[641.2438125,"misleading (the \"very uneven\" part) is emphasized."],[648.1868125,"misleading (the \"Baseline outcome rates\" part) is no longer emphasized."],[648.1868125,"misleading (the \"very uneven\" part) is no longer emphasized."]]},{"start":648.7868125,"say":"A responsible analysis would report the raw total and the severity-specific results, then use stratification, standardization, regression, matching, or a suitable study design to compare like with like.","live":null,"does":[[662.6378125,"appropriate is hidden from the screen — left the board."],[662.6378125,"appropriate_label is hidden from the screen — left the board."],[662.6378125,"heading_decision is hidden from the screen — left the board."],[662.6378125,"misleading is hidden from the screen — left the board."],[662.6378125,"misleading_label is hidden from the screen — left the board."]]},{"start":663.2378125,"say":"One final caution. Finding a Simpson reversal does not automatically prove that the subgroup result is causal. Severity may be measured imperfectly, other confounders may remain, and the chosen subgroups may themselves be consequences of earlier decisions.","live":[],"does":[[663.2378125,"heading_cause is shown on the screen, written out."],[664.1898125,"warning is shown on the screen, written out."]]},{"start":681.0438125000001,"say":"The reversal is a diagnostic signal. It tells us that aggregation and composition matter, and that we should investigate how patients entered each group before drawing a treatment conclusion.","live":["warning","heading_cause"],"does":[[683.2848125,"warning (the \"signal\" part) is emphasized."],[688.0438125000001,"warning (the \"signal\" part) is no longer emphasized."]]},{"start":693.1258124999999,"say":"So ask what the total literally measures. In this study, thirty-six and seventy-four percent describe outcomes in two observed populations. They are valid summaries of those populations.","live":null,"does":[]},{"start":706.8798125,"say":"Then ask what comparison the decision requires. If the goal is a treatment effect, compare treatment and no treatment at the same severity mix. If the goal is operational forecasting, use the mix expected in the population where the decision will be applied.","live":null,"does":[[714.3448125,"final_rule is shown on the screen, written out."]]},{"start":723.5478125,"say":"The lesson of Simpson's paradox is not simply to distrust averages. It is to identify their weights, understand what creates those weights, and make sure both sides of a comparison answer the same question for the same target population.","live":["warning","final_rule","heading_cause"],"does":[[736.1568125000001,"A box is drawn around final_rule."],[738.0115208333333,"final_rule is hidden from the screen — left the board."],[738.0115208333333,"heading_cause is hidden from the screen — left the board."],[738.0115208333333,"warning is hidden from the screen — left the board."]]}]}]},"durationSeconds":739,"chapters":[{"title":"The Reversal in the Counts","startSeconds":0,"narration":"Suppose two hospitals are comparing a treatment with ordinary care. The treatment will perform better among mild cases, and better among severe cases. Then we will pool everyone, and the apparent winner will reverse. That reversal is Simpson's paradox. Here is the question in its sharpest form. Can every clinically relevant subgroup favor the treatment while the total favors no treatment? Yes. The important task is not merely to be surprised. It is to see exactly how the reversal is made. We will use one thousand treated patients and one thousand untreated patients. The entries are successes over total cases, followed by the corresponding success rate. Start with mild cases. Among the treated patients, ninety of one hundred succeed. That is a ninety percent success rate. Without treatment, seven hundred twenty of nine hundred mild cases succeed. That is eighty percent. Within mild cases, treatment is ahead by ten percentage points. Now take severe cases. Among treated patients, two hundred seventy of nine hundred succeed, giving thirty percent. Without treatment, twenty of one hundred severe cases succeed, giving twenty percent. Treatment is again ahead by ten percentage points. Read the mild row once more. Ninety percent with treatment is ten points above eighty percent without treatment. Read the severe row. Thirty percent with treatment is ten points above twenty percent without treatment. There is no subgroup here where ordinary care has the higher success rate. Before revealing the total, make a prediction. Treatment wins by the same margin in both rows. It is natural to expect treatment to win when the rows are added. Now add the actual counts. Treatment has three hundred sixty successes among one thousand patients, for thirty-six percent. No treatment has seven hundred forty successes among one thousand, for seventy-four percent. The pooled comparison therefore says the opposite. Thirty-six minus seventy-four is negative thirty-eight percentage points. Within mild cases, treatment wins by ten points. Within severe cases, treatment wins by ten points. After pooling, treatment loses by thirty-eight points. Every fraction is correct, and the reversal is not an arithmetic error. What changed was the comparison itself. The subgroup rows compare patients with similar severity. The total compares two populations with radically different mixtures of mild and severe cases. That mixture is where the paradox lives."},{"title":"Why Pooling Flips the Result","startSeconds":167.9937083333333,"narration":"The pooled rates become understandable as soon as we draw who entered each comparison. Blue means mild cases and red means severe cases. Each strip contains one thousand patients. The treatment strip is almost entirely severe cases. Nine hundred are severe, and only one hundred are mild. The no-treatment strip has exactly the opposite composition. Only one hundred cases are severe, while nine hundred are mild. These labels name the two populations. Treatment is being judged mostly on severe patients. No treatment is being judged mostly on mild patients. Severity matters enormously even before considering treatment. Without treatment, mild cases succeed eighty percent of the time, while severe cases succeed only twenty percent. That is a sixty-point baseline-risk gap. It is much larger than the treatment's ten-point advantage inside either severity group. A pooled rate is a weighted average. For treatment, mild cases have weight one tenth, while severe cases have weight nine tenths. Those weights place ninety percent beside a small group and thirty percent beside a large group. The weighted average is thirty-six percent. For no treatment, the weights switch. Mild cases now receive weight nine tenths, and severe cases receive only one tenth. That weighted average is seventy-four percent. It is high because the no-treatment population contains many more of the patients who had a high baseline chance of success. Pooling has not averaged the same ingredients in the same proportions. It has compared one average dominated by severe cases with another average dominated by mild cases. We can separate the reversal into two forces. First comes the treatment effect visible within either group: a positive ten percentage points. Second comes the composition effect. The treatment population has eighty percentage points fewer mild cases. Mild and severe baseline success rates are sixty points apart. Multiplying those differences gives negative forty-eight points. The unfavorable case mix costs forty-eight points in the pooled comparison. Combine a positive ten-point treatment effect with a negative forty-eight point composition effect. The result is negative thirty-eight points, exactly the aggregate reversal we observed. The total answers a real question: what fraction succeeded in each observed population? It does not answer the like-for-like question: what would happen if treatment and no treatment were applied to populations with the same severity mix?"},{"title":"The Geometry of Unequal Weights","startSeconds":339.2772083333333,"narration":"Now turn the arithmetic into geometry. The horizontal axis records the fraction of cases that are mild. The vertical axis records the resulting success rate. At the far left, the population is entirely severe. Treatment succeeds at thirty percent, while no treatment succeeds at twenty percent. At the far right, the population is entirely mild. Treatment succeeds at ninety percent, while no treatment succeeds at eighty percent. Every mixture between those endpoints is a weighted average. Moving right gives more weight to mild cases, so both overall success rates rise. Most importantly, the red treatment line stays ten percentage points above the green control line at every common mixture. Start with a population that is ten percent mild. At this same mixture, treatment is thirty-six percent and no treatment is twenty-six percent. Sweep the common mixture toward ninety percent mild. The two points move together, and the yellow gap remains ten points from one end to the other. The same-mix formula confirms the picture. Both lines gain sixty points for a full change in mild share, so those matching terms cancel and leave a ten-point treatment advantage. If both groups had the same case mix, every vertical comparison would favor treatment. That is the geometric form of the two subgroup results. But the study did not observe both groups at one horizontal position. Remove the common-mixture comparison and mark the two populations that were actually observed. The treated population sits here, at ten percent mild and thirty-six percent successful. The untreated population sits far to the right, at ninety percent mild and seventy-four percent successful. Those two points are not vertically aligned. The control point benefits from a much easier case mix, so it can sit higher even though its entire green line lies below the red line. The observed comparison reads diagonally across the picture: left-hand red against right-hand green. The like-for-like comparison reads vertically: red against green at one shared mild-case fraction. That is the reversal made visible. Weighted averages are points between subgroup endpoints. Different weights select different horizontal positions, and a large enough horizontal shift can overwhelm the vertical treatment advantage."},{"title":"When Pooling Answers the Question","startSeconds":498.0758125,"narration":"The practical question is not whether totals are forbidden. Totals are useful summaries. The question is whether the pooled total answers the decision you believe it answers. Begin with severity. In our example, severity strongly predicts the outcome. Mild cases have a much higher baseline chance of success than severe cases. Severity also predicts which comparison group a patient enters. The treatment group received mostly severe cases, while the untreated group received mostly mild cases. Treatment choice may itself affect the outcome. But severity points toward both treatment choice and outcome, creating a backdoor route that mixes treatment effects with case selection. This is why the first question matters. Does the subgroup variable predict the outcome even without treatment? If it does, changing the subgroup mix can change an aggregate rate. Second, are treatment and comparison cases distributed differently across those subgroups? A risk factor cannot create this particular distortion if both groups carry the same mix. Third, which population is the decision about? A hospital may need the expected result for its actual patient mix. A causal comparison may instead need both options evaluated on one common mix. Pooling is informative when treatment assignment is balanced across the important risk groups. Then each overall rate combines comparable ingredients in comparable proportions. Pooling can also be appropriate when the target mixture is deliberately specified. If a health system expects seventy percent mild and thirty percent severe cases, it can standardize both options to that same mix. In that setting, the pooled answer is not pretending to be universal. It is an answer for a named population, using weights chosen to represent that population. Pooling becomes dangerous when baseline risks differ greatly, treatment exposure is uneven across the risk groups, and the crude total is read as though it were a like-for-like treatment effect. Our example satisfies all three warning conditions. Severity changes baseline success by sixty points. The mild share changes from ten percent to ninety percent. The pooled rates then compare unlike populations. A responsible analysis would report the raw total and the severity-specific results, then use stratification, standardization, regression, matching, or a suitable study design to compare like with like. One final caution. Finding a Simpson reversal does not automatically prove that the subgroup result is causal. Severity may be measured imperfectly, other confounders may remain, and the chosen subgroups may themselves be consequences of earlier decisions. The reversal is a diagnostic signal. It tells us that aggregation and composition matter, and that we should investigate how patients entered each group before drawing a treatment conclusion. So ask what the total literally measures. In this study, thirty-six and seventy-four percent describe outcomes in two observed populations. They are valid summaries of those populations. Then ask what comparison the decision requires. If the goal is a treatment effect, compare treatment and no treatment at the same severity mix. If the goal is operational forecasting, use the mix expected in the population where the decision will be applied. The lesson of Simpson's paradox is not simply to distrust averages. It is to identify their weights, understand what creates those weights, and make sure both sides of a comparison answer the same question for the same target population."}]}}
