Why a 99% Accurate Test Can Still Be Wrong
- 0 views
- Last updated
- Statistics
A visual introduction to base rates and Bayes' theorem for viewers without statistics training. Starting with 10,000 people and a rare disease, the lecture separates true positives from false positives, reads the probability directly from those populations, and only then introduces Bayes' formula. It shows how prevalence and symmetric test accuracy change the meaning of a positive result, then follows the same population through a second conditionally independent test.
Suppose a medical test is described as ninety-nine percent accurate, and it comes back positive. That sounds almost conclusive. But if the disease is rare, the positive result can still be more likely wrong than right. We are going to see why by counting people before writing any probability formula. Here is the question in its most personal form. A positive result says you have a rare disease. Does ninety-nine percent accurate mean there is a ninety-nine percent chance you have it? No. That number describes how the test behaves inside known groups. It does not yet answer what group a positive person probably came from. Take ten thousand people. The large gray field represents the people in this cohort who do not have the disease. I have magnified the affected people so we can actually see them. Let the disease affect one person in a thousand. That is a prevalence of zero point one percent. In ten thousand people, only ten actually have the disease. The remaining nine thousand nine hundred ninety are healthy. Now we must say exactly what ninety-nine percent accurate means. For this lecture, it means two things. Sensitivity is ninety-nine percent, so among people who truly have the disease, the test is positive ninety-nine percent of the time. Specificity is also ninety-nine percent, so among healthy people, the test is negative ninety-nine percent of the time. Apply sensitivity to the ten affected people. Ninety-nine percent of ten is nine point nine. So across many cohorts like this one, we expect about nine point nine true positive results. Now turn to the healthy majority. Ninety-nine percent specificity leaves a one percent false-positive rate. One percent sounds tiny, but it acts on nine thousand nine hundred ninety people. One percent of that enormous healthy group is ninety-nine point nine false positives. The decimal counts are expected counts, averages over many equally sized cohorts. In one real cohort we would see whole people, very close to these values.
The test has now run on all ten thousand people. But after a positive result, most of that original cohort is no longer relevant. We need a new reference group: everyone whose result was positive. The green pile contains the positive results from people who truly have the disease. Its expected size is nine point nine. These are the true positives. The yellow pile contains positive results from healthy people. Its expected size is ninety-nine point nine. These are false positives, contributed by the enormous healthy majority. Pause on the picture. The test is excellent inside either group. Yet the false-positive pile is about ten times taller, because the healthy group supplying it began nine hundred ninety-nine times larger than the disease group. Now gather the two piles. Write nine point nine true positives above ninety-nine point nine false positives. Rule beneath them and add. The positive-test group contains about one hundred nine point eight people. To answer our question, ask what fraction of that positive group came from the green pile. The numerator is nine point nine true positives. The denominator is every positive result, one hundred nine point eight. That fraction is about nine percent. So after one positive test, the chance of actually having the disease is only about nine percent under our assumptions. The complementary probability is about ninety-one percent. In other words, this positive result is probably wrong, even though the test has ninety-nine percent sensitivity and ninety-nine percent specificity. Nothing paradoxical happened. The test made errors on only one percent of healthy people. There were simply so many healthy people that their small error rate produced far more positive results than the rare disease did.
Now that the populations are visible, we can compress the same reasoning into Bayes' theorem. Start with the rule we already used: true positives divided by all positive results. The vertical bar means given. P of D given positive asks: among people known to have a positive result, what fraction have disease? The phrase after the bar names the reference group. Before building the formula, name its ingredients. P of D is prevalence, the fraction who have the disease before testing. Here it is zero point zero zero one. P of positive given D is sensitivity, the positive rate inside the disease group. Here it is zero point nine nine. P of D complement is the healthy share, zero point nine nine nine. And P of positive given D complement is the false-positive rate, zero point zero one. Now replace each pile by the probability that creates it. Sensitivity times prevalence creates the true-positive share. False-positive rate times healthy share creates the false-positive share. The numerator keeps the true-positive route. The denominator adds both routes into the positive group. This is exactly what the two visible piles did, with the common population size canceled out. Substitute our values. The true-positive route is zero point nine nine times zero point zero zero one. The false-positive route is zero point zero one times zero point nine nine nine. The result is about zero point zero nine zero, or nine percent. Bayes' theorem has not introduced a new argument. It has merely named the count in a form that works for any cohort size. The common mistake is to reverse the condition. Ninety-nine percent sensitivity describes positive results among people already known to have disease. We wanted disease among people already known to have a positive result. Those are different questions.
Bayes' formula lets us change one ingredient at a time. First keep the test fixed at ninety-nine percent sensitivity and specificity, and vary only the disease prevalence. The horizontal coordinate is prevalence as a percentage. The vertical coordinate is the chance of disease after a positive result. At zero point one percent prevalence, our yellow point reads about nine percent. Raise prevalence to one percent. Now one person in a hundred has the disease before testing. The point climbs to fifty percent, because the expected true-positive and false-positive piles are equal. Raise prevalence to five percent. The test has not improved at all, but the positive result now means about eighty-three point nine percent. At ten percent prevalence, the posterior reaches about ninety-one point seven percent. The same test result means something very different in a high-risk population than in a low-risk population. Prevalence is the starting information, sometimes called the prior probability. A positive test updates that starting point. It does not erase it. Now restore the very rare prevalence of zero point one percent and change the test itself. To keep the phrase accuracy unambiguous, a will mean both sensitivity and specificity. At ninety-nine percent accuracy, the point again sits near nine percent. Its label gives accuracy first and the posterior probability second. Drop accuracy to ninety-five percent. The posterior falls below two percent. A five percent false-positive rate applied to nearly ten thousand healthy people overwhelms the true-positive pile. Return to ninety-nine percent, and we recover about nine percent. Now push the accuracy to ninety-nine point nine percent. The false-positive rate falls from one percent to one tenth of one percent. That extra nine in the accuracy raises the posterior to about fifty percent. For an extremely rare disease, tiny changes in the false-positive rate can matter enormously because that rate acts on the healthy majority. So the phrase ninety-nine percent accurate is incomplete on its own. We need sensitivity, specificity, and prevalence. Change any one of them and the meaning of a positive result can swing dramatically.
Suppose the same person is tested again and the second result is also positive. Begin with the group that survived the first test: about nine point nine true positives and ninety-nine point nine false positives. Assume the second test is conditionally independent of the first. Among the people who truly have disease, it again detects ninety-nine percent. Ninety-nine percent of nine point nine is nine point eight zero one. Among the healthy people who produced the first false positive, only one percent produce another false positive independently. One percent of ninety-nine point nine is zero point nine nine nine. Now read the two surviving piles. About nine point eight people are true positives twice, while about one person is falsely positive twice. The green pile is finally much larger than the yellow pile. The probability of disease after two positive results is the green count divided by the two surviving counts together. That is about ninety point eight percent. One independent repeat test has moved the answer from about nine percent to about ninety-one percent by filtering both piles again. There is a compact way to understand that jump. Start with disease odds of ten to nine thousand nine hundred ninety, which reduce to one to nine hundred ninety-nine. A positive result is ninety-nine times more likely when disease is present than when it is absent. That factor, ninety-nine, is called the positive likelihood ratio. The first positive result multiplies the prior odds by ninety-nine. That produces the same roughly nine percent probability we found from the first two piles. Under conditional independence, the second positive multiplies by the same factor again. Two positives contribute ninety-nine times ninety-nine, changing the odds by a factor of nine thousand eight hundred one. Convert those final odds back to a probability and we recover ninety point eight percent. The count method and the odds method are two views of the same update. The independence assumption matters. If both tests use the same sample, the same instrument, or the same biological signal, their errors may be correlated. A repeated error can then be more likely than this calculation assumes, so the second positive may add less evidence. The lesson is not to distrust accurate tests. It is to ask the complete question. How rare is the disease? What are the sensitivity and specificity? And is new evidence genuinely independent? With those facts, a surprising positive result becomes a count we can understand.
Loading discussion…