The Central Limit Theorem, Shown by Sampling
- 0 views
- Last updated
- Statistics
A skewed population, samples of four days at a time, and the wall of gray bricks their averages build: this lecture watches the Central Limit Theorem happen before stating it. The tail dies because averaging is a tug of war; the bell's width comes out as sigma over root n; growing the sample from four to sixteen to sixty-four halves the width twice; and the closing chapters read the fine print and cash the theorem in as the error bar on every reported measurement.
Every field that measures anything keeps meeting the same curve. Biologists find it in heights, physicists in noise, pollsters in polls: a symmetric bell, over and over, in data that has no business agreeing about anything. Today we find out where that bell comes from, and the answer is one theorem about averages. So let me build a world with no bell in it anywhere. Here is a population: the number of sales a small shop makes in a day. Most days are quiet, and one sale is the commonest day of all, at thirty-two percent. Two sales happen on twenty-two percent of days, and three on sixteen. Then the tail begins: four, five, six, and every so often a wild day of seven sales. Two numbers will follow us all lecture. The mean, mu, is two point eight, the balance point of the whole shape. And the standard deviation, sigma, is one point eight, the typical distance a single day lands from that mean. And look at the shape itself. It is lopsided: the tail runs off to the right, and a seven sits far from everything else. Nothing about this picture is symmetric, and nothing about it is a bell. Now the destination. This curve on the right is the normal curve: the bell that every one of those fields keeps finding, symmetric about its middle and thin in both tails. Keep the two pictures side by side, because together they are the whole lecture. The claim is that the lopsided shop on the left will manufacture the bell on the right, and the machine that does it is averaging. Take four days at random, average them, and keep the average. Do it again and again, and watch what the averages do.
Here is the machine we ended on, and its one rule: draw four days at random, average them, and keep nothing but the average. On the left, the population we draw from. On the right, an empty frame where the averages will live. First sample. The shop hands us a two, a one, a five, and a three, and there they are, sitting on the population. Now the arithmetic, in full. Two plus one plus five plus three is eleven, and eleven over four is two point seven five. That is the sample mean, written x bar, and it is the only thing we keep. Over it goes. The tick marks two point seven five on the new axis, and then the whole sample becomes one gray brick, standing in the bin it landed in. The four days that made it are gone. Again. A one, a four, and a pair of twos: the sum is nine, nine over four is two point two five, and a second brick lands just to the left of the first. Once more: a six, a one, a three, and a two. Twelve over four is three exactly, and a third brick lands. Now a lucky draw. A seven comes up, with a one, a one, and a two. Add them: eleven again, so the average is two point seven five, again. Even a wild day, averaged with three quiet ones, gets dragged back to the middle of the pile. From here, let it run. Sample, average, brick. Sample, average, brick. Five more spins of that loop, five more bricks, and already the pile stands tallest in the middle and thin at both ends. Nine averages hint at a shape without showing it. To see it properly we need not nine but dozens, so let the machine take over and run while we watch.
Sixty-four averages now: the machine kept sampling while we talked, and every average left its brick. Look at the wall they built. It stands tallest just short of three, and it falls away on both sides. Set the wall against where it came from. The population leans left and drags its tail out to seven. The averages have no tail worth the name: past four and a half, the bins are simply empty. Averaging erased the skew, and the question is how. So watch a lucky sample try to build a tail. A seven comes up, and with it a one, a two, and a three. Their average: thirteen over four, which is three and a quarter. The seven hauled it to the right, and the quiet days hauled it straight back. Averaging is a tug of war, and the middle wins. An average out at seven would need all four draws to be sevens at once, and each has probability zero point zero five. Multiply the four together: about six chances in a million. That is why the wall is dead long before seven. Now bring back the curve from the opening minute and lay it over the wall. This is the normal curve, and it hugs the bricks. A lopsided population, pushed through nothing but averaging, has manufactured the bell. What we built deserves its name. The sample mean is itself random: a different sample gives a different average. Its distribution, across every sample you could possibly draw, is called the sampling distribution of the mean, and our histogram is a portrait of it. And measure the portrait. Its centre sits at two point eight, the population mean, exactly. Its width is zero point nine, and zero point nine is one point eight divided by two: sigma, divided by the square root of four. The sample size has crept into the width, and that is the theorem knocking.
Here it is then, in full: the Central Limit Theorem. Draw samples of size n from any population at all, provided it has a mean mu and a finite standard deviation sigma. As n grows, the distribution of the sample mean approaches the normal curve, centred on mu, with width sigma over the square root of n. That one line is the quantitative heart. The width of the bell of averages is sigma over root n. The population's shape is not in it, and the skew is not in it. Only mu, sigma, and n survive the averaging. Let us watch that square root earn its keep. Here is the bell for n equals four, centred on the dashed line at two point eight. The bracket underneath is one width of it: sigma of x bar, the quantity our formula computes. The computation sits beside it. Sigma is one point eight, the spread we measured back at the start. Root four is two. And one point eight over two is zero point nine. Now buy more data. Quadruple the sample to sixteen: root sixteen is four, and the width drops to one point eight over four, zero point four five. Watch the bracket shrink to match the new bell. Mark the price of that. Four times the data bought exactly half the width, not a quarter of it. The square root is the reason, and it never relents. Quadruple again, to sixty-four. Root sixty-four is eight, and the width is zero point two two five: the bell has sharpened into a spike. It grows taller as it narrows because the area under every one of these curves is one, so the probability has nowhere to go but up. So this is what averaging buys, and what it costs. It forgets the population's shape entirely, and it shrinks the noise by exactly root n. What remains is the fine print, because a theorem this strong has conditions.
A theorem this useful earns its fine print, and there are three clauses. Beside them, the shop again, to test each clause against. First: the draws must be independent, and all from one population. Our machine had that by construction. Real data, measured over time or across neighbours, has to earn it. Second: sigma must be finite. Here is a population that breaks the rule: this red tail dies away so slowly that the spread is infinite, and averages drawn from it never settle into any bell. No sigma, no theorem. Third: how large n must be depends on the shape you start from. Our shop's mild tail was nearly gone by n equals four; a harsher skew takes more averaging. The theorem promises the bell in the limit, and the shape decides the price of getting there. Now spend the theorem on the thing you meet everywhere: the error bar. A lab measures something noisy, averages n readings, and reports the average: three point one. An honest report also says how far off that number is likely to be. The theorem is the answer. That reported average was one draw from a bell whose width is sigma over root n, and here that width is zero point two. So draw the bar: from two point nine to three point three, one width each side of the reading. That is the error bar, drawn. So a quoted plus or minus is not decoration, and it is not a guess. It is this theorem, applied: the width of the normal curve the average came from. Wherever an error bar stands, a bell stands behind it. One curve carries the whole lecture. Here is the noise in the average, sigma over root n, plotted against the sample size. At n equals four it stood at zero point nine. At sixteen, zero point four five. At sixty-four, zero point two two five. The curve falls like one over root n, so each halving of the noise costs a quadrupling of the data. That is why large experiments are large, and why the last decimal place is always the expensive one. And that is the Central Limit Theorem. Averages forget the shape they were drawn from, remember the mean, and close in on it like root n. The bell was never hiding in the population. Averaging builds it, every time.
Loading discussion…