Concentration of Measure on the Sphere
- 0 views
- Last updated
- Mathematics
On a high-dimensional sphere, a band you would barely notice on a globe holds essentially all of the area. This lecture shows that phenomenon before explaining it, then builds the explanation in full: the epsilon-extension of a set, Levy's spherical isoperimetric inequality argued by comparing one cap against two, and the concentration inequality with its universal constants, cashed in at measure one half. The spheres are then named a normal Levy family, Lipschitz functions are forced to their medians in three lines, and the payoff is Milman's proof of Dvoretzky's theorem: every n-dimensional normed space has an almost Euclidean subspace of dimension on the order of log n, with the intuition carried by the straight and diagonal slices of a cube.
Take a sphere. Not the one in front of you, but its cousin in a thousand dimensions. Almost everything you know about the round one quietly fails up there, and it fails in a useful direction: by the end of this lecture, one fact about where a sphere keeps its area will hand us a theorem about every normed space there is. So start in three dimensions, where we can look. Here is the unit sphere, and there, drawn around its middle, is its equator. Let it turn once, so you can see it is a ball and not a disc. Now mark off a band around that equator: every point within zero point two radians of it, about twelve degrees of latitude either way. On this sphere the band is nothing special. It holds about twenty percent of the area, and you can see that: most of the surface is out in the two caps. Keep that picture, and raise the dimension. I cannot draw the sphere in dimension one thousand, but I can tell you what happens to this same band up there: it swallows essentially all of the area, and the two caps, which look like most of the sphere here, hold almost nothing at all. The reason is a picture we can draw. Slice the sphere by latitude, and ask how much area sits at each one. In dimension n the answer is proportional to sine of theta, raised to the power n minus one. For our sphere that exponent is just one, and the profile is this gentle mound. Now raise the exponent. At n equals eleven the mound sharpens. At n around fifty it is a spike, standing on the equator. The area has nowhere else to go: any latitude away from the equator has sine less than one, and a number below one, raised to the fiftieth power, is nothing. Here are the guarantees, for the very same band. On our sphere, dimension two, it holds about twenty percent. By dimension one hundred it already holds at least seventy two percent. By dimension one thousand, at least ninety nine point nine nine nine nine nine percent. And by ten thousand, everything except one part in ten to the eighty six. A band you would barely notice on a globe is, up there, the whole sphere. That is concentration of measure, and two questions come with it. Why the equator, of all places, and what guarantees those numbers? Keep one eye on where this is heading: the same phenomenon will, at the end, buy Dvoretzky's theorem. But first the why, and the why is an inequality about caps.
A cap is the sphere's own disc. Pick a point x zero, and take everything within spherical distance R of it. There it is, poured around the pole. Let it turn, so you can see the cap lying on the surface, and there is its formula. Now grow it. Take any set A and fatten it by epsilon: everything within distance epsilon of A. That is the epsilon extension, written A sub epsilon, and for a cap it means the cap plus this collar. Growing a set costs area, and the bill is paid along the boundary: the longer the edge, the more the collar adds. So here is Levy's question. Among all sets of one fixed area, which shape, when you fatten it by epsilon, grows the least? Watch what happens to a bad candidate. Here is the same total area, split into two half-size caps, one at each pole. Nothing about the area changed. What changed is the edge: this set has two boundary circles where the single cap had one. Now fatten both sides by the same epsilon. The single cap grows one collar. The split set grows two, and two collars around two circles is more new area than one. Same area in, more area out: splitting was a bad trade. And that is the pattern in general. Wiggly sets, stretched sets, scattered sets: at a fixed area, every one of them carries at least as much boundary as the cap, and grows at least as fast. The cap is the minimizer. Say it precisely. If A and the cap C cover the same fraction of the sphere, then for every positive epsilon, the extension of A is at least as big as the extension of C. That is Levy's spherical isoperimetric inequality, and it is the engine under everything that follows. One more reading before we use it. The cap is solving the same problem the disc solves in the plane and the ball solves in space: least boundary for the area enclosed. On the sphere, least boundary becomes least growth. Now let us spend it.
Here is the payoff move, and it is short. Take any set A covering exactly half the sphere: half is where a set and its complement balance, and it is the case that powers everything. Which cap has measure one half? The one of radius pi over two: a hemisphere. So the hemisphere is the shape to beat. Let it turn once, so that blue reads as half of a surface rather than as a lid laid across one. Fatten the hemisphere by epsilon and you get this: everything above the equator, plus a collar of angular width epsilon below it. By the isoperimetric inequality, whatever A actually was, its extension covers at least as much of the sphere as this does. So what does the grown hemisphere miss? Only this: the cap of radius pi over two minus epsilon, around the opposite pole. Call it B. The extension covers one minus the measure of B, and everything now rides on how small B is. And B is exactly the kind of region the spike killed. Every point of it sits at latitude at least epsilon away from the equator, where the sine to the n minus one profile has already collapsed. Integrate that profile and the measure of B comes out at most a constant times e to the minus c n epsilon squared. Put the chain together. Any set of measure one half, grown by epsilon, covers all of the sphere except an exponentially small remainder. The constants C and c are universal: they depend on nothing, not the dimension, not the set, not epsilon. That is the concentration inequality. How small is exponentially small? Let me plot the escaping mass, e to the minus n epsilon squared, against the dimension. Here it is for epsilon equal to zero point two, and above it, decaying slower, the thinner margin zero point one. Now ride the steep one. At n around fifty the escaping mass is already near a tenth. By one hundred and fifty it is a quarter of one percent, and off the right edge, in dimension one thousand, it is ten to the minus eighteen. And compare the two curves: halving epsilon quarters the exponent, so the thin margin decays four times slower. The square on epsilon is real. This behavior earned the spheres a title. A sequence of spaces satisfying exactly this bound, uniformly in n, is called a normal Levy family, and the spheres are the founding example. The name matters because the mechanism travels: any family with the bound inherits everything we do next.
Concentration so far is about sets. The version that gets used is about functions. Take any real function f on the sphere that is Lipschitz with constant one, so it moves by at most the distance you moved, and let M be its median: half the sphere sits at or below M. Now grow that half. Every point of the extension is within epsilon of somewhere f was at most M, and a Lipschitz function cannot climb faster than distance. So on the whole extension, f is at most M plus epsilon, and the extension is nearly everything. Run the same argument from above, with the set where f is at least M, and the mirror bound comes out. Put the two together: apart from an exception of size at most twice the old one, every point of the sphere has f within epsilon of the median. There it is: essentially the entire sphere lands in this little window. Read that slowly, because it is strange. A Lipschitz function on a high-dimensional sphere is, for every practical purpose, a constant. You may design f as cleverly as you like; the geometry flattens it. Milman's insight was that this is not a curiosity but a tool. Here is the tool at work, first where we can see it. This is the cube, and everything that follows is one flat cut through its centre, taken two different ways. Let it turn first, so you know where its corners are. Cut it straight across, parallel to a face. The knife goes in flat, like this, and the face it leaves behind is a square. That much you would have guessed without me. Now tilt the knife, until it stands square to the long diagonal of the cube, the line joining one corner to the corner furthest away from it. Same cube, same centre, different angle. And the face it leaves is a hexagon. A regular one, with six equal sides, cut out of a body built entirely from squares. This is the step nobody believes until they watch it. So turn it again, and follow the rim. It crosses six of the cube's twelve edges, and it crosses every one of them at its midpoint. Six midpoints, six equal sides. Now lay the two sections flat and measure them. Draw the biggest circle that fits inside each one, and the smallest that wraps around it. For the square, the inner circle has radius one and the outer reaches the corners at root two, so the outer is forty one percent bigger than the inner. The hexagon is measurably rounder. Its inner circle is the larger of the two, while its outer circle is the same root two, because its corners are still edge midpoints of the same cube. The ratio drops to two over root three, about one point one five. Same cube, better direction, and the section has forgotten most of the cube's corners. In dimension n the same cube has central sections round to within any tolerance you name, and, more surprisingly, you do not have to hunt for the direction. The norm of the cube, like any norm, is a Lipschitz function on the sphere, so it concentrates at its median. On a random subspace of the right dimension it is nearly constant, and nearly constant is nearly round. That is Dvoretzky's theorem, in the sharp form Milman proved with exactly this concentration. For every tolerance epsilon there is a constant c of epsilon, and every n-dimensional normed space, no matter how jagged its unit ball, contains a subspace of dimension at least c of epsilon times log n on which the norm is within one plus epsilon of Euclidean. And the proof is this lecture run in reverse. Caps grow slowest, so sets of half measure grow to everything, so Lipschitz functions sit at their medians, so norms are constant on random subspaces, so round sections exist. One inequality about the oldest shape on the sphere, pushed four steps, reaches every norm in every dimension. Log n is small, and that is honest: the cube itself shows you cannot do better than logarithmic in general. But it is not zero, and that is the miracle. In high dimension, roundness is not the exception. It is what is left when there is nowhere else for the measure to go.
Loading discussion…