# The Mathematics of Neural Networks and Gradient Descent

> A first course in the mathematics behind neural networks. We build one unit out of inputs, weights and a bias, show why a purely linear network collapses into a single straight line, and introduce the sigmoid and ReLU activations that bend it. We then define mean squared error, watch the loss become a function of the weights, and walk downhill: derivatives as slopes, the learning rate, partial derivatives and the gradient. The final section assembles the chain rule into backpropagation and carries one complete training step through in numbers, from the forward pass to the updated weight and the loss that fell.

- Canonical watch page: [The Mathematics of Neural Networks and Gradient Descent](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent)
- Publisher: [Academa, Inc.](https://academa.ai)
- Subject: Machine Learning
- Published: 2026-10-01T20:02:52.912Z
- Updated: 2026-10-01T20:02:52.912Z
- Duration: PT871S (14 minutes 31 seconds)
- Chapters: 5
- Views: 1
- Language: en-US
- Access: Free
- Video stream: [HLS content](https://academa.ai/media/l/01M3WDT9XXSAEX76WGMC8M29GQ/0/dark/master.m3u8)
- Embed: [Player](https://academa.ai/embed/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent)
- Audiovisual record: [Semantic JSON](https://academa.ai/media/l/01M3WDT9XXSAEX76WGMC8M29GQ/0/semantic.json)
- Thumbnail: [Image](https://academa.ai/media/l/01M3WDT9XXSAEX76WGMC8M29GQ/0/dark/poster.jpg)

## Description

How neural networks compute, how loss measures error, and how derivatives, the gradient and the chain rule update every weight.

## Chapters

- [00:00–03:5.759 · From Numbers to a Prediction](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=0)
- [03:5.759–05:45.684 · Why a Straight Line Is Not Enough](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=185.75941666666662)
- [05:45.684–07:41.83 · A Number for Being Wrong](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=345.68372916666664)
- [07:41.83–10:57.565 · Walking Downhill](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333)
- [10:57.565–14:31 · Backpropagation: The Chain Rule Doing the Work](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333)

## Transcript

### [00:00 · From Numbers to a Prediction](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=0)

A neural network is a piece of arithmetic with adjustable numbers inside it. Suppose two measurements describe a student: hours of study, and hours of sleep. We want one number out, a predicted exam score. So there are two questions. What arithmetic should we do, and how can the numbers inside it fix themselves when the answer comes out wrong? Draw that arithmetic as a picture. The two inputs sit on the left, and I will call them x one and x two. For our student, x one is three hours of study and x two is eight hours of sleep. Over here is a single unit, and its job is to turn those two numbers into one. Every arrow carries a number of its own, called a weight, and the weight decides how much that input matters. The first weight is zero point five. The second is zero point two. The unit does two things with those numbers. First it adds up the weighted inputs. Then it adds one more number of its own, which belongs to the unit rather than to any input. That number is called the bias, and here it is minus one. It lets the unit shift its answer up or down whatever the inputs happen to be. Written out, that is the whole of one unit. z equals w one x one, plus w two x two, plus b. A weighted sum of the inputs, and then the bias. Put our numbers in: zero point five times three, plus zero point two times eight, minus one. One point five, plus one point six, minus one. The unit's output is two point one. Now clear the arithmetic and watch the weight itself. Raise the first weight to one, and the output climbs to three point six, because study now counts for twice as much. Lower it to zero point one, and the output drops to zero point nine. Those weights are the dials, and everything that follows is about how to set them. One unit is not a network. Start again from the two inputs, and this time give them three units instead of one. Every unit reads both of them, with its own pair of weights and its own bias. Each unit computes its own weighted sum, so the layer turns two numbers into three. Feed those three into one last unit, and out comes the single prediction, y hat. Written as algebra, the whole layer is one line: the vector z equals a matrix W times the vector x, plus the vector b. W holds every weight in the layer, one row for each unit, and b holds every bias. Change a single entry of W and the prediction changes. So the machinery is in place, and two things are missing. We need a number, call it L, that says how wrong the prediction is. And we need a rule that changes every weight and every bias to make that number smaller. Those two are the rest of this lecture.

### [03:5.759 · Why a Straight Line Is Not Enough](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=185.75941666666662)

A weighted sum is a straight-line formula, and straight lines have a limit. Here are five measurements. As x grows the output falls to zero and then climbs again. Try matching them with one straight line. Tilt the line one way, and the left half is wrong. Tilt it the other way, and the right half is wrong. There is no slope and no intercept that passes through all five points, because the data bends and the line cannot. You might hope that stacking two layers fixes that. Let us check. The first layer computes v x plus c. The second layer takes that answer and computes w times it, plus b. Substitute, and expand. The answer is w v times x, plus w c plus b. That is a straight line again, with a new slope and a new intercept. Two straight layers are one straight layer, and a hundred of them would still be one. So something has to bend. Take the simplest bend there is. It is called ReLU, and it returns its input when the input is positive, and zero otherwise. ReLU of two minus x is the blue piece. It slopes down until x reaches two, and after that it is flat at zero. ReLU of x minus two is the green piece: flat at zero until two, and then rising. Wherever one of them is positive the other is zero, so adding them gives exactly the shape through all five points. One bend per unit, and a layer of units can fold a straight line into very nearly any shape you want. Two bends are standard, and the first is the sigmoid. It squeezes any input at all into the range from zero to one. A large positive input gives almost one, a large negative input gives almost zero, and an input of zero gives exactly one half. That is useful when the output should read as a probability. We will also need its slope later, and the slope is unusually tidy: sigma prime equals sigma times one minus sigma. Where the curve is steep that number is large, and out at the flat ends it is almost nothing. The other standard bend is the one we just used, ReLU. It is not smooth at zero, but it is cheap to compute, and its slope is simply one on the right and zero on the left. Either way, a unit now does two things: it forms a weighted sum, and then it bends the result.

### [05:45.684 · A Number for Being Wrong](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=345.68372916666664)

A network has to know when it is wrong, and by how much. Here are three training examples. The red dots are the correct answers, recorded in advance, and the blue line is what the model currently predicts. For each example, subtract the target from the prediction. That difference is the error. The first prediction is half a unit too low, the second is a whole unit too high, and the third is half a unit too low again. Errors come with signs, and if we simply added them a prediction that is too high would cancel one that is too low. So square each error first. Squaring makes every contribution positive, and it punishes a big miss much more than a small one. Add the three squares. A quarter, plus one, plus a quarter, is one and a half. Then divide by the number of examples. The mean squared error is zero point five. In general, then: average the squared difference between prediction and target over all n examples. One number, for the whole data set, and it is the number we are going to make as small as we can. Now here is the change of view that the rest of the lecture rests on. The data are fixed. We cannot alter a single target. The only things we can move are the weights, so the loss is really a function of the weights. To see that clearly, keep one single example: input equal to one, target equal to one. With one weight and no bias the prediction is just w, so the loss is w minus one, all squared. Sweep the weight and watch the loss. At w equals two the loss is one. Bring the weight down and the loss falls, reaches zero at w equals one, and climbs the far wall again. So training means finding the bottom of that bowl.

### [07:41.83 · Walking Downhill](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333)

We have a bowl, and we need to walk to the bottom of it. But we are not allowed to see the whole bowl. At any moment the network knows only the weight it currently has, and how steeply the loss is rising or falling right there. That steepness is exactly what a derivative is. The derivative of w minus one squared is two times w minus one. At w equals two it comes to two, a positive number, and the tangent line there rises to the right. A positive slope means the loss grows as the weight grows. So to make the loss smaller, move the weight the other way. Subtract something proportional to the slope, and that is the whole of gradient descent in one line. New weight equals old weight, minus eta times the derivative. Eta is a small positive number called the learning rate, and it decides how far each step carries us. Take eta equal to zero point three. The step is zero point three times two, which is zero point six, so the weight goes from two to one point four. Watch the tangent flatten as it lands. Do it again. At one point four the slope is only zero point eight, so the step is smaller and the weight moves to one point one six. Once more, and it reaches one point zero six. The steps shrink by themselves, because near the bottom there is hardly any slope left to multiply. Had the slope been negative, the minus sign would have pushed the weight up instead. Either way we go downhill. But eta matters. Suppose at w equals two we had used one point one. The step would be two point two, the weight would land at minus zero point two, and the loss there is worse than where we began. A real network has thousands of weights, not one, and nothing changes except that there is now a slope for each of them. Here is a loss with two weights, drawn as a contour map. Every ring is a set of weights giving the same loss, and the bottom of the bowl is inside the smallest ring. Hold w two still and ask how the loss changes as w one moves on its own. That is the partial derivative with respect to w one, and taking a step against it gives the blue arrow. Hold w one still instead, and you get the green arrow. Collect the partial derivatives into one list and you have the gradient. Here it is minus four and six. The gradient points in the direction the loss increases fastest, straight across the contour lines. So step the opposite way. Add the two component steps, and the red arrow is the diagonal of their parallelogram: that is where we land. The rule reads as before, with the gradient standing in for the single derivative. And repeat. Each step crosses to an inner ring, the arrows shorten as the ground flattens out, and the weights settle near the bottom. That is gradient descent. What is left is the hard part. In a real network, with layers feeding layers, how do we actually compute those partial derivatives?

### [10:57.565 · Backpropagation: The Chain Rule Doing the Work](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333)

We need the slope of the loss with respect to one weight. But the weight does not touch the loss directly. Follow the arithmetic forward: the weight and the input make z. The bend turns z into the activation a. And a, compared with the target y, produces the loss. So nudging the weight nudges z, and that nudges a, and that nudges the loss. Three links in one chain. Read the chain backwards, which is where the name backpropagation comes from. First, how fast does the loss respond to a? Next, how fast does a respond to z? Finally, how fast does z respond to the weight? Multiply those three local rates together and you have the slope you wanted. Each of those three rates is a one-line derivative. The loss is a minus y, all squared, so differentiating it with respect to a gives two times a minus y. The activation is sigma of z, and the sigmoid's slope is the tidy expression we met earlier: a times one minus a. And z is w x plus b, so differentiating with respect to w leaves just x. With respect to b it leaves one, because the bias is added on its own. Multiply the three together and there it is. The slope of the loss with respect to this weight is two times a minus y, times a times one minus a, times x. For the bias it is the same product without the x on the end. Now put numbers in. One input equal to one, a target of one, a weight of a half, and a bias of zero. The weighted sum is a half. The sigmoid of a half is zero point six two two, so the network answers zero point six two two when it should answer one. That red gap is the error. a minus y is minus zero point three seven eight, and twice it is minus zero point seven five five. There is the first factor. The second factor is the steepness of the sigmoid at this point, a times one minus a, which works out as zero point two three five. You can see it in the tangent line: a gentle slope, so a gentle factor. Multiply them, together with x, which is one. The slope of the loss with respect to this weight is minus zero point one seven seven. It is negative, so raising the weight lowers the loss. With a learning rate of a half the weight moves from zero point five to zero point five eight nine, and the activation climbs the curve toward its target. And the loss falls, from zero point one four three to zero point one two seven. The gap is narrower than it was. That is one training step for one weight, and a network does it for every weight it has, over and over again. So that is the whole loop. Forward, weighted sums and bends turn the inputs into a prediction. The loss collapses all of that into one number. The chain rule turns that number into a slope for every weight and every bias in the network. Then each weight takes a small step against its own slope, and we go round again. Everything else in deep learning is this loop, done at scale.

## About Academa, Inc.

Academa makes technical knowledge easier to understand through visual lectures and lets learners request new lecture videos on the topics they need.

## Complete audiovisual record

Immutable source: [semantic.json](https://academa.ai/media/l/01M3WDT9XXSAEX76WGMC8M29GQ/0/semantic.json)

Record version: 1. Render attempt: 0.

### How to read this timeline

Each scene owns its object identifiers. A beat's board is the complete board when listed, empty when marked empty, and unchanged from the nearest earlier listed board in the same scene when marked unchanged. Action times are absolute positions in the published video.

### Scene 1: [From Numbers to a Prediction](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=0)

Span: 00:00–03:5.759 (0s–185.75941666666662s).

#### Objects

- bias\_link: an Arrow \[green\] labelled "b = -1" drawn in net (start=(3.4, 1.0), end=(3.4, 1.85))
- cell: a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14)
- edges\_in: a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 3.5))
- edges\_in\_2: a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 2.1))
- edges\_in\_3: a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 0.7))
- edges\_in\_4: a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 3.5))
- edges\_in\_5: a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 2.1))
- edges\_in\_6: a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 0.7))
- edges\_out: a Line \[gray\] drawn in stack (start=(3.25, 3.5), end=(5.35, 2.1))
- edges\_out\_2: a Line \[gray\] drawn in stack (start=(3.25, 2.1), end=(5.35, 2.1))
- edges\_out\_3: a Line \[gray\] drawn in stack (start=(3.25, 0.7), end=(5.35, 2.1))
- form: a Math \[text\] that says "$z = w\_1 x\_1 + w\_2 x\_2 + b$"
- heading\_layer: a Heading that says "A Layer of Units"
- hidden: a Point \[yellow\] labelled "z\_1" drawn in stack (location=(3.1, 3.5))
- hidden\_2: a Point \[yellow\] labelled "z\_2" drawn in stack (location=(3.1, 2.1))
- hidden\_3: a Point \[yellow\] labelled "z\_3" drawn in stack (location=(3.1, 0.7))
- in\_1: a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0))
- in\_2: a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3))
- layer\_math: a Math \[text\] that says "$arrow(z) = W arrow(x) + arrow(b)$"
- link\_1: an Arrow \[gray\] labelled "w\_1 = 0.5" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35))
- link\_2: an Arrow \[gray\] labelled "w\_2 = 0.2" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95))
- net: a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5))
- numbers: a Math \[text\] that says "$z = 0.5 dot.op 3 + 0.2 dot.op 8 - 1$"
- out\_link: an Arrow \[red\] labelled "z = 2.1" drawn in net (start=(3.7, 2.15), end=(5.3, 2.15))
- promises: a Block \[text\] that says "A number $L$ that measures how wrong $hat(y)$ is. A rule that changes every weight to make $L$ smaller."
- question: a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"
- s\_in\_1: a Point \[blue\] labelled "x\_1" drawn in stack (location=(0.8, 2.9))
- s\_in\_2: a Point \[blue\] labelled "x\_2" drawn in stack (location=(0.8, 1.3))
- s\_out: a Point \[red\] labelled "hat(y)" drawn in stack (location=(5.5, 2.1))
- shapes: a Math \[text\] that says "$W: 3 times 2, quad arrow(b): 3 times 1$"
- stack: a Figure (x\_range=(0.0, 6.4), y\_range=(0.2, 4.2), aspect=(6.4, 4.0))
- w1: a VariableNumber (initial\_value=0.5, format\_spec='.1f')
- z\_num: a VariableNumber (initial\_value=2.1, format\_spec='.1f')

#### Beats

##### [00:00](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=0)

Narration: A neural network is a piece of arithmetic with adjustable numbers inside it. Suppose two measurements describe a student: hours of study, and hours of sleep. We want one number out, a predicted exam score. So there are two questions. What arithmetic should we do, and how can the numbers inside it fix themselves when the answer comes out wrong?

Board: Empty.

Actions:
- [00:00](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=0): question is shown on the screen, written out.
- [00:22.942](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=22.9415): question moves to a new place on the board.

##### [00:23.542](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=23.541500000000003)

Narration: Draw that arithmetic as a picture. The two inputs sit on the left, and I will call them x one and x two. For our student, x one is three hours of study and x two is eight hours of sleep. Over here is a single unit, and its job is to turn those two numbers into one.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"

Actions:
- [00:23.542](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=23.541500000000003): net is shown on the screen, written out.
- [00:27.698](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=27.698): in\_1 is shown on the screen, written out.
- [00:28.132](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=28.131558149589278): in\_2 is shown on the screen, written out.
- [00:38.855](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=38.855): cell is shown on the screen, written out.

##### [00:43.309](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=43.309000000000005)

Narration: Every arrow carries a number of its own, called a weight, and the weight decides how much that input matters. The first weight is zero point five. The second is zero point two.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"; net — a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5)); in\_1 — a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0)); in\_2 — a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3)); cell — a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14)

Actions:
- [00:44.505](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=44.505): link\_1 is shown on the screen, written out.
- [00:53.131](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=53.131): link\_2 is shown on the screen, written out.

##### [00:55.461](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=55.4605)

Narration: The unit does two things with those numbers. First it adds up the weighted inputs. Then it adds one more number of its own, which belongs to the unit rather than to any input. That number is called the bias, and here it is minus one. It lets the unit shift its answer up or down whatever the inputs happen to be.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"; net — a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5)); in\_1 — a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0)); in\_2 — a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3)); cell — a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14); link\_1 — an Arrow \[gray\] labelled "w\_1 = 0.5" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35)); link\_2 — an Arrow \[gray\] labelled "w\_2 = 0.2" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95))

Actions:
- [00:58.746](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=58.746): cell is indicated — a transient flash.
- [01:8.022](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=68.02199999999999): bias\_link is shown on the screen, written out.

##### [01:16.076](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=76.07600000000001)

Narration: Written out, that is the whole of one unit. z equals w one x one, plus w two x two, plus b. A weighted sum of the inputs, and then the bias. Put our numbers in: zero point five times three, plus zero point two times eight, minus one.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"; net — a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5)); in\_1 — a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0)); in\_2 — a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3)); cell — a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14); link\_1 — an Arrow \[gray\] labelled "w\_1 = 0.5" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35)); link\_2 — an Arrow \[gray\] labelled "w\_2 = 0.2" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95)); bias\_link — an Arrow \[green\] labelled "b = -1" drawn in net (start=(3.4, 1.0), end=(3.4, 1.85))

Actions:
- [01:16.378](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=76.378): net moves to a new place on the board.
- [01:16.378](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=76.378): form is shown on the screen, written out.
- [01:28.174](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=88.174): numbers is shown on the screen, written out.

##### [01:35.026](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=95.02550000000001)

Narration: One point five, plus one point six, minus one. The unit's output is two point one.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"; net — a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5)); form — a Math \[text\] that says "$z = w\_1 x\_1 + w\_2 x\_2 + b$"; numbers — a Math \[text\] that says "$z = 0.5 dot.op 3 + 0.2 dot.op 8 - 1$"; in\_1 — a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0)); in\_2 — a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3)); cell — a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14); link\_1 — an Arrow \[gray\] labelled "w\_1 = 0.5" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35)); link\_2 — an Arrow \[gray\] labelled "w\_2 = 0.2" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95)); bias\_link — an Arrow \[green\] labelled "b = -1" drawn in net (start=(3.4, 1.0), end=(3.4, 1.85))

Actions:
- [01:35.026](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=95.02550000000001): numbers becomes "$z = 1.5 + 1.6 - 1$".
- [01:39.327](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=99.32700000000001): numbers becomes "$z = 1.5 + 1.6 - 1 = 2.1$".
- [01:39.327](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=99.32700000000001): out\_link is shown on the screen, written out.

##### [01:41.704](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=101.7035)

Narration: Now clear the arithmetic and watch the weight itself. Raise the first weight to one, and the output climbs to three point six, because study now counts for twice as much. Lower it to zero point one, and the output drops to zero point nine. Those weights are the dials, and everything that follows is about how to set them.

Board: question — a Panel that says "Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?"; net — a Figure (x\_range=(0.0, 6.6), y\_range=(0.4, 3.9), aspect=(6.6, 3.5)); form — a Math \[text\] that says "$z = w\_1 x\_1 + w\_2 x\_2 + b$"; numbers — a Math \[text\] that says "$z = 0.5 dot.op 3 + 0.2 dot.op 8 - 1$"; in\_1 — a Point \[blue\] labelled "x\_1" drawn in net (location=(0.9, 3.0)); in\_2 — a Point \[blue\] labelled "x\_2" drawn in net (location=(0.9, 1.3)); cell — a Point \[yellow\] labelled "Sigma" drawn in net (location=(3.4, 2.15), marker\_radius=0.14); link\_1 — an Arrow \[gray\] labelled "w\_1 = 0.5" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35)); link\_2 — an Arrow \[gray\] labelled "w\_2 = 0.2" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95)); bias\_link — an Arrow \[green\] labelled "b = -1" drawn in net (start=(3.4, 1.0), end=(3.4, 1.85)); out\_link — an Arrow \[red\] labelled "z = 2.1" drawn in net (start=(3.7, 2.15), end=(5.3, 2.15))

Actions:
- [01:42.226](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=102.226): numbers is hidden from the screen.
- [01:45.163](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=105.163): link\_1 is redrawn as the numbers it depends on change.
- [01:45.163](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=105.163): out\_link is redrawn as the numbers it depends on change.
- [01:45.163](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=105.163): w1 ticks to 1.0.
- [01:45.163](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=105.163): z\_num ticks to 3.6.
- [01:52.408](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=112.408): link\_1 is redrawn as the numbers it depends on change.
- [01:52.408](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=112.408): out\_link is redrawn as the numbers it depends on change.
- [01:52.408](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=112.408): w1 ticks to 0.1.
- [01:52.408](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=112.408): z\_num ticks to 0.9.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): form is hidden from the screen — left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): net is hidden from the screen — left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): in\_1 is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): in\_2 is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): cell is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): link\_1 is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): link\_2 is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): bias\_link is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): out\_link is hidden from the screen — net left the board.
- [02:1.917](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=121.91650000000001): question is hidden from the screen — left the board.

##### [02:3.117](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=123.1165)

Narration: One unit is not a network. Start again from the two inputs, and this time give them three units instead of one. Every unit reads both of them, with its own pair of weights and its own bias.

Board: Empty.

Actions:
- [02:3.117](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=123.1165): heading\_layer is shown on the screen, written out.
- [02:3.117](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=123.1165): stack is shown on the screen, written out.
- [02:6.75](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=126.75): s\_in\_1 is shown on the screen, written out.
- [02:7.061](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=127.06057028695969): s\_in\_2 is shown on the screen, written out.
- [02:8.457](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=128.457): hidden is shown on the screen, written out.
- [02:8.592](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=128.5924431529241): hidden\_2 is shown on the screen, written out.
- [02:8.728](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=128.7278863058482): hidden\_3 is shown on the screen, written out.
- [02:11.209](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.209): edges\_in is shown on the screen, written out.
- [02:11.27](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.27042390119868): edges\_in\_2 is shown on the screen, written out.
- [02:11.332](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.3318478023974): edges\_in\_3 is shown on the screen, written out.
- [02:11.393](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.39327170359607): edges\_in\_4 is shown on the screen, written out.
- [02:11.455](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.45469560479475): edges\_in\_5 is shown on the screen, written out.
- [02:11.52](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=131.51999622230295): edges\_in\_6 is shown on the screen, written out.

##### [02:15.93](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=135.9305)

Narration: Each unit computes its own weighted sum, so the layer turns two numbers into three. Feed those three into one last unit, and out comes the single prediction, y hat.

Board: stack — a Figure (x\_range=(0.0, 6.4), y\_range=(0.2, 4.2), aspect=(6.4, 4.0)); heading\_layer — a Heading that says "A Layer of Units"; s\_in\_1 — a Point \[blue\] labelled "x\_1" drawn in stack (location=(0.8, 2.9)); s\_in\_2 — a Point \[blue\] labelled "x\_2" drawn in stack (location=(0.8, 1.3)); hidden — a Point \[yellow\] labelled "z\_1" drawn in stack (location=(3.1, 3.5)); hidden\_2 — a Point \[yellow\] labelled "z\_2" drawn in stack (location=(3.1, 2.1)); hidden\_3 — a Point \[yellow\] labelled "z\_3" drawn in stack (location=(3.1, 0.7)); edges\_in — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 3.5)); edges\_in\_2 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 2.1)); edges\_in\_3 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 0.7)); edges\_in\_4 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 3.5)); edges\_in\_5 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 2.1)); edges\_in\_6 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 0.7))

Actions:
- [02:22.072](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=142.072): edges\_out is shown on the screen, written out.
- [02:22.155](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=142.1553800460155): edges\_out\_2 is shown on the screen, written out.
- [02:22.239](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=142.23876009203096): edges\_out\_3 is shown on the screen, written out.
- [02:24.65](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=144.65): s\_out is shown on the screen, written out.

##### [02:27.804](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=147.804)

Narration: Written as algebra, the whole layer is one line: the vector z equals a matrix W times the vector x, plus the vector b. W holds every weight in the layer, one row for each unit, and b holds every bias. Change a single entry of W and the prediction changes.

Board: stack — a Figure (x\_range=(0.0, 6.4), y\_range=(0.2, 4.2), aspect=(6.4, 4.0)); heading\_layer — a Heading that says "A Layer of Units"; s\_in\_1 — a Point \[blue\] labelled "x\_1" drawn in stack (location=(0.8, 2.9)); s\_in\_2 — a Point \[blue\] labelled "x\_2" drawn in stack (location=(0.8, 1.3)); hidden — a Point \[yellow\] labelled "z\_1" drawn in stack (location=(3.1, 3.5)); hidden\_2 — a Point \[yellow\] labelled "z\_2" drawn in stack (location=(3.1, 2.1)); hidden\_3 — a Point \[yellow\] labelled "z\_3" drawn in stack (location=(3.1, 0.7)); edges\_in — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 3.5)); edges\_in\_2 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 2.1)); edges\_in\_3 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 0.7)); edges\_in\_4 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 3.5)); edges\_in\_5 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 2.1)); edges\_in\_6 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 0.7)); edges\_out — a Line \[gray\] drawn in stack (start=(3.25, 3.5), end=(5.35, 2.1)); edges\_out\_2 — a Line \[gray\] drawn in stack (start=(3.25, 2.1), end=(5.35, 2.1)); edges\_out\_3 — a Line \[gray\] drawn in stack (start=(3.25, 0.7), end=(5.35, 2.1)); s\_out — a Point \[red\] labelled "hat(y)" drawn in stack (location=(5.5, 2.1))

Actions:
- [02:30.463](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=150.46300000000002): stack moves to a new place on the board.
- [02:30.463](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=150.46300000000002): layer\_math is shown on the screen, written out.
- [02:39.762](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=159.762): shapes is shown on the screen, written out.

##### [02:47.874](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=167.874)

Narration: So the machinery is in place, and two things are missing. We need a number, call it L, that says how wrong the prediction is. And we need a rule that changes every weight and every bias to make that number smaller. Those two are the rest of this lecture.

Board: layer\_math — a Math \[text\] that says "$arrow(z) = W arrow(x) + arrow(b)$"; shapes — a Math \[text\] that says "$W: 3 times 2, quad arrow(b): 3 times 1$"; stack — a Figure (x\_range=(0.0, 6.4), y\_range=(0.2, 4.2), aspect=(6.4, 4.0)); heading\_layer — a Heading that says "A Layer of Units"; s\_in\_1 — a Point \[blue\] labelled "x\_1" drawn in stack (location=(0.8, 2.9)); s\_in\_2 — a Point \[blue\] labelled "x\_2" drawn in stack (location=(0.8, 1.3)); hidden — a Point \[yellow\] labelled "z\_1" drawn in stack (location=(3.1, 3.5)); hidden\_2 — a Point \[yellow\] labelled "z\_2" drawn in stack (location=(3.1, 2.1)); hidden\_3 — a Point \[yellow\] labelled "z\_3" drawn in stack (location=(3.1, 0.7)); edges\_in — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 3.5)); edges\_in\_2 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 2.1)); edges\_in\_3 — a Line \[gray\] drawn in stack (start=(0.95, 2.9), end=(2.95, 0.7)); edges\_in\_4 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 3.5)); edges\_in\_5 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 2.1)); edges\_in\_6 — a Line \[gray\] drawn in stack (start=(0.95, 1.3), end=(2.95, 0.7)); edges\_out — a Line \[gray\] drawn in stack (start=(3.25, 3.5), end=(5.35, 2.1)); edges\_out\_2 — a Line \[gray\] drawn in stack (start=(3.25, 2.1), end=(5.35, 2.1)); edges\_out\_3 — a Line \[gray\] drawn in stack (start=(3.25, 0.7), end=(5.35, 2.1)); s\_out — a Point \[red\] labelled "hat(y)" drawn in stack (location=(5.5, 2.1))

Actions:
- [02:51.056](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=171.056): promises is shown on the screen, written out.
- [02:54.956](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=174.95600000000002): promises (the "how wrong" part) is emphasized.
- [02:57.533](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=177.53300000000002): promises (the "changes every weight" part) is emphasized.
- [02:57.533](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=177.53300000000002): promises (the "how wrong" part) is no longer emphasized.
- [03:3.211](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=183.211): promises (the "changes every weight" part) is no longer emphasized.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): heading\_layer is hidden from the screen — left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): layer\_math is hidden from the screen — left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): promises is hidden from the screen — left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): shapes is hidden from the screen — left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): stack is hidden from the screen — left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): s\_in\_1 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): s\_in\_2 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): hidden is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): hidden\_2 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): hidden\_3 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in\_2 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in\_3 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in\_4 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in\_5 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_in\_6 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_out is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_out\_2 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): edges\_out\_3 is hidden from the screen — stack left the board.
- [03:4.718](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=184.71775): s\_out is hidden from the screen — stack left the board.

### Scene 2: [Why a Straight Line Is Not Enough](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=185.75941666666662)

Span: 03:5.759–05:45.684 (185.75941666666662s–345.68372916666664s).

#### Objects

- ceiling: a Line \[gray\] drawn in sig\_axes (start=(-4.0, 1.0), end=(4.0, 1.0), dashed=True)
- collapse: a Derivation \[text\] that says "$h &= v x + c \\ hat(y) &= w h + b \\ &= w (v x + c) + b \\ &= (w v) x + (w c + b)$"
- fit: an Axes (x\_range=(-0.4, 4.6), y\_range=(-0.6, 3.0), x\_ticks\_every=1.0)
- fit\_line: a Math \[text\] that says "$hat(y) = upright("ReLU")(2 - x) + upright("ReLU")(x - 2)$"
- guess: a Line \[yellow\] labelled "v x + c" drawn in fit (start=(0.0, \<VariableNumber intercept = 0.4\>), end=(4.4, ((slope \* 4.4) + intercept)))
- half: a Point \[yellow\] labelled "(0, 0.5)" drawn in sig\_axes (location=(0.0, 0.5))
- heading\_bend: a Heading that says "One Bend Is Enough"
- heading\_collapse: a Heading that says "Two Straight Layers Collapse"
- heading\_two: a Heading that says "Two Standard Bends"
- intercept: a VariableNumber (initial\_value=1.0)
- left\_piece: a FunctionPlot \[blue\] drawn in fit (function=\<function\>, x\_range=(-0.3, 4.4))
- piece\_left: a Math \[blue\] that says "$upright("ReLU")(2 - x)$"
- piece\_right: a Math \[green\] that says "$upright("ReLU")(x - 2)$"
- point: a Point \[yellow\] drawn in sig\_axes (location=(0.0, 0.5))
- point\_2: a Point \[yellow\] drawn in sig\_axes (location=(3.0, 0.953))
- relu\_axes: an Axes (x\_range=(-3.0, 3.0), y\_range=(-0.4, 3.0), x\_ticks\_every=1.0)
- relu\_curve: a FunctionPlot \[green\] drawn in relu\_axes (function=\<function\>, x\_range=(-3.0, 3.0))
- relu\_def: a Math \[text\] that says "$upright("ReLU")(t) = max(0, t)$"
- relu\_formula: a Math \[text\] that says "$upright("ReLU")(z) = max(0, z)$"
- relu\_slope: a Text \[text\] that says "Slope $1$ right of zero, $0$ left."
- right\_piece: a FunctionPlot \[green\] drawn in fit (function=\<function\>, x\_range=(-0.3, 4.4))
- samples: a Point \[red\] drawn in fit (location=(0.0, 2.0))
- samples\_2: a Point \[red\] drawn in fit (location=(1.0, 1.0))
- samples\_3: a Point \[red\] drawn in fit (location=(2.0, 0.0))
- samples\_4: a Point \[red\] drawn in fit (location=(3.0, 1.0))
- samples\_5: a Point \[red\] drawn in fit (location=(4.0, 2.0))
- sig\_axes: an Axes (x\_range=(-4.0, 4.0), y\_range=(-0.2, 1.2), x\_ticks\_every=2.0)
- sig\_curve: a FunctionPlot \[blue\] drawn in sig\_axes (function=\<function\>, x\_range=(-4.0, 4.0))
- sig\_deriv: a Math \[text\] that says "$sigma'(z) = sigma(z) (1 - sigma(z))$"
- sig\_formula: a Math \[text\] that says "$sigma(z) = frac(1, 1 + e^(-z))$"
- slope: a VariableNumber (initial\_value=0.2)

#### Beats

##### [03:5.759](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=185.75941666666662)

Narration: A weighted sum is a straight-line formula, and straight lines have a limit. Here are five measurements. As x grows the output falls to zero and then climbs again. Try matching them with one straight line.

Board: Empty.

Actions:
- [03:5.759](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=185.75941666666662): heading\_collapse is shown on the screen, written out.
- [03:11.877](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=191.87741666666662): fit is shown on the screen, written out.
- [03:11.877](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=191.87741666666662): samples is shown on the screen, written out.
- [03:12.046](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=192.04570723823946): samples\_2 is shown on the screen, written out.
- [03:12.214](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=192.21399780981233): samples\_3 is shown on the screen, written out.
- [03:12.382](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=192.38228838138517): samples\_4 is shown on the screen, written out.
- [03:12.551](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=192.550578952958): samples\_5 is shown on the screen, written out.
- [03:17.532](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=197.53241666666662): guess is shown on the screen, written out.

##### [03:20.361](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=200.3609166666666)

Narration: Tilt the line one way, and the left half is wrong. Tilt it the other way, and the right half is wrong. There is no slope and no intercept that passes through all five points, because the data bends and the line cannot.

Board: fit — an Axes (x\_range=(-0.4, 4.6), y\_range=(-0.6, 3.0), x\_ticks\_every=1.0); heading\_collapse — a Heading that says "Two Straight Layers Collapse"; samples — a Point \[red\] drawn in fit (location=(0.0, 2.0)); samples\_2 — a Point \[red\] drawn in fit (location=(1.0, 1.0)); samples\_3 — a Point \[red\] drawn in fit (location=(2.0, 0.0)); samples\_4 — a Point \[red\] drawn in fit (location=(3.0, 1.0)); samples\_5 — a Point \[red\] drawn in fit (location=(4.0, 2.0)); guess — a Line \[yellow\] labelled "v x + c" drawn in fit (start=(0.0, \<VariableNumber intercept = 0.4\>), end=(4.4, ((slope \* 4.4) + intercept)))

Actions:
- [03:20.709](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=200.7094166666666): guess is redrawn as the numbers it depends on change.
- [03:20.709](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=200.7094166666666): slope ticks to -0.45.
- [03:20.709](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=200.7094166666666): intercept ticks to 2.0.
- [03:24.146](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=204.14641666666662): guess is redrawn as the numbers it depends on change.
- [03:24.146](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=204.14641666666662): slope ticks to 0.45.
- [03:24.146](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=204.14641666666662): intercept ticks to 0.4.

##### [03:35.334](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=215.3344166666666)

Narration: You might hope that stacking two layers fixes that. Let us check. The first layer computes v x plus c. The second layer takes that answer and computes w times it, plus b.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [03:40.617](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=220.6174166666666): fit moves to a new place on the board.
- [03:40.617](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=220.6174166666666): collapse is shown on the screen, written out.
- [03:44.1](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=224.1004166666666): collapse is shown on the screen, written out.

##### [03:48.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=228.64191666666662)

Narration: Substitute, and expand. The answer is w v times x, plus w c plus b. That is a straight line again, with a new slope and a new intercept. Two straight layers are one straight layer, and a hundred of them would still be one.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [03:48.833](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=228.8334166666666): collapse is shown on the screen, written out.
- [03:49.878](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=229.87841666666662): collapse is shown on the screen, written out.
- [03:58.237](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=238.2374166666666): collapse (the "(w v) x" part) is emphasized.
- [03:59.085](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=239.08541666666662): collapse (the "(w c + b)" part) is emphasized.
- [03:59.085](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=239.08541666666662): collapse (the "(w v) x" part) is no longer emphasized.
- [04:3.195](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=243.19541666666663): collapse (the "(w c + b)" part) is no longer emphasized.
- [04:4.879](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=244.8789166666666): collapse is hidden from the screen — left the board.
- [04:4.879](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=244.8789166666666): heading\_collapse is hidden from the screen — left the board.

##### [04:6.079](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=246.07891666666663)

Narration: So something has to bend. Take the simplest bend there is. It is called ReLU, and it returns its input when the input is positive, and zero otherwise.

Board: fit — an Axes (x\_range=(-0.4, 4.6), y\_range=(-0.6, 3.0), x\_ticks\_every=1.0); samples — a Point \[red\] drawn in fit (location=(0.0, 2.0)); samples\_2 — a Point \[red\] drawn in fit (location=(1.0, 1.0)); samples\_3 — a Point \[red\] drawn in fit (location=(2.0, 0.0)); samples\_4 — a Point \[red\] drawn in fit (location=(3.0, 1.0)); samples\_5 — a Point \[red\] drawn in fit (location=(4.0, 2.0)); guess — a Line \[yellow\] labelled "v x + c" drawn in fit (start=(0.0, \<VariableNumber intercept = 0.4\>), end=(4.4, ((slope \* 4.4) + intercept)))

Actions:
- [04:6.079](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=246.07891666666663): heading\_bend is shown on the screen, written out.
- [04:6.079](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=246.07891666666663): guess is hidden from the screen.
- [04:10.873](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=250.87341666666663): relu\_def is shown on the screen, written out.

##### [04:17.058](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=257.05791666666664)

Narration: ReLU of two minus x is the blue piece. It slopes down until x reaches two, and after that it is flat at zero. ReLU of x minus two is the green piece: flat at zero until two, and then rising.

Board: fit — an Axes (x\_range=(-0.4, 4.6), y\_range=(-0.6, 3.0), x\_ticks\_every=1.0); samples — a Point \[red\] drawn in fit (location=(0.0, 2.0)); samples\_2 — a Point \[red\] drawn in fit (location=(1.0, 1.0)); samples\_3 — a Point \[red\] drawn in fit (location=(2.0, 0.0)); samples\_4 — a Point \[red\] drawn in fit (location=(3.0, 1.0)); samples\_5 — a Point \[red\] drawn in fit (location=(4.0, 2.0)); relu\_def — a Math \[text\] that says "$upright("ReLU")(t) = max(0, t)$"; heading\_bend — a Heading that says "One Bend Is Enough"

Actions:
- [04:19.728](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=259.7284166666666): piece\_left is shown on the screen, written out.
- [04:19.728](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=259.7284166666666): left\_piece is shown on the screen, written out.
- [04:28.157](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=268.1574166666666): piece\_right is shown on the screen, written out.
- [04:28.157](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=268.1574166666666): right\_piece is shown on the screen, written out.

##### [04:33.169](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=273.1694166666666)

Narration: Wherever one of them is positive the other is zero, so adding them gives exactly the shape through all five points. One bend per unit, and a layer of units can fold a straight line into very nearly any shape you want.

Board: fit — an Axes (x\_range=(-0.4, 4.6), y\_range=(-0.6, 3.0), x\_ticks\_every=1.0); samples — a Point \[red\] drawn in fit (location=(0.0, 2.0)); samples\_2 — a Point \[red\] drawn in fit (location=(1.0, 1.0)); samples\_3 — a Point \[red\] drawn in fit (location=(2.0, 0.0)); samples\_4 — a Point \[red\] drawn in fit (location=(3.0, 1.0)); samples\_5 — a Point \[red\] drawn in fit (location=(4.0, 2.0)); relu\_def — a Math \[text\] that says "$upright("ReLU")(t) = max(0, t)$"; piece\_left — a Math \[blue\] that says "$upright("ReLU")(2 - x)$"; piece\_right — a Math \[green\] that says "$upright("ReLU")(x - 2)$"; heading\_bend — a Heading that says "One Bend Is Enough"; left\_piece — a FunctionPlot \[blue\] drawn in fit (function=\<function\>, x\_range=(-0.3, 4.4)); right\_piece — a FunctionPlot \[green\] drawn in fit (function=\<function\>, x\_range=(-0.3, 4.4))

Actions:
- [04:36.652](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=276.6524166666666): fit\_line is shown on the screen, written out.
- [04:37.464](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=277.46441666666664): left\_piece is indicated — a transient flash.
- [04:37.735](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=277.73484639126303): right\_piece is indicated — a transient flash.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): fit is hidden from the screen — left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): samples is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): samples\_2 is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): samples\_3 is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): samples\_4 is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): samples\_5 is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): left\_piece is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): right\_piece is hidden from the screen — fit left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): fit\_line is hidden from the screen — left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): heading\_bend is hidden from the screen — left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): piece\_left is hidden from the screen — left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): piece\_right is hidden from the screen — left the board.
- [04:47.032](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=287.03191666666663): relu\_def is hidden from the screen — left the board.

##### [04:48.232](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=288.2319166666666)

Narration: Two bends are standard, and the first is the sigmoid. It squeezes any input at all into the range from zero to one. A large positive input gives almost one, a large negative input gives almost zero, and an input of zero gives exactly one half.

Board: Empty.

Actions:
- [04:48.232](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=288.2319166666666): heading\_two is shown on the screen, written out.
- [04:51.343](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=291.3434166666666): sig\_axes is shown on the screen, written out.
- [04:51.343](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=291.3434166666666): sig\_formula is shown on the screen, written out.
- [04:52.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=292.88741666666664): sig\_curve is shown on the screen, written out.
- [04:54.861](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=294.8614166666666): ceiling is shown on the screen, written out.
- [05:5.124](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=305.1244166666666): half is shown on the screen, written out.

##### [05:6.456](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=306.45641666666666)

Narration: That is useful when the output should read as a probability. We will also need its slope later, and the slope is unusually tidy: sigma prime equals sigma times one minus sigma. Where the curve is steep that number is large, and out at the flat ends it is almost nothing.

Board: sig\_formula — a Math \[text\] that says "$sigma(z) = frac(1, 1 + e^(-z))$"; sig\_axes — an Axes (x\_range=(-4.0, 4.0), y\_range=(-0.2, 1.2), x\_ticks\_every=2.0); heading\_two — a Heading that says "Two Standard Bends"; sig\_curve — a FunctionPlot \[blue\] drawn in sig\_axes (function=\<function\>, x\_range=(-4.0, 4.0)); ceiling — a Line \[gray\] drawn in sig\_axes (start=(-4.0, 1.0), end=(4.0, 1.0), dashed=True); half — a Point \[yellow\] labelled "(0, 0.5)" drawn in sig\_axes (location=(0.0, 0.5))

Actions:
- [05:13.863](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=313.86341666666664): sig\_deriv is shown on the screen, written out.
- [05:20.109](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=320.1094166666666): point is shown on the screen, grown.
- [05:22.353](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=322.35305829620194): point is hidden from the screen.
- [05:22.64](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=322.6404166666666): point\_2 is shown on the screen, grown.
- [05:24.941](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=324.9413613161131): point\_2 is hidden from the screen.

##### [05:25.446](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=325.44641666666666)

Narration: The other standard bend is the one we just used, ReLU. It is not smooth at zero, but it is cheap to compute, and its slope is simply one on the right and zero on the left. Either way, a unit now does two things: it forms a weighted sum, and then it bends the result.

Board: sig\_formula — a Math \[text\] that says "$sigma(z) = frac(1, 1 + e^(-z))$"; sig\_deriv — a Math \[text\] that says "$sigma'(z) = sigma(z) (1 - sigma(z))$"; sig\_axes — an Axes (x\_range=(-4.0, 4.0), y\_range=(-0.2, 1.2), x\_ticks\_every=2.0); heading\_two — a Heading that says "Two Standard Bends"; sig\_curve — a FunctionPlot \[blue\] drawn in sig\_axes (function=\<function\>, x\_range=(-4.0, 4.0)); ceiling — a Line \[gray\] drawn in sig\_axes (start=(-4.0, 1.0), end=(4.0, 1.0), dashed=True); half — a Point \[yellow\] labelled "(0, 0.5)" drawn in sig\_axes (location=(0.0, 0.5))

Actions:
- [05:28.488](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=328.4884166666667): relu\_axes is shown on the screen, written out.
- [05:28.488](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=328.4884166666667): relu\_formula is shown on the screen, written out.
- [05:29.72](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=329.72008742516255): relu\_curve is shown on the screen, written out.
- [05:33.956](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=333.95641666666666): relu\_slope is shown on the screen, written out.
- [05:43.256](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=343.2564166666666): sig\_formula is indicated — a transient flash.
- [05:43.423](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=343.4231694771073): relu\_formula is indicated — a transient flash.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): heading\_two is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): relu\_axes is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): relu\_curve is hidden from the screen — relu\_axes left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): relu\_formula is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): relu\_slope is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): sig\_axes is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): sig\_curve is hidden from the screen — sig\_axes left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): ceiling is hidden from the screen — sig\_axes left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): half is hidden from the screen — sig\_axes left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): sig\_deriv is hidden from the screen — left the board.
- [05:44.642](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=344.6420625): sig\_formula is hidden from the screen — left the board.

### Scene 3: [A Number for Being Wrong](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=345.68372916666664)

Span: 05:45.684–07:41.83 (345.68372916666664s–461.83008333333333s).

#### Objects

- bowl: an Axes (x\_range=(-0.4, 2.6), y\_range=(-0.3, 2.3), x\_ticks\_every=0.5)
- bowl\_curve: a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.2, 2.4))
- e1: a Line \[yellow\] labelled "-0.5" drawn in pred (start=(1.0, 1.0), end=(1.0, 1.5))
- e2: a Line \[yellow\] labelled "+1.0" drawn in pred (start=(2.0, 2.0), end=(2.0, 1.0))
- e3: a Line \[yellow\] labelled "-0.5" drawn in pred (start=(3.0, 3.0), end=(3.0, 3.5))
- err\_def: a Math \[text\] that says "$e\_i = hat(y)\_i - y\_i$"
- heading\_loss: a Heading that says "The Loss, and What It Depends On"
- loss\_w: a Math \[text\] that says "$L(w) = (w - 1)^2$"
- model: a FunctionPlot \[blue\] drawn in pred (function=\<function\>, x\_range=(0.0, 3.6))
- mse\_def: a Panel that says "The mean squared error averages the squared difference between prediction and target over all $n$ training examples."
- mse\_formula: a Math \[text\] that says "$L = frac(1, n) sum\_(i = 1)^n (hat(y)\_i - y\_i)^2$"
- mse\_value: a Math \[text\] that says "$L = frac(1, 3) dot.op 1.50 = 0.50$"
- pred: an Axes (x\_range=(-0.3, 4.0), y\_range=(-0.3, 4.2), x\_ticks\_every=1.0)
- problem: a Tex \[text\] that says "How wrong is this model?"
- spot: a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='bowl\_curve', x=\<VariableNumber w\_live = 1.0\>)
- square\_sum: an Arithmetic \[text\] that says "$(-0.5)^2 = 0.25 (+1.0)^2 = 1.00 (-0.5)^2 = 0.25 upright("total") = 1.50$" (operator='+', operands=('(-0.5)^2 = 0.25', '(+1.0)^2 = 1.00', '(-0.5)^2 = 0.25'), result='upright("total") = 1.50')
- t1: a Point \[red\] labelled "y\_1" drawn in pred (location=(1.0, 1.5))
- t2: a Point \[red\] labelled "y\_2" drawn in pred (location=(2.0, 1.0))
- t3: a Point \[red\] labelled "y\_3" drawn in pred (location=(3.0, 3.5))
- w\_live: a VariableNumber (initial\_value=2.0, format\_spec='.2f')

#### Beats

##### [05:45.684](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=345.68372916666664)

Narration: A network has to know when it is wrong, and by how much. Here are three training examples. The red dots are the correct answers, recorded in advance, and the blue line is what the model currently predicts.

Board: Empty.

Actions:
- [05:45.684](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=345.68372916666664): problem is shown on the screen, written out.
- [05:51.106](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=351.10572916666666): pred is shown on the screen, written out.
- [05:52.813](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=352.81272916666666): t1 is shown on the screen, written out.
- [05:52.983](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=352.98267054020096): t2 is shown on the screen, written out.
- [05:53.421](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=353.42058595058626): t3 is shown on the screen, written out.
- [05:56.864](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=356.86372916666664): model is shown on the screen, written out.

##### [06:0.355](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=360.35472916666663)

Narration: For each example, subtract the target from the prediction. That difference is the error. The first prediction is half a unit too low, the second is a whole unit too high, and the third is half a unit too low again.

Board: pred — an Axes (x\_range=(-0.3, 4.0), y\_range=(-0.3, 4.2), x\_ticks\_every=1.0); problem — a Tex \[text\] that says "How wrong is this model?"; t1 — a Point \[red\] labelled "y\_1" drawn in pred (location=(1.0, 1.5)); t2 — a Point \[red\] labelled "y\_2" drawn in pred (location=(2.0, 1.0)); t3 — a Point \[red\] labelled "y\_3" drawn in pred (location=(3.0, 3.5)); model — a FunctionPlot \[blue\] drawn in pred (function=\<function\>, x\_range=(0.0, 3.6))

Actions:
- [06:4.581](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=364.58072916666663): pred moves to a new place on the board.
- [06:4.581](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=364.58072916666663): err\_def is shown on the screen, written out.
- [06:6.288](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=366.2877291666666): e1 is shown on the screen, written out.
- [06:8.633](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=368.63272916666665): e2 is shown on the screen, written out.
- [06:10.967](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=370.96672916666665): e3 is shown on the screen, written out.

##### [06:13.935](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=373.9347291666666)

Narration: Errors come with signs, and if we simply added them a prediction that is too high would cancel one that is too low. So square each error first. Squaring makes every contribution positive, and it punishes a big miss much more than a small one.

Board: err\_def — a Math \[text\] that says "$e\_i = hat(y)\_i - y\_i$"; pred — an Axes (x\_range=(-0.3, 4.0), y\_range=(-0.3, 4.2), x\_ticks\_every=1.0); problem — a Tex \[text\] that says "How wrong is this model?"; t1 — a Point \[red\] labelled "y\_1" drawn in pred (location=(1.0, 1.5)); t2 — a Point \[red\] labelled "y\_2" drawn in pred (location=(2.0, 1.0)); t3 — a Point \[red\] labelled "y\_3" drawn in pred (location=(3.0, 3.5)); model — a FunctionPlot \[blue\] drawn in pred (function=\<function\>, x\_range=(0.0, 3.6)); e1 — a Line \[yellow\] labelled "-0.5" drawn in pred (start=(1.0, 1.0), end=(1.0, 1.5)); e2 — a Line \[yellow\] labelled "+1.0" drawn in pred (start=(2.0, 2.0), end=(2.0, 1.0)); e3 — a Line \[yellow\] labelled "-0.5" drawn in pred (start=(3.0, 3.0), end=(3.0, 3.5))

Actions:
- [06:21.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=381.5857291666666): square\_sum is shown on the screen, written out.
- [06:21.861](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=381.8607621274585): square\_sum is shown on the screen, written out.
- [06:22.393](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=382.39272377123854): square\_sum is shown on the screen, written out.

##### [06:29.965](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=389.96472916666664)

Narration: Add the three squares. A quarter, plus one, plus a quarter, is one and a half. Then divide by the number of examples. The mean squared error is zero point five.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [06:30.313](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=390.3127291666666): square\_sum is shown on the screen, drawn.
- [06:30.931](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=390.93143870578774): square\_sum is shown on the screen, drawn.
- [06:35.387](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=395.3867291666666): square\_sum is shown on the screen, written out.
- [06:36.989](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=396.98872916666664): mse\_value is shown on the screen, written out.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): err\_def is hidden from the screen — left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): mse\_value is hidden from the screen — left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): pred is hidden from the screen — left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): t1 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): t2 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): t3 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): model is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): e1 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): e2 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): e3 is hidden from the screen — pred left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): problem is hidden from the screen — left the board.
- [06:43.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.0837291666666): square\_sum is hidden from the screen — left the board.

##### [06:43.684](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.68372916666664)

Narration: In general, then: average the squared difference between prediction and target over all n examples. One number, for the whole data set, and it is the number we are going to make as small as we can.

Board: Empty.

Actions:
- [06:43.684](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=403.68372916666664): heading\_loss is shown on the screen, written out.
- [06:44.485](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=404.4847291666666): mse\_def is shown on the screen, written out.
- [06:45.739](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=405.73872916666664): mse\_formula is shown on the screen, written out.

##### [06:56.498](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=416.49772916666666)

Narration: Now here is the change of view that the rest of the lecture rests on. The data are fixed. We cannot alter a single target. The only things we can move are the weights, so the loss is really a function of the weights.

Board: mse\_def — a Panel that says "The mean squared error averages the squared difference between prediction and target over all $n$ training examples."; mse\_formula — a Math \[text\] that says "$L = frac(1, n) sum\_(i = 1)^n (hat(y)\_i - y\_i)^2$"; heading\_loss — a Heading that says "The Loss, and What It Depends On"

Actions:
- [07:1.374](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=421.37372916666664): mse\_formula (the "(hat(y)\_i - y\_i)^2" part) is emphasized.
- [07:6.053](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=426.0527291666666): mse\_formula (the "(hat(y)\_i - y\_i)^2" part) is no longer emphasized.

##### [07:9.892](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=429.89222916666665)

Narration: To see that clearly, keep one single example: input equal to one, target equal to one. With one weight and no bias the prediction is just w, so the loss is w minus one, all squared.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [07:9.892](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=429.89222916666665): bowl is shown on the screen, written out.
- [07:19.505](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=439.5047291666666): loss\_w is shown on the screen, written out.
- [07:21.247](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=441.2467291666666): bowl\_curve is shown on the screen, written out.
- [07:23.221](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=443.2207291666666): spot is shown on the screen, written out.

##### [07:24.831](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=444.83072916666663)

Narration: Sweep the weight and watch the loss. At w equals two the loss is one. Bring the weight down and the loss falls, reaches zero at w equals one, and climbs the far wall again. So training means finding the bottom of that bowl.

Board: mse\_def — a Panel that says "The mean squared error averages the squared difference between prediction and target over all $n$ training examples."; mse\_formula — a Math \[text\] that says "$L = frac(1, n) sum\_(i = 1)^n (hat(y)\_i - y\_i)^2$"; loss\_w — a Math \[text\] that says "$L(w) = (w - 1)^2$"; bowl — an Axes (x\_range=(-0.4, 2.6), y\_range=(-0.3, 2.3), x\_ticks\_every=0.5); heading\_loss — a Heading that says "The Loss, and What It Depends On"; bowl\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.2, 2.4)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='bowl\_curve', x=\<VariableNumber w\_live = 1.0\>)

Actions:
- [07:31.483](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=451.4827291666666): spot is redrawn as the numbers it depends on change.
- [07:31.483](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=451.4827291666666): w\_live ticks to 1.0.
- [07:35.744](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=455.74372916666664): spot is redrawn as the numbers it depends on change.
- [07:35.744](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=455.74372916666664): w\_live ticks to 0.0.
- [07:39.239](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=459.23872916666664): spot is redrawn as the numbers it depends on change.
- [07:39.239](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=459.23872916666664): w\_live ticks to 1.0.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): bowl is hidden from the screen — left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): bowl\_curve is hidden from the screen — bowl left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): spot is hidden from the screen — bowl left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): heading\_loss is hidden from the screen — left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): loss\_w is hidden from the screen — left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): mse\_def is hidden from the screen — left the board.
- [07:40.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=460.78841666666665): mse\_formula is hidden from the screen — left the board.

### Scene 4: [Walking Downhill](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333)

Span: 07:41.83–10:57.565 (461.83008333333333s–657.5645208333333s).

#### Objects

- bowl: an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0)
- comp\_1: a Vector \[blue\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>))
- comp\_2: a Vector \[green\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5)))))
- contours: a LevelCurves \[gray\] drawn in surf (function=\<function\>, values=(0.25, 1.0, 2.0, 3.0))
- d\_at\_two: a Math \[text\] that says "$upright("at") w = 2: quad frac(dif L, dif w) = 2$"
- d\_general: a Math \[text\] that says "$frac(dif L, dif w) = 2 (w - 1)$"
- eta\_note: a Text \[text\] that says "$eta$ is the learning rate: it sets how big a step to take."
- grad\_def: a Math \[text\] that says "$nabla L = (frac(partial L, partial w\_1), frac(partial L, partial w\_2))$"
- grad\_val: a Math \[text\] that says "$nabla L = (-4, thin 6)$"
- heading\_one: a Heading that says "One Weight, One Slope"
- heading\_two: a Heading that says "Two Weights at Once"
- here: a Point \[yellow\] labelled "w" drawn in surf (location=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>))
- loss\_2d: a Math \[text\] that says "$L = (w\_1 - 0.5)^2 + 2 (w\_2 - 0.5)^2$"
- loss\_curve: a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5))
- p1: a VariableNumber (initial\_value=-1.5, format\_spec='.2f')
- p2: a VariableNumber (initial\_value=2.0, format\_spec='.2f')
- rule: a Math \[text\] that says "$w arrow.l w - eta frac(dif L, dif w)$"
- rule\_vec: a Math \[text\] that says "$(w\_1, w\_2) arrow.l (w\_1, w\_2) - eta nabla L$"
- side\_1: a Line \[gray\] drawn in surf (start=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))), dashed=True)
- side\_2: a Line \[gray\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5)))), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))), dashed=True)
- spot: a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>)
- step\_arrow: a Vector \[red\] labelled "- eta nabla L" drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))))
- step\_big: a Math \[text\] that says "$eta = 1.1: quad w arrow.l 2 - 1.1 dot.op 2 = -0.2$"
- step\_num: a Math \[text\] that says "$(-1.5, thin 2) - 0.15 (-4, thin 6) = (-0.9, thin 1.1)$"
- step\_one: a Math \[text\] that says "$w arrow.l 2 - 0.3 dot.op 2 = 1.4$"
- surf: an Axes (x\_range=(-2.2, 2.4), y\_range=(-1.0, 2.5), aspect=(4.6, 3.5))
- tangent: a TangentLine \[yellow\] drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>, length=1.0)
- w: a VariableNumber (initial\_value=2.0, format\_spec='.2f')

#### Beats

##### [07:41.83](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333)

Narration: We have a bowl, and we need to walk to the bottom of it. But we are not allowed to see the whole bowl. At any moment the network knows only the weight it currently has, and how steeply the loss is rising or falling right there.

Board: Empty.

Actions:
- [07:41.83](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333): heading\_one is shown on the screen, written out.
- [07:41.83](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=461.83008333333333): bowl is shown on the screen, written out.
- [07:42.329](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=462.32908333333336): loss\_curve is shown on the screen, written out.
- [07:50.224](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=470.22408333333334): spot is shown on the screen, written out.

##### [07:55.631](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=475.63058333333333)

Narration: That steepness is exactly what a derivative is. The derivative of w minus one squared is two times w minus one. At w equals two it comes to two, a positive number, and the tangent line there rises to the right.

Board: bowl — an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0); heading\_one — a Heading that says "One Weight, One Slope"; loss\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>)

Actions:
- [07:57.685](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=477.68508333333335): bowl moves to a new place on the board.
- [07:57.685](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=477.68508333333335): d\_general is shown on the screen, written out.
- [08:6.021](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=486.0210833333333): d\_at\_two is shown on the screen, written out.
- [08:8.32](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=488.32008333333334): tangent is shown on the screen, written out.

##### [08:11.312](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=491.3115833333333)

Narration: A positive slope means the loss grows as the weight grows. So to make the loss smaller, move the weight the other way. Subtract something proportional to the slope, and that is the whole of gradient descent in one line.

Board: d\_general — a Math \[text\] that says "$frac(dif L, dif w) = 2 (w - 1)$"; d\_at\_two — a Math \[text\] that says "$upright("at") w = 2: quad frac(dif L, dif w) = 2$"; bowl — an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0); heading\_one — a Heading that says "One Weight, One Slope"; loss\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>); tangent — a TangentLine \[yellow\] drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>, length=1.0)

Actions:
- [08:19.821](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=499.8210833333333): rule is shown on the screen, written out.

##### [08:26.192](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=506.1915833333333)

Narration: New weight equals old weight, minus eta times the derivative. Eta is a small positive number called the learning rate, and it decides how far each step carries us.

Board: d\_general — a Math \[text\] that says "$frac(dif L, dif w) = 2 (w - 1)$"; d\_at\_two — a Math \[text\] that says "$upright("at") w = 2: quad frac(dif L, dif w) = 2$"; rule — a Math \[text\] that says "$w arrow.l w - eta frac(dif L, dif w)$"; bowl — an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0); heading\_one — a Heading that says "One Weight, One Slope"; loss\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>); tangent — a TangentLine \[yellow\] drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>, length=1.0)

Actions:
- [08:28.885](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=508.88508333333334): rule (the "eta" part) is emphasized.
- [08:33.692](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=513.6920833333334): eta\_note is shown on the screen, written out.
- [08:35.143](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=515.1430833333334): rule (the "eta" part) is no longer emphasized.

##### [08:38.541](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=518.5410833333333)

Narration: Take eta equal to zero point three. The step is zero point three times two, which is zero point six, so the weight goes from two to one point four. Watch the tangent flatten as it lands.

Board: d\_general — a Math \[text\] that says "$frac(dif L, dif w) = 2 (w - 1)$"; d\_at\_two — a Math \[text\] that says "$upright("at") w = 2: quad frac(dif L, dif w) = 2$"; rule — a Math \[text\] that says "$w arrow.l w - eta frac(dif L, dif w)$"; eta\_note — a Text \[text\] that says "$eta$ is the learning rate: it sets how big a step to take."; bowl — an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0); heading\_one — a Heading that says "One Weight, One Slope"; loss\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>); tangent — a TangentLine \[yellow\] drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>, length=1.0)

Actions:
- [08:41.873](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=521.8730833333333): step\_one is shown on the screen, written out.
- [08:49.222](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=529.2220833333333): spot is redrawn as the numbers it depends on change.
- [08:49.222](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=529.2220833333333): tangent is redrawn as the numbers it depends on change.
- [08:49.222](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=529.2220833333333): w ticks to 1.4.

##### [08:52.226](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=532.2255833333334)

Narration: Do it again. At one point four the slope is only zero point eight, so the step is smaller and the weight moves to one point one six. Once more, and it reaches one point zero six. The steps shrink by themselves, because near the bottom there is hardly any slope left to multiply.

Board: d\_general — a Math \[text\] that says "$frac(dif L, dif w) = 2 (w - 1)$"; d\_at\_two — a Math \[text\] that says "$upright("at") w = 2: quad frac(dif L, dif w) = 2$"; rule — a Math \[text\] that says "$w arrow.l w - eta frac(dif L, dif w)$"; eta\_note — a Text \[text\] that says "$eta$ is the learning rate: it sets how big a step to take."; step\_one — a Math \[text\] that says "$w arrow.l 2 - 0.3 dot.op 2 = 1.4$"; bowl — an Axes (x\_range=(-0.5, 2.8), y\_range=(-0.3, 2.8), x\_ticks\_every=1.0); heading\_one — a Heading that says "One Weight, One Slope"; loss\_curve — a FunctionPlot \[blue\] drawn in bowl (function=\<function\>, x\_range=(-0.3, 2.5)); spot — a PlotPoint \[yellow\] labelled "w = 2.00" drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>); tangent — a TangentLine \[yellow\] drawn in bowl (target='loss\_curve', x=\<VariableNumber w = -0.2\>, length=1.0)

Actions:
- [08:58.483](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=538.4830833333333): spot is redrawn as the numbers it depends on change.
- [08:58.483](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=538.4830833333333): tangent is redrawn as the numbers it depends on change.
- [08:58.483](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=538.4830833333333): w ticks to 1.16.
- [09:2.095](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=542.0950833333334): spot is redrawn as the numbers it depends on change.
- [09:2.095](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=542.0950833333334): tangent is redrawn as the numbers it depends on change.
- [09:2.095](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=542.0950833333334): w ticks to 1.06.

##### [09:11.031](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=551.0305833333333)

Narration: Had the slope been negative, the minus sign would have pushed the weight up instead. Either way we go downhill. But eta matters. Suppose at w equals two we had used one point one. The step would be two point two, the weight would land at minus zero point two, and the loss there is worse than where we began.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [09:19.784](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=559.7840833333333): spot is redrawn as the numbers it depends on change.
- [09:19.784](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=559.7840833333333): tangent is redrawn as the numbers it depends on change.
- [09:19.784](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=559.7840833333333): w ticks to 2.0.
- [09:23.941](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=563.9410833333334): step\_big is shown on the screen, written out.
- [09:26.147](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=566.1470833333333): spot is redrawn as the numbers it depends on change.
- [09:26.147](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=566.1470833333333): tangent is redrawn as the numbers it depends on change.
- [09:26.147](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=566.1470833333333): w ticks to -0.2.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): bowl is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): loss\_curve is hidden from the screen — bowl left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): spot is hidden from the screen — bowl left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): tangent is hidden from the screen — bowl left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): d\_at\_two is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): d\_general is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): eta\_note is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): heading\_one is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): rule is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): step\_big is hidden from the screen — left the board.
- [09:30.884](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=570.8835833333334): step\_one is hidden from the screen — left the board.

##### [09:32.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=572.0835833333333)

Narration: A real network has thousands of weights, not one, and nothing changes except that there is now a slope for each of them. Here is a loss with two weights, drawn as a contour map. Every ring is a set of weights giving the same loss, and the bottom of the bowl is inside the smallest ring.

Board: Empty.

Actions:
- [09:32.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=572.0835833333333): heading\_two is shown on the screen, written out.
- [09:32.084](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=572.0835833333333): surf is shown on the screen, written out.
- [09:40.118](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=580.1180833333333): surf moves to a new place on the board.
- [09:40.118](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=580.1180833333333): loss\_2d is shown on the screen, written out.
- [09:40.791](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=580.7910833333333): here is shown on the screen, written out.
- [09:41.871](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=581.8710833333333): contours is shown on the screen, written out.

##### [09:50.006](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=590.0055833333333)

Narration: Hold w two still and ask how the loss changes as w one moves on its own. That is the partial derivative with respect to w one, and taking a step against it gives the blue arrow. Hold w one still instead, and you get the green arrow.

Board: loss\_2d — a Math \[text\] that says "$L = (w\_1 - 0.5)^2 + 2 (w\_2 - 0.5)^2$"; surf — an Axes (x\_range=(-2.2, 2.4), y\_range=(-1.0, 2.5), aspect=(4.6, 3.5)); heading\_two — a Heading that says "Two Weights at Once"; contours — a LevelCurves \[gray\] drawn in surf (function=\<function\>, values=(0.25, 1.0, 2.0, 3.0)); here — a Point \[yellow\] labelled "w" drawn in surf (location=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>))

Actions:
- [09:56.728](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=596.7280833333333): grad\_def is shown on the screen, written out.
- [10:1.256](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=601.2560833333333): comp\_1 is shown on the screen, written out.
- [10:5.644](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=605.6440833333334): comp\_2 is shown on the screen, written out.

##### [10:7.371](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=607.3705833333333)

Narration: Collect the partial derivatives into one list and you have the gradient. Here it is minus four and six. The gradient points in the direction the loss increases fastest, straight across the contour lines.

Board: loss\_2d — a Math \[text\] that says "$L = (w\_1 - 0.5)^2 + 2 (w\_2 - 0.5)^2$"; grad\_def — a Math \[text\] that says "$nabla L = (frac(partial L, partial w\_1), frac(partial L, partial w\_2))$"; surf — an Axes (x\_range=(-2.2, 2.4), y\_range=(-1.0, 2.5), aspect=(4.6, 3.5)); heading\_two — a Heading that says "Two Weights at Once"; contours — a LevelCurves \[gray\] drawn in surf (function=\<function\>, values=(0.25, 1.0, 2.0, 3.0)); here — a Point \[yellow\] labelled "w" drawn in surf (location=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>)); comp\_1 — a Vector \[blue\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>)); comp\_2 — a Vector \[green\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5)))))

Actions:
- [10:12.096](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=612.0960833333334): grad\_val is shown on the screen, written out.
- [10:18.62](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=618.6200833333334): contours is indicated — a transient flash.

##### [10:21.043](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=621.0430833333334)

Narration: So step the opposite way. Add the two component steps, and the red arrow is the diagonal of their parallelogram: that is where we land. The rule reads as before, with the gradient standing in for the single derivative.

Board: loss\_2d — a Math \[text\] that says "$L = (w\_1 - 0.5)^2 + 2 (w\_2 - 0.5)^2$"; grad\_def — a Math \[text\] that says "$nabla L = (frac(partial L, partial w\_1), frac(partial L, partial w\_2))$"; grad\_val — a Math \[text\] that says "$nabla L = (-4, thin 6)$"; surf — an Axes (x\_range=(-2.2, 2.4), y\_range=(-1.0, 2.5), aspect=(4.6, 3.5)); heading\_two — a Heading that says "Two Weights at Once"; contours — a LevelCurves \[gray\] drawn in surf (function=\<function\>, values=(0.25, 1.0, 2.0, 3.0)); here — a Point \[yellow\] labelled "w" drawn in surf (location=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>)); comp\_1 — a Vector \[blue\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>)); comp\_2 — a Vector \[green\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5)))))

Actions:
- [10:23.551](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=623.5510833333334): side\_1 is shown on the screen, written out.
- [10:23.801](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=623.8012408880836): side\_2 is shown on the screen, written out.
- [10:25.71](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=625.7100833333334): step\_arrow is shown on the screen, written out.
- [10:29.878](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=629.8780833333334): step\_num is shown on the screen, written out.
- [10:31.353](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=631.3530833333334): rule\_vec is shown on the screen, written out.

##### [10:36.284](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.2835833333334)

Narration: And repeat. Each step crosses to an inner ring, the arrows shorten as the ground flattens out, and the weights settle near the bottom. That is gradient descent. What is left is the hard part. In a real network, with layers feeding layers, how do we actually compute those partial derivatives?

Board: loss\_2d — a Math \[text\] that says "$L = (w\_1 - 0.5)^2 + 2 (w\_2 - 0.5)^2$"; grad\_def — a Math \[text\] that says "$nabla L = (frac(partial L, partial w\_1), frac(partial L, partial w\_2))$"; grad\_val — a Math \[text\] that says "$nabla L = (-4, thin 6)$"; rule\_vec — a Math \[text\] that says "$(w\_1, w\_2) arrow.l (w\_1, w\_2) - eta nabla L$"; step\_num — a Math \[text\] that says "$(-1.5, thin 2) - 0.15 (-4, thin 6) = (-0.9, thin 1.1)$"; surf — an Axes (x\_range=(-2.2, 2.4), y\_range=(-1.0, 2.5), aspect=(4.6, 3.5)); heading\_two — a Heading that says "Two Weights at Once"; contours — a LevelCurves \[gray\] drawn in surf (function=\<function\>, values=(0.25, 1.0, 2.0, 3.0)); here — a Point \[yellow\] labelled "w" drawn in surf (location=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>)); comp\_1 — a Vector \[blue\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>)); comp\_2 — a Vector \[green\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5))))); side\_1 — a Line \[gray\] drawn in surf (start=((p1 - (0.3 \* (p1 - 0.5))), \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))), dashed=True); side\_2 — a Line \[gray\] drawn in surf (start=(\<VariableNumber p1 = -0.19\>, (p2 - (0.6 \* (p2 - 0.5)))), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))), dashed=True); step\_arrow — a Vector \[red\] labelled "- eta nabla L" drawn in surf (start=(\<VariableNumber p1 = -0.19\>, \<VariableNumber p2 = 0.6\>), end=((p1 - (0.3 \* (p1 - 0.5))), (p2 - (0.6 \* (p2 - 0.5)))))

Actions:
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): here is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): comp\_1 is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): comp\_2 is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): side\_1 is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): side\_2 is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): step\_arrow is redrawn as the numbers it depends on change.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): p1 ticks to -0.9.
- [10:36.887](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=636.8870833333334): p2 ticks to 1.1.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): here is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): comp\_1 is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): comp\_2 is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): side\_1 is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): side\_2 is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): step\_arrow is redrawn as the numbers it depends on change.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): p1 ticks to -0.48.
- [10:40.881](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=640.8810833333334): p2 ticks to 0.74.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): here is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): comp\_1 is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): comp\_2 is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): side\_1 is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): side\_2 is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): step\_arrow is redrawn as the numbers it depends on change.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): p1 ticks to -0.19.
- [10:43.586](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=643.5860833333334): p2 ticks to 0.6.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): grad\_def is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): grad\_val is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): heading\_two is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): loss\_2d is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): rule\_vec is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): step\_num is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): surf is hidden from the screen — left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): contours is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): here is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): comp\_1 is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): comp\_2 is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): side\_1 is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): side\_2 is hidden from the screen — surf left the board.
- [10:56.523](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=656.5228541666667): step\_arrow is hidden from the screen — surf left the board.

### Scene 5: [Backpropagation: The Chain Rule Doing the Work](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333)

Span: 10:57.565–14:30.512 (657.5645208333333s–870.5116874999999s).

#### Objects

- b\_1: an Arrow \[red\] labelled "frac(partial z, partial w)" drawn in chain (start=(2.75, 1.0), end=(0.95, 1.0))
- b\_2: an Arrow \[red\] labelled "frac(dif a, dif z)" drawn in chain (start=(4.95, 1.0), end=(3.25, 1.0))
- b\_3: an Arrow \[red\] labelled "frac(partial L, partial a)" drawn in chain (start=(7.05, 1.0), end=(5.45, 1.0))
- chain: a Figure (x\_range=(0.0, 8.0), y\_range=(0.2, 3.0), aspect=(8.0, 2.8))
- chain\_rule: a Math \[text\] that says "$frac(partial L, partial w) = frac(partial L, partial a) dot.op frac(dif a, dif z) dot.op frac(partial z, partial w)$"
- f\_1: an Arrow \[gray\] labelled "w x + b" drawn in chain (start=(0.95, 2.2), end=(2.75, 2.2))
- f\_2: an Arrow \[gray\] labelled "sigma" drawn in chain (start=(3.25, 2.2), end=(4.95, 2.2))
- f\_3: an Arrow \[gray\] labelled "(a - y)^2" drawn in chain (start=(5.45, 2.2), end=(7.05, 2.2))
- gap\_seg: a Line \[red\] labelled "a - y" drawn in sig (start=(\<VariableNumber z\_live = 0.589\>, (1.0 / (1.0 + exp((-1.0 \* z\_l…, end=(\<VariableNumber z\_live = 0.589\>, 1.0))
- heading\_chain: a Heading that says "From the Weight to the Loss"
- heading\_loop: a Heading that says "The Whole Loop"
- heading\_numbers: a Heading that says "One Step, in Numbers"
- heading\_pieces: a Heading that says "Three Local Derivatives"
- loss\_change: a Math \[text\] that says "$L: 0.143 arrow.r 0.127$"
- n\_a: a Point \[green\] labelled "a" drawn in chain (location=(5.2, 2.2))
- n\_l: a Point \[red\] labelled "L" drawn in chain (location=(7.3, 2.2))
- n\_x: a Point \[blue\] labelled "x" drawn in chain (location=(0.7, 2.2))
- n\_z: a Point \[yellow\] labelled "z" drawn in chain (location=(3.0, 2.2))
- num\_0: a Math \[text\] that says "$x = 1, quad y = 1, quad w = 0.5, quad b = 0$"
- num\_1: a Math \[text\] that says "$z = 0.5 dot.op 1 + 0 = 0.5$"
- num\_2: a Math \[text\] that says "$a = sigma(0.5) = 0.622$"
- num\_3: a Math \[text\] that says "$frac(partial L, partial a) = 2 (0.622 - 1) = -0.755$"
- num\_4: a Math \[text\] that says "$frac(dif a, dif z) = 0.622 (1 - 0.622) = 0.235$"
- num\_5: a Math \[text\] that says "$frac(partial L, partial w) = -0.755 dot.op 0.235 dot.op 1 = -0.177$"
- num\_6: a Math \[text\] that says "$w arrow.l 0.5 - 0.5 dot.op (-0.177) = 0.589$"
- pieces: a Derivation \[text\] that says "$L &= (a - y)^2 \\ frac(partial L, partial a) &= 2 (a - y) \\ a &= sigma(z) \\ frac(dif a, dif z) &= a (1 - a) \\ z &= w x + b, quad frac(partial z, partial w) = x$"
- recap: a Block \[text\] that says "Forward: weighted sums and bends turn inputs into a prediction. Loss: one number, $L$, says how wrong that prediction is. Backward: the chain rule turns $L$ into a slope for every weight. Step: each weight moves against its own slope, and …"
- result\_b: a Math \[text\] that says "$frac(partial L, partial b) = 2 (a - y) dot.op a (1 - a)$"
- result\_w: a Math \[text\] that says "$frac(partial L, partial w) = 2 (a - y) dot.op a (1 - a) dot.op x$"
- riding: a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>)
- sig: an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0)
- sig\_curve: a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0))
- slope\_here: a TangentLine \[yellow\] drawn in sig (target='sig\_curve', x=0.5, show\_dot=False)
- target: a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True)
- z\_live: a VariableNumber (initial\_value=0.5, format\_spec='.3f')

#### Beats

##### [10:57.565](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333)

Narration: We need the slope of the loss with respect to one weight. But the weight does not touch the loss directly. Follow the arithmetic forward: the weight and the input make z. The bend turns z into the activation a. And a, compared with the target y, produces the loss.

Board: Empty.

Actions:
- [10:57.565](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333): heading\_chain is shown on the screen, written out.
- [10:57.565](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=657.5645208333333): chain is shown on the screen, written out.
- [11:4.357](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=664.3565208333333): n\_x is shown on the screen, written out.
- [11:7.724](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=667.7235208333333): f\_1 is shown on the screen, written out.
- [11:8.153](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=668.1525208333333): n\_z is shown on the screen, written out.
- [11:9.732](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=669.7315208333333): f\_2 is shown on the screen, written out.
- [11:11.497](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=671.4965208333333): n\_a is shown on the screen, written out.
- [11:15.12](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=675.1195208333332): f\_3 is shown on the screen, written out.
- [11:16.327](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=676.3265208333332): n\_l is shown on the screen, written out.

##### [11:18.448](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=678.4475208333333)

Narration: So nudging the weight nudges z, and that nudges a, and that nudges the loss. Three links in one chain.

Board: chain — a Figure (x\_range=(0.0, 8.0), y\_range=(0.2, 3.0), aspect=(8.0, 2.8)); heading\_chain — a Heading that says "From the Weight to the Loss"; n\_x — a Point \[blue\] labelled "x" drawn in chain (location=(0.7, 2.2)); f\_1 — an Arrow \[gray\] labelled "w x + b" drawn in chain (start=(0.95, 2.2), end=(2.75, 2.2)); n\_z — a Point \[yellow\] labelled "z" drawn in chain (location=(3.0, 2.2)); f\_2 — an Arrow \[gray\] labelled "sigma" drawn in chain (start=(3.25, 2.2), end=(4.95, 2.2)); n\_a — a Point \[green\] labelled "a" drawn in chain (location=(5.2, 2.2)); f\_3 — an Arrow \[gray\] labelled "(a - y)^2" drawn in chain (start=(5.45, 2.2), end=(7.05, 2.2)); n\_l — a Point \[red\] labelled "L" drawn in chain (location=(7.3, 2.2))

Actions:
- [11:19.76](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=679.7595208333333): f\_1 is indicated — a transient flash.
- [11:20.052](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=680.0517841348993): f\_2 is indicated — a transient flash.
- [11:21.059](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=681.0587362540507): f\_3 is indicated — a transient flash.

##### [11:26.553](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=686.5530208333333)

Narration: Read the chain backwards, which is where the name backpropagation comes from. First, how fast does the loss respond to a? Next, how fast does a respond to z? Finally, how fast does z respond to the weight? Multiply those three local rates together and you have the slope you wanted.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [11:31.622](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=691.6215208333333): b\_3 is shown on the screen, written out.
- [11:35.546](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=695.5455208333333): b\_2 is shown on the screen, written out.
- [11:38.391](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=698.3905208333333): b\_1 is shown on the screen, written out.
- [11:42.512](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=702.5115208333333): chain\_rule is shown on the screen, written out.
- [11:46.842](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=706.8420208333333): chain moves to a new place on the board.
- [11:46.842](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=706.8420208333333): chain\_rule is hidden from the screen — left the board.
- [11:46.842](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=706.8420208333333): heading\_chain is hidden from the screen — left the board.

##### [11:48.042](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=708.0420208333333)

Narration: Each of those three rates is a one-line derivative. The loss is a minus y, all squared, so differentiating it with respect to a gives two times a minus y.

Board: chain — a Figure (x\_range=(0.0, 8.0), y\_range=(0.2, 3.0), aspect=(8.0, 2.8)); n\_x — a Point \[blue\] labelled "x" drawn in chain (location=(0.7, 2.2)); f\_1 — an Arrow \[gray\] labelled "w x + b" drawn in chain (start=(0.95, 2.2), end=(2.75, 2.2)); n\_z — a Point \[yellow\] labelled "z" drawn in chain (location=(3.0, 2.2)); f\_2 — an Arrow \[gray\] labelled "sigma" drawn in chain (start=(3.25, 2.2), end=(4.95, 2.2)); n\_a — a Point \[green\] labelled "a" drawn in chain (location=(5.2, 2.2)); f\_3 — an Arrow \[gray\] labelled "(a - y)^2" drawn in chain (start=(5.45, 2.2), end=(7.05, 2.2)); n\_l — a Point \[red\] labelled "L" drawn in chain (location=(7.3, 2.2)); b\_3 — an Arrow \[red\] labelled "frac(partial L, partial a)" drawn in chain (start=(7.05, 1.0), end=(5.45, 1.0)); b\_2 — an Arrow \[red\] labelled "frac(dif a, dif z)" drawn in chain (start=(4.95, 1.0), end=(3.25, 1.0)); b\_1 — an Arrow \[red\] labelled "frac(partial z, partial w)" drawn in chain (start=(2.75, 1.0), end=(0.95, 1.0))

Actions:
- [11:48.042](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=708.0420208333333): heading\_pieces is shown on the screen, written out.
- [11:51.77](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=711.7695208333333): pieces is shown on the screen, written out.
- [11:56.611](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=716.6105208333333): pieces is shown on the screen, written out.
- [11:56.611](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=716.6105208333333): b\_3 is indicated — a transient flash.

##### [11:59.196](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=719.1960208333333)

Narration: The activation is sigma of z, and the sigmoid's slope is the tidy expression we met earlier: a times one minus a.

Board: chain — a Figure (x\_range=(0.0, 8.0), y\_range=(0.2, 3.0), aspect=(8.0, 2.8)); n\_x — a Point \[blue\] labelled "x" drawn in chain (location=(0.7, 2.2)); f\_1 — an Arrow \[gray\] labelled "w x + b" drawn in chain (start=(0.95, 2.2), end=(2.75, 2.2)); n\_z — a Point \[yellow\] labelled "z" drawn in chain (location=(3.0, 2.2)); f\_2 — an Arrow \[gray\] labelled "sigma" drawn in chain (start=(3.25, 2.2), end=(4.95, 2.2)); n\_a — a Point \[green\] labelled "a" drawn in chain (location=(5.2, 2.2)); f\_3 — an Arrow \[gray\] labelled "(a - y)^2" drawn in chain (start=(5.45, 2.2), end=(7.05, 2.2)); n\_l — a Point \[red\] labelled "L" drawn in chain (location=(7.3, 2.2)); b\_3 — an Arrow \[red\] labelled "frac(partial L, partial a)" drawn in chain (start=(7.05, 1.0), end=(5.45, 1.0)); b\_2 — an Arrow \[red\] labelled "frac(dif a, dif z)" drawn in chain (start=(4.95, 1.0), end=(3.25, 1.0)); b\_1 — an Arrow \[red\] labelled "frac(partial z, partial w)" drawn in chain (start=(2.75, 1.0), end=(0.95, 1.0)); heading\_pieces — a Heading that says "Three Local Derivatives"

Actions:
- [11:59.731](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=719.7305208333333): pieces is shown on the screen, written out.
- [12:3.515](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=723.5145208333333): pieces is shown on the screen, written out.
- [12:3.515](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=723.5145208333333): b\_2 is indicated — a transient flash.

##### [12:8.527](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=728.5265208333333)

Narration: And z is w x plus b, so differentiating with respect to w leaves just x. With respect to b it leaves one, because the bias is added on its own.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [12:10.547](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=730.5465208333333): pieces is shown on the screen, written out.
- [12:13.902](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=733.9015208333333): b\_1 is indicated — a transient flash.

##### [12:20.47](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=740.4695208333333)

Narration: Multiply the three together and there it is. The slope of the loss with respect to this weight is two times a minus y, times a times one minus a, times x. For the bias it is the same product without the x on the end.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [12:20.818](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=740.8175208333333): result\_w is shown on the screen, written out.
- [12:32.347](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=752.3465208333333): result\_b is shown on the screen, written out.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): chain is hidden from the screen — left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): n\_x is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): f\_1 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): n\_z is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): f\_2 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): n\_a is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): f\_3 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): n\_l is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): b\_3 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): b\_2 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): b\_1 is hidden from the screen — chain left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): heading\_pieces is hidden from the screen — left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): pieces is hidden from the screen — left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): result\_b is hidden from the screen — left the board.
- [12:35.725](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=755.7250208333332): result\_w is hidden from the screen — left the board.

##### [12:36.925](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=756.9250208333333)

Narration: Now put numbers in. One input equal to one, a target of one, a weight of a half, and a bias of zero. The weighted sum is a half. The sigmoid of a half is zero point six two two, so the network answers zero point six two two when it should answer one.

Board: Empty.

Actions:
- [12:36.925](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=756.9250208333333): heading\_numbers is shown on the screen, written out.
- [12:36.925](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=756.9250208333333): sig is shown on the screen, written out.
- [12:37.727](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=757.7265208333333): sig\_curve is shown on the screen, written out.
- [12:37.727](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=757.7265208333333): num\_0 is shown on the screen, written out.
- [12:40.594](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=760.5935208333333): target is shown on the screen, written out.
- [12:45.39](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=765.3895208333333): num\_1 is shown on the screen, written out.
- [12:47.34](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=767.3395208333333): riding is shown on the screen, written out.
- [12:47.34](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=767.3395208333333): num\_2 is shown on the screen, written out.

##### [12:54.941](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=774.9405208333333)

Narration: That red gap is the error. a minus y is minus zero point three seven eight, and twice it is minus zero point seven five five. There is the first factor.

Board: num\_0 — a Math \[text\] that says "$x = 1, quad y = 1, quad w = 0.5, quad b = 0$"; num\_1 — a Math \[text\] that says "$z = 0.5 dot.op 1 + 0 = 0.5$"; num\_2 — a Math \[text\] that says "$a = sigma(0.5) = 0.622$"; sig — an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0); heading\_numbers — a Heading that says "One Step, in Numbers"; sig\_curve — a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0)); target — a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True); riding — a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>)

Actions:
- [12:55.788](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=775.7875208333332): gap\_seg is shown on the screen, written out.
- [13:1.384](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=781.3835208333333): num\_3 is shown on the screen, written out.

##### [13:6.854](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=786.8540208333333)

Narration: The second factor is the steepness of the sigmoid at this point, a times one minus a, which works out as zero point two three five. You can see it in the tangent line: a gentle slope, so a gentle factor.

Board: num\_0 — a Math \[text\] that says "$x = 1, quad y = 1, quad w = 0.5, quad b = 0$"; num\_1 — a Math \[text\] that says "$z = 0.5 dot.op 1 + 0 = 0.5$"; num\_2 — a Math \[text\] that says "$a = sigma(0.5) = 0.622$"; num\_3 — a Math \[text\] that says "$frac(partial L, partial a) = 2 (0.622 - 1) = -0.755$"; sig — an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0); heading\_numbers — a Heading that says "One Step, in Numbers"; sig\_curve — a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0)); target — a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True); riding — a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>); gap\_seg — a Line \[red\] labelled "a - y" drawn in sig (start=(\<VariableNumber z\_live = 0.589\>, (1.0 / (1.0 + exp((-1.0 \* z\_l…, end=(\<VariableNumber z\_live = 0.589\>, 1.0))

Actions:
- [13:12.7](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=792.6995208333333): num\_4 is shown on the screen, written out.
- [13:16.799](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=796.7985208333333): slope\_here is shown on the screen, written out.
- [13:20.85](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=800.8500208333332): num\_4 moves to a new place on the board.
- [13:20.85](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=800.8500208333332): num\_0 is hidden from the screen — left the board.
- [13:20.85](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=800.8500208333332): num\_1 is hidden from the screen — left the board.
- [13:20.85](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=800.8500208333332): num\_2 is hidden from the screen — left the board.
- [13:20.85](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=800.8500208333332): num\_3 is hidden from the screen — left the board.

##### [13:21.45](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=801.4500208333333)

Narration: Multiply them, together with x, which is one. The slope of the loss with respect to this weight is minus zero point one seven seven.

Board: num\_4 — a Math \[text\] that says "$frac(dif a, dif z) = 0.622 (1 - 0.622) = 0.235$"; sig — an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0); heading\_numbers — a Heading that says "One Step, in Numbers"; sig\_curve — a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0)); target — a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True); riding — a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>); gap\_seg — a Line \[red\] labelled "a - y" drawn in sig (start=(\<VariableNumber z\_live = 0.589\>, (1.0 / (1.0 + exp((-1.0 \* z\_l…, end=(\<VariableNumber z\_live = 0.589\>, 1.0)); slope\_here — a TangentLine \[yellow\] drawn in sig (target='sig\_curve', x=0.5, show\_dot=False)

Actions:
- [13:25.549](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=805.5485208333333): num\_5 is shown on the screen, written out.

##### [13:32.093](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=812.0925208333333)

Narration: It is negative, so raising the weight lowers the loss. With a learning rate of a half the weight moves from zero point five to zero point five eight nine, and the activation climbs the curve toward its target.

Board: num\_4 — a Math \[text\] that says "$frac(dif a, dif z) = 0.622 (1 - 0.622) = 0.235$"; sig — an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0); heading\_numbers — a Heading that says "One Step, in Numbers"; sig\_curve — a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0)); target — a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True); riding — a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>); gap\_seg — a Line \[red\] labelled "a - y" drawn in sig (start=(\<VariableNumber z\_live = 0.589\>, (1.0 / (1.0 + exp((-1.0 \* z\_l…, end=(\<VariableNumber z\_live = 0.589\>, 1.0)); slope\_here — a TangentLine \[yellow\] drawn in sig (target='sig\_curve', x=0.5, show\_dot=False); num\_5 — a Math \[text\] that says "$frac(partial L, partial w) = -0.755 dot.op 0.235 dot.op 1 = -0.177$"

Actions:
- [13:38.258](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=818.2575208333333): num\_6 is shown on the screen, written out.
- [13:43.088](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=823.0875208333333): riding is redrawn as the numbers it depends on change.
- [13:43.088](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=823.0875208333333): gap\_seg is redrawn as the numbers it depends on change.
- [13:43.088](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=823.0875208333333): z\_live ticks to 0.589.

##### [13:45.859](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=825.8585208333333)

Narration: And the loss falls, from zero point one four three to zero point one two seven. The gap is narrower than it was. That is one training step for one weight, and a network does it for every weight it has, over and over again.

Board: num\_4 — a Math \[text\] that says "$frac(dif a, dif z) = 0.622 (1 - 0.622) = 0.235$"; sig — an Axes (x\_range=(-3.2, 3.2), y\_range=(-0.1, 1.2), x\_ticks\_every=1.0); heading\_numbers — a Heading that says "One Step, in Numbers"; sig\_curve — a FunctionPlot \[blue\] drawn in sig (function=\<function\>, x\_range=(-3.0, 3.0)); target — a Line \[red\] labelled "y = 1" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True); riding — a PlotPoint \[yellow\] labelled "a" drawn in sig (target='sig\_curve', x=\<VariableNumber z\_live = 0.589\>); gap\_seg — a Line \[red\] labelled "a - y" drawn in sig (start=(\<VariableNumber z\_live = 0.589\>, (1.0 / (1.0 + exp((-1.0 \* z\_l…, end=(\<VariableNumber z\_live = 0.589\>, 1.0)); slope\_here — a TangentLine \[yellow\] drawn in sig (target='sig\_curve', x=0.5, show\_dot=False); num\_5 — a Math \[text\] that says "$frac(partial L, partial w) = -0.755 dot.op 0.235 dot.op 1 = -0.177$"; num\_6 — a Math \[text\] that says "$w arrow.l 0.5 - 0.5 dot.op (-0.177) = 0.589$"

Actions:
- [13:46.904](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=826.9035208333333): loss\_change is shown on the screen, written out.
- [13:53.127](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=833.1265208333333): gap\_seg is indicated — a transient flash.
- [13:56.064](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=836.0635208333333): A box is drawn around loss\_change.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): heading\_numbers is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): loss\_change is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): num\_4 is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): num\_5 is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): num\_6 is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): sig is hidden from the screen — left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): sig\_curve is hidden from the screen — sig left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): target is hidden from the screen — sig left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): riding is hidden from the screen — sig left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): gap\_seg is hidden from the screen — sig left the board.
- [14:1.358](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=841.3580208333333): slope\_here is hidden from the screen — sig left the board.

##### [14:2.558](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=842.5580208333333)

Narration: So that is the whole loop. Forward, weighted sums and bends turn the inputs into a prediction.

Board: Empty.

Actions:
- [14:2.558](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=842.5580208333333): heading\_loop is shown on the screen, written out.
- [14:3.836](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=843.8355208333334): recap is shown on the screen, written out.
- [14:4.927](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=844.9265208333334): recap (the "weighted sums and bends" part) is emphasized.

##### [14:10.136](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=850.1355208333333)

Narration: The loss collapses all of that into one number. The chain rule turns that number into a slope for every weight and every bias in the network.

Board: recap — a Block \[text\] that says "Forward: weighted sums and bends turn inputs into a prediction. Loss: one number, $L$, says how wrong that prediction is. Backward: the chain rule turns $L$ into a slope for every weight. Step: each weight moves against its own slope, and …"; heading\_loop — a Heading that says "The Whole Loop"

Actions:
- [14:11.204](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=851.2035208333333): recap (the "one number" part) is emphasized.
- [14:11.204](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=851.2035208333333): recap (the "weighted sums and bends" part) is no longer emphasized.
- [14:14.13](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=854.1295208333333): recap (the "one number" part) is no longer emphasized.
- [14:14.13](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=854.1295208333333): recap (the "the chain rule" part) is emphasized.

##### [14:20.163](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=860.1625208333332)

Narration: Then each weight takes a small step against its own slope, and we go round again. Everything else in deep learning is this loop, done at scale.

Board: Unchanged from the preceding beat in this scene.

Actions:
- [14:22.137](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=862.1365208333333): recap (the "against its own slope" part) is emphasized.
- [14:22.137](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=862.1365208333333): recap (the "the chain rule" part) is no longer emphasized.
- [14:29.22](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=869.2200208333333): recap (the "against its own slope" part) is no longer emphasized.
- [14:29.47](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=869.4700208333333): heading\_loop is hidden from the screen — left the board.
- [14:29.47](https://academa.ai/@engineer/lectures/the-mathematics-of-neural-networks-and-gradient-descent?t=869.4700208333333): recap is hidden from the screen — left the board.
