{"version":1,"lectureId":"01M3WDT9XXSAEX76WGMC8M29GQ","attempt":0,"publication":{"slug":"the-mathematics-of-neural-networks-and-gradient-descent","title":"The Mathematics of Neural Networks and Gradient Descent","subject":"machine-learning","summary":"A first course in the mathematics behind neural networks. We build one unit out of inputs, weights and a bias, show why a purely linear network collapses into a single straight line, and introduce the sigmoid and ReLU activations that bend it. We then define mean squared error, watch the loss become a function of the weights, and walk downhill: derivatives as slopes, the learning rate, partial derivatives and the gradient. The final section assembles the chain rule into backpropagation and carries one complete training step through in numbers, from the forward pass to the updated weight and the loss that fell.","metaDescription":"How neural networks compute, how loss measures error, and how derivatives, the gradient and the chain rule update every weight.","transcript":"A neural network is a piece of arithmetic with adjustable numbers inside it. Suppose two measurements describe a student: hours of study, and hours of sleep. We want one number out, a predicted exam score. So there are two questions. What arithmetic should we do, and how can the numbers inside it fix themselves when the answer comes out wrong? Draw that arithmetic as a picture. The two inputs sit on the left, and I will call them x one and x two. For our student, x one is three hours of study and x two is eight hours of sleep. Over here is a single unit, and its job is to turn those two numbers into one. Every arrow carries a number of its own, called a weight, and the weight decides how much that input matters. The first weight is zero point five. The second is zero point two. The unit does two things with those numbers. First it adds up the weighted inputs. Then it adds one more number of its own, which belongs to the unit rather than to any input. That number is called the bias, and here it is minus one. It lets the unit shift its answer up or down whatever the inputs happen to be. Written out, that is the whole of one unit. z equals w one x one, plus w two x two, plus b. A weighted sum of the inputs, and then the bias. Put our numbers in: zero point five times three, plus zero point two times eight, minus one. One point five, plus one point six, minus one. The unit's output is two point one. Now clear the arithmetic and watch the weight itself. Raise the first weight to one, and the output climbs to three point six, because study now counts for twice as much. Lower it to zero point one, and the output drops to zero point nine. Those weights are the dials, and everything that follows is about how to set them. One unit is not a network. Start again from the two inputs, and this time give them three units instead of one. Every unit reads both of them, with its own pair of weights and its own bias. Each unit computes its own weighted sum, so the layer turns two numbers into three. Feed those three into one last unit, and out comes the single prediction, y hat. Written as algebra, the whole layer is one line: the vector z equals a matrix W times the vector x, plus the vector b. W holds every weight in the layer, one row for each unit, and b holds every bias. Change a single entry of W and the prediction changes. So the machinery is in place, and two things are missing. We need a number, call it L, that says how wrong the prediction is. And we need a rule that changes every weight and every bias to make that number smaller. Those two are the rest of this lecture. A weighted sum is a straight-line formula, and straight lines have a limit. Here are five measurements. As x grows the output falls to zero and then climbs again. Try matching them with one straight line. Tilt the line one way, and the left half is wrong. Tilt it the other way, and the right half is wrong. There is no slope and no intercept that passes through all five points, because the data bends and the line cannot. You might hope that stacking two layers fixes that. Let us check. The first layer computes v x plus c. The second layer takes that answer and computes w times it, plus b. Substitute, and expand. The answer is w v times x, plus w c plus b. That is a straight line again, with a new slope and a new intercept. Two straight layers are one straight layer, and a hundred of them would still be one. So something has to bend. Take the simplest bend there is. It is called ReLU, and it returns its input when the input is positive, and zero otherwise. ReLU of two minus x is the blue piece. It slopes down until x reaches two, and after that it is flat at zero. ReLU of x minus two is the green piece: flat at zero until two, and then rising. Wherever one of them is positive the other is zero, so adding them gives exactly the shape through all five points. One bend per unit, and a layer of units can fold a straight line into very nearly any shape you want. Two bends are standard, and the first is the sigmoid. It squeezes any input at all into the range from zero to one. A large positive input gives almost one, a large negative input gives almost zero, and an input of zero gives exactly one half. That is useful when the output should read as a probability. We will also need its slope later, and the slope is unusually tidy: sigma prime equals sigma times one minus sigma. Where the curve is steep that number is large, and out at the flat ends it is almost nothing. The other standard bend is the one we just used, ReLU. It is not smooth at zero, but it is cheap to compute, and its slope is simply one on the right and zero on the left. Either way, a unit now does two things: it forms a weighted sum, and then it bends the result. A network has to know when it is wrong, and by how much. Here are three training examples. The red dots are the correct answers, recorded in advance, and the blue line is what the model currently predicts. For each example, subtract the target from the prediction. That difference is the error. The first prediction is half a unit too low, the second is a whole unit too high, and the third is half a unit too low again. Errors come with signs, and if we simply added them a prediction that is too high would cancel one that is too low. So square each error first. Squaring makes every contribution positive, and it punishes a big miss much more than a small one. Add the three squares. A quarter, plus one, plus a quarter, is one and a half. Then divide by the number of examples. The mean squared error is zero point five. In general, then: average the squared difference between prediction and target over all n examples. One number, for the whole data set, and it is the number we are going to make as small as we can. Now here is the change of view that the rest of the lecture rests on. The data are fixed. We cannot alter a single target. The only things we can move are the weights, so the loss is really a function of the weights. To see that clearly, keep one single example: input equal to one, target equal to one. With one weight and no bias the prediction is just w, so the loss is w minus one, all squared. Sweep the weight and watch the loss. At w equals two the loss is one. Bring the weight down and the loss falls, reaches zero at w equals one, and climbs the far wall again. So training means finding the bottom of that bowl. We have a bowl, and we need to walk to the bottom of it. But we are not allowed to see the whole bowl. At any moment the network knows only the weight it currently has, and how steeply the loss is rising or falling right there. That steepness is exactly what a derivative is. The derivative of w minus one squared is two times w minus one. At w equals two it comes to two, a positive number, and the tangent line there rises to the right. A positive slope means the loss grows as the weight grows. So to make the loss smaller, move the weight the other way. Subtract something proportional to the slope, and that is the whole of gradient descent in one line. New weight equals old weight, minus eta times the derivative. Eta is a small positive number called the learning rate, and it decides how far each step carries us. Take eta equal to zero point three. The step is zero point three times two, which is zero point six, so the weight goes from two to one point four. Watch the tangent flatten as it lands. Do it again. At one point four the slope is only zero point eight, so the step is smaller and the weight moves to one point one six. Once more, and it reaches one point zero six. The steps shrink by themselves, because near the bottom there is hardly any slope left to multiply. Had the slope been negative, the minus sign would have pushed the weight up instead. Either way we go downhill. But eta matters. Suppose at w equals two we had used one point one. The step would be two point two, the weight would land at minus zero point two, and the loss there is worse than where we began. A real network has thousands of weights, not one, and nothing changes except that there is now a slope for each of them. Here is a loss with two weights, drawn as a contour map. Every ring is a set of weights giving the same loss, and the bottom of the bowl is inside the smallest ring. Hold w two still and ask how the loss changes as w one moves on its own. That is the partial derivative with respect to w one, and taking a step against it gives the blue arrow. Hold w one still instead, and you get the green arrow. Collect the partial derivatives into one list and you have the gradient. Here it is minus four and six. The gradient points in the direction the loss increases fastest, straight across the contour lines. So step the opposite way. Add the two component steps, and the red arrow is the diagonal of their parallelogram: that is where we land. The rule reads as before, with the gradient standing in for the single derivative. And repeat. Each step crosses to an inner ring, the arrows shorten as the ground flattens out, and the weights settle near the bottom. That is gradient descent. What is left is the hard part. In a real network, with layers feeding layers, how do we actually compute those partial derivatives? We need the slope of the loss with respect to one weight. But the weight does not touch the loss directly. Follow the arithmetic forward: the weight and the input make z. The bend turns z into the activation a. And a, compared with the target y, produces the loss. So nudging the weight nudges z, and that nudges a, and that nudges the loss. Three links in one chain. Read the chain backwards, which is where the name backpropagation comes from. First, how fast does the loss respond to a? Next, how fast does a respond to z? Finally, how fast does z respond to the weight? Multiply those three local rates together and you have the slope you wanted. Each of those three rates is a one-line derivative. The loss is a minus y, all squared, so differentiating it with respect to a gives two times a minus y. The activation is sigma of z, and the sigmoid's slope is the tidy expression we met earlier: a times one minus a. And z is w x plus b, so differentiating with respect to w leaves just x. With respect to b it leaves one, because the bias is added on its own. Multiply the three together and there it is. The slope of the loss with respect to this weight is two times a minus y, times a times one minus a, times x. For the bias it is the same product without the x on the end. Now put numbers in. One input equal to one, a target of one, a weight of a half, and a bias of zero. The weighted sum is a half. The sigmoid of a half is zero point six two two, so the network answers zero point six two two when it should answer one. That red gap is the error. a minus y is minus zero point three seven eight, and twice it is minus zero point seven five five. There is the first factor. The second factor is the steepness of the sigmoid at this point, a times one minus a, which works out as zero point two three five. You can see it in the tangent line: a gentle slope, so a gentle factor. Multiply them, together with x, which is one. The slope of the loss with respect to this weight is minus zero point one seven seven. It is negative, so raising the weight lowers the loss. With a learning rate of a half the weight moves from zero point five to zero point five eight nine, and the activation climbs the curve toward its target. And the loss falls, from zero point one four three to zero point one two seven. The gap is narrower than it was. That is one training step for one weight, and a network does it for every weight it has, over and over again. So that is the whole loop. Forward, weighted sums and bends turn the inputs into a prediction. The loss collapses all of that into one number. The chain rule turns that number into a slope for every weight and every bias in the network. Then each weight takes a small step against its own slope, and we go round again. Everything else in deep learning is this loop, done at scale.","watch":{"version":1,"scenes":[{"title":"From Numbers to a Prediction","start":0,"end":185.75941666666662,"objects":{"bias_link":"an Arrow [green] labelled \"b = -1\" drawn in net (start=(3.4, 1.0), end=(3.4, 1.85))","cell":"a Point [yellow] labelled \"Sigma\" drawn in net (location=(3.4, 2.15), marker_radius=0.14)","edges_in":"a Line [gray] drawn in stack (start=(0.95, 2.9), end=(2.95, 3.5))","edges_in_2":"a Line [gray] drawn in stack (start=(0.95, 2.9), end=(2.95, 2.1))","edges_in_3":"a Line [gray] drawn in stack (start=(0.95, 2.9), end=(2.95, 0.7))","edges_in_4":"a Line [gray] drawn in stack (start=(0.95, 1.3), end=(2.95, 3.5))","edges_in_5":"a Line [gray] drawn in stack (start=(0.95, 1.3), end=(2.95, 2.1))","edges_in_6":"a Line [gray] drawn in stack (start=(0.95, 1.3), end=(2.95, 0.7))","edges_out":"a Line [gray] drawn in stack (start=(3.25, 3.5), end=(5.35, 2.1))","edges_out_2":"a Line [gray] drawn in stack (start=(3.25, 2.1), end=(5.35, 2.1))","edges_out_3":"a Line [gray] drawn in stack (start=(3.25, 0.7), end=(5.35, 2.1))","form":"a Math [text] that says \"$z = w_1 x_1 + w_2 x_2 + b$\"","heading_layer":"a Heading that says \"A Layer of Units\"","hidden":"a Point [yellow] labelled \"z_1\" drawn in stack (location=(3.1, 3.5))","hidden_2":"a Point [yellow] labelled \"z_2\" drawn in stack (location=(3.1, 2.1))","hidden_3":"a Point [yellow] labelled \"z_3\" drawn in stack (location=(3.1, 0.7))","in_1":"a Point [blue] labelled \"x_1\" drawn in net (location=(0.9, 3.0))","in_2":"a Point [blue] labelled \"x_2\" drawn in net (location=(0.9, 1.3))","layer_math":"a Math [text] that says \"$arrow(z) = W arrow(x) + arrow(b)$\"","link_1":"an Arrow [gray] labelled \"w_1 = 0.5\" drawn in net (start=(1.15, 2.95), end=(3.1, 2.35))","link_2":"an Arrow [gray] labelled \"w_2 = 0.2\" drawn in net (start=(1.15, 1.35), end=(3.1, 1.95))","net":"a Figure (x_range=(0.0, 6.6), y_range=(0.4, 3.9), aspect=(6.6, 3.5))","numbers":"a Math [text] that says \"$z = 0.5 dot.op 3 + 0.2 dot.op 8 - 1$\"","out_link":"an Arrow [red] labelled \"z = 2.1\" drawn in net (start=(3.7, 2.15), end=(5.3, 2.15))","promises":"a Block [text] that says \"A number $L$ that measures how wrong $hat(y)$ is. A rule that changes every weight to make $L$ smaller.\"","question":"a Panel that says \"Two numbers go in and one number comes out. What arithmetic turns the inputs into the prediction, and how can that arithmetic correct itself when the prediction is wrong?\"","s_in_1":"a Point [blue] labelled \"x_1\" drawn in stack (location=(0.8, 2.9))","s_in_2":"a Point [blue] labelled \"x_2\" drawn in stack (location=(0.8, 1.3))","s_out":"a Point [red] labelled \"hat(y)\" drawn in stack (location=(5.5, 2.1))","shapes":"a Math [text] that says \"$W: 3 times 2, quad arrow(b): 3 times 1$\"","stack":"a Figure (x_range=(0.0, 6.4), y_range=(0.2, 4.2), aspect=(6.4, 4.0))","w1":"a VariableNumber (initial_value=0.5, format_spec='.1f')","z_num":"a VariableNumber (initial_value=2.1, format_spec='.1f')"},"beats":[{"start":0,"say":"A neural network is a piece of arithmetic with adjustable numbers inside it. Suppose two measurements describe a student: hours of study, and hours of sleep. We want one number out, a predicted exam score. So there are two questions. What arithmetic should we do, and how can the numbers inside it fix themselves when the answer comes out wrong?","live":[],"does":[[0,"question is shown on the screen, written out."],[22.9415,"question moves to a new place on the board."]]},{"start":23.541500000000003,"say":"Draw that arithmetic as a picture. The two inputs sit on the left, and I will call them x one and x two. For our student, x one is three hours of study and x two is eight hours of sleep. Over here is a single unit, and its job is to turn those two numbers into one.","live":["question"],"does":[[23.541500000000003,"net is shown on the screen, written out."],[27.698,"in_1 is shown on the screen, written out."],[28.131558149589278,"in_2 is shown on the screen, written out."],[38.855,"cell is shown on the screen, written out."]]},{"start":43.309000000000005,"say":"Every arrow carries a number of its own, called a weight, and the weight decides how much that input matters. The first weight is zero point five. The second is zero point two.","live":["question","net","in_1","in_2","cell"],"does":[[44.505,"link_1 is shown on the screen, written out."],[53.131,"link_2 is shown on the screen, written out."]]},{"start":55.4605,"say":"The unit does two things with those numbers. First it adds up the weighted inputs. Then it adds one more number of its own, which belongs to the unit rather than to any input. That number is called the bias, and here it is minus one. It lets the unit shift its answer up or down whatever the inputs happen to be.","live":["question","net","in_1","in_2","cell","link_1","link_2"],"does":[[58.746,"cell is indicated — a transient flash."],[68.02199999999999,"bias_link is shown on the screen, written out."]]},{"start":76.07600000000001,"say":"Written out, that is the whole of one unit. z equals w one x one, plus w two x two, plus b. A weighted sum of the inputs, and then the bias. Put our numbers in: zero point five times three, plus zero point two times eight, minus one.","live":["question","net","in_1","in_2","cell","link_1","link_2","bias_link"],"does":[[76.378,"net moves to a new place on the board."],[76.378,"form is shown on the screen, written out."],[88.174,"numbers is shown on the screen, written out."]]},{"start":95.02550000000001,"say":"One point five, plus one point six, minus one. The unit's output is two point one.","live":["question","net","form","numbers","in_1","in_2","cell","link_1","link_2","bias_link"],"does":[[95.02550000000001,"numbers becomes \"$z = 1.5 + 1.6 - 1$\"."],[99.32700000000001,"numbers becomes \"$z = 1.5 + 1.6 - 1 = 2.1$\"."],[99.32700000000001,"out_link is shown on the screen, written out."]]},{"start":101.7035,"say":"Now clear the arithmetic and watch the weight itself. Raise the first weight to one, and the output climbs to three point six, because study now counts for twice as much. Lower it to zero point one, and the output drops to zero point nine. Those weights are the dials, and everything that follows is about how to set them.","live":["question","net","form","numbers","in_1","in_2","cell","link_1","link_2","bias_link","out_link"],"does":[[102.226,"numbers is hidden from the screen."],[105.163,"link_1 is redrawn as the numbers it depends on change."],[105.163,"out_link is redrawn as the numbers it depends on change."],[105.163,"w1 ticks to 1.0."],[105.163,"z_num ticks to 3.6."],[112.408,"link_1 is redrawn as the numbers it depends on change."],[112.408,"out_link is redrawn as the numbers it depends on change."],[112.408,"w1 ticks to 0.1."],[112.408,"z_num ticks to 0.9."],[121.91650000000001,"form is hidden from the screen — left the board."],[121.91650000000001,"net is hidden from the screen — left the board."],[121.91650000000001,"in_1 is hidden from the screen — net left the board."],[121.91650000000001,"in_2 is hidden from the screen — net left the board."],[121.91650000000001,"cell is hidden from the screen — net left the board."],[121.91650000000001,"link_1 is hidden from the screen — net left the board."],[121.91650000000001,"link_2 is hidden from the screen — net left the board."],[121.91650000000001,"bias_link is hidden from the screen — net left the board."],[121.91650000000001,"out_link is hidden from the screen — net left the board."],[121.91650000000001,"question is hidden from the screen — left the board."]]},{"start":123.1165,"say":"One unit is not a network. Start again from the two inputs, and this time give them three units instead of one. Every unit reads both of them, with its own pair of weights and its own bias.","live":[],"does":[[123.1165,"heading_layer is shown on the screen, written out."],[123.1165,"stack is shown on the screen, written out."],[126.75,"s_in_1 is shown on the screen, written out."],[127.06057028695969,"s_in_2 is shown on the screen, written out."],[128.457,"hidden is shown on the screen, written out."],[128.5924431529241,"hidden_2 is shown on the screen, written out."],[128.7278863058482,"hidden_3 is shown on the screen, written out."],[131.209,"edges_in is shown on the screen, written out."],[131.27042390119868,"edges_in_2 is shown on the screen, written out."],[131.3318478023974,"edges_in_3 is shown on the screen, written out."],[131.39327170359607,"edges_in_4 is shown on the screen, written out."],[131.45469560479475,"edges_in_5 is shown on the screen, written out."],[131.51999622230295,"edges_in_6 is shown on the screen, written out."]]},{"start":135.9305,"say":"Each unit computes its own weighted sum, so the layer turns two numbers into three. Feed those three into one last unit, and out comes the single prediction, y hat.","live":["stack","heading_layer","s_in_1","s_in_2","hidden","hidden_2","hidden_3","edges_in","edges_in_2","edges_in_3","edges_in_4","edges_in_5","edges_in_6"],"does":[[142.072,"edges_out is shown on the screen, written out."],[142.1553800460155,"edges_out_2 is shown on the screen, written out."],[142.23876009203096,"edges_out_3 is shown on the screen, written out."],[144.65,"s_out is shown on the screen, written out."]]},{"start":147.804,"say":"Written as algebra, the whole layer is one line: the vector z equals a matrix W times the vector x, plus the vector b. W holds every weight in the layer, one row for each unit, and b holds every bias. Change a single entry of W and the prediction changes.","live":["stack","heading_layer","s_in_1","s_in_2","hidden","hidden_2","hidden_3","edges_in","edges_in_2","edges_in_3","edges_in_4","edges_in_5","edges_in_6","edges_out","edges_out_2","edges_out_3","s_out"],"does":[[150.46300000000002,"stack moves to a new place on the board."],[150.46300000000002,"layer_math is shown on the screen, written out."],[159.762,"shapes is shown on the screen, written out."]]},{"start":167.874,"say":"So the machinery is in place, and two things are missing. We need a number, call it L, that says how wrong the prediction is. And we need a rule that changes every weight and every bias to make that number smaller. Those two are the rest of this lecture.","live":["layer_math","shapes","stack","heading_layer","s_in_1","s_in_2","hidden","hidden_2","hidden_3","edges_in","edges_in_2","edges_in_3","edges_in_4","edges_in_5","edges_in_6","edges_out","edges_out_2","edges_out_3","s_out"],"does":[[171.056,"promises is shown on the screen, written out."],[174.95600000000002,"promises (the \"how wrong\" part) is emphasized."],[177.53300000000002,"promises (the \"changes every weight\" part) is emphasized."],[177.53300000000002,"promises (the \"how wrong\" part) is no longer emphasized."],[183.211,"promises (the \"changes every weight\" part) is no longer emphasized."],[184.71775,"heading_layer is hidden from the screen — left the board."],[184.71775,"layer_math is hidden from the screen — left the board."],[184.71775,"promises is hidden from the screen — left the board."],[184.71775,"shapes is hidden from the screen — left the board."],[184.71775,"stack is hidden from the screen — left the board."],[184.71775,"s_in_1 is hidden from the screen — stack left the board."],[184.71775,"s_in_2 is hidden from the screen — stack left the board."],[184.71775,"hidden is hidden from the screen — stack left the board."],[184.71775,"hidden_2 is hidden from the screen — stack left the board."],[184.71775,"hidden_3 is hidden from the screen — stack left the board."],[184.71775,"edges_in is hidden from the screen — stack left the board."],[184.71775,"edges_in_2 is hidden from the screen — stack left the board."],[184.71775,"edges_in_3 is hidden from the screen — stack left the board."],[184.71775,"edges_in_4 is hidden from the screen — stack left the board."],[184.71775,"edges_in_5 is hidden from the screen — stack left the board."],[184.71775,"edges_in_6 is hidden from the screen — stack left the board."],[184.71775,"edges_out is hidden from the screen — stack left the board."],[184.71775,"edges_out_2 is hidden from the screen — stack left the board."],[184.71775,"edges_out_3 is hidden from the screen — stack left the board."],[184.71775,"s_out is hidden from the screen — stack left the board."]]}]},{"title":"Why a Straight Line Is Not Enough","start":185.75941666666662,"end":345.68372916666664,"objects":{"ceiling":"a Line [gray] drawn in sig_axes (start=(-4.0, 1.0), end=(4.0, 1.0), dashed=True)","collapse":"a Derivation [text] that says \"$h &= v x + c \\ hat(y) &= w h + b \\ &= w (v x + c) + b \\ &= (w v) x + (w c + b)$\"","fit":"an Axes (x_range=(-0.4, 4.6), y_range=(-0.6, 3.0), x_ticks_every=1.0)","fit_line":"a Math [text] that says \"$hat(y) = upright(\"ReLU\")(2 - x) + upright(\"ReLU\")(x - 2)$\"","guess":"a Line [yellow] labelled \"v x + c\" drawn in fit (start=(0.0, <VariableNumber intercept = 0.4>), end=(4.4, ((slope * 4.4) + intercept)))","half":"a Point [yellow] labelled \"(0, 0.5)\" drawn in sig_axes (location=(0.0, 0.5))","heading_bend":"a Heading that says \"One Bend Is Enough\"","heading_collapse":"a Heading that says \"Two Straight Layers Collapse\"","heading_two":"a Heading that says \"Two Standard Bends\"","intercept":"a VariableNumber (initial_value=1.0)","left_piece":"a FunctionPlot [blue] drawn in fit (function=<function>, x_range=(-0.3, 4.4))","piece_left":"a Math [blue] that says \"$upright(\"ReLU\")(2 - x)$\"","piece_right":"a Math [green] that says \"$upright(\"ReLU\")(x - 2)$\"","point":"a Point [yellow] drawn in sig_axes (location=(0.0, 0.5))","point_2":"a Point [yellow] drawn in sig_axes (location=(3.0, 0.953))","relu_axes":"an Axes (x_range=(-3.0, 3.0), y_range=(-0.4, 3.0), x_ticks_every=1.0)","relu_curve":"a FunctionPlot [green] drawn in relu_axes (function=<function>, x_range=(-3.0, 3.0))","relu_def":"a Math [text] that says \"$upright(\"ReLU\")(t) = max(0, t)$\"","relu_formula":"a Math [text] that says \"$upright(\"ReLU\")(z) = max(0, z)$\"","relu_slope":"a Text [text] that says \"Slope $1$ right of zero, $0$ left.\"","right_piece":"a FunctionPlot [green] drawn in fit (function=<function>, x_range=(-0.3, 4.4))","samples":"a Point [red] drawn in fit (location=(0.0, 2.0))","samples_2":"a Point [red] drawn in fit (location=(1.0, 1.0))","samples_3":"a Point [red] drawn in fit (location=(2.0, 0.0))","samples_4":"a Point [red] drawn in fit (location=(3.0, 1.0))","samples_5":"a Point [red] drawn in fit (location=(4.0, 2.0))","sig_axes":"an Axes (x_range=(-4.0, 4.0), y_range=(-0.2, 1.2), x_ticks_every=2.0)","sig_curve":"a FunctionPlot [blue] drawn in sig_axes (function=<function>, x_range=(-4.0, 4.0))","sig_deriv":"a Math [text] that says \"$sigma'(z) = sigma(z) (1 - sigma(z))$\"","sig_formula":"a Math [text] that says \"$sigma(z) = frac(1, 1 + e^(-z))$\"","slope":"a VariableNumber (initial_value=0.2)"},"beats":[{"start":185.75941666666662,"say":"A weighted sum is a straight-line formula, and straight lines have a limit. Here are five measurements. As x grows the output falls to zero and then climbs again. Try matching them with one straight line.","live":[],"does":[[185.75941666666662,"heading_collapse is shown on the screen, written out."],[191.87741666666662,"fit is shown on the screen, written out."],[191.87741666666662,"samples is shown on the screen, written out."],[192.04570723823946,"samples_2 is shown on the screen, written out."],[192.21399780981233,"samples_3 is shown on the screen, written out."],[192.38228838138517,"samples_4 is shown on the screen, written out."],[192.550578952958,"samples_5 is shown on the screen, written out."],[197.53241666666662,"guess is shown on the screen, written out."]]},{"start":200.3609166666666,"say":"Tilt the line one way, and the left half is wrong. Tilt it the other way, and the right half is wrong. There is no slope and no intercept that passes through all five points, because the data bends and the line cannot.","live":["fit","heading_collapse","samples","samples_2","samples_3","samples_4","samples_5","guess"],"does":[[200.7094166666666,"guess is redrawn as the numbers it depends on change."],[200.7094166666666,"slope ticks to -0.45."],[200.7094166666666,"intercept ticks to 2.0."],[204.14641666666662,"guess is redrawn as the numbers it depends on change."],[204.14641666666662,"slope ticks to 0.45."],[204.14641666666662,"intercept ticks to 0.4."]]},{"start":215.3344166666666,"say":"You might hope that stacking two layers fixes that. Let us check. The first layer computes v x plus c. The second layer takes that answer and computes w times it, plus b.","live":null,"does":[[220.6174166666666,"fit moves to a new place on the board."],[220.6174166666666,"collapse is shown on the screen, written out."],[224.1004166666666,"collapse is shown on the screen, written out."]]},{"start":228.64191666666662,"say":"Substitute, and expand. The answer is w v times x, plus w c plus b. That is a straight line again, with a new slope and a new intercept. Two straight layers are one straight layer, and a hundred of them would still be one.","live":null,"does":[[228.8334166666666,"collapse is shown on the screen, written out."],[229.87841666666662,"collapse is shown on the screen, written out."],[238.2374166666666,"collapse (the \"(w v) x\" part) is emphasized."],[239.08541666666662,"collapse (the \"(w c + b)\" part) is emphasized."],[239.08541666666662,"collapse (the \"(w v) x\" part) is no longer emphasized."],[243.19541666666663,"collapse (the \"(w c + b)\" part) is no longer emphasized."],[244.8789166666666,"collapse is hidden from the screen — left the board."],[244.8789166666666,"heading_collapse is hidden from the screen — left the board."]]},{"start":246.07891666666663,"say":"So something has to bend. Take the simplest bend there is. It is called ReLU, and it returns its input when the input is positive, and zero otherwise.","live":["fit","samples","samples_2","samples_3","samples_4","samples_5","guess"],"does":[[246.07891666666663,"heading_bend is shown on the screen, written out."],[246.07891666666663,"guess is hidden from the screen."],[250.87341666666663,"relu_def is shown on the screen, written out."]]},{"start":257.05791666666664,"say":"ReLU of two minus x is the blue piece. It slopes down until x reaches two, and after that it is flat at zero. ReLU of x minus two is the green piece: flat at zero until two, and then rising.","live":["fit","samples","samples_2","samples_3","samples_4","samples_5","relu_def","heading_bend"],"does":[[259.7284166666666,"piece_left is shown on the screen, written out."],[259.7284166666666,"left_piece is shown on the screen, written out."],[268.1574166666666,"piece_right is shown on the screen, written out."],[268.1574166666666,"right_piece is shown on the screen, written out."]]},{"start":273.1694166666666,"say":"Wherever one of them is positive the other is zero, so adding them gives exactly the shape through all five points. One bend per unit, and a layer of units can fold a straight line into very nearly any shape you want.","live":["fit","samples","samples_2","samples_3","samples_4","samples_5","relu_def","piece_left","piece_right","heading_bend","left_piece","right_piece"],"does":[[276.6524166666666,"fit_line is shown on the screen, written out."],[277.46441666666664,"left_piece is indicated — a transient flash."],[277.73484639126303,"right_piece is indicated — a transient flash."],[287.03191666666663,"fit is hidden from the screen — left the board."],[287.03191666666663,"samples is hidden from the screen — fit left the board."],[287.03191666666663,"samples_2 is hidden from the screen — fit left the board."],[287.03191666666663,"samples_3 is hidden from the screen — fit left the board."],[287.03191666666663,"samples_4 is hidden from the screen — fit left the board."],[287.03191666666663,"samples_5 is hidden from the screen — fit left the board."],[287.03191666666663,"left_piece is hidden from the screen — fit left the board."],[287.03191666666663,"right_piece is hidden from the screen — fit left the board."],[287.03191666666663,"fit_line is hidden from the screen — left the board."],[287.03191666666663,"heading_bend is hidden from the screen — left the board."],[287.03191666666663,"piece_left is hidden from the screen — left the board."],[287.03191666666663,"piece_right is hidden from the screen — left the board."],[287.03191666666663,"relu_def is hidden from the screen — left the board."]]},{"start":288.2319166666666,"say":"Two bends are standard, and the first is the sigmoid. It squeezes any input at all into the range from zero to one. A large positive input gives almost one, a large negative input gives almost zero, and an input of zero gives exactly one half.","live":[],"does":[[288.2319166666666,"heading_two is shown on the screen, written out."],[291.3434166666666,"sig_axes is shown on the screen, written out."],[291.3434166666666,"sig_formula is shown on the screen, written out."],[292.88741666666664,"sig_curve is shown on the screen, written out."],[294.8614166666666,"ceiling is shown on the screen, written out."],[305.1244166666666,"half is shown on the screen, written out."]]},{"start":306.45641666666666,"say":"That is useful when the output should read as a probability. We will also need its slope later, and the slope is unusually tidy: sigma prime equals sigma times one minus sigma. Where the curve is steep that number is large, and out at the flat ends it is almost nothing.","live":["sig_formula","sig_axes","heading_two","sig_curve","ceiling","half"],"does":[[313.86341666666664,"sig_deriv is shown on the screen, written out."],[320.1094166666666,"point is shown on the screen, grown."],[322.35305829620194,"point is hidden from the screen."],[322.6404166666666,"point_2 is shown on the screen, grown."],[324.9413613161131,"point_2 is hidden from the screen."]]},{"start":325.44641666666666,"say":"The other standard bend is the one we just used, ReLU. It is not smooth at zero, but it is cheap to compute, and its slope is simply one on the right and zero on the left. Either way, a unit now does two things: it forms a weighted sum, and then it bends the result.","live":["sig_formula","sig_deriv","sig_axes","heading_two","sig_curve","ceiling","half"],"does":[[328.4884166666667,"relu_axes is shown on the screen, written out."],[328.4884166666667,"relu_formula is shown on the screen, written out."],[329.72008742516255,"relu_curve is shown on the screen, written out."],[333.95641666666666,"relu_slope is shown on the screen, written out."],[343.2564166666666,"sig_formula is indicated — a transient flash."],[343.4231694771073,"relu_formula is indicated — a transient flash."],[344.6420625,"heading_two is hidden from the screen — left the board."],[344.6420625,"relu_axes is hidden from the screen — left the board."],[344.6420625,"relu_curve is hidden from the screen — relu_axes left the board."],[344.6420625,"relu_formula is hidden from the screen — left the board."],[344.6420625,"relu_slope is hidden from the screen — left the board."],[344.6420625,"sig_axes is hidden from the screen — left the board."],[344.6420625,"sig_curve is hidden from the screen — sig_axes left the board."],[344.6420625,"ceiling is hidden from the screen — sig_axes left the board."],[344.6420625,"half is hidden from the screen — sig_axes left the board."],[344.6420625,"sig_deriv is hidden from the screen — left the board."],[344.6420625,"sig_formula is hidden from the screen — left the board."]]}]},{"title":"A Number for Being Wrong","start":345.68372916666664,"end":461.83008333333333,"objects":{"bowl":"an Axes (x_range=(-0.4, 2.6), y_range=(-0.3, 2.3), x_ticks_every=0.5)","bowl_curve":"a FunctionPlot [blue] drawn in bowl (function=<function>, x_range=(-0.2, 2.4))","e1":"a Line [yellow] labelled \"-0.5\" drawn in pred (start=(1.0, 1.0), end=(1.0, 1.5))","e2":"a Line [yellow] labelled \"+1.0\" drawn in pred (start=(2.0, 2.0), end=(2.0, 1.0))","e3":"a Line [yellow] labelled \"-0.5\" drawn in pred (start=(3.0, 3.0), end=(3.0, 3.5))","err_def":"a Math [text] that says \"$e_i = hat(y)_i - y_i$\"","heading_loss":"a Heading that says \"The Loss, and What It Depends On\"","loss_w":"a Math [text] that says \"$L(w) = (w - 1)^2$\"","model":"a FunctionPlot [blue] drawn in pred (function=<function>, x_range=(0.0, 3.6))","mse_def":"a Panel that says \"The mean squared error averages the squared difference between prediction and target over all $n$ training examples.\"","mse_formula":"a Math [text] that says \"$L = frac(1, n) sum_(i = 1)^n (hat(y)_i - y_i)^2$\"","mse_value":"a Math [text] that says \"$L = frac(1, 3) dot.op 1.50 = 0.50$\"","pred":"an Axes (x_range=(-0.3, 4.0), y_range=(-0.3, 4.2), x_ticks_every=1.0)","problem":"a Tex [text] that says \"How wrong is this model?\"","spot":"a PlotPoint [yellow] labelled \"w = 2.00\" drawn in bowl (target='bowl_curve', x=<VariableNumber w_live = 1.0>)","square_sum":"an Arithmetic [text] that says \"$(-0.5)^2 = 0.25 (+1.0)^2 = 1.00 (-0.5)^2 = 0.25 upright(\"total\") = 1.50$\" (operator='+', operands=('(-0.5)^2 = 0.25', '(+1.0)^2 = 1.00', '(-0.5)^2 = 0.25'), result='upright(\"total\") = 1.50')","t1":"a Point [red] labelled \"y_1\" drawn in pred (location=(1.0, 1.5))","t2":"a Point [red] labelled \"y_2\" drawn in pred (location=(2.0, 1.0))","t3":"a Point [red] labelled \"y_3\" drawn in pred (location=(3.0, 3.5))","w_live":"a VariableNumber (initial_value=2.0, format_spec='.2f')"},"beats":[{"start":345.68372916666664,"say":"A network has to know when it is wrong, and by how much. Here are three training examples. The red dots are the correct answers, recorded in advance, and the blue line is what the model currently predicts.","live":[],"does":[[345.68372916666664,"problem is shown on the screen, written out."],[351.10572916666666,"pred is shown on the screen, written out."],[352.81272916666666,"t1 is shown on the screen, written out."],[352.98267054020096,"t2 is shown on the screen, written out."],[353.42058595058626,"t3 is shown on the screen, written out."],[356.86372916666664,"model is shown on the screen, written out."]]},{"start":360.35472916666663,"say":"For each example, subtract the target from the prediction. That difference is the error. The first prediction is half a unit too low, the second is a whole unit too high, and the third is half a unit too low again.","live":["pred","problem","t1","t2","t3","model"],"does":[[364.58072916666663,"pred moves to a new place on the board."],[364.58072916666663,"err_def is shown on the screen, written out."],[366.2877291666666,"e1 is shown on the screen, written out."],[368.63272916666665,"e2 is shown on the screen, written out."],[370.96672916666665,"e3 is shown on the screen, written out."]]},{"start":373.9347291666666,"say":"Errors come with signs, and if we simply added them a prediction that is too high would cancel one that is too low. So square each error first. Squaring makes every contribution positive, and it punishes a big miss much more than a small one.","live":["err_def","pred","problem","t1","t2","t3","model","e1","e2","e3"],"does":[[381.5857291666666,"square_sum is shown on the screen, written out."],[381.8607621274585,"square_sum is shown on the screen, written out."],[382.39272377123854,"square_sum is shown on the screen, written out."]]},{"start":389.96472916666664,"say":"Add the three squares. A quarter, plus one, plus a quarter, is one and a half. Then divide by the number of examples. The mean squared error is zero point five.","live":null,"does":[[390.3127291666666,"square_sum is shown on the screen, drawn."],[390.93143870578774,"square_sum is shown on the screen, drawn."],[395.3867291666666,"square_sum is shown on the screen, written out."],[396.98872916666664,"mse_value is shown on the screen, written out."],[403.0837291666666,"err_def is hidden from the screen — left the board."],[403.0837291666666,"mse_value is hidden from the screen — left the board."],[403.0837291666666,"pred is hidden from the screen — left the board."],[403.0837291666666,"t1 is hidden from the screen — pred left the board."],[403.0837291666666,"t2 is hidden from the screen — pred left the board."],[403.0837291666666,"t3 is hidden from the screen — pred left the board."],[403.0837291666666,"model is hidden from the screen — pred left the board."],[403.0837291666666,"e1 is hidden from the screen — pred left the board."],[403.0837291666666,"e2 is hidden from the screen — pred left the board."],[403.0837291666666,"e3 is hidden from the screen — pred left the board."],[403.0837291666666,"problem is hidden from the screen — left the board."],[403.0837291666666,"square_sum is hidden from the screen — left the board."]]},{"start":403.68372916666664,"say":"In general, then: average the squared difference between prediction and target over all n examples. One number, for the whole data set, and it is the number we are going to make as small as we can.","live":[],"does":[[403.68372916666664,"heading_loss is shown on the screen, written out."],[404.4847291666666,"mse_def is shown on the screen, written out."],[405.73872916666664,"mse_formula is shown on the screen, written out."]]},{"start":416.49772916666666,"say":"Now here is the change of view that the rest of the lecture rests on. The data are fixed. We cannot alter a single target. The only things we can move are the weights, so the loss is really a function of the weights.","live":["mse_def","mse_formula","heading_loss"],"does":[[421.37372916666664,"mse_formula (the \"(hat(y)_i - y_i)^2\" part) is emphasized."],[426.0527291666666,"mse_formula (the \"(hat(y)_i - y_i)^2\" part) is no longer emphasized."]]},{"start":429.89222916666665,"say":"To see that clearly, keep one single example: input equal to one, target equal to one. With one weight and no bias the prediction is just w, so the loss is w minus one, all squared.","live":null,"does":[[429.89222916666665,"bowl is shown on the screen, written out."],[439.5047291666666,"loss_w is shown on the screen, written out."],[441.2467291666666,"bowl_curve is shown on the screen, written out."],[443.2207291666666,"spot is shown on the screen, written out."]]},{"start":444.83072916666663,"say":"Sweep the weight and watch the loss. At w equals two the loss is one. Bring the weight down and the loss falls, reaches zero at w equals one, and climbs the far wall again. So training means finding the bottom of that bowl.","live":["mse_def","mse_formula","loss_w","bowl","heading_loss","bowl_curve","spot"],"does":[[451.4827291666666,"spot is redrawn as the numbers it depends on change."],[451.4827291666666,"w_live ticks to 1.0."],[455.74372916666664,"spot is redrawn as the numbers it depends on change."],[455.74372916666664,"w_live ticks to 0.0."],[459.23872916666664,"spot is redrawn as the numbers it depends on change."],[459.23872916666664,"w_live ticks to 1.0."],[460.78841666666665,"bowl is hidden from the screen — left the board."],[460.78841666666665,"bowl_curve is hidden from the screen — bowl left the board."],[460.78841666666665,"spot is hidden from the screen — bowl left the board."],[460.78841666666665,"heading_loss is hidden from the screen — left the board."],[460.78841666666665,"loss_w is hidden from the screen — left the board."],[460.78841666666665,"mse_def is hidden from the screen — left the board."],[460.78841666666665,"mse_formula is hidden from the screen — left the board."]]}]},{"title":"Walking Downhill","start":461.83008333333333,"end":657.5645208333333,"objects":{"bowl":"an Axes (x_range=(-0.5, 2.8), y_range=(-0.3, 2.8), x_ticks_every=1.0)","comp_1":"a Vector [blue] drawn in surf (start=(<VariableNumber p1 = -0.19>, <VariableNumber p2 = 0.6>), end=((p1 - (0.3 * (p1 - 0.5))), <VariableNumber p2 = 0.6>))","comp_2":"a Vector [green] drawn in surf (start=(<VariableNumber p1 = -0.19>, <VariableNumber p2 = 0.6>), end=(<VariableNumber p1 = -0.19>, (p2 - (0.6 * (p2 - 0.5)))))","contours":"a LevelCurves [gray] drawn in surf (function=<function>, values=(0.25, 1.0, 2.0, 3.0))","d_at_two":"a Math [text] that says \"$upright(\"at\") w = 2: quad frac(dif L, dif w) = 2$\"","d_general":"a Math [text] that says \"$frac(dif L, dif w) = 2 (w - 1)$\"","eta_note":"a Text [text] that says \"$eta$ is the learning rate: it sets how big a step to take.\"","grad_def":"a Math [text] that says \"$nabla L = (frac(partial L, partial w_1), frac(partial L, partial w_2))$\"","grad_val":"a Math [text] that says \"$nabla L = (-4, thin 6)$\"","heading_one":"a Heading that says \"One Weight, One Slope\"","heading_two":"a Heading that says \"Two Weights at Once\"","here":"a Point [yellow] labelled \"w\" drawn in surf (location=(<VariableNumber p1 = -0.19>, <VariableNumber p2 = 0.6>))","loss_2d":"a Math [text] that says \"$L = (w_1 - 0.5)^2 + 2 (w_2 - 0.5)^2$\"","loss_curve":"a FunctionPlot [blue] drawn in bowl (function=<function>, x_range=(-0.3, 2.5))","p1":"a VariableNumber (initial_value=-1.5, format_spec='.2f')","p2":"a VariableNumber (initial_value=2.0, format_spec='.2f')","rule":"a Math [text] that says \"$w arrow.l w - eta frac(dif L, dif w)$\"","rule_vec":"a Math [text] that says \"$(w_1, w_2) arrow.l (w_1, w_2) - eta nabla L$\"","side_1":"a Line [gray] drawn in surf (start=((p1 - (0.3 * (p1 - 0.5))), <VariableNumber p2 = 0.6>), end=((p1 - (0.3 * (p1 - 0.5))), (p2 - (0.6 * (p2 - 0.5)))), dashed=True)","side_2":"a Line [gray] drawn in surf (start=(<VariableNumber p1 = -0.19>, (p2 - (0.6 * (p2 - 0.5)))), end=((p1 - (0.3 * (p1 - 0.5))), (p2 - (0.6 * (p2 - 0.5)))), dashed=True)","spot":"a PlotPoint [yellow] labelled \"w = 2.00\" drawn in bowl (target='loss_curve', x=<VariableNumber w = -0.2>)","step_arrow":"a Vector [red] labelled \"- eta nabla L\" drawn in surf (start=(<VariableNumber p1 = -0.19>, <VariableNumber p2 = 0.6>), end=((p1 - (0.3 * (p1 - 0.5))), (p2 - (0.6 * (p2 - 0.5)))))","step_big":"a Math [text] that says \"$eta = 1.1: quad w arrow.l 2 - 1.1 dot.op 2 = -0.2$\"","step_num":"a Math [text] that says \"$(-1.5, thin 2) - 0.15 (-4, thin 6) = (-0.9, thin 1.1)$\"","step_one":"a Math [text] that says \"$w arrow.l 2 - 0.3 dot.op 2 = 1.4$\"","surf":"an Axes (x_range=(-2.2, 2.4), y_range=(-1.0, 2.5), aspect=(4.6, 3.5))","tangent":"a TangentLine [yellow] drawn in bowl (target='loss_curve', x=<VariableNumber w = -0.2>, length=1.0)","w":"a VariableNumber (initial_value=2.0, format_spec='.2f')"},"beats":[{"start":461.83008333333333,"say":"We have a bowl, and we need to walk to the bottom of it. But we are not allowed to see the whole bowl. At any moment the network knows only the weight it currently has, and how steeply the loss is rising or falling right there.","live":[],"does":[[461.83008333333333,"heading_one is shown on the screen, written out."],[461.83008333333333,"bowl is shown on the screen, written out."],[462.32908333333336,"loss_curve is shown on the screen, written out."],[470.22408333333334,"spot is shown on the screen, written out."]]},{"start":475.63058333333333,"say":"That steepness is exactly what a derivative is. The derivative of w minus one squared is two times w minus one. At w equals two it comes to two, a positive number, and the tangent line there rises to the right.","live":["bowl","heading_one","loss_curve","spot"],"does":[[477.68508333333335,"bowl moves to a new place on the board."],[477.68508333333335,"d_general is shown on the screen, written out."],[486.0210833333333,"d_at_two is shown on the screen, written out."],[488.32008333333334,"tangent is shown on the screen, written out."]]},{"start":491.3115833333333,"say":"A positive slope means the loss grows as the weight grows. So to make the loss smaller, move the weight the other way. Subtract something proportional to the slope, and that is the whole of gradient descent in one line.","live":["d_general","d_at_two","bowl","heading_one","loss_curve","spot","tangent"],"does":[[499.8210833333333,"rule is shown on the screen, written out."]]},{"start":506.1915833333333,"say":"New weight equals old weight, minus eta times the derivative. Eta is a small positive number called the learning rate, and it decides how far each step carries us.","live":["d_general","d_at_two","rule","bowl","heading_one","loss_curve","spot","tangent"],"does":[[508.88508333333334,"rule (the \"eta\" part) is emphasized."],[513.6920833333334,"eta_note is shown on the screen, written out."],[515.1430833333334,"rule (the \"eta\" part) is no longer emphasized."]]},{"start":518.5410833333333,"say":"Take eta equal to zero point three. The step is zero point three times two, which is zero point six, so the weight goes from two to one point four. Watch the tangent flatten as it lands.","live":["d_general","d_at_two","rule","eta_note","bowl","heading_one","loss_curve","spot","tangent"],"does":[[521.8730833333333,"step_one is shown on the screen, written out."],[529.2220833333333,"spot is redrawn as the numbers it depends on change."],[529.2220833333333,"tangent is redrawn as the numbers it depends on change."],[529.2220833333333,"w ticks to 1.4."]]},{"start":532.2255833333334,"say":"Do it again. At one point four the slope is only zero point eight, so the step is smaller and the weight moves to one point one six. Once more, and it reaches one point zero six. The steps shrink by themselves, because near the bottom there is hardly any slope left to multiply.","live":["d_general","d_at_two","rule","eta_note","step_one","bowl","heading_one","loss_curve","spot","tangent"],"does":[[538.4830833333333,"spot is redrawn as the numbers it depends on change."],[538.4830833333333,"tangent is redrawn as the numbers it depends on change."],[538.4830833333333,"w ticks to 1.16."],[542.0950833333334,"spot is redrawn as the numbers it depends on change."],[542.0950833333334,"tangent is redrawn as the numbers it depends on change."],[542.0950833333334,"w ticks to 1.06."]]},{"start":551.0305833333333,"say":"Had the slope been negative, the minus sign would have pushed the weight up instead. Either way we go downhill. But eta matters. Suppose at w equals two we had used one point one. The step would be two point two, the weight would land at minus zero point two, and the loss there is worse than where we began.","live":null,"does":[[559.7840833333333,"spot is redrawn as the numbers it depends on change."],[559.7840833333333,"tangent is redrawn as the numbers it depends on change."],[559.7840833333333,"w ticks to 2.0."],[563.9410833333334,"step_big is shown on the screen, written out."],[566.1470833333333,"spot is redrawn as the numbers it depends on change."],[566.1470833333333,"tangent is redrawn as the numbers it depends on change."],[566.1470833333333,"w ticks to -0.2."],[570.8835833333334,"bowl is hidden from the screen — left the board."],[570.8835833333334,"loss_curve is hidden from the screen — bowl left the board."],[570.8835833333334,"spot is hidden from the screen — bowl left the board."],[570.8835833333334,"tangent is hidden from the screen — bowl left the board."],[570.8835833333334,"d_at_two is hidden from the screen — left the board."],[570.8835833333334,"d_general is hidden from the screen — left the board."],[570.8835833333334,"eta_note is hidden from the screen — left the board."],[570.8835833333334,"heading_one is hidden from the screen — left the board."],[570.8835833333334,"rule is hidden from the screen — left the board."],[570.8835833333334,"step_big is hidden from the screen — left the board."],[570.8835833333334,"step_one is hidden from the screen — left the board."]]},{"start":572.0835833333333,"say":"A real network has thousands of weights, not one, and nothing changes except that there is now a slope for each of them. Here is a loss with two weights, drawn as a contour map. Every ring is a set of weights giving the same loss, and the bottom of the bowl is inside the smallest ring.","live":[],"does":[[572.0835833333333,"heading_two is shown on the screen, written out."],[572.0835833333333,"surf is shown on the screen, written out."],[580.1180833333333,"surf moves to a new place on the board."],[580.1180833333333,"loss_2d is shown on the screen, written out."],[580.7910833333333,"here is shown on the screen, written out."],[581.8710833333333,"contours is shown on the screen, written out."]]},{"start":590.0055833333333,"say":"Hold w two still and ask how the loss changes as w one moves on its own. That is the partial derivative with respect to w one, and taking a step against it gives the blue arrow. Hold w one still instead, and you get the green arrow.","live":["loss_2d","surf","heading_two","contours","here"],"does":[[596.7280833333333,"grad_def is shown on the screen, written out."],[601.2560833333333,"comp_1 is shown on the screen, written out."],[605.6440833333334,"comp_2 is shown on the screen, written out."]]},{"start":607.3705833333333,"say":"Collect the partial derivatives into one list and you have the gradient. Here it is minus four and six. The gradient points in the direction the loss increases fastest, straight across the contour lines.","live":["loss_2d","grad_def","surf","heading_two","contours","here","comp_1","comp_2"],"does":[[612.0960833333334,"grad_val is shown on the screen, written out."],[618.6200833333334,"contours is indicated — a transient flash."]]},{"start":621.0430833333334,"say":"So step the opposite way. Add the two component steps, and the red arrow is the diagonal of their parallelogram: that is where we land. The rule reads as before, with the gradient standing in for the single derivative.","live":["loss_2d","grad_def","grad_val","surf","heading_two","contours","here","comp_1","comp_2"],"does":[[623.5510833333334,"side_1 is shown on the screen, written out."],[623.8012408880836,"side_2 is shown on the screen, written out."],[625.7100833333334,"step_arrow is shown on the screen, written out."],[629.8780833333334,"step_num is shown on the screen, written out."],[631.3530833333334,"rule_vec is shown on the screen, written out."]]},{"start":636.2835833333334,"say":"And repeat. Each step crosses to an inner ring, the arrows shorten as the ground flattens out, and the weights settle near the bottom. That is gradient descent. What is left is the hard part. In a real network, with layers feeding layers, how do we actually compute those partial derivatives?","live":["loss_2d","grad_def","grad_val","rule_vec","step_num","surf","heading_two","contours","here","comp_1","comp_2","side_1","side_2","step_arrow"],"does":[[636.8870833333334,"here is redrawn as the numbers it depends on change."],[636.8870833333334,"comp_1 is redrawn as the numbers it depends on change."],[636.8870833333334,"comp_2 is redrawn as the numbers it depends on change."],[636.8870833333334,"side_1 is redrawn as the numbers it depends on change."],[636.8870833333334,"side_2 is redrawn as the numbers it depends on change."],[636.8870833333334,"step_arrow is redrawn as the numbers it depends on change."],[636.8870833333334,"p1 ticks to -0.9."],[636.8870833333334,"p2 ticks to 1.1."],[640.8810833333334,"here is redrawn as the numbers it depends on change."],[640.8810833333334,"comp_1 is redrawn as the numbers it depends on change."],[640.8810833333334,"comp_2 is redrawn as the numbers it depends on change."],[640.8810833333334,"side_1 is redrawn as the numbers it depends on change."],[640.8810833333334,"side_2 is redrawn as the numbers it depends on change."],[640.8810833333334,"step_arrow is redrawn as the numbers it depends on change."],[640.8810833333334,"p1 ticks to -0.48."],[640.8810833333334,"p2 ticks to 0.74."],[643.5860833333334,"here is redrawn as the numbers it depends on change."],[643.5860833333334,"comp_1 is redrawn as the numbers it depends on change."],[643.5860833333334,"comp_2 is redrawn as the numbers it depends on change."],[643.5860833333334,"side_1 is redrawn as the numbers it depends on change."],[643.5860833333334,"side_2 is redrawn as the numbers it depends on change."],[643.5860833333334,"step_arrow is redrawn as the numbers it depends on change."],[643.5860833333334,"p1 ticks to -0.19."],[643.5860833333334,"p2 ticks to 0.6."],[656.5228541666667,"grad_def is hidden from the screen — left the board."],[656.5228541666667,"grad_val is hidden from the screen — left the board."],[656.5228541666667,"heading_two is hidden from the screen — left the board."],[656.5228541666667,"loss_2d is hidden from the screen — left the board."],[656.5228541666667,"rule_vec is hidden from the screen — left the board."],[656.5228541666667,"step_num is hidden from the screen — left the board."],[656.5228541666667,"surf is hidden from the screen — left the board."],[656.5228541666667,"contours is hidden from the screen — surf left the board."],[656.5228541666667,"here is hidden from the screen — surf left the board."],[656.5228541666667,"comp_1 is hidden from the screen — surf left the board."],[656.5228541666667,"comp_2 is hidden from the screen — surf left the board."],[656.5228541666667,"side_1 is hidden from the screen — surf left the board."],[656.5228541666667,"side_2 is hidden from the screen — surf left the board."],[656.5228541666667,"step_arrow is hidden from the screen — surf left the board."]]}]},{"title":"Backpropagation: The Chain Rule Doing the Work","start":657.5645208333333,"end":870.5116874999999,"objects":{"b_1":"an Arrow [red] labelled \"frac(partial z, partial w)\" drawn in chain (start=(2.75, 1.0), end=(0.95, 1.0))","b_2":"an Arrow [red] labelled \"frac(dif a, dif z)\" drawn in chain (start=(4.95, 1.0), end=(3.25, 1.0))","b_3":"an Arrow [red] labelled \"frac(partial L, partial a)\" drawn in chain (start=(7.05, 1.0), end=(5.45, 1.0))","chain":"a Figure (x_range=(0.0, 8.0), y_range=(0.2, 3.0), aspect=(8.0, 2.8))","chain_rule":"a Math [text] that says \"$frac(partial L, partial w) = frac(partial L, partial a) dot.op frac(dif a, dif z) dot.op frac(partial z, partial w)$\"","f_1":"an Arrow [gray] labelled \"w x + b\" drawn in chain (start=(0.95, 2.2), end=(2.75, 2.2))","f_2":"an Arrow [gray] labelled \"sigma\" drawn in chain (start=(3.25, 2.2), end=(4.95, 2.2))","f_3":"an Arrow [gray] labelled \"(a - y)^2\" drawn in chain (start=(5.45, 2.2), end=(7.05, 2.2))","gap_seg":"a Line [red] labelled \"a - y\" drawn in sig (start=(<VariableNumber z_live = 0.589>, (1.0 / (1.0 + exp((-1.0 * z_l…, end=(<VariableNumber z_live = 0.589>, 1.0))","heading_chain":"a Heading that says \"From the Weight to the Loss\"","heading_loop":"a Heading that says \"The Whole Loop\"","heading_numbers":"a Heading that says \"One Step, in Numbers\"","heading_pieces":"a Heading that says \"Three Local Derivatives\"","loss_change":"a Math [text] that says \"$L: 0.143 arrow.r 0.127$\"","n_a":"a Point [green] labelled \"a\" drawn in chain (location=(5.2, 2.2))","n_l":"a Point [red] labelled \"L\" drawn in chain (location=(7.3, 2.2))","n_x":"a Point [blue] labelled \"x\" drawn in chain (location=(0.7, 2.2))","n_z":"a Point [yellow] labelled \"z\" drawn in chain (location=(3.0, 2.2))","num_0":"a Math [text] that says \"$x = 1, quad y = 1, quad w = 0.5, quad b = 0$\"","num_1":"a Math [text] that says \"$z = 0.5 dot.op 1 + 0 = 0.5$\"","num_2":"a Math [text] that says \"$a = sigma(0.5) = 0.622$\"","num_3":"a Math [text] that says \"$frac(partial L, partial a) = 2 (0.622 - 1) = -0.755$\"","num_4":"a Math [text] that says \"$frac(dif a, dif z) = 0.622 (1 - 0.622) = 0.235$\"","num_5":"a Math [text] that says \"$frac(partial L, partial w) = -0.755 dot.op 0.235 dot.op 1 = -0.177$\"","num_6":"a Math [text] that says \"$w arrow.l 0.5 - 0.5 dot.op (-0.177) = 0.589$\"","pieces":"a Derivation [text] that says \"$L &= (a - y)^2 \\ frac(partial L, partial a) &= 2 (a - y) \\ a &= sigma(z) \\ frac(dif a, dif z) &= a (1 - a) \\ z &= w x + b, quad frac(partial z, partial w) = x$\"","recap":"a Block [text] that says \"Forward: weighted sums and bends turn inputs into a prediction. Loss: one number, $L$, says how wrong that prediction is. Backward: the chain rule turns $L$ into a slope for every weight. Step: each weight moves against its own slope, and …\"","result_b":"a Math [text] that says \"$frac(partial L, partial b) = 2 (a - y) dot.op a (1 - a)$\"","result_w":"a Math [text] that says \"$frac(partial L, partial w) = 2 (a - y) dot.op a (1 - a) dot.op x$\"","riding":"a PlotPoint [yellow] labelled \"a\" drawn in sig (target='sig_curve', x=<VariableNumber z_live = 0.589>)","sig":"an Axes (x_range=(-3.2, 3.2), y_range=(-0.1, 1.2), x_ticks_every=1.0)","sig_curve":"a FunctionPlot [blue] drawn in sig (function=<function>, x_range=(-3.0, 3.0))","slope_here":"a TangentLine [yellow] drawn in sig (target='sig_curve', x=0.5, show_dot=False)","target":"a Line [red] labelled \"y = 1\" drawn in sig (start=(-3.0, 1.0), end=(3.0, 1.0), dashed=True)","z_live":"a VariableNumber (initial_value=0.5, format_spec='.3f')"},"beats":[{"start":657.5645208333333,"say":"We need the slope of the loss with respect to one weight. But the weight does not touch the loss directly. Follow the arithmetic forward: the weight and the input make z. The bend turns z into the activation a. And a, compared with the target y, produces the loss.","live":[],"does":[[657.5645208333333,"heading_chain is shown on the screen, written out."],[657.5645208333333,"chain is shown on the screen, written out."],[664.3565208333333,"n_x is shown on the screen, written out."],[667.7235208333333,"f_1 is shown on the screen, written out."],[668.1525208333333,"n_z is shown on the screen, written out."],[669.7315208333333,"f_2 is shown on the screen, written out."],[671.4965208333333,"n_a is shown on the screen, written out."],[675.1195208333332,"f_3 is shown on the screen, written out."],[676.3265208333332,"n_l is shown on the screen, written out."]]},{"start":678.4475208333333,"say":"So nudging the weight nudges z, and that nudges a, and that nudges the loss. Three links in one chain.","live":["chain","heading_chain","n_x","f_1","n_z","f_2","n_a","f_3","n_l"],"does":[[679.7595208333333,"f_1 is indicated — a transient flash."],[680.0517841348993,"f_2 is indicated — a transient flash."],[681.0587362540507,"f_3 is indicated — a transient flash."]]},{"start":686.5530208333333,"say":"Read the chain backwards, which is where the name backpropagation comes from. First, how fast does the loss respond to a? Next, how fast does a respond to z? Finally, how fast does z respond to the weight? Multiply those three local rates together and you have the slope you wanted.","live":null,"does":[[691.6215208333333,"b_3 is shown on the screen, written out."],[695.5455208333333,"b_2 is shown on the screen, written out."],[698.3905208333333,"b_1 is shown on the screen, written out."],[702.5115208333333,"chain_rule is shown on the screen, written out."],[706.8420208333333,"chain moves to a new place on the board."],[706.8420208333333,"chain_rule is hidden from the screen — left the board."],[706.8420208333333,"heading_chain is hidden from the screen — left the board."]]},{"start":708.0420208333333,"say":"Each of those three rates is a one-line derivative. The loss is a minus y, all squared, so differentiating it with respect to a gives two times a minus y.","live":["chain","n_x","f_1","n_z","f_2","n_a","f_3","n_l","b_3","b_2","b_1"],"does":[[708.0420208333333,"heading_pieces is shown on the screen, written out."],[711.7695208333333,"pieces is shown on the screen, written out."],[716.6105208333333,"pieces is shown on the screen, written out."],[716.6105208333333,"b_3 is indicated — a transient flash."]]},{"start":719.1960208333333,"say":"The activation is sigma of z, and the sigmoid's slope is the tidy expression we met earlier: a times one minus a.","live":["chain","n_x","f_1","n_z","f_2","n_a","f_3","n_l","b_3","b_2","b_1","heading_pieces"],"does":[[719.7305208333333,"pieces is shown on the screen, written out."],[723.5145208333333,"pieces is shown on the screen, written out."],[723.5145208333333,"b_2 is indicated — a transient flash."]]},{"start":728.5265208333333,"say":"And z is w x plus b, so differentiating with respect to w leaves just x. With respect to b it leaves one, because the bias is added on its own.","live":null,"does":[[730.5465208333333,"pieces is shown on the screen, written out."],[733.9015208333333,"b_1 is indicated — a transient flash."]]},{"start":740.4695208333333,"say":"Multiply the three together and there it is. The slope of the loss with respect to this weight is two times a minus y, times a times one minus a, times x. For the bias it is the same product without the x on the end.","live":null,"does":[[740.8175208333333,"result_w is shown on the screen, written out."],[752.3465208333333,"result_b is shown on the screen, written out."],[755.7250208333332,"chain is hidden from the screen — left the board."],[755.7250208333332,"n_x is hidden from the screen — chain left the board."],[755.7250208333332,"f_1 is hidden from the screen — chain left the board."],[755.7250208333332,"n_z is hidden from the screen — chain left the board."],[755.7250208333332,"f_2 is hidden from the screen — chain left the board."],[755.7250208333332,"n_a is hidden from the screen — chain left the board."],[755.7250208333332,"f_3 is hidden from the screen — chain left the board."],[755.7250208333332,"n_l is hidden from the screen — chain left the board."],[755.7250208333332,"b_3 is hidden from the screen — chain left the board."],[755.7250208333332,"b_2 is hidden from the screen — chain left the board."],[755.7250208333332,"b_1 is hidden from the screen — chain left the board."],[755.7250208333332,"heading_pieces is hidden from the screen — left the board."],[755.7250208333332,"pieces is hidden from the screen — left the board."],[755.7250208333332,"result_b is hidden from the screen — left the board."],[755.7250208333332,"result_w is hidden from the screen — left the board."]]},{"start":756.9250208333333,"say":"Now put numbers in. One input equal to one, a target of one, a weight of a half, and a bias of zero. The weighted sum is a half. The sigmoid of a half is zero point six two two, so the network answers zero point six two two when it should answer one.","live":[],"does":[[756.9250208333333,"heading_numbers is shown on the screen, written out."],[756.9250208333333,"sig is shown on the screen, written out."],[757.7265208333333,"sig_curve is shown on the screen, written out."],[757.7265208333333,"num_0 is shown on the screen, written out."],[760.5935208333333,"target is shown on the screen, written out."],[765.3895208333333,"num_1 is shown on the screen, written out."],[767.3395208333333,"riding is shown on the screen, written out."],[767.3395208333333,"num_2 is shown on the screen, written out."]]},{"start":774.9405208333333,"say":"That red gap is the error. a minus y is minus zero point three seven eight, and twice it is minus zero point seven five five. There is the first factor.","live":["num_0","num_1","num_2","sig","heading_numbers","sig_curve","target","riding"],"does":[[775.7875208333332,"gap_seg is shown on the screen, written out."],[781.3835208333333,"num_3 is shown on the screen, written out."]]},{"start":786.8540208333333,"say":"The second factor is the steepness of the sigmoid at this point, a times one minus a, which works out as zero point two three five. You can see it in the tangent line: a gentle slope, so a gentle factor.","live":["num_0","num_1","num_2","num_3","sig","heading_numbers","sig_curve","target","riding","gap_seg"],"does":[[792.6995208333333,"num_4 is shown on the screen, written out."],[796.7985208333333,"slope_here is shown on the screen, written out."],[800.8500208333332,"num_4 moves to a new place on the board."],[800.8500208333332,"num_0 is hidden from the screen — left the board."],[800.8500208333332,"num_1 is hidden from the screen — left the board."],[800.8500208333332,"num_2 is hidden from the screen — left the board."],[800.8500208333332,"num_3 is hidden from the screen — left the board."]]},{"start":801.4500208333333,"say":"Multiply them, together with x, which is one. The slope of the loss with respect to this weight is minus zero point one seven seven.","live":["num_4","sig","heading_numbers","sig_curve","target","riding","gap_seg","slope_here"],"does":[[805.5485208333333,"num_5 is shown on the screen, written out."]]},{"start":812.0925208333333,"say":"It is negative, so raising the weight lowers the loss. With a learning rate of a half the weight moves from zero point five to zero point five eight nine, and the activation climbs the curve toward its target.","live":["num_4","sig","heading_numbers","sig_curve","target","riding","gap_seg","slope_here","num_5"],"does":[[818.2575208333333,"num_6 is shown on the screen, written out."],[823.0875208333333,"riding is redrawn as the numbers it depends on change."],[823.0875208333333,"gap_seg is redrawn as the numbers it depends on change."],[823.0875208333333,"z_live ticks to 0.589."]]},{"start":825.8585208333333,"say":"And the loss falls, from zero point one four three to zero point one two seven. The gap is narrower than it was. That is one training step for one weight, and a network does it for every weight it has, over and over again.","live":["num_4","sig","heading_numbers","sig_curve","target","riding","gap_seg","slope_here","num_5","num_6"],"does":[[826.9035208333333,"loss_change is shown on the screen, written out."],[833.1265208333333,"gap_seg is indicated — a transient flash."],[836.0635208333333,"A box is drawn around loss_change."],[841.3580208333333,"heading_numbers is hidden from the screen — left the board."],[841.3580208333333,"loss_change is hidden from the screen — left the board."],[841.3580208333333,"num_4 is hidden from the screen — left the board."],[841.3580208333333,"num_5 is hidden from the screen — left the board."],[841.3580208333333,"num_6 is hidden from the screen — left the board."],[841.3580208333333,"sig is hidden from the screen — left the board."],[841.3580208333333,"sig_curve is hidden from the screen — sig left the board."],[841.3580208333333,"target is hidden from the screen — sig left the board."],[841.3580208333333,"riding is hidden from the screen — sig left the board."],[841.3580208333333,"gap_seg is hidden from the screen — sig left the board."],[841.3580208333333,"slope_here is hidden from the screen — sig left the board."]]},{"start":842.5580208333333,"say":"So that is the whole loop. Forward, weighted sums and bends turn the inputs into a prediction.","live":[],"does":[[842.5580208333333,"heading_loop is shown on the screen, written out."],[843.8355208333334,"recap is shown on the screen, written out."],[844.9265208333334,"recap (the \"weighted sums and bends\" part) is emphasized."]]},{"start":850.1355208333333,"say":"The loss collapses all of that into one number. The chain rule turns that number into a slope for every weight and every bias in the network.","live":["recap","heading_loop"],"does":[[851.2035208333333,"recap (the \"one number\" part) is emphasized."],[851.2035208333333,"recap (the \"weighted sums and bends\" part) is no longer emphasized."],[854.1295208333333,"recap (the \"one number\" part) is no longer emphasized."],[854.1295208333333,"recap (the \"the chain rule\" part) is emphasized."]]},{"start":860.1625208333332,"say":"Then each weight takes a small step against its own slope, and we go round again. Everything else in deep learning is this loop, done at scale.","live":null,"does":[[862.1365208333333,"recap (the \"against its own slope\" part) is emphasized."],[862.1365208333333,"recap (the \"the chain rule\" part) is no longer emphasized."],[869.2200208333333,"recap (the \"against its own slope\" part) is no longer emphasized."],[869.4700208333333,"heading_loop is hidden from the screen — left the board."],[869.4700208333333,"recap is hidden from the screen — left the board."]]}]}]},"durationSeconds":871,"chapters":[{"title":"From Numbers to a Prediction","startSeconds":0,"narration":"A neural network is a piece of arithmetic with adjustable numbers inside it. Suppose two measurements describe a student: hours of study, and hours of sleep. We want one number out, a predicted exam score. So there are two questions. What arithmetic should we do, and how can the numbers inside it fix themselves when the answer comes out wrong? Draw that arithmetic as a picture. The two inputs sit on the left, and I will call them x one and x two. For our student, x one is three hours of study and x two is eight hours of sleep. Over here is a single unit, and its job is to turn those two numbers into one. Every arrow carries a number of its own, called a weight, and the weight decides how much that input matters. The first weight is zero point five. The second is zero point two. The unit does two things with those numbers. First it adds up the weighted inputs. Then it adds one more number of its own, which belongs to the unit rather than to any input. That number is called the bias, and here it is minus one. It lets the unit shift its answer up or down whatever the inputs happen to be. Written out, that is the whole of one unit. z equals w one x one, plus w two x two, plus b. A weighted sum of the inputs, and then the bias. Put our numbers in: zero point five times three, plus zero point two times eight, minus one. One point five, plus one point six, minus one. The unit's output is two point one. Now clear the arithmetic and watch the weight itself. Raise the first weight to one, and the output climbs to three point six, because study now counts for twice as much. Lower it to zero point one, and the output drops to zero point nine. Those weights are the dials, and everything that follows is about how to set them. One unit is not a network. Start again from the two inputs, and this time give them three units instead of one. Every unit reads both of them, with its own pair of weights and its own bias. Each unit computes its own weighted sum, so the layer turns two numbers into three. Feed those three into one last unit, and out comes the single prediction, y hat. Written as algebra, the whole layer is one line: the vector z equals a matrix W times the vector x, plus the vector b. W holds every weight in the layer, one row for each unit, and b holds every bias. Change a single entry of W and the prediction changes. So the machinery is in place, and two things are missing. We need a number, call it L, that says how wrong the prediction is. And we need a rule that changes every weight and every bias to make that number smaller. Those two are the rest of this lecture."},{"title":"Why a Straight Line Is Not Enough","startSeconds":185.75941666666662,"narration":"A weighted sum is a straight-line formula, and straight lines have a limit. Here are five measurements. As x grows the output falls to zero and then climbs again. Try matching them with one straight line. Tilt the line one way, and the left half is wrong. Tilt it the other way, and the right half is wrong. There is no slope and no intercept that passes through all five points, because the data bends and the line cannot. You might hope that stacking two layers fixes that. Let us check. The first layer computes v x plus c. The second layer takes that answer and computes w times it, plus b. Substitute, and expand. The answer is w v times x, plus w c plus b. That is a straight line again, with a new slope and a new intercept. Two straight layers are one straight layer, and a hundred of them would still be one. So something has to bend. Take the simplest bend there is. It is called ReLU, and it returns its input when the input is positive, and zero otherwise. ReLU of two minus x is the blue piece. It slopes down until x reaches two, and after that it is flat at zero. ReLU of x minus two is the green piece: flat at zero until two, and then rising. Wherever one of them is positive the other is zero, so adding them gives exactly the shape through all five points. One bend per unit, and a layer of units can fold a straight line into very nearly any shape you want. Two bends are standard, and the first is the sigmoid. It squeezes any input at all into the range from zero to one. A large positive input gives almost one, a large negative input gives almost zero, and an input of zero gives exactly one half. That is useful when the output should read as a probability. We will also need its slope later, and the slope is unusually tidy: sigma prime equals sigma times one minus sigma. Where the curve is steep that number is large, and out at the flat ends it is almost nothing. The other standard bend is the one we just used, ReLU. It is not smooth at zero, but it is cheap to compute, and its slope is simply one on the right and zero on the left. Either way, a unit now does two things: it forms a weighted sum, and then it bends the result."},{"title":"A Number for Being Wrong","startSeconds":345.68372916666664,"narration":"A network has to know when it is wrong, and by how much. Here are three training examples. The red dots are the correct answers, recorded in advance, and the blue line is what the model currently predicts. For each example, subtract the target from the prediction. That difference is the error. The first prediction is half a unit too low, the second is a whole unit too high, and the third is half a unit too low again. Errors come with signs, and if we simply added them a prediction that is too high would cancel one that is too low. So square each error first. Squaring makes every contribution positive, and it punishes a big miss much more than a small one. Add the three squares. A quarter, plus one, plus a quarter, is one and a half. Then divide by the number of examples. The mean squared error is zero point five. In general, then: average the squared difference between prediction and target over all n examples. One number, for the whole data set, and it is the number we are going to make as small as we can. Now here is the change of view that the rest of the lecture rests on. The data are fixed. We cannot alter a single target. The only things we can move are the weights, so the loss is really a function of the weights. To see that clearly, keep one single example: input equal to one, target equal to one. With one weight and no bias the prediction is just w, so the loss is w minus one, all squared. Sweep the weight and watch the loss. At w equals two the loss is one. Bring the weight down and the loss falls, reaches zero at w equals one, and climbs the far wall again. So training means finding the bottom of that bowl."},{"title":"Walking Downhill","startSeconds":461.83008333333333,"narration":"We have a bowl, and we need to walk to the bottom of it. But we are not allowed to see the whole bowl. At any moment the network knows only the weight it currently has, and how steeply the loss is rising or falling right there. That steepness is exactly what a derivative is. The derivative of w minus one squared is two times w minus one. At w equals two it comes to two, a positive number, and the tangent line there rises to the right. A positive slope means the loss grows as the weight grows. So to make the loss smaller, move the weight the other way. Subtract something proportional to the slope, and that is the whole of gradient descent in one line. New weight equals old weight, minus eta times the derivative. Eta is a small positive number called the learning rate, and it decides how far each step carries us. Take eta equal to zero point three. The step is zero point three times two, which is zero point six, so the weight goes from two to one point four. Watch the tangent flatten as it lands. Do it again. At one point four the slope is only zero point eight, so the step is smaller and the weight moves to one point one six. Once more, and it reaches one point zero six. The steps shrink by themselves, because near the bottom there is hardly any slope left to multiply. Had the slope been negative, the minus sign would have pushed the weight up instead. Either way we go downhill. But eta matters. Suppose at w equals two we had used one point one. The step would be two point two, the weight would land at minus zero point two, and the loss there is worse than where we began. A real network has thousands of weights, not one, and nothing changes except that there is now a slope for each of them. Here is a loss with two weights, drawn as a contour map. Every ring is a set of weights giving the same loss, and the bottom of the bowl is inside the smallest ring. Hold w two still and ask how the loss changes as w one moves on its own. That is the partial derivative with respect to w one, and taking a step against it gives the blue arrow. Hold w one still instead, and you get the green arrow. Collect the partial derivatives into one list and you have the gradient. Here it is minus four and six. The gradient points in the direction the loss increases fastest, straight across the contour lines. So step the opposite way. Add the two component steps, and the red arrow is the diagonal of their parallelogram: that is where we land. The rule reads as before, with the gradient standing in for the single derivative. And repeat. Each step crosses to an inner ring, the arrows shorten as the ground flattens out, and the weights settle near the bottom. That is gradient descent. What is left is the hard part. In a real network, with layers feeding layers, how do we actually compute those partial derivatives?"},{"title":"Backpropagation: The Chain Rule Doing the Work","startSeconds":657.5645208333333,"narration":"We need the slope of the loss with respect to one weight. But the weight does not touch the loss directly. Follow the arithmetic forward: the weight and the input make z. The bend turns z into the activation a. And a, compared with the target y, produces the loss. So nudging the weight nudges z, and that nudges a, and that nudges the loss. Three links in one chain. Read the chain backwards, which is where the name backpropagation comes from. First, how fast does the loss respond to a? Next, how fast does a respond to z? Finally, how fast does z respond to the weight? Multiply those three local rates together and you have the slope you wanted. Each of those three rates is a one-line derivative. The loss is a minus y, all squared, so differentiating it with respect to a gives two times a minus y. The activation is sigma of z, and the sigmoid's slope is the tidy expression we met earlier: a times one minus a. And z is w x plus b, so differentiating with respect to w leaves just x. With respect to b it leaves one, because the bias is added on its own. Multiply the three together and there it is. The slope of the loss with respect to this weight is two times a minus y, times a times one minus a, times x. For the bias it is the same product without the x on the end. Now put numbers in. One input equal to one, a target of one, a weight of a half, and a bias of zero. The weighted sum is a half. The sigmoid of a half is zero point six two two, so the network answers zero point six two two when it should answer one. That red gap is the error. a minus y is minus zero point three seven eight, and twice it is minus zero point seven five five. There is the first factor. The second factor is the steepness of the sigmoid at this point, a times one minus a, which works out as zero point two three five. You can see it in the tangent line: a gentle slope, so a gentle factor. Multiply them, together with x, which is one. The slope of the loss with respect to this weight is minus zero point one seven seven. It is negative, so raising the weight lowers the loss. With a learning rate of a half the weight moves from zero point five to zero point five eight nine, and the activation climbs the curve toward its target. And the loss falls, from zero point one four three to zero point one two seven. The gap is narrower than it was. That is one training step for one weight, and a network does it for every weight it has, over and over again. So that is the whole loop. Forward, weighted sums and bends turn the inputs into a prediction. The loss collapses all of that into one number. The chain rule turns that number into a slope for every weight and every bias in the network. Then each weight takes a small step against its own slope, and we go round again. Everything else in deep learning is this loop, done at scale."}]}}
