The Transformer: From Recurrence Failure to Attention

About this lecture

Recurrent networks read a sentence one position at a time, which makes them slow to train and poor at relating words that sit far apart. This lecture builds the Transformer out of that failure, following the 2017 paper Attention Is All You Need. We begin with queries, keys and values, assemble scaled dot-product attention term by term, and see why the scores are divided by the square root of the key dimension. Multi-head attention follows as several parallel projections into narrower subspaces, sinusoidal positional encodings restore the order that attention throws away, and the encoder and decoder stacks are then assembled from those parts with residual connections, layer normalisation and a causal mask. We close on the paper's own comparison of path length, parallelism and training cost, and the translation results that made the design stick. Matrix multiplication and a first course in neural networks are assumed; nothing else is.

Transcript

Loading discussion…