Home › Lectures › 3Blue1Brown — Neural networks

Transformers, the tech behind LLMs

An overview of how transformers process text, predict the next token, and turn those predictions into generated language.

3Blue1Brown⏱ 27 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

The lecture introduces transformers as neural networks used in language, image, and audio models. It explains how a language model generates text by repeatedly predicting a probability distribution for the next token, sampling a token, adding it to the input, and predicting again. A chatbot can use the same process after receiving instructions and a user prompt.

The input is split into tokens, which are converted into vectors by an embedding matrix. Attention blocks let those vectors exchange context, while feed-forward layers transform them independently. Repeated layers build richer representations; a fixed context size limits how much text the model can consider. At the end, an unembedding matrix maps the final vector to scores for possible next tokens, and softmax turns the scores into probabilities.

The lecture also reviews the deep-learning foundations behind this process: models learn weights from examples through backpropagation, and much of their computation can be expressed as matrix multiplication. It introduces embeddings as vectors whose directions can capture semantic relationships, dot products as a measure of alignment, and temperature as a control over how concentrated or varied the sampling distribution is. These concepts prepare the ground for studying attention in more detail.

Key ideas

Sample questions

In a language model, each word is represented by a vector in a high-dimensional space. What relationship would you expect between the vectors for two words with similar meanings?

  1. ATheir distance has no connection to their meanings.
  2. BThey are always farther apart than vectors for unrelated words.
  3. CThey must have exactly the same coordinates.
  4. DThey tend to be close to each other in that space.
Show answer

Correct answer: D. The vectors are intended to encode meaning, so semantic similarity is reflected in spatial proximity.

In a language-modeling approach that uses a flexible structure with adjustable parameters, why are many examples used instead of writing explicit code for every step of the task?

  1. AThe examples provide experience from which the adjustable parameters can be tuned, rather than requiring every procedure to be specified as a rule.
  2. BThe examples are used only to convert words into vectors, while task behavior is fully specified by hand.
  3. CThe examples eliminate the need for any adjustable parameters in the model.
  4. DThe examples guarantee that the model will follow a fixed procedure that was explicitly coded in advance.
Show answer

Correct answer: A. This approach relies on tuning a flexible model using many examples, rather than spelling out the task as a hand-written procedure.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →