Part 1 left you with a model full of random numbers. It can take an image and confidently tell you it's the digit 4 when it's a 9. Training is the process of dragging those random numbers towards ones that are usually right — and every single step of that process needs a gradient: a number for each parameter that says "increase me to make the loss worse, decrease me to make it better, and by roughly this much."
For a model with a hundred million parameters, computing those gradients by hand is not on the table. Autograd is the part of PyTorch that does it for you — automatically, exactly, and fast enough that you never think about it. Understanding how it works is the difference between training models and debugging them by superstition.
The calculus you were promised you'd never need
Picture the loss as a landscape. Every parameter is an axis; the height at any point is how wrong the model is with those parameter values. Training is walking downhill. The gradient is the arrow pointing in the steepest uphill direction, so you step the opposite way:
Do that a few hundred thousand times and, if the landscape is kind and the step size is sane, you end up somewhere low. The entire job of autograd is to hand you ∇L — the gradient of the loss with respect to every parameter — after each forward pass.
The graph builds itself as your code runs
The moment a tensor has requires_grad=True, PyTorch starts watching. Every operation you perform on it creates a new node in a computational graph: the result tensor, a reference to the function that produced it (its grad_fn), and pointers back to the inputs. There is no compile step and no graph definition phase — the graph is a side effect of running your forward pass, which is exactly why you can put if statements, loops, and print calls inside a model and it all just works.
import torchx = torch.tensor(2.0, requires_grad=True)y = x ** 3 + 2 * xprint(y) # tensor(12., grad_fn=<AddBackward0>)print(y.grad_fn) # <AddBackward0 object ...># Walk one step back into the graph by hand:print(y.grad_fn.next_functions) # the pow and mul nodes feeding the add
For y = x³ + 2x, the graph PyTorch just recorded looks like this — a small pipeline of operations, each remembering what it needs to compute its own local derivative later:
backward() walks it in reverse
Calling y.backward() starts at y with a seed gradient of 1 and traverses the graph towards the leaves. At each node it multiplies the incoming gradient by that operation's local derivative — the chain rule, applied mechanically — and when two paths reach the same tensor, their contributions add. The final number lands in each leaf tensor's .grad.
import torchx = torch.tensor(2.0, requires_grad=True)y = x ** 3 + 2 * xy.backward()print(x.grad) # tensor(14.)# Check it: d/dx (x³ + 2x) = 3x² + 2 -> 3·4 + 2 = 14print(3 * x.item() ** 2 + 2) # 14.0
That's the trick in full. It scales from this three-node graph to a Transformer with billions of parameters without any change in principle — just more nodes, and the arithmetic running on a GPU.
Gradient descent, in motion
Below is the simplest possible optimisation — minimising L(θ) = θ² starting from θ = 10 — run with a sensible learning rate and with one that's slightly too large. Hover the lines. The takeaway that will save you the most time: a diverging loss is almost never a broken model, it's a learning rate that's too high.
Minimising θ² from θ₀ = 10
η = 0.1 slides smoothly to zero. η = 1.1 overshoots the minimum by more each step and explodes. Same model, same code — only the step size changed.
| step | η = 0.1 (stable) | η = 1.1 (diverges) |
|---|---|---|
| 0 | 100 | 100 |
| 1 | 64 | 144 |
| 2 | 40.96 | 207.4 |
| 3 | 26.2 | 298.6 |
| 4 | 16.8 | 430 |
| 5 | 10.7 | 619.2 |
| 6 | 6.87 | 891.6 |
| 7 | 4.4 | 1284 |
The training loop, line by line
Every supervised training loop in PyTorch — from a tutorial notebook to a foundation-model run — is a variation on these five lines. The optimizer holds a reference to the model's parameters and knows one update rule (SGD, Adam, …); you hand it gradients, it takes the step.
import torch.nn as nnfrom torch.optim import Adammodel = MLP(784, 128, 10)opt = Adam(model.parameters(), lr=1e-3)loss_fn = nn.CrossEntropyLoss()for images, labels in loader:opt.zero_grad() # 1. clear last step's gradientslogits = model(images) # 2. forward pass — builds the graphloss = loss_fn(logits, labels)loss.backward() # 3. backward pass — fills every .gradopt.step() # 4. update parameters using .grad# 5. repeat a few hundred thousand times
opt.zero_grad()— gradients accumulate by default. Skip this and step N's gradient is added on top of step N−1's, and your training quietly falls apart.loss.backward()— frees the graph as it walks it. Call it twice on one forward pass and PyTorch raises an error unless you asked forretain_graph=True.opt.step()— reads.grad, mutates the parameters in place, never touches the graph.with torch.no_grad():— switch graph-building off entirely for evaluation and inference. Saves memory and time, and you weren't going to callbackward()anyway.
Turning tracking off on purpose
Two tools stop autograd following a tensor. tensor.detach() returns a view that shares the same storage but sits outside the graph — reach for it when you want a value for logging or a metric but not its history. with torch.no_grad(): disables recording for its whole block, which is the standard wrapper around a validation loop and any inference code.
And that's the core of how PyTorch learns: a forward pass quietly records a graph, backward() walks it in reverse to produce an exact gradient for every parameter, and the optimizer turns those gradients into a step downhill. Everything above this in the stack — schedulers, mixed precision, distributed training, torch.compile — is an optimisation of this loop, never a replacement for it. Part 3 goes up a level: how to structure the model itself so it survives contact with a real project.
References
- [1]Autograd mechanics · PyTorch documentationHow the graph is recorded, when it's freed, and the in-place-operation rules.
- [2]A Gentle Introduction to torch.autograd · PyTorch tutorials
- [3]torch.autograd.backward · PyTorch documentation
- [4]torch.optim — optimizers and the update step · PyTorch documentation
- [5]Learning representations by back-propagating errors · Rumelhart, Hinton & Williams, Nature 1986The original backpropagation paper.


