1 article, newest first.
The five-line loop from Part 2 trains a toy. A real run adds a device, mixed precision for a near-free 2×, a learning-rate schedule, gradient clipping, and checkpoints you can actually resume from. Here's the whole thing.