Opening the archive
Opening the archive
Everything published here so far. Filter by topic, or sort by what gets read most.
Series · PyTorchClear
You bought a fast GPU and it's sitting at 30% utilisation. The model isn't the bottleneck — the input pipeline is. Here's how Dataset and DataLoader work, and how to stop starving the accelerator.
The five-line loop from Part 2 trains a toy. A real run adds a device, mixed precision for a near-free 2×, a learning-rate schedule, gradient clipping, and checkpoints you can actually resume from. Here's the whole thing.
nn.Sequential is training wheels. Real models are custom Modules — with skip connections, shared weights, and branches. Here's how nn.Module actually works, and every footgun it hides.
Every training loop on Earth is basically the same four lines. The one that does the real work — loss.backward() — is a graph PyTorch builds while your code runs and then walks backwards. Let's take it apart.
People think PyTorch is enormous. It isn't. It's one data structure and a pile of operations on it — and once that clicks, every model you'll ever read looks obvious. Here's the ground floor.