Everything tagged #pytorch
5 articles, newest first.
Datasets and DataLoaders: Keeping the GPU Fed
You bought a fast GPU and it's sitting at 30% utilisation. The model isn't the bottleneck — the input pipeline is. Here's how Dataset and DataLoader work, and how to stop starving the accelerator.
Training for Real: GPUs, Mixed Precision, and Schedules
The five-line loop from Part 2 trains a toy. A real run adds a device, mixed precision for a near-free 2×, a learning-rate schedule, gradient clipping, and checkpoints you can actually resume from. Here's the whole thing.
nn.Module: Building Networks That Hold Together
nn.Sequential is training wheels. Real models are custom Modules — with skip connections, shared weights, and branches. Here's how nn.Module actually works, and every footgun it hides.
Autograd: How PyTorch Learns
Every training loop on Earth is basically the same four lines. The one that does the real work — loss.backward() — is a graph PyTorch builds while your code runs and then walks backwards. Let's take it apart.
PyTorch: From Tensor to Neural Network
People think PyTorch is enormous. It isn't. It's one data structure and a pile of operations on it — and once that clicks, every model you'll ever read looks obvious. Here's the ground floor.