The first time you open a real PyTorch codebase, it looks like a lot. Modules importing modules, forward methods, .cuda() sprinkled everywhere, a training loop with cryptic four-letter calls. It feels like there must be a hundred concepts to learn before any of it makes sense.
There aren't. PyTorch is a numerical computing library that turned out to be very good at deep learning. Take away the pretrained models and the training helpers and what's left is a single data structure — the tensor — and a large, well-organised set of operations on it. A neural network is just a function built from those operations. Training is just calculus on that function. Get comfortable with the tensor and the rest of the library stops being intimidating and starts being predictable.
This series builds the whole picture from that starting point. This first part is the tensor itself and how a network is assembled from it.
Everything is a tensor
A tensor is an n-dimensional array where every element has the same type. That's the entire definition. The dimensionality is the only thing that changes as you go up:
- A scalar — a single loss value — is a 0-D tensor.
- A vector — the 10 class scores for one image — is 1-D, shape
(10,). - A matrix — a batch of those score vectors — is 2-D, shape
(32, 10). - A batch of RGB images is 4-D, shape
(32, 3, 224, 224): batch, channels, height, width. - A batch of token embeddings is 3-D: batch, sequence length, embedding dimension.
If you've used NumPy, this is familiar ground — the API was deliberately kept close, and most np. code has a one-to-one torch. translation. Try the snippet below; every line runs.
import torchx = torch.tensor([[1.0, 2.0], [3.0, 4.0]])print(x.shape) # torch.Size([2, 2])print(x.dtype) # torch.float32print(x.mean()) # tensor(2.5000)print(x @ x) # matrix multiplyprint(x.sum(dim=0)) # column sums -> tensor([4., 6.])print(x.reshape(4)) # tensor([1., 2., 3., 4.])print(x.T) # transpose
The two attributes that make it deep learning
A NumPy array has a shape and a dtype. A PyTorch tensor has two more attributes, and they are the entire reason the library exists.
device says where the data physically lives: cpu (your RAM) or cuda:0 (the first GPU's memory). Operations run wherever their inputs live, so moving a tensor to the GPU with .to("cuda") is literally what makes the maths run on the accelerator. Mix devices in one operation and PyTorch stops you with an error — a rite of passage covered properly in part 5.
requires_grad is a boolean flag that says "track every operation you do to me, because I'm going to want gradients later." Set it on a tensor and PyTorch quietly starts recording a graph of everything that happens downstream. That graph is how a network learns, and it's the whole subject of part 2.
import torchw = torch.randn(3, 3, requires_grad=True)print(w.device) # cpuprint(w.requires_grad) # True# A parameter you'll optimise: float, on the right device, tracked.# A batch of input data: float, same device, NOT tracked.batch = torch.randn(16, 3)print(batch.requires_grad) # False
The device attribute isn't a formality. The same operation — a matrix multiply — runs orders of magnitude faster on a GPU once the matrices are big enough to hide the cost of shipping them there. Below is A @ B for square matrices of growing size; hover the lines. For small tensors the CPU wins (no transfer overhead); the GPU pulls ahead fast and then runs away.
Illustrative — exact numbers depend on hardware
Below ~256, the GPU's transfer and launch overhead makes it slower. Past ~1024 it isn't close. This gap is the entire reason `.to("cuda")` exists.
| matrix size N | CPU | GPU (cuda) |
|---|---|---|
| 128 | 0.2 | 0.35 |
| 256 | 1.4 | 0.4 |
| 512 | 9.8 | 0.7 |
| 1024 | 78 | 2.1 |
| 2048 | 620 | 9.5 |
| 4096 | 4900 | 62 |
Shapes are where you'll actually spend your time
Here's the honest truth about writing models: the maths is rarely the hard part. The hard part is getting a (32, 3, 224, 224) tensor to line up with a layer that expected (32, 150528). Shape errors are the mosquitoes of deep learning — not dangerous, endlessly annoying, and you will get bitten daily until you learn the patterns.
The key rule is broadcasting: when two tensors have different shapes, PyTorch tries to make them compatible by aligning dimensions from the right and stretching any dimension of size 1 to match its partner. It never copies data to do this — it's a view trick — so it's free.
| Operation | Result | What happened |
|---|---|---|
(32, 10) + (10,) | (32, 10) | the vector is added to every row |
(32, 10) * (32, 1) | (32, 10) | each row scaled by its own value |
(8, 1, 6) + (1, 5, 6) | (8, 5, 6) | both size-1 dims stretch |
(32, 10) + (32,) | error | trailing dims 10 and 32 don't match |
Building a layer from scratch
A neural network layer is nothing more exotic than a function from tensors to tensors that carries some learnable numbers of its own. The workhorse is the linear (or fully-connected, or dense) layer, which multiplies its input by a weight matrix and adds a bias:
That's it. You can write the whole thing with the operators you already have:
import torchbatch, in_features, out_features = 4, 3, 2x = torch.randn(batch, in_features)W = torch.randn(out_features, in_features)b = torch.randn(out_features)y = x @ W.T + b # broadcasting adds b to every rowprint(y.shape) # torch.Size([4, 2])# A non-linearity turns a stack of linear layers into something# more expressive than one big linear layer.h = torch.relu(y)print(h.min() >= 0) # tensor(True)
Stack two of these with a non-linearity between them and you have a multi-layer perceptron — a universal function approximator, in principle. The only missing ingredient is a way to choose good values for W and b instead of random ones.
The same thing, the grown-up way
Writing layers by hand gets old fast: you'd be tracking every weight tensor manually, moving each one to the GPU, remembering which ones to save. torch.nn does the bookkeeping. nn.Linear holds its weight and bias as nn.Parameter tensors (which flip requires_grad=True for you), and wrapping your model in nn.Module gets you .parameters(), .to(device), .train() / .eval(), and checkpoint saving for free.
import torch.nn as nnclass MLP(nn.Module):def __init__(self, in_dim, hidden, out_dim):super().__init__()self.net = nn.Sequential(nn.Linear(in_dim, hidden),nn.ReLU(),nn.Linear(hidden, out_dim),)def forward(self, x):return self.net(x)model = MLP(784, 128, 10) # 28x28 image in, 10 classes outn_params = sum(p.numel() for p in model.parameters())print(f"{n_params:,} parameters") # 101,770 parameters
Every serious model — a ResNet, a Transformer, Stable Diffusion's U-Net — is this same idea scaled up: nn.Module objects holding other nn.Module objects, with a forward method that says how tensors flow through them. The object graph is small:
What you actually built
- Tensors hold the numbers, plus a shape, a dtype, a device, and a gradient-tracking flag.
- Operations on tensors — matmul,
relu,sum— are the arithmetic a network is made of. - Broadcasting is how tensors of different shapes combine without copying; shape errors name the dimension that failed.
nn.Parameteris a tensor a layer owns and wants gradients for.nn.Modulebundles parameters with aforwardmethod into a reusable unit — and a model is just Modules nested inside Modules.
References
- [1]torch.Tensor — API reference · PyTorch documentation
- [2]Tensors — Learn the Basics · PyTorch tutorials
- [3]Broadcasting semantics · PyTorch documentationThe exact rules for how mismatched shapes are reconciled.
- [4]Building models with nn.Module — Introduction to PyTorch · PyTorch tutorials
- [5]PyTorch: An Imperative Style, High-Performance Deep Learning Library · Paszke et al., NeurIPS 2019The design paper — why PyTorch is built the way it is.


