A function, not a brain
The neural network is a function — a parameterized mapping from inputs to outputs. Given an input vector x, it produces an output vector y, with the exact shape of the mapping determined by a set of numerical weights, often billions of them. Training a neural network is the process of finding values for those weights that make the function behave the way we want.
That is the entire skeleton. The "neural" metaphor is evocative but not load-bearing. A convolutional network is not, despite the name, imitating the visual cortex. A Transformer is not, despite the occasional slogan, imitating neuroscience. Both are functions with specific architectural biases, chosen because they make certain problems easy to learn.
The loss landscape
To say a network "learns" is to say we adjust its weights to make a particular number — the loss — small. The loss is a scalar that measures, on a given training example, how far the network's output is from the answer we wanted. For a classifier, it might be the cross-entropy; for a regressor, the squared error. It does not matter which. What matters is that the loss is differentiable with respect to every weight in the network.
That differentiability is what unlocks the entire enterprise. It means we can ask, for every single weight in the network, "if I nudged this up a little, would the loss go up or down, and by how much?" That gradient — a vector with one entry per weight — points in the direction of steepest decrease in the loss. We then take a small step in that direction, recompute, and repeat. This is gradient descent.
θ ← θ − η · ∇L(θ)
Read that line carefully. It is, with very few embellishments, the entire training loop of every neural network ever trained. The cleverness is in the details — what is L, what is η, how do we estimate the gradient cheaply, what architectural choices make the landscape nice — but the core idea fits in one equation.
Backprop, intuitively
The trouble with the line above is that L depends on the weights through a composition of millions of operations. Computing the gradient naively would require, for each weight, a separate forward pass through the network. That is impossibly expensive.
Backpropagation solves this by working backwards. Starting from the loss at the output, we use the chain rule of calculus to propagate the gradient one layer at a time. Each layer's gradient depends only on the layer's local computation and the gradient arriving from the layer above. The cost is roughly twice a forward pass — a remarkable bargain.
Backprop is not an algorithm. It is a direct consequence of the chain rule applied to a layered function — and the layered structure of neural networks exists, primarily, to make that application efficient. — DR. MAYA CHEN · TORONTO
Why depth?
Why do "deep" networks work better than shallow ones? The honest answer is that we do not fully know, but we have a few clues. Depth gives a network the ability to compose features: early layers learn edges, middle layers learn textures and parts, late layers learn objects. Each composition step refines the representation. Theoretically, a sufficiently wide one-layer network can approximate any function, but it would need an astronomical number of neurons to do so.
Practically, depth turns out to be the most efficient prior we have ever found for learning from natural data. We do not fully understand why. That gap — between the empirical success of depth and our theoretical grasp of it — is one of the great open questions in the field.
The limits of the picture
The view of deep learning as "just optimization" is necessary, but not sufficient. It tells you how the weights move, but not what they represent. It tells you that training converges, but not why some architectures generalize to new data and others memorize. It tells you how to compute a gradient, but not what to do when the gradient is zero and the loss is still high.
The frontier of deep learning research is, increasingly, the frontier of understanding what these systems learn. Not just how to train them faster, but what it means that they learn what they learn — and what it would take to make them learn something different.
Filed under: Deep Learning · Optimization · Theory
Cite as: Chen, M. (2026). Deep Learning: how neural networks actually learn. Signal, Vol. 1, Issue 11.