From pixels to meaning
An image, to a computer, is a three-dimensional array of numbers: height, width, and three color channels. There is no "cat" in the array, no "tree", no "horizon". Just numbers. The entire problem of computer vision is to map these numbers to something semantically useful — a category, a bounding box, a depth map, a caption, or, increasingly, another image.
This mapping is the oldest sub-problem of modern AI. It is also, for a long time, the hardest. The raw pixel space of a single 1024×1024 photograph contains on the order of three million numbers, and the mapping from those numbers to "this is a golden retriever catching a frisbee at sunset" is non-linear, hierarchical, and full of context.
The convolutions that won
The breakthrough that broke the field open was the convolutional neural network — an architecture that, almost by accident, encoded the right prior for visual data. A convolution slides a small filter across the image, applying the same weights everywhere. This gives the network two crucial properties: translation equivariance (a cat in the top-left is the same as a cat in the bottom-right) and locality (neighboring pixels are more related than distant ones).
Stack many of these convolutions, intersperse them with pooling operations that shrink the spatial dimensions, and you get the canonical architecture that dominated the 2010s: AlexNet, VGG, ResNet, EfficientNet. Trained on millions of labeled images, these networks learned to classify, localize, and segment the visual world with superhuman accuracy.
Attention sees too
Around 2020, vision researchers began to ask an uncomfortable question: do we really need convolutions? The Transformer, originally designed for sequences, treats its input as a set of tokens that attend to each other. If you slice an image into 16×16 patches and treat each patch as a token, the same machinery applies.
The result was the Vision Transformer (ViT) — and, to many researchers' surprise, it worked as well as the best convolutional networks, often better, when given enough data. Today, almost every frontier vision system is a hybrid: convolutional inductive biases for early layers, attention for global reasoning, and increasingly, the two are merged into unified architectures.
The lesson of the 2020s was not that convolutions were wrong, or that attention was right. It was that the right answer depends on how much data you have, and how willing you are to let the network learn the inductive bias itself. — LIAM PARK · BERLIN
Beyond classification
Classification — "is there a cat in this image" — was the canonical task that drove early progress. But it is also the most trivial. Today's vision systems do far more: they detect objects (where is the cat?), segment them (which pixels are cat?), estimate depth and pose, track them across video, and increasingly, generate them from text.
Diffusion models extended the visual revolution beyond understanding into creation. Systems like Stable Diffusion, DALL·E, and Sora turned text into photoreal images and coherent video. The vision model is no longer just an analyzer — it is a synthesizer. The boundary between "vision" and "generation" is dissolving.
What vision is for
The applications keep multiplying. Medical imaging reads X-rays and pathology slides. Autonomous vehicles read streets. Agricultural drones read crop health. Satellite systems read deforestation. Industrial inspection reads manufacturing defects. Robotics reads the physical world in real time.
But the deepest shift is the most subtle: vision is becoming the primary input modality for AI in the physical world. Most of what we want machines to do — drive, deliver, manufacture, operate, assist — happens in a visual environment. Computer vision is, increasingly, the sense organ of every other AI capability.
Filed under: Computer Vision · CNN · Vision Transformers · Multimodal
Cite as: Park, L. (2026). Computer Vision: teaching machines to see. Signal, Vol. 1, Issue 10.