The Vision Transformer (ViT) is a neural network architecture that applies the transformer self-attention mechanism directly to sequences of fixed-size image patches, treating each patch embedding as a token analogous to a word in natural-language processing. Introduced by Dosovitskiy et al. (2020), ViT demonstrated that pure attention-based models can match or exceed convolutional networks on image classification benchmarks when pre-trained on sufficiently large datasets.
Content
- Prior to ViT, convolutional neural networks dominated computer vision. The key insight of Dosovitskiy et al.’s “An Image is Worth 16x16 Words” (2020, ICLR 2021) was that the NLP transformer architecture, when applied to image patches, could match CNN performance on ImageNet with far less domain-specific engineering — but only when pre-trained on very large datasets. Earlier works (Image Transformer, AxialTransformer) applied attention to images but retained convolutional elements; ViT was the first clean pure-attention vision model to scale competitively.
- The ViT pipeline begins by splitting an H×W image into N patches of size P×P, giving N = HW/P² tokens. Each patch is flattened and projected to dimension D via a learned linear embedding. A learnable [CLS] token is prepended, and 1D positional embeddings are added. The sequence then passes through L transformer encoder layers, each containing multi-head self-attention (complexity O(N²D)) and a feed-forward network. The [CLS] token representation is used for classification. Variants include DeiT (data-efficient training with distillation), Swin Transformer (hierarchical local windows), and BEiT (masked image modelling pre-training).
- ViT matters because it unifies the architecture for vision and language under the transformer framework, enabling parameter sharing and joint pre-training that were impractical with CNNs. The Swin Transformer’s hierarchical design made ViT competitive for dense prediction tasks (detection, segmentation) that require multi-scale features. ViT-based encoders power most modern text-to-image models (Stable Diffusion uses a CLIP ViT image encoder), video understanding models (ViViT), and medical imaging systems.
- By 2024–2025, ViT has become the default image encoder in frontier AI systems. Meta’s DINOv2 demonstrated high-quality self-supervised ViT features usable across dozens of vision tasks without fine-tuning. Google’s SoViT investigated optimal patch sizes and training objectives. The Diffusion Transformer (DiT) replaced the U-Net in Stable Diffusion 3 and OpenAI Sora’s video generation pipeline. Scaling laws for ViT models have been published confirming that compute-optimal training follows similar power laws to language models.