A visual representation is a learned encoding of image or video content into a feature vector or embedding that captures semantic and structural properties useful for downstream tasks. Such representations, produced by convolutional or transformer-based encoders, support classification, retrieval, detection, and multimodal alignment. The quality of a visual representation determines transferability and sample efficiency across vision applications.

Content

  • Visual representations are learned through supervised, self-supervised (e.g. contrastive or masked-image modelling), or multimodal objectives. Good representations are linearly separable, disentangle nuisance factors, and transfer to unseen domains; they form the latent backbone shared between perception, generation, and cross-modal retrieval pipelines.