Image captioning is the artificial-intelligence task of generating a natural-language description of the content of an image. It sits at the intersection of computer vision and natural-language generation, typically pairing a visual encoder that extracts image features with a language decoder that produces a fluent sentence. Modern systems use attention mechanisms and large vision-language models to ground the generated text in salient regions of the image.

Overview

  • Image captioning bridges perception and language, requiring a model both to recognise objects, attributes and relationships in a scene and to express them coherently. It evolved from encoder-decoder pipelines with convolutional vision backbones to attention-based and large multimodal vision-language architectures.

Mechanisms

  • Visual feature extraction with convolutional or transformer encoders
  • Sequence generation by a language decoder
  • Attention that grounds words in image regions
  • Evaluation with metrics comparing generated and reference captions

Applications

  • Alt-text generation for web accessibility
  • Media indexing and image search
  • Assistive technology for visually impaired users
  • Dataset annotation and content moderation

Provenance