Image captioning is the artificial-intelligence task of generating a natural-language description of the content of an image. It sits at the intersection of computer vision and natural-language generation, typically pairing a visual encoder that extracts image features with a language decoder that produces a fluent sentence. Modern systems use attention mechanisms and large vision-language models to ground the generated text in salient regions of the image.
Overview
- Image captioning bridges perception and language, requiring a model both to recognise objects, attributes and relationships in a scene and to express them coherently. It evolved from encoder-decoder pipelines with convolutional vision backbones to attention-based and large multimodal vision-language architectures.
Mechanisms
- Visual feature extraction with convolutional or transformer encoders
- Sequence generation by a language decoder
- Attention that grounds words in image regions
- Evaluation with metrics comparing generated and reference captions
Applications
- Alt-text generation for web accessibility
- Media indexing and image search
- Assistive technology for visually impaired users
- Dataset annotation and content moderation