Video understanding is the field of artificial intelligence concerned with extracting semantic meaning from video, including recognising objects, actions, events, and their temporal relationships across frames. Unlike single-image analysis, it must model motion, temporal context, and long-range dependencies to interpret what is happening over time. Modern approaches combine spatial feature extraction with temporal modelling using recurrent, 3D-convolutional, and transformer-based architectures, increasingly fused with language for captioning, retrieval, and question answering.
Overview
- Video understanding extends computer vision into the temporal dimension, where the central difficulty is jointly modelling appearance and motion across many frames while remaining computationally tractable. Early methods relied on hand-crafted motion features and optical flow; the deep-learning era introduced two-stream networks, 3D convolutions, and, more recently, video transformers that attend across space and time. The growing fusion with language models has enabled open-vocabulary recognition, dense captioning, and natural-language video retrieval.
Key aspects
- Joint spatial-temporal modelling across frames
- Action and activity recognition over time
- Temporal localisation of events within long videos
- Multimodal fusion of vision with audio and language
- Efficient handling of high-dimensional video data
Applications
- Content moderation and video search
- Surveillance and anomaly detection
- Sports and broadcast analytics
- Autonomous driving perception
- Video captioning and question answering
Provenance
- This class was materialised to resolve inbound references from existing classes in the knowledge graph.