Video object segmentation is the computer-vision task of delineating and tracking the pixel-level boundaries of one or more objects across the frames of a video sequence. It extends single-image segmentation with temporal coherence, propagating masks while handling motion, occlusion and appearance change. It underpins video editing, autonomous perception, surveillance and content analysis.
Content
- Methods range from semi-supervised mask propagation from a first-frame annotation to unsupervised and promptable models. Key challenges are temporal consistency, occlusion handling and real-time performance, addressed with memory networks, optical flow and transformer architectures.