Semantic scene understanding is the computer-vision task of parsing a visual environment into labelled, structured representations of objects, surfaces, and their spatial and functional relationships. It goes beyond object detection by assigning meaning to regions, inferring affordances, and building a coherent model of the scene that downstream systems can reason over. It is foundational to spatial computing, where digital content must be anchored to real-world geometry and semantics.
Content
- Pipelines typically combine semantic segmentation, instance segmentation, and depth estimation with relational reasoning to produce a scene graph. In augmented and mixed reality, this enables persistent placement, occlusion handling, and physics-aware interaction between virtual and physical objects.