A Computer Vision Task is a specific computational problem solved using visual input data, encompassing image classification, object detection, semantic segmentation, instance segmentation, and pose estimation. These tasks form the building blocks of downstream vision applications such as scene understanding, autonomous navigation, and visual question answering, typically implemented via convolutional or transformer-based neural architectures.
Semantic Classification
Content
Task Categories
-
Image Classification: Categorizing entire images based on dominant content (e.g., cat vs. dog)
-
Object Detection: Locating objects with bounding boxes using architectures like YOLO, RetinaNet, Faster RCNN
-
Semantic Segmentation: Assigning class labels to each pixel in an image
-
Instance Segmentation: Distinguishing individual object instances with precise boundaries
-
Pose Estimation: Detecting human body keypoints and skeletal structure
Modern Frameworks
-
Ultralytics YOLO11: Supports detection, segmentation, OBB, classification, and pose estimation
-
MATLAB Computer Vision Toolbox: Deep learning and CNN-based approaches