Unlabeled data is a collection of raw observations, such as images, text, or sensor readings, that lacks the target annotations or ground-truth labels needed for supervised learning. It is typically abundant and inexpensive to collect relative to labelled data, since it requires no manual annotation effort. Unlabeled data is the input substrate for unsupervised learning, self-supervised pretraining, and active learning, which selectively queries labels for the most informative examples.

Provenance