Annotated Training Data is a dataset whose examples have been augmented with ground-truth labels, bounding boxes, segmentation masks, or other targets that supervised models learn to predict. The annotations are produced by humans, programmatic rules, or model-assisted labeling and define the task the model is trained to solve. Its quality, coverage, and label consistency are primary determinants of supervised-model performance.
Content
- Labels may be class tags, boxes, pixel masks, keypoints, or text spans, created by human annotators, weak supervision, or model-in-the-loop pipelines. Because models inherit the biases and errors of their labels, annotation guidelines, inter-annotator agreement, and quality control are as important as raw volume, and labeled data is often the scarcest and most expensive resource in a project.