Human annotation data is the output of people labelling, ranking, or judging raw content, such as text, images, audio, or model responses, to create ground-truth signals for training and evaluating machine learning systems. It is produced through structured annotation workflows involving guidelines, multiple annotators, and inter-annotator agreement checks such as Cohen’s kappa. Human annotation data underpins supervised learning, evaluation metrics such as COMET, and reinforcement learning from human feedback, where annotators’ preference judgements directly shape model behaviour.