Inter-annotator agreement is a measure of the degree to which independent human annotators assign consistent labels to the same data, commonly quantified using statistics such as Cohen’s kappa or Krippendorff’s alpha. It is used to assess the reliability of human evaluation and human feedback used to train or benchmark AI systems, since low agreement signals ambiguous guidelines or task definitions. High inter-annotator agreement is generally treated as a precondition for trusting labelled data as ground truth.

Provenance