Fairness Constraints are mathematical formalizations of equitable treatment requirements in AI systems, expressed as conditions that model predictions must satisfy with respect to protected attributes such as race, gender, or age. The three canonical constraint families are Independence (demographic parity: predictions are statistically independent of protected attributes), Separation (equalized odds: predictions are independent of protected attributes conditional on the true label), and Sufficiency (calibration: true labels are independent of protected attributes conditional on predictions). These constraints are incorporated into model training as regularisation penalties or constrained optimisation objectives, and are subject to fundamental incompatibility theorems when base rates differ across protected groups.

Semantic Classification

Content

Fairness constraints operationalise the intuitive social value of equitable treatment into computable mathematical objects that can be incorporated into machine learning training and evaluation pipelines. The foundational taxonomy, formalised in Hardt et al. (2016) and surveyed extensively in Barocas et al. (2019), distinguishes three mutually exclusive constraint families: Independence, Separation, and Sufficiency — each encoding a different moral intuition about what fairness between groups requires.

The Independence criterion (demographic parity) demands that a model’s predictions be statistically uncorrelated with protected group membership. A hiring algorithm satisfying demographic parity would accept the same proportion of applicants from each demographic group. This is intuitive as a baseline equality measure but can require predicting outcomes that contradict ground-truth base rate differences, potentially undermining predictive accuracy. The Separation criterion (equalized odds) conditions on the true label: it requires that true positive rates and false positive rates be equal across groups. This is appropriate when the ground truth labels are considered reliable and unbiased — a condition rarely fully satisfied in practice. The Sufficiency criterion (calibration across groups) requires that a model’s predicted probabilities be equally well-calibrated for all groups, so that a prediction of 70% risk means the same actual risk regardless of group membership.

The Chouldechova (2017) impossibility theorem proves mathematically that when base rates differ between groups, it is impossible for a predictor to simultaneously satisfy separation and sufficiency unless it has zero error. This has profound practical implications: regulators and practitioners must choose which fairness criterion to prioritise based on domain context and stakeholder values, accepting that other criteria will be violated.

Implementation of fairness constraints in model training typically uses constrained optimisation with Lagrange multipliers, adversarial debiasing (training an auxiliary adversary to prevent a model from encoding protected attribute information), or post-processing threshold adjustments (calibrating decision thresholds separately per group after training). Each approach carries accuracy costs and requires empirical validation using dedicated fairness auditing tools to verify that constraints are met not just in aggregate but across intersectional subgroups.

Provenance