Data curation is the process of collecting, filtering, cleaning, annotating and organising raw data into a coherent, high-quality dataset suitable for training or evaluating machine learning models. It includes deduplication, removal of low-quality or harmful content, balancing of class or source distributions, and documentation of provenance and licensing. Rigorous data curation has a substantial effect on downstream model quality and is increasingly recognised as being as important as model architecture in large-scale AI systems.