A training corpus is the body of text or other data used to fit the parameters of a machine learning model, whose scale, diversity and quality directly shape what the resulting model can learn. It is the input from which subword vocabularies are derived by algorithms such as byte pair encoding during tokenisation, prior to any model training taking place. Curation choices around a training corpus, including deduplication and filtering, materially affect downstream model behaviour.