A large-scale corpus is a training dataset comprising an extremely large volume of text or other sequential data, typically gathered from web crawls, digitised books, or code repositories, and used to pretrain large neural language models. Its scale is a primary determinant of model capability under empirically observed scaling laws, alongside model parameter count and compute budget. Curation and deduplication of a large-scale corpus materially affect downstream model quality and the presence of memorised content.