A speech corpus is a curated collection of recorded audio, usually paired with transcriptions, speaker metadata or other annotations, used to train and evaluate speech and voice systems. Applications such as speech recognition and voice cloning both require a sufficiently large and diverse speech corpus to learn accurate acoustic and linguistic models. Corpus quality factors, including recording conditions, speaker diversity and transcription accuracy, directly bound the performance achievable by models trained on it.

Provenance