A Parallel Corpus is a collection of texts paired with their translations in one or more other languages, aligned at the sentence or segment level. It provides the supervised training signal for statistical and neural machine-translation systems by exemplifying how meaning maps across languages. The size, quality, and domain coverage of a parallel corpus strongly influence the accuracy of trained translation models.

Content

  • Alignment at the segment level lets models learn how meaning maps between languages. Corpus size, translation quality, and domain coverage directly bound achievable accuracy, so curation and cleaning of parallel data are central concerns in machine-translation development.