A video codec (coder-decoder) is an algorithm or hardware implementation that compresses and decompresses digital video by exploiting spatial redundancy within frames (intra-prediction), temporal redundancy across frames (inter-prediction with motion compensation), and transform coding of residuals, enabling practical storage and transmission of video at bitrates orders of magnitude lower than uncompressed formats whilst maintaining perceptual quality.
Content
- The history of video coding standards traces from CCITT H.120 (1984) through MPEG-1 (1993) and MPEG-2/H.262 (1995, enabling DVD and digital television) to H.263 for video conferencing, and H.264/AVC (2003), which became the dominant internet video format. The ITU-T and ISO/IEC JTC1 standards bodies collaborate through the Joint Collaborative Team on Video Coding (JCT-VC) and Joint Video Expert Team (JVET) to develop successive generations roughly doubling compression efficiency every five to seven years.
- H.265/HEVC (2013) doubled coding efficiency over H.264 but attracted patent licensing disputes that fractured the ecosystem. Google, Mozilla, Cisco, Amazon, Intel, and others formed the Alliance for Open Media (AOM) in 2015 to develop AV1 as a royalty-free alternative, ratified in 2018. AV1 achieves 30–40% bitrate savings over HEVC at the cost of substantially higher encoder complexity. H.266/VVC (2020) offers similar efficiency to AV1 with potentially simpler encoder implementations but carries HEVC-like licensing risk. EVC (MPEG-5 Part 1) provides a royalty-free baseline profile.
- Hardware codec support is critical for deployment: Apple A-series, Qualcomm Snapdragon, Google Tensor, and AMD/Nvidia GPUs include dedicated AV1 decode and, increasingly, encode engines. Intel Arc GPUs added AV1 hardware encoding in 2022. YouTube, Netflix, and Meta have deployed AV1 for the majority of their streaming traffic by 2024, using open-source encoders (SVT-AV1, libaom) at cloud scale with GPU- and ASIC-accelerated decode on client devices.
- Neural video coding has emerged as a research frontier: end-to-end learned codecs (e.g., Scale-Space Flow, DCVC) use convolutional and transformer networks for all coding stages, and hybrid learned/traditional approaches (neural in-loop filters, ML motion estimation) are being standardised in MPEG’s LCEVC and Neural Enhancement extensions. By 2025, learned video coding has closed the gap to AV1 on perceptual metrics at low bitrates, with deployment beginning in video conferencing and surveillance applications where encoder latency constraints are relaxed.