Image and Video Restoration is a computational imaging discipline within the artificial-intelligence and computer-vision domain concerned with recovering high-quality, perceptually faithful visual content from degraded observations, where degradation encompasses noise corruption (Gaussian/Poisson…
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:SuperResolutionModel))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:DenoisingNetwork))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:InpaintingModel))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:DeblurringNetwork))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:FaceRestorationModule))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:DegradationModel))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:PerceptualLossFunction))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:hasPart ai:QualityAssessmentMetric))
## Dependency Relationships
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:requires ai:TrainingDataset))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:requires ai:GPUCompute))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:requires ai:DegradationPrior))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:requires ai:DeepLearningFramework))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:requires ai:ImageQualityMetric))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:dependsOn ai:ConvolutionalNeuralNetworks))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:dependsOn ai:TransformerArchitecture))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:dependsOn ai:DiffusionModels))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:dependsOn ai:PerceptualLoss))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:dependsOn ai:GenerativeAdversarialNetworks))
## Capability Relationships
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:MediaArchivalRestoration))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:MedicalImageEnhancement))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:SatelliteImageAnalysis))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:ComputationalPhotography))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:VideoProductionPipeline))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:enables ai:ForensicImageEnhancement))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:supports ai:FilmAndTelevisionProduction))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:supports ai:CulturalHeritagePreservation))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:supports ai:ImmersiveMedia))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:supports ai:RemoteSensing))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:supports ai:ForensicScience))
## Implementation Relationships
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:RealESRGAN))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:SwinIR))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:NAFNet))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:GFPGAN))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:CodeFormer))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:LamaInpainting))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:StableDiffusionInpaint))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:implements ai:DnCNN))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:ResidualLearning))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:AttentionMechanism))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:VectorQuantisation))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:ContrastiveLearning))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:PSNRMetric))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:uses ai:LPIPSMetric))
## Reduction Relationships
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:ImageDegradation))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:NoisePower))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:MotionBlur))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:CompressionArtefacts))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:SpatialResolutionLoss))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:reduces ai:ArchivalDegradationRisk))
## Association Relationships
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:relatedTo ai:GenerativeAI))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:relatedTo ai:ComputationalImaging))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:relatedTo ai:ImageEnhancement))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:relatedTo ai:VideoProcessing))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:contrastsWith ai:ImageSynthesis))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:contrastsWith ai:ImageCompression))
SubClassOf(ai:ImageAndVideoRestoration
ObjectSomeValuesFrom(ai:contrastsWith ai:Image3DReconstruction))
## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:ImageAndVideoRestoration "AI-1091"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:ImageAndVideoRestoration "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:marketSizeUSD2024 ai:ImageAndVideoRestoration "4200000000"^^xsd:integer)
DataPropertyAssertion(ai:marketSizeUSD2030 ai:ImageAndVideoRestoration "12800000000"^^xsd:integer)
DataPropertyAssertion(ai:marketCAGR ai:ImageAndVideoRestoration "0.204"^^xsd:decimal)
DataPropertyAssertion(ai:benchmarkPSNR_CBSD68_sigma25 ai:ImageAndVideoRestoration "39.59"^^xsd:decimal)
DataPropertyAssertion(ai:benchmarkPSNR_GoPro_deblur ai:ImageAndVideoRestoration "40.30"^^xsd:decimal)
## Property Constraints
SubClassOf(ai:ImageAndVideoRestoration
DataMinCardinality(1 ai:hasDegradationModel xsd:string))
SubClassOf(ai:ImageAndVideoRestoration
DataMinCardinality(1 ai:hasQualityMetric xsd:string))
SubClassOf(ai:ImageAndVideoRestoration
DataAllValuesFrom(ai:isBlindRestoration xsd:boolean))
SubClassOf(ai:ImageAndVideoRestoration
DataSomeValuesFrom(ai:targetPSNR xsd:decimal))
## Annotations
AnnotationAssertion(rdfs:label ai:ImageAndVideoRestoration "Image and Video Restoration"@en)
AnnotationAssertion(rdfs:comment ai:ImageAndVideoRestoration "Computational imaging discipline recovering high-quality visual content from degraded observations (noise, blur, compression, downscaling, missing regions), deploying CNNs (DnCNN, Real-ESRGAN, BSRGAN), transformers (SwinIR ICCV 2021, NAFNet ECCV 2022 Best Paper, HAT), diffusion models (StableSR, DifFace, ResShift), and face-specialised methods (GFPGAN, CodeFormer NeurIPS 2022, RestoreFormer), implemented in commercial pipelines (Topaz Photo AI, Topaz Video AI, ChaiNNer, Magnific AI, Adobe Firefly Generative Fill), assessed via PSNR/SSIM/LPIPS metrics, deployed across medical imaging, astronomy, cultural heritage archival, satellite remote sensing, and immersive media, constituting a $4.2B 2024 market growing to $12.8B 2030 at CAGR 20.4%."@en)
AnnotationAssertion(dcterms:identifier ai:ImageAndVideoRestoration "AI-1091"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:ImageAndVideoRestoration "Computer Vision, Image Processing, Super-Resolution, Denoising, Inpainting, Face Restoration, Deep Learning"@en)
)
Property Characteristics
AsymmetricObjectProperty(ai:requires) AsymmetricObjectProperty(ai:enables) AsymmetricObjectProperty(ai:implements) AsymmetricObjectProperty(ai:contrastsWith) TransitiveObjectProperty(ai:dependsOn) FunctionalDataProperty(ai:benchmarkPSNR_CBSD68_sigma25) FunctionalDataProperty(ai:marketCAGR)
About Image and Video Restoration
- Image and Video Restoration is the branch of computer vision and computational photography concerned with recovering perceptually faithful, high-quality visual content from degraded observations. Unlike image synthesis, which generates content from scratch, or image segmentation, which classifies existing content, restoration operates under an ill-posed inverse problem framework: many degraded images could have arisen from the same underlying clean image, and the goal is to recover the most plausible (or most perceptually pleasing) clean result given prior knowledge about the statistics of natural images and the degradation process.
- Restoration is distinguished from adjacent fields by its input-constrained nature: a restoration system always begins with a degraded observation and must respect the information present in that observation, modifying pixels to remove noise or recover detail rather than synthesising arbitrary new content. This constraint distinguishes restoration from text-to-image generation (unconstrained) and positions it between pure signal recovery (reconstructing known signals from measurements) and perceptual enhancement (improving subjective quality beyond physical ground truth).
- The discipline spans six primary task categories: super-resolution (recovering spatial detail from low-resolution inputs), denoising (removing stochastic or structured noise from sensor-contaminated images), deblurring (reversing motion or defocus blur), inpainting (filling missing, masked, or damaged regions with semantically coherent content), compression artefact removal (eliminating blocking and ringing from lossy codecs), and face restoration (identity-consistent recovery of degraded facial images using face-specific priors). In video, temporal consistency across frames adds an additional constraint absent from still-image restoration.
- The field has undergone four paradigm shifts since the 1960s. Classical signal-processing methods (Wiener filters, total variation regularisation, sparse coding with K-SVD learned dictionaries, non-local means self-similarity exploitation) provided mathematically rigorous priors grounded in signal statistics but struggled to model the rich non-linear complexity of natural images at scale. Convolutional neural networks from 2012 onwards learned powerful implicit priors from large paired datasets (LR/HR image pairs synthesised by synthetic degradation, clean/noisy pairs, sharp/blurry pairs), dramatically outperforming classical methods on every benchmark. Transformer architectures from 2021 (SwinIR, Restormer, HAT) extended the effective receptive field to global image context through self-attention, enabling better handling of long-range dependencies critical for structural coherence in restoration. Diffusion models from 2022 introduced a fourth paradigm: leveraging generative priors from models trained on billions of images to hallucinate plausible high-frequency detail rather than merely minimising pixel-loss objectives, accepting a fidelity-perception trade-off in exchange for dramatically improved perceptual quality scores and enabling open-ended text-guided restoration.
- The market reflects this maturity trajectory. Professional restoration tooling (Topaz Labs, Adobe Firefly, Magnific AI, ChaiNNer) generated an estimated 12.8B by 2030 at 20.4% CAGR driven by democratisation of computational photography, streaming-driven remastering demand, and the growing archival digitisation industry globally.
Core Problem Formulation
- Restoration tasks share a mathematical core: given a degraded observation y = D(x) + n, recover the clean image x, where D is a degradation operator (blur kernel convolution, downsampling, masking) and n is additive noise. The ill-posed nature means the mapping from y to x is one-to-many, requiring regularisation priors P(x). Classical MAP (Maximum A Posteriori) estimation maximises log P(x|y) = log P(y|x) + log P(x). Modern neural approaches implicitly learn these priors from training data pairs {(x_i, y_i)} or from unpaired distributions. Diffusion-based methods explicitly model P(x) as a learned score function ∇_x log P(x), using the degraded y as a conditioning signal during reverse diffusion sampling.
- Degradation modelling is the central technical challenge in training generalisable restoration systems. Bicubic downsampling was used exclusively in early SR training (SRCNN, VDSR) but produces models that only handle this specific downsampling kernel—real camera images are degraded by a cascade of optical aberrations, motion blur, sensor noise, demosaicing errors, and JPEG compression that simple bicubic degradation fails to simulate. Higher-order degradation pipelines (Real-ESRGAN, BSRGAN) apply multiple degradation stages in random order, covering 30+ degradation types, dramatically improving real-world generalisation. The ongoing research frontier involves jointly modelling the entire camera image signal processing (ISP) pipeline from photon capture through demosaicing, gain application, noise profiling, and JPEG encoding.
- Training objectives divide into pixel-wise losses (MSE/L1/L2 optimising PSNR but producing blurry over-smoothed results by averaging plausible reconstructions), perceptual losses (VGG feature-space distances capturing texture and structure better than pixel distance, used in ESRGAN), adversarial losses (GANs driving the generator to produce outputs indistinguishable from real images to a trained discriminator, enabling high-frequency detail recovery at cost of potential artefacts), and score-matching objectives (diffusion models training the denoising score ∇_x log p(x|y) iteratively). The perception-distortion trade-off (Blau and Michaeli CVPR 2018) formally proves these objectives sit on a Pareto frontier: reducing distortion (improving PSNR) necessarily increases perceptual distance (worsening LPIPS), and this bound is tight—no method can simultaneously achieve optimal PSNR and optimal perceptual quality.
- The field distinguishes non-blind restoration (where the degradation parameters, e.g. noise standard deviation σ or blur kernel k, are known at inference time) from blind restoration (where degradation parameters are unknown and must be estimated or marginalised). Real-world restoration is almost universally blind because practical images suffer from complex, spatially-varying and unknown combinations of degradations—hence the importance of models like Real-ESRGAN and BSRGAN trained on randomised degradation pipelines covering 30+ degradation types.
- Evaluation metrics reflect the dual nature of restoration quality. Reference-based metrics include PSNR (peak signal-to-noise ratio, measuring log MSE, higher is better, state-of-art SR ~33 dB on BSD100 4×), SSIM (structural similarity capturing luminance/contrast/structure, 0-1 scale, high quality >0.92), and LPIPS (learned perceptual image patch similarity using VGG/AlexNet features, lower is better, face restoration good results 0.2-0.4). No-reference metrics include NIQE (natural image quality evaluator using multivariate Gaussian model of natural statistics), BRISQUE (blind/referenceless image spatial quality evaluator), MUSIQ (multi-scale image quality transformer), and CLIP-IQA (using CLIP features for open-ended quality questions). Diffusion-based restoration also uses FID (Fréchet Inception Distance measuring distributional realism). No single metric captures all quality aspects, motivating the use of metric suites rather than single benchmarks.
Components/Architecture of Restoration Systems
- A complete restoration system typically comprises: (1) Degradation estimator — identifying the type and severity of degradation present in the input, either via blind estimation networks (e.g. predicting noise level map σ for FFDNet) or by training on diverse degradation distributions and relying on implicit routing; (2) Feature extraction backbone — CNN (ResNet, U-Net, DenseNet), transformer (Swin, ViT), or hybrid backbone mapping the input image to latent feature representations at multiple scales; (3) Restoration network — the core model architecture performing the degradation reversal, which may be a feed-forward network (DnCNN, NAFNet, SwinIR producing a single deterministic output), a GAN (with generator and discriminator trained adversarially), or a diffusion model (iteratively denoising from Gaussian to clean distribution conditioned on the degraded input); (4) Loss supervision — combination of pixel-wise, perceptual, adversarial, and frequency-domain losses guiding training; (5) Post-processing — face detection and GFPGAN/CodeFormer face enhancement, colour correction, sharpening mask application; (6) Quality assessment — automated PSNR/SSIM/LPIPS metrics on test sets and optional no-reference metrics for deployment monitoring.
- U-Net architectures dominate encoder-decoder restoration designs because skip connections between corresponding encoder and decoder levels preserve spatial detail that bottleneck representations discard. NAFNet uses a U-shaped design with 4 encoder stages and 4 decoder stages, progressively halving spatial resolution while doubling channel count, with NAFBlocks (LayerNorm + depthwise convolution + gating + channel attention) at each level. SwinIR uses a residual group design with Swin Transformer blocks rather than convolutions, achieving global receptive field through shifted windows without the quadratic cost of full self-attention.
- Frequency-domain processing has emerged as a complementary approach. LaMa’s Fourier Convolution Layers achieve global receptive fields in the frequency domain, crucial for coherent inpainting of large regions. FFTformer applies frequency-domain attention for deblurring. Wavelet-domain methods (MWCNN, DWDN) separate image content into multi-scale frequency bands, applying task-specific processing at each scale. These approaches offer computational efficiency advantages for high-resolution images where spatial-domain convolutions are expensive.
- Conditioning mechanisms for generative restoration include: spatial feature transform (SFT) layers injecting semantic priors into the restoration network activation maps (used in GFPGAN to inject StyleGAN2 face priors), cross-attention (used in text-guided inpainting to condition on CLIP text embeddings), ControlNet-style conditioning (copying frozen backbone weights and adding trainable conditioning branches for spatial control), and classifier-free guidance in diffusion models (amplifying the conditioning signal by linearly extrapolating between conditional and unconditional score predictions, enabling creativity slider effects in Magnific AI).
Super-Resolution Family
- Super-resolution (SR) recovers spatial detail lost through optical downsampling, sensor limitations, or compression. Scale factors 2×, 4×, and 8× are standard benchmarks; 16× and beyond enters the hallucination regime where generative priors dominate. SR is the highest-volume restoration task by deployment count, implemented in every modern smartphone and television, and driving the largest commercial product category (Topaz Gigapixel AI, Adobe Super Resolution, NVIDIA DLSS, Samsung Super Resolution).
- SRCNN (Dong et al. 2014, TPAMI 2016) pioneered the CNN approach with a three-layer network mapping bicubic-upsampled LR images to HR counterparts, already surpassing all classical methods and sparking the deep SR revolution. The architecture’s simplicity—patch extraction, non-linear mapping, reconstruction—established the basic pipeline template.
- VDSR (Kim et al. CVPR 2016) introduced residual learning to SR—the network predicts the high-frequency residual rather than the full image—enabling much deeper networks (20 layers) without vanishing gradient problems. Residual connections are now standard in virtually every restoration backbone.
- SRGAN (Ledig et al. CVPR 2017) introduced adversarial training for SR, demonstrating that perceptual losses combined with a GAN discriminator produce textures that human observers prefer over PSNR-optimal MSE-trained models, even when PSNR is marginally lower. This paper formally introduced the perceptual-fidelity trade-off as a practical concern.
- RCAN (Zhang et al. ECCV 2018) applied channel attention (Squeeze-and-Excitation mechanisms) within residual groups for SR, achieving 32.63 dB on BSD100 4× SR at the time, demonstrating that attention mechanisms effectively concentrate on high-frequency detail channels relevant to restoration.
- ESRGAN (Wang et al. ECCV 2018 Workshops) added a relativistic discriminator (predicting relative realness between real and fake) and perceptual loss computed at VGG features rather than pixel space, producing textures that are sharper and more natural than PSNR-optimised networks even when PSNR metrics are marginally lower. This trade-off—perceptual quality vs. pixel fidelity—became a central tension of the field.
- RDN (Zhang et al. CVPR 2018) introduced residual dense networks connecting every layer to every subsequent layer within residual blocks, achieving extremely rich feature reuse for SR and reaching 32.47 dB on BSD100 4× at time of publication.
- Real-ESRGAN (Wang et al. ICCV 2021 Workshops) extended ESRGAN with a higher-order degradation pipeline: multiple rounds of blur → downscale → noise → JPEG compression applied in random order, producing a training distribution matching real-world camera images far better than the synthetic Gaussian+bicubic pipelines used by predecessors. Real-ESRGAN generalises to in-the-wild photographs without fine-tuning, enabling mass consumer deployment. The model achieved GitHub adoption exceeding 28,000 stars and was integrated into Topaz Gigapixel AI, ChaiNNer, and dozens of consumer tools.
- BSRGAN (Zhang et al. ICCV 2021) addressed blind SR similarly, proposing a more systematic randomisation of 30+ degradation types including motion blur, lens blur, downsampling kernels, noise (Gaussian/Poisson/JPEG/JPEG2000), with random permutation of the application order, achieving better generalisation to diverse real degradations.
- SwinIR (Liang et al. ICCV 2021) replaced the CNN backbone with Swin Transformer, exploiting shifted window self-attention to model long-range dependencies across the image. SwinIR dominated the NTIRE 2022 and AIM 2022 restoration challenges, achieving PSNR 32.72 dB on BSD100 4× SR vs. ESRGAN 32.73 dB but with substantially better SSIM 0.9020, and won the JPEG artefact reduction track with 40.80 dB at QF=10. Its computational cost was higher than CNNs but manageable on 8GB+ VRAM.
- HAT (Chen et al. CVPR 2023) introduced a Hybrid Attention Transformer combining channel attention (compressing channel correlations) with window attention (modelling spatial correlations), achieving 33.04 dB on BSD100 4× SR—roughly 1 dB above SwinIR—and establishing a new state of the art at 224×224 window training resolution.
- DRCT (Wang et al. 2024) applied dense residual connections within the transformer framework, achieving 33.11 dB on BSD100 4× and further pushing the PSNR ceiling on established SR benchmarks while maintaining reasonable inference speed.
- Diffusion-based SR from 2022 introduced a fundamentally different philosophy. StableSR (Wang et al. 2023) adapted Stable Diffusion v1.4 for 4× SR via a time-aware encoder and spatial feature transform layers, producing photorealistic texture details beyond the training data’s true high-frequency content, measured by improved LPIPS and FID at the cost of higher PSNR variance. OSEDiff (Wu et al. NeurIPS 2024) reduced SD inference to a single step via online score estimation, achieving competitive perceptual quality at 50× lower latency than full DDIM sampling. ResShift (Yue et al. NeurIPS 2023) modelled SR as residual shift between LR and HR distributions, needing only 15 denoising steps for high-quality 4× SR vs. 1000 for standard DDPM.
- SUPIR (Yu et al. CVPR 2024) scaled restoration to the SDXL (Stable Diffusion XL) backbone, achieving 8K-resolution in-the-wild image restoration with prompted semantic guidance—users can specify “restore this portrait to look like a professional DSLR shot” and the model complies using its massive generative prior. SUPIR represents the current frontier of AI image restoration at the cost of extremely high computational requirements (A100 80GB recommended).
Denoising Family
- Denoising is the most mature sub-task of image restoration, with deep learning methods saturating standard benchmarks. The central challenge has shifted from absolute PSNR performance to real-world generalisation (handling non-Gaussian real sensor noise), speed (achieving <100ms denoising for camera preview), and perceptual quality beyond PSNR.
- DnCNN (Zhang et al. TIP 2017) established the modern deep denoising baseline: a 17-layer fully convolutional network predicting the noise residual n from the noisy observation y via residual learning, trained on Gaussian noise at multiple σ values (blind denoising) or single σ (non-blind). DnCNN achieved 31.73 dB PSNR on BSD68 at σ=25 and became the standard comparison baseline, with variants achieving 40+ dB at low noise levels.
- FFDNet (Zhang et al. TIP 2018) improved DnCNN by accepting a noise level map as an additional input, enabling spatially-varying and user-controlled denoising in a single unified model without retraining for each σ. This spatial noise level conditioning is particularly useful for images with non-uniform noise (bright regions vs. dark regions of a photograph have different noise characteristics).
- CBDNet (Guo et al. CVPR 2019) addressed real-world camera noise by modelling the full camera ISP pipeline noise (Poisson–Gaussian mixed noise from photon shot noise and read noise, amplified and spatially correlated by the ISP). CBDNet jointly estimates a noise map and performs non-blind denoising, achieving substantially better results on real noisy datasets like SIDD and DND (Darmstadt Noise Dataset) than models trained only on synthetic Gaussian noise.
- NAFNet (Chen et al. ECCV 2022 Best Paper) achieved state-of-the-art denoising (39.99 dB on SIDD benchmark for sRGB denoising) and deblurring (40.30 dB on GoPro) using an elegantly simple non-linear activation-free block comprising LayerNorm, pointwise convolution, gating via element-wise product (replacing GELU), and channel attention. The key insight: many nonlinear activations previously assumed essential (GELU, Sigmoid, Softmax) could be replaced by simple gating, reducing computation while improving accuracy. NAFNet won the NTIRE 2022 deblurring and denoising tracks.
- Restormer (Zamir et al. CVPR 2022) applied transformer attention across channels rather than spatial positions for denoising, achieving competitive results while maintaining O(N) computational cost in image pixels rather than O(N²) for full spatial attention. Restormer achieves 40.02 dB on SIDD for colour denoising, demonstrating transformer applicability to high-resolution denoising where spatial attention is computationally prohibitive.
- Diffusion-based denoising (Gao et al. 2023, Score-based denoising) applies score-matching frameworks where a pretrained score network guides iterative denoising, achieving perceptually excellent results at cost of 50-100 NFE (number of function evaluations) vs. single-forward-pass feed-forward methods. For scientific imaging where distributional realism matters more than inference speed, diffusion denoising is increasingly preferred.
- Video denoising adds temporal coherence requirements: adjacent frames must be consistently denoised to avoid flickering artefacts in playback. FastDVDnet (Tassano et al. CVPR 2020) and RVRT (Liang et al. NeurIPS 2022) use optical flow-guided temporal alignment before joint denoising. The DAVIS Video Denoising Benchmark provides standardised evaluation for temporal methods.
Inpainting Family
- Inpainting recovers missing, masked, or damaged pixels in an image. The challenge is generating semantically plausible and texture-consistent fill for arbitrarily shaped missing regions.
- Early deep inpainting used context encoders (Pathak et al. CVPR 2016) with an adversarial loss to predict centre regions from surrounding context — the first demonstration that a network could generate semantically plausible content rather than blurring.
- Partial Convolution (Liu et al. ECCV 2018) masked the convolution operation to prevent invalid pixels (in the masked region) from influencing output, avoiding propagation of damaged information through the network and enabling accurate reconstruction for irregularly-shaped masks.
- GatedConvolution (Yu et al. ICCV 2019) refined partial convolution by learning soft gates for each spatial location and channel, providing more flexible handling of irregular masks and achieving better results on Places2 and CELEBA-HQ.
- LaMa (Suvorov et al. WACV 2022) introduced Fourier Convolution Layers (FFC) as the backbone for inpainting, exploiting the property that Fourier-domain convolutions have global receptive fields from the first layer, enabling the model to reason about global image structure when filling large missing regions.
- LaMa achieved state-of-the-art on Places2 and CelebA-HQ inpainting, generating coherent fills for masks covering up to 60% of image area, with inference in under 1 second per image on consumer GPUs.
- LaMa became the default high-quality inpainting backbone used in production pipelines including Stable Diffusion’s masked inpainting conditioning, and powers multiple commercial products including the Lama Cleaner open-source tool.
- Stable Diffusion Inpainting (Rombach et al. 2022, Runway ML / StabilityAI) fine-tuned the Stable Diffusion v1.5 latent diffusion model on inpainting tasks, accepting image, mask, and text prompt as conditioning inputs.
- SD Inpaint enabled semantically guided inpainting: removing objects and replacing them with prompted content, extending image backgrounds, or editing specific regions while preserving surroundings.
- SD Inpaint achieved photorealistic results for complex scene understanding tasks where LaMa-style feed-forward networks struggled with coherent content generation beyond simple texture filling.
- Adobe Firefly Generative Fill (2023-2026) integrated diffusion-based inpainting into Adobe Photoshop with a consumer-facing interface: users lasso a region, optionally describe the desired content via text prompt, and the system generates multiple plausible fills.
- By 2025, Adobe reported 12 billion Firefly-generated images, with Generative Fill the most-used feature across Creative Cloud.
- Firefly Generative Fill uses Content Credentials (C2PA standard) to cryptographically tag generated content for provenance tracking.
- Firefly 3 (May 2024) introduced Structure Reference and Style Reference conditioning, enabling matching composition and aesthetic treatment across inpainted regions.
- Stable Diffusion XL Inpainting (2024) brought higher-resolution inpainting to 1024×1024 native resolution, dramatically improving coherence for large inpainted regions in portrait and landscape photography.
- Diffusion-based Outpainting extends inpainting to regions outside the original image boundary — effectively extending the canvas without distorting existing content.
- Outpainting is used in creative tools (Adobe Photoshop Generative Expand, Midjourney /zoom out, DALL-E 3 outpaint mode) and production workflows where images must be reframed for different aspect ratios (16:9 → 9:16 for vertical social formats).
Face Restoration Family
- Blind face restoration presents unique challenges: faces carry strong semantic priors (symmetry, specific feature placement, skin texture) that generic SR/denoising models cannot exploit. Face-specific methods leverage identity-consistent priors from large-scale face datasets.
- GFPGAN (Wang et al. CVPR 2021) integrated pretrained StyleGAN2 as a generative prior for face restoration. A U-Net degradation removal backbone extracts latent features which are spatially transformed into StyleGAN2’s feature space, leveraging its distribution over human faces to generate identity-consistent, high-quality facial features. GFPGAN achieved LPIPS 0.365 on the CelebA-HQ blind face restoration benchmark—dramatically better than pixel-loss methods (LPIPS ~0.51)—and was integrated into Real-ESRGAN pipelines for human photograph restoration. GitHub stars exceeded 35,000 by 2026.
- CodeFormer (Zhou et al. NeurIPS 2022) proposed a vector-quantised codebook approach where face tokens are looked up in a codebook of canonical face parts learned during self-supervised pretraining. A controllable fidelity-quality trade-off parameter w ∈ [0,1] enables users to balance between high-fidelity (w→1, preserving input appearance at cost of quality) and high-quality (w→0, generating idealised facial features at cost of identity fidelity). CodeFormer achieved LPIPS 0.398 overall with better identity preservation than GFPGAN on severe degradations, and was adopted by Stable Diffusion face fixers (ADetailer, sd-webui-roop) and professional photo restoration workflows.
- RestoreFormer (Wang et al. CVPR 2022) used multi-head cross-attention between degraded face features and a high-quality face dictionary, enabling feature-level restoration without StyleGAN’s generative sampling randomness, achieving more deterministic results suitable for identity-critical applications.
- DifFace (Yue et al. 2023) applied diffusion-based face restoration with a transition model that shifts between degraded and clean face distributions, achieving LPIPS 0.2936 on CelebA-HQ blind restoration—the best among diffusion-based face restoration methods. The model requires 100 diffusion steps but produces richer texture details than GAN-based approaches.
- BFR via Diffusion — broader blind face restoration using diffusion includes DPR-Face (Zhao et al. 2024), DAEFR (Tsai et al. 2023), and SUPIR (Yu et al. 2024) which extended diffusion restoration to arbitrary content beyond faces using SDXL as the prior, enabling 8K-resolution natural image restoration with prompted semantic control.
Deblurring
- Motion and defocus blur are common photographic degradations. Traditional blind deblurring estimated the blur kernel then applied Wiener deconvolution, but was brittle for complex spatially-varying motion blur.
- NAFNet (Chen et al. ECCV 2022) achieved 40.30 dB PSNR on GoPro motion deblurring, and 33.69 dB on HIDE dataset, without explicit kernel estimation by directly learning the clean↔blurry mapping end-to-end. The simple architecture trained with Adam for 400K iterations on paired GoPro data remains competitive with far more complex models in 2026.
- Restormer (Zamir et al. CVPR 2022) achieved 40.02 dB on GoPro deblurring via its efficient channel-attention transformer, within 0.3 dB of NAFNet at higher computational cost but with better generalization across defocus and motion blur types.
- FFTformer (Kong et al. CVPR 2023) leveraged frequency-domain self-attention for deblurring, achieving 40.11 dB on GoPro while reducing memory consumption via the compactness of the Fourier representation.
- FFTformer (Kong et al. CVPR 2023) leveraged frequency-domain self-attention for deblurring, achieving 40.11 dB on GoPro while reducing memory consumption via the compactness of the Fourier representation.
- Deblur-GS (Lee et al. 2024, referenced in original stub) applies 3D Gaussian Splatting for novel view synthesis deblurring—distinct from 2D image deblurring—recovering sharp radiance fields from blurry multi-view captures.
- Defocus deblurring — recovering images blurred by shallow depth-of-field optical defocus — differs from motion deblurring in that the point spread function is spatially varying (disk-like near edges) and must be estimated from focus cues. DPDNet (Abuolaim et al. ECCV 2020) used dual-pixel stereo pairs available in modern DSLR/mirrorless cameras for depth-aided defocus deblurring, achieving better results than single-image methods.
- Blind deconvolution approaches (Krishnan et al. NIPS 2011, Michaeli-Irani ECCV 2014) attempted to jointly estimate the blur kernel and clean image via MAP optimisation with sparse priors. These classical approaches are still used in scientific imaging (microscopy, astronomy) where the physical point spread function can be partly characterised from instrument models.
- Video deblurring extends image deblurring to temporal sequences, using neighbouring frames as references to recover detail from blurry frames via deformable alignment and temporal aggregation (EDVR, PVDNet, ESTRNN architectures).
Commercial Workflow Tools
- Topaz Photo AI (versions 3.0–3.5, 2024-2026) unifies three previously separate Topaz products into one AI-native RAW-processing application. Noise removal uses a proprietary model (successor to DeNoise AI v3) achieving competitive performance with DnCNN/NAFNet on DSLR raw noise. Sharpening uses a motion-blur estimation and deconvolution pipeline. Upscaling uses a proprietary variant of Real-ESRGAN with additional face enhancement layers. Processing speed on NVIDIA RTX 4090: 24 MP images in 2-3 seconds; Apple M2 Pro: 4-6 seconds via Core ML acceleration. Topaz Photo AI v3.3 (January 2025) added subject detection for selective regional processing. Pricing: 99/year.
- Topaz Video AI (versions 4.0–5.0, 2024-2026) offers multiple specialised video models: Iris Pro for face-aware upscaling, Proteus for general footage, Dione for deinterlacing archival video, Chronos for frame interpolation (24fps→60fps, 60fps→120fps), and Nyx for noise reduction on cinema footage. Topaz Video AI 5.0 (March 2026) introduced real-time preview on RTX 4080 and batch processing of multi-camera footage. 4K→8K upscaling requires 6-8 seconds per frame on RTX 4090. Widely used in film restoration workflows (home video, 8mm/16mm digitisation) and streaming platform remastering.
- ChaiNNer (open source, GitHub 7K+ stars 2025) is a node-based visual pipeline builder supporting Real-ESRGAN, ESRGAN, SwinIR, GFPGAN, CodeFormer, BSRGAN, and custom models in ONNX/PyTorch format. Nodes for tiling, face detection, colour correction, and format conversion enable complex multi-model pipelines. ChaiNNer supports CPU, CUDA, DirectML, and Apple MPS backends, making it accessible across hardware. Preferred by prosumer restorers handling archival film digitisation and photography.
- Magnific AI (Jaime de los Rios, launched November 2023) takes a generative hallucination approach: instead of recovering existing detail, it generates plausible high-frequency texture via a diffusion model conditioned on the input image, text prompt, style slider (photographic to ultra-detailed), and creativity slider. At 2× upscale it functions similarly to SR; at 8×–16× it generates substantial invented detail, which is artistically valuable but forensically inappropriate. Magnific achieved viral adoption among commercial photographers, concept artists, and advertising agencies; pricing 299/month. In 2025 Magnific was acquired by Freepik.
- Adobe Firefly Generative Fill and Generative Expand (integrated into Photoshop and Lightroom since May 2023) brought professional-grade inpainting to 33M+ Creative Cloud subscribers without requiring ML expertise. Firefly 3 (May 2024) introduced Structure Reference for controlling composition, and Style Reference for matching aesthetic treatment. By early 2025 Adobe reported 12 billion Firefly image generations. Content Credentials (C2PA) are embedded in all Firefly outputs for provenance tracking.
Use Cases / Major Families
- Archival Heritage Restoration: BFI (British Film Institute) National Archive holds 1.3 million film cans, of which large proportions are in danger of deterioration.
- AI-assisted restoration pipelines using Real-ESRGAN for grain removal and upscaling, combined with LaMa-based scratch/damage inpainting, reduce manual digital restoration costs from £200-500/minute of footage to £50-100/minute.
- BBC Archive (1.5 million hours of footage) restoration and upscaling from SD/HD to 4K/8K uses Topaz Video AI and proprietary BBC R&D pipelines.
- Imperial War Museum (IWM) applies face restoration to 120 million historical photographs using GFPGAN/CodeFormer.
- The BFI Film Forever programme (£250M UK Government commitment) funds AI-automated digitisation pipelines for nitrate and acetate film restoration at scale.
- Peter Jackson’s WingNut Films used deep restoration and colourisation for World War I footage in “They Shall Not Grow Old” (2018) — a precursor to the systematic AI pipelines now used in broadcast archival.
- Medical Imaging Enhancement: MRI super-resolution reduces scan acquisition time by 2-4× while maintaining diagnostic resolution—crucial for paediatric and claustrophobic patients.
- CT denoising enables ultra-low-dose protocols reducing patient radiation exposure 50-70% without sacrificing diagnostic accuracy.
- Retinal OCT (Optical Coherence Tomography) restoration improves drusen quantification in AMD (Age-related Macular Degeneration) screening, enabling earlier intervention.
- Pathology whole-slide image (WSI) SR improves cell nucleus visibility in cancer grading when scanner optics are insufficient.
- Cryo-electron microscopy (cryo-EM) denoising and particle picking enhancement enables higher-resolution protein structure determination.
- These applications require non-blind characterisation of known scanner noise models, use of PSNR/SSIM metrics on paired data, and regulatory validation under MHRA (UK) and FDA (US) Software as a Medical Device (SaMD) frameworks.
- Satellite and Remote Sensing SR: Planet Labs’ 3-5m spatial resolution imagery can be super-resolved to simulated 0.75-1m resolution for crop monitoring, urban mapping, and disaster assessment.
- Maxar’s 30cm WorldView imagery benefits from denoising and deblurring under atmospheric turbulence, critical for defence reconnaissance and precision agriculture.
- EU Copernicus Sentinel-2’s 10m multispectral imagery is routinely super-resolved for precision agriculture, flood mapping, and deforestation monitoring applications.
- Smartphone Computational Photography: Every major smartphone maker (Apple, Samsung, Google, Huawei) uses neural image processing pipelines incorporating SR, denoising, and HDR fusion.
- Apple Deep Fusion (A13+) and Photonic Engine (A16+) apply ML denoising before demosaicing, extracting maximum signal from sensor data before any pixel is committed to RGB.
- Google Tensor G3 Night Sight uses transformer-based long-exposure fusion and denoising at 30fps on-device.
- Samsung Galaxy S24 Ultra uses AI Upscale for video recording with real-time temporal SR on Exynos/Snapdragon NPU.
- Video Game and Real-Time SR: NVIDIA DLSS 3 (2023) and DLSS 3.5 (Ray Reconstruction, 2024) use SR networks to upscale from 0.5× native resolution to display resolution, achieving 2-4× performance gains.
- AMD FSR 3 and Intel XeSS 2 compete in this real-time upscaling segment, using spatial upsampling and temporal anti-aliasing to produce competitive quality at <33ms frame budget.
- Meta Quest 3’s Fixed Foveated Rendering and App SpaceWarp use temporal SR to maintain 72-90fps in XR at manageable GPU load.
- Forensic Enhancement: UK police and intelligence services apply video super-resolution and face restoration to CCTV footage for identification under Home Office CAST (Centre for Applied Science and Technology) forensic validation frameworks.
- Admissibility constraints require documented, validated, and reproducible processing pipelines, limiting deployment to certified vendors such as Amped Software and Ocean Systems.
- ENFSI (European Network of Forensic Science Institutes) published guidance in 2024 on AI-assisted CCTV enhancement, requiring explicit uncertainty quantification attached to any enhanced image used as evidence.
- Astronomy and Scientific Imaging: Ground-based telescope images are routinely degraded by atmospheric seeing (turbulence), limiting effective resolution to 1-2 arcseconds even for 8-meter telescopes.
- Adaptive optics (AO) systems partially correct atmospheric turbulence; AI post-processing further recovers detail beyond AO correction limits, reaching near diffraction-limited resolution.
- Hubble Space Telescope legacy images have been reprocessed with deep denoising and SR to extract detail from archival data without new observations.
- JWST (James Webb Space Telescope) infrared imaging uses on-board noise characterisation with ground-based super-resolution to maximise effective angular resolution in galaxy morphology studies.
Academic Context
- Image restoration has a multi-decade academic lineage rooted in inverse problems and statistical signal processing. Tikhonov regularisation (1963) introduced the concept of regularising ill-posed inverse problems via penalising solution norm, foundational to all subsequent restoration frameworks. Total Variation (TV) regularisation (Rudin-Osher-Fatemi 1992) captured the piecewise-smooth prior of natural images, enabling effective noise removal and edge preservation without blurring. Sparse coding with K-SVD learned dictionaries (Elad-Aharon TPAMI 2006) learned compact overcomplete representations capturing image structure, bridging classical and learning-based approaches. Non-local means (Buades-Coll-Morel CVPR 2005) pioneered self-similarity exploitation: noisy pixels are denoised by weighted averaging with similar patches from across the full image, achieving remarkable denoising without learned parameters. BM3D (Dabov et al. TIP 2007) combined non-local grouping with collaborative sparse filtering in the wavelet domain, remaining competitive with early CNNs and still used as a baseline.
- The deep learning revolution in restoration crystallised with SRCNN (Dong et al. ECCV 2014) demonstrating that even a shallow CNN trained end-to-end on paired synthetic data surpassed all classical SR methods. This result was decisive: it established that implicit learned priors from large datasets outperform manually engineered priors for natural image restoration. The subsequent decade saw rapid benchmark progression driven by residual learning (VDSR 2016), perceptual losses (SRGAN 2017, ESRGAN 2018), attention mechanisms (RCAN 2018), dense connections (RDN 2018), and transformer architectures (SwinIR 2021, HAT 2023).
- The field crystallised around standardised benchmarks enabling rigorous comparison: BSD68 (68 images from Berkeley Segmentation Dataset, gold standard for Gaussian denoising), Set5/Set14/BSD100/Urban100/Manga109 (SR benchmarks at 2×/4× scale), GoPro (motion deblurring, Park et al. CVPR 2017 with 3,214 blurry-sharp pairs captured at 240fps and frame-averaged to different blur levels), SIDD (Smartphone Image Denoising Dataset, Abdelhamed et al. CVPR 2018 with real sRGB noise from 10 smartphones in 10 lighting conditions), HIDE (Hiding Identities while Enhancing Dataset for deblurring with pedestrians), and CelebA-HQ/LFW/WebFace (face restoration benchmarks with identity-level evaluation). These benchmarks while essential have known limitations: BSD100 is too simple for modern methods (differences below human JND threshold), synthetic degradation benchmarks do not capture real-world degradation diversity, and PSNR/SSIM metrics do not correlate well with human perceptual preferences above a certain baseline quality level.
- The NTIRE Workshop (New Trends in Image Restoration and Enhancement, since 2017 at CVPR) hosts annual challenges that have defined state-of-the-art milestones: NTIRE 2017 introduced the DIV2K dataset (1000 2K-resolution training images); NTIRE 2018 introduced enhanced upscaling methods and realistic SR tracks; NTIRE 2021 saw ESRGAN-based methods swept by SwinIR; NTIRE 2022 saw NAFNet win deblurring (ECCV Best Paper) and SwinIR win SR; NTIRE 2023 introduced diffusion SR tracks and video restoration challenges; NTIRE 2024 included video SR and realistic noise synthesis tracks. The PIRM Challenge at ECCV 2018 first formalised the perceptual-quality vs. distortion trade-off curve, demonstrating that methods optimising PSNR necessarily sacrifice perceptual quality (LPIPS) and vice versa.
- The perception-distortion trade-off (Blau and Michaeli, CVPR 2018) is a foundational theoretical result: for a fixed distortion level (PSNR/SSIM), the minimum achievable perceptual distance (FID/LPIPS) is bounded below by a function of the distortion, and this bound is tight. This explains why GAN-based and diffusion-based methods achieving better perceptual scores necessarily worsen PSNR—they are operating on the Pareto frontier of the perception-distortion curve. Practical consequence: the choice of restoration method must be matched to application requirements. Medical imaging and forensics require high distortion fidelity (maximise PSNR/SSIM); consumer photography and creative tools require high perceptual quality (minimise LPIPS/FID).
- Self-supervised and zero-shot restoration represents a major academic frontier from 2020 onwards. Methods like Noise2Noise (Lehtinen et al. ICML 2018) demonstrated that denoisers can be trained on pairs of noisy images without clean ground truth by exploiting the statistical independence of noise realizations. Noise2Self and Noise2Void (Batson-Royer NeurIPS 2019, Krull et al. CVPR 2019) further removed the requirement for paired noisy images, enabling blind-spot denoising from single noisy observations. Zero-shot SR (Shocher et al. CVPR 2018) observed that SR can be achieved by training on crops from the test image itself (exploiting scale-space self-similarity), achieving competitive results without any external training data. These approaches are particularly valuable for scientific imaging (microscopy, MRI, astronomy) where paired training data is unavailable.
- Implicit Neural Representations (INR) have been explored for restoration since 2021. NeRF-inspired continuous scene representations applied to image restoration (SCI, Zero-DCE++) use coordinate-based networks to represent images as continuous functions, enabling arbitrary-scale SR without discrete upsampling. Limitations in training speed and explicit control have constrained practical adoption compared to feed-forward discriminative approaches.
- The NTIRE Workshop (New Trends in Image Restoration and Enhancement, since 2017 at CVPR) hosts annual challenges that have defined state-of-the-art milestones: NTIRE 2021 saw ESRGAN-based methods swept by SwinIR; NTIRE 2022 saw NAFNet win deblurring; NTIRE 2023 introduced diffusion SR tracks; NTIRE 2024 included video SR and realistic noise synthesis tracks. The PIRM Challenge at ECCV 2018 first formalised the perceptual-quality vs. distortion trade-off curve, demonstrating that methods optimising PSNR necessarily sacrifice perceptual quality (LPIPS) and vice versa—the so-called perception-distortion trade-off of Blau and Michaeli CVPR 2018.
- The perception-distortion trade-off (Blau and Michaeli, CVPR 2018) is a foundational theoretical result: for a fixed distortion level (PSNR/SSIM), the minimum achievable perceptual distance (FID/LPIPS) is bounded below by a function of the distortion, and this bound is tight. This explains why GAN-based and diffusion-based methods achieving better perceptual scores necessarily worsen PSNR—they are operating on the Pareto frontier of the perception-distortion curve.
Current Landscape (2026)
- Benchmark saturation is an increasing concern across classical benchmarks. On BSD68 denoising at σ=50, state-of-the-art methods (NAFNet, Restormer, DnCNN variants) all cluster between 29.9 and 30.2 dB—differences indistinguishable to human observers. On GoPro deblurring, the top 5 methods (NAFNet 40.30, FFTformer 40.11, Restormer 40.02, MPRNet 39.71, MIMO-UNet+ 39.45) differ by less than 1 dB. This convergence motivates a shift from benchmark-chasing toward real-world evaluation, robustness to out-of-distribution degradations, and task-specific application metrics.
- By 2026, image and video restoration has reached a production-ready plateau across all major task types. The dominant paradigm for super-resolution has shifted from purely GAN-based (ESRGAN era, 2018-2021) through transformer-based (SwinIR/HAT era, 2021-2023) to hybrid generative (diffusion-conditioned, 2023-2026). For quality-critical, pixel-faithful applications (medical imaging, satellite analysis, forensics), transformer-based deterministic methods (NAFNet, SwinIR, HAT) remain preferred for their reproducibility and calibrated PSNR. For creative and consumer applications where perceptual richness matters more than pixel-fidelity (photo enhancement, art upscaling, heritage presentation), diffusion-based methods (StableSR, SUPIR, OSEDiff, Magnific) dominate.
- Unified restoration models have emerged as a key 2024-2026 trend. PromptIR (Potlapalli et al. NeurIPS 2023) learned a single model across all degradation types (denoising+deblurring+deraining+haze removal) via prompt-conditioned routing. AirNet (Li et al. CVPR 2022) used contrastive learning to dynamically route degraded images to appropriate restoration branches. InstructIR (Conde et al. 2024) followed natural-language instructions for restoration, enabling prompts like “remove moderate Gaussian noise” or “enhance sharpness without hallucinating detail.”
- Video restoration has matured with temporal coherence as the central challenge. Video SR methods including EDVR (Wang et al. CVPR 2019), BasicVSR++ (Chan et al. CVPR 2022), and VRT (Liang et al. 2022) use deformable alignment and propagation to maintain temporal consistency. Real-time video SR on RTX 4090 achieves 1080p→4K at 30fps as of 2025.
- Efficiency has become a central research axis: OSEDiff, ResShift, and DRCT (Wang et al. 2024) all target reducing diffusion inference steps while preserving quality. TinyDiffusion approaches target mobile deployment of diffusion-based SR on NPU (Neural Processing Unit) at 1-2 second latency for 12MP images.
- The synthetic-to-real domain gap remains an active challenge. Models trained on synthetic degradation pipelines (even sophisticated ones like Real-ESRGAN) still misfire on unusual real-world degradation combinations (severe motion blur + extreme noise + non-standard lenses). Continued progress on dataset diversity, augmentation strategies, and domain adaptation addresses this.
UK Context
- The United Kingdom holds a structurally significant position in image and video restoration through the combination of world-class archival institutions (BFI, IWM, National Archives), leading broadcast technology organisations (BBC R&D, ITV Technology), strong academic research groups (Imperial, UCL, Edinburgh, Manchester, Oxford), and a thriving creative industries sector (film, VFX, advertising).
- Imperial College London Image and Video Communications group (Prof. Yiannis Andreopoulos and collaborators) conducts research in video coding, super-resolution for immersive media, and neural image compression, with specific work on spatially adaptive SR for 360° video relevant to XR platforms.
- Imperial’s involvement in the EPSRC-funded Visual Media Lab addresses restoration for broadcast and streaming applications, including collaboration with Sky UK on HDR content remastering from SDR archives.
- Imperial’s Department of Electrical and Electronic Engineering has specific expertise in compressed sensing and sparse signal recovery foundational to MRI acceleration and single-pixel camera SR.
- BBC Research and Development (Salford, Greater Manchester, and White City, London) operates one of Europe’s largest media technology research programmes with approximately 250 researchers across broadcast, audio, and visual engineering.
- BBC R&D’s Visual Experience team works on AI-assisted archive restoration, super-resolution for BBC iPlayer streaming, and automated video quality assessment for UHD delivery.
- The BBC R&D Remaster project (2022-2026) developed automated pipelines combining DnCNN-class denoising, Real-ESRGAN upscaling, and temporal stabilisation to bring SD-era BBC content to HD/4K for streaming platforms including iPlayer and BritBox.
- BBC R&D contributed to the EBU (European Broadcasting Union) technical standards for AI-assisted remastering, including EBU Tech 3386 on video quality metrics applicable to restored content.
- BBC R&D’s Artificial Intelligence team published open-source tools for video quality assessment and frame interpolation, including the BBC iPlayer video quality monitoring pipeline.
- University of Edinburgh Institute for Digital Communications (IDCOM) conducts research on sparse signal recovery, compressed sensing reconstruction (relevant to MRI SR), and neural image compression.
- Edinburgh’s collaboration with NHS Lothian and Edinburgh Imaging Facility applies SR and denoising to clinical MRI workflows, with pilot deployments accelerating fetal MRI at Little France hospital.
- Edinburgh’s School of Informatics has contributed theoretical work on score-based diffusion models foundational to DifFace and StableSR — key restoration methods.
- UCL Department of Computer Science Advanced Image Processing group applies denoising and SR in medical contexts.
- UCL’s collaboration with Moorfields Eye Hospital (the world’s largest eye hospital, treating 450,000+ patients/year) applies retinal OCT super-resolution to improve drusen quantification and diabetic retinopathy screening sensitivity.
- UCL’s pathology group enhances whole-slide images for the NHS Digital Pathology programme, which is rolling out AI-assisted pathology across 15 NHS trusts by 2026.
- UCL’s cryo-EM SR work at the UK National Electron Bio-imaging Centre (eBIC, Diamond Light Source) contributes to protein structure determination at near-atomic resolution.
- University of Manchester Photon Science group (School of Physics and Astronomy) develops SR and denoising algorithms for X-ray ptychography and synchrotron imaging at Diamond Light Source (Harwell, Oxfordshire).
- Manchester’s CMSP (Centre for Mathematical Sciences and Physics) contributes to compressed sensing theory and Bayesian inference underpinning MRI SR.
- Manchester’s collaboration with Christie Hospital applies CT denoising protocols for ultra-low-dose radiotherapy planning — directly reducing patient radiation burden.
- University of Cambridge Machine Intelligence Lab and Computer Vision Group contributes to video understanding and restoration, with specific work on self-supervised video denoising for autonomous vehicle camera systems.
- Cambridge’s collaboration with ARM Holdings applies restoration-aware neural processing unit (NPU) design for mobile SR deployment on Cambridge-designed processor cores.
- Alan Turing Institute (London) hosts the Image Reconstruction programme applying denoising and SR in astronomy (SKA Square Kilometre Array radio telescope data, which will generate 300 petabytes per year requiring aggressive denoising), materials science (electron microscopy), and geoscience (seismic imaging).
- The Turing’s AI for Science programme funded five restoration-adjacent projects in 2024-2025, including cryo-EM SR (Cambridge), synchrotron denoising (Diamond Light Source), and radio astronomy image reconstruction (Jodrell Bank Observatory, Macclesfield).
- Northern England industrial and academic applications:
- Sheffield Hallam University’s Sports Technology Institute applies video SR and motion blur correction for performance analysis in football, cricket, and athletics coaching.
- Newcastle-upon-Tyne’s Centre for Life (Biomedical Research Centre) applies retinal imaging SR for population-scale ophthalmology screening in the North East, which has above-average rates of diabetic retinopathy due to dietary and demographic factors.
- Hartlepool and Teesside manufacturing clusters apply SR-enhanced machine vision for defect detection on steel, polymer, and chemical production lines, with Teesside University’s machine vision laboratory (Prof. Tao Xiang group) providing technical support under Innovate UK Smart Factory programmes.
- Leeds Teaching Hospitals NHS Trust applies CT and MRI denoising as part of the Yorkshire and Humber AI in Health programme.
- Newcastle University’s Computing Science department has specific expertise in video restoration for surveillance applications, with EPSRC-funded research on uncertainty-aware face restoration for forensic scenarios.
- Creative industries sector:
- Pinewood Studios (Buckinghamshire) and Warner Bros. Leavesden (Hertfordshire) use Topaz Video AI and proprietary pipelines for VFX plate enhancement and film grain management in digital intermediate post-production.
- ILM (Industrial Light and Magic, London, established 2014) applies neural SR for rendering acceleration in virtual production environments, reducing render cost while maintaining 4K output quality.
- Framestore (London, 650+ artists) uses custom AI restoration tools for film remastering and commercial advertising workflows where archival footage must match modern camera captures.
- Double Negative (DNEG, London, 3000+ artists) has developed proprietary temporal SR models for visual effects compositing workflows where traditional upscaling introduced unacceptable artefacts.
- The UK VFX industry generates approximately £2.1B annually (BFI 2024 screen industries report), with AI-assisted restoration and enhancement representing a significant and growing efficiency driver across commercial production pipelines.
Tooling Ecosystem and Open-Source Landscape
- The restoration tooling ecosystem spans academic reference implementations, prosumer desktop applications, commercial cloud APIs, and integrated creative cloud workflows.
- Stable Diffusion WebUI (A1111) — Automatic1111’s AUTOMATIC1111/stable-diffusion-webui has native support for img2img inpainting, restoration via ControlNet (Tile Resample for SR, Inpaint for masked restoration), and face restoration extension integrating GFPGAN and CodeFormer as post-processing steps. The extension ecosystem includes ADetailer (face detection + CodeFormer), sd-webui-extras (multiple upscalers), and Ultimate SD Upscale (tiled SR for any resolution).
- ComfyUI — node-based workflow interface for Stable Diffusion with native support for restoration workflows including tiled SR via VAE-decode patching, ControlNet conditioning, and custom model loading. ComfyUI’s JSON workflow sharing has enabled community distribution of complex restoration pipelines.
- Real-ESRGAN — official inference code from xinntao/Real-ESRGAN (28K+ GitHub stars) supporting command-line batch processing, face enhancement with GFPGAN, video restoration with FFMPEG integration, and model zoo with general/anime/face-specific variants.
- BasicSR — comprehensive super-resolution and restoration training and inference framework from XPixelGroup supporting DnCNN, ESRGAN, RealESRGAN, SwinIR, HAT, GFPGAN, CodeFormer. Used by most academic restoration papers as the standard training codebase.
- IQA-PyTorch — comprehensive image quality assessment library supporting 40+ reference and no-reference metrics (PSNR, SSIM, LPIPS, DISTS, NIQE, BRISQUE, MUSIQ) with consistent APIs for restoration evaluation.
- Lama Cleaner — open-source web application built on LaMa/SD inpainting for object removal and inpainting, with 17K+ GitHub stars, supporting CPU and CUDA backends. Widely used for background cleanup in photography.
- OpenModelDB — community database of 1000+ restoration model files in ONNX and PyTorch formats compatible with ChaiNNer and other tools, including Real-ESRGAN variants, face restoration models, and specialised anime/illustration upscalers.
- Topaz Labs — professional-grade desktop applications with GPU-optimised proprietary models: Topaz Photo AI (denoising, sharpening, upscaling, 299 perpetual), Topaz DeNoise AI (standalone, legacy), with perpetual + annual subscription pricing models.
- Adobe Firefly API — programmatic access to Firefly Generative Fill and Generative Expand for enterprise and developer integration, enabling automated inpainting workflows in content management systems and e-commerce product photography pipelines.
- Magnific AI API — developer API for diffusion-based SR at 2×-16× scale with style/creativity controls, used by agencies and marketplaces for automated image quality uplifting.
- Python ecosystem: OpenCV (classical restoration functions), scikit-image (quality metrics, basic filters), PyTorch hub (pretrained restoration models), Hugging Face diffusers (SD inpainting, pipeline abstractions), ONNX Runtime (cross-platform model deployment for production).
Risks, Limitations, and Ethical Considerations
- Hallucination and fabrication risk: Diffusion-based and GAN-based restoration methods generate details that were not present in the original image. At 16× super-resolution, most texture details are invented. For forensic analysis (CCTV enhancement, medical diagnosis, materials inspection), invented detail is worse than absence of detail—it is actionable misinformation. The systematic reporting of LPIPS improvements without PSNR context in commercial marketing obscures this risk from non-expert users.
- Face identity risks: Face restoration models (GFPGAN, CodeFormer) may alter identity features of the restored subject, replacing their specific facial characteristics with population-average features from the GAN or codebook prior. This is documented in GFPGAN’s own ablation studies where severe degradations produce results that are perceptually high-quality but identity-inconsistent. For biometric applications (identification from CCTV, family photo restoration, medical records), identity drift is a safety-critical failure mode.
- Deepfake misuse potential: High-quality face restoration combined with face swapping and expression manipulation tools creates a pipeline for producing photorealistic synthetic faces from low-quality source imagery. CodeFormer’s 35,000+ GitHub stars reflect both legitimate restoration demand and this dual-use concern. Platform detection systems (C2PA content credentials, watermarking, detection classifiers) are being deployed but remain imperfect.
- Synthetic-to-real generalisation failures: Models trained on carefully curated synthetic degradation pipelines (even sophisticated ones like Real-ESRGAN) systematically fail on degradation combinations outside their training distribution. Known failure modes include: extremely old film grain combined with water damage, non-standard lens bokeh patterns, digital camera banding from electronics interference, and archival photographic fading with non-uniform colour channel degradation. Production deployment requires degradation-specific validation datasets beyond standard benchmarks.
- Computational cost inequality: State-of-the-art restoration (SwinIR, NAFNet, HAT) requires NVIDIA RTX 3080+ or equivalent for practical inference speeds (>1 fps at 12MP). Diffusion-based restoration (StableSR, DifFace, SUPIR) requires RTX 4090 or cloud GPUs for sub-minute inference. This effectively restricts access to high-quality restoration to well-resourced individuals and organisations, creating a quality gap between users with and without access to recent GPU hardware. Consumer tools like Topaz mitigate this through optimised model deployment and CPU fallback, but processing times on CPU-only systems can exceed 10 minutes per image.
- Regulatory uncertainty for medical AI restoration: AI-enhanced medical images (MRI SR, CT denoising) face regulatory uncertainty in many jurisdictions. MHRA (UK) and FDA (US) require registration as a medical device for software that modifies diagnostic images in ways that could influence clinical decisions, but the threshold for what constitutes a “material” modification is unclear. Several healthcare AI companies have pursued Software as a Medical Device (SaMD) clearance for their restoration pipelines; others operate under claimed “enhancements tools” exemptions that may not withstand regulatory scrutiny.
Future Directions (2026-2030)
- The trajectory of image and video restoration over 2026-2030 is shaped by three converging forces: the continued improvement of generative models (diffusion, flow matching) providing richer priors; the democratisation of GPU acceleration bringing professional-grade restoration to mobile devices; and the formalisation of regulatory frameworks for AI-enhanced imagery in forensic, medical, and archival contexts.
- Agentic restoration pipelines: Orchestrating multi-step restoration with LLM-based quality assessment and adaptive model selection—identifying degradation type, severity, and semantically important regions before selecting appropriate model chains—will replace single-model universal approaches for professional workflows.
- An agentic pipeline for archival film restoration might: (1) classify degradation types via a classifier (grain, scratches, fading, interlacing); (2) route each degradation to a specialised restoration model; (3) apply temporal coherence post-processing; (4) assess output quality via no-reference metrics; (5) flag low-confidence outputs for human review — all orchestrated by an LLM reasoning about quality constraints.
- Physical-world priors via camera-aware restoration: Integrating camera-specific lens characterisation, sensor noise profiles, and ISP simulation into degradation models will narrow the synthetic-to-real generalisation gap.
- Joint demosaicing+SR+denoising processing (as opposed to sequential application of separate models) will become feasible as unified camera-ISP models mature, processing RAW sensor data directly to high-quality output without intermediate lossy conversion steps.
- Flow matching-based restoration (building on Rectified Flow / Stable Diffusion 3 architecture) may replace DDPM-based diffusion for restoration by 2028, offering fewer NFE (1-4 steps) while maintaining comparable quality — enabling real-time diffusion-based SR on consumer hardware.
- Standardised perceptual metrics: The field lacks a single universally accepted perceptual metric. LPIPS, DISTS, FID, MUSIQ, NIQE all measure different aspects of perceptual quality without consensus on weighting.
- An ISO/IEC standard for image restoration quality evaluation would enable interoperability between vendor claims and regulatory submissions for medical and forensic applications.
- JPEG AI (ISO/IEC 15444-17) and JVET (Joint Video Experts Team) are developing perceptual quality metrics standardisation as part of next-generation codec evaluation frameworks.
- Regulatory frameworks for AI in forensic enhancement: ENFSI published 2024 guidance on AI-assisted CCTV enhancement requiring explicit uncertainty quantification and reproducibility documentation.
- UK Home Office CAST (Centre for Applied Science and Technology) has initiated a validation programme for AI-enhanced forensic imagery, expected to produce accreditation requirements by 2027.
- The College of Policing guidance on digital forensics (updated 2025) now explicitly requires that any AI processing of digital evidence be documented with model version, parameters, and validation evidence for admissibility.
- Real-time 8K video restoration: By 2028-2030, dedicated NPU acceleration on consumer graphics cards and mobile SoCs will enable real-time 4K→8K SR at 30fps for streaming, home video, and XR headsets.
- NVIDIA DLSS 5, AMD FSR 4, and Intel XeSS 2 represent commercial milestones on this trajectory.
- Apple Vision Pro generation 3 (expected 2027-2028) is anticipated to incorporate real-time neural SR for pass-through video, reducing display resolution requirements while maintaining perceived sharpness.
- Joint compression and restoration: Neural image codecs (JPEG XL, VVC supplemented by neural post-filtering) will increasingly integrate restoration as a standard post-decoding enhancement step.
- The boundary between codec artefact removal and super-resolution is already blurring: HEVC and AVC in-loop filters remove blocking, and learned post-processing codecs like Mean Scale Hyperprior (Ballé et al. 2018) integrate reconstruction and perceptual quality jointly.
- By 2028, end-to-end learned image codecs (from sensor to display via a joint compression+restoration pipeline) may become competitive with separate codec+restoration stacks for streaming applications.
- Personalised face restoration: Deploying face restoration explicitly conditioned on a verified identity reference will preserve individual-specific features rather than regressing to population averages.
- Identity reference-conditioned restoration is technically achievable (an encoder maps reference face embeddings into StyleGAN-style conditioning); regulatory frameworks preventing non-consensual face modification at scale will constrain deployment.
- The UK ICO (Information Commissioner’s Office) has flagged face restoration in archival research contexts as requiring explicit consent under UK GDPR Article 9 (biometric data processing), creating legal uncertainty for large-scale heritage digitisation projects.
- Multimodal restoration guidance: Text-guided restoration (InstructIR, SUPIR) will be extended to reference image guidance, video description guidance, and depth/semantic map guidance, enabling increasingly precise control over restoration outcomes.
- By 2030, restoration APIs integrated into creative cloud platforms will accept natural-language intent descriptions alongside images, with the restoration system automatically selecting models, parameters, and post-processing steps to achieve the specified aesthetic or technical goal.
Research & Literature
- Key foundational and contemporary references spanning the full restoration taxonomy. The restoration literature is dominated by CVPR, ICCV, and ECCV proceedings, with important theory contributions at NeurIPS and IEEE TIP/TPAMI journals. NTIRE workshop proceedings at CVPR provide the most comprehensive annual state-of-the-art snapshots.
- The field’s publication velocity is extremely high: PapersWithCode tracks 200+ active restoration leaderboards across task/benchmark combinations, with new state-of-the-art claims published weekly.
- Key academic groups leading the field: XPixelGroup (Chao Dong, Jiantao Zhou — ESRGAN/BasicSR), VGG Oxford (Karen Simonyan — perceptual features/LPIPS), CMU/CMU-Africa (Alyosha Efros — GAN/perceptual work), KAIST (Jaejun Yoo — DifFace/diffusion restoration), NTU (Chen Change Loy — GFPGAN/CodeFormer), HKUST (Ping Luo — NAFNet backbone), KAUST/Inception (Syed Waqas Zamir — Restormer).
- Industry research groups: Adobe Research (Firefly, Super Resolution), NVIDIA Research (DLSS, perceptual quality), Google DeepMind (image quality, WaDIQaM), Meta AI Research (image reconstruction, compression).
- Dong, C., Loy, C.C., He, K., Tang, X. (2014). “Learning a Deep Convolutional Network for Image Super-Resolution.” ECCV 2014. [SRCNN foundational work]
- Zhang, K., Zuo, W., Chen, Y., Meng, D., Zhang, L. (2017). “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising.” IEEE TIP 26(7). [DnCNN]
- Wang, X., Yu, K., Wu, S., et al. (2018). “ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks.” ECCV Workshops 2018. [ESRGAN]
- Wang, X., et al. (2021). “Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data.” ICCV 2021 Workshops. [Real-ESRGAN]
- Zhang, K., et al. (2021). “Designing a Practical Degradation Model for Deep Blind Image Super-Resolution.” ICCV 2021. [BSRGAN]
- Liang, J., et al. (2021). “SwinIR: Image Restoration Using Swin Transformer.” ICCV 2021 Workshops. [SwinIR]
- Chen, L., et al. (2022). “Simple Baselines for Image Restoration.” ECCV 2022. [NAFNet — Best Paper Award]
- Zamir, S.W., et al. (2022). “Restormer: Efficient Transformer for High-Resolution Image Restoration.” CVPR 2022. [Restormer]
- Suvorov, R., et al. (2022). “Resolution-robust Large Mask Inpainting with Fourier Convolutions.” WACV 2022. [LaMa]
- Wang, X., et al. (2021). “Towards Real-World Blind Face Restoration with Generative Facial Prior.” CVPR 2021. [GFPGAN]
- Zhou, S., et al. (2022). “Towards Robust Blind Face Restoration with Codebook Lookup Transformer.” NeurIPS 2022. [CodeFormer]
- Wang, T., et al. (2022). “RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs.” CVPR 2022. [RestoreFormer]
- Wang, J., et al. (2023). “Exploiting Diffusion Prior for Real-World Image Super-Resolution.” arXiv 2305.07015. [StableSR]
- Yue, Z., et al. (2023). “DifFace: Blind Face Restoration with Diffused Error Contraction.” arXiv 2212.06512. [DifFace]
- Yue, Z., et al. (2023). “ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting.” NeurIPS 2023. [ResShift]
- Chen, X., et al. (2023). “HAT: Hybrid Attention Transformer for Image Restoration.” CVPR 2023. [HAT]
- Blau, Y., Michaeli, T. (2018). “The Perception-Distortion Tradeoff.” CVPR 2018. [Theoretical foundation]
- Zhang, R., et al. (2018). “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric.” CVPR 2018. [LPIPS metric]
- Rombach, R., et al. (2022). “High-Resolution Image Synthesis with Latent Diffusion Models.” CVPR 2022. [Stable Diffusion — foundation for inpainting]
- Wang, Z., et al. (2004). “Image Quality Assessment: From Error Visibility to Structural Similarity.” IEEE TIP 13(4). [SSIM metric]
- Abdelhamed, A., et al. (2018). “A High-Quality Denoising Dataset for Smartphone Cameras.” CVPR 2018. [SIDD benchmark]
- Park, S., et al. (2017). “Deep Multi-Scale Convolutional Neural Network for Dynamic Scene Deblurring.” CVPR 2017. [GoPro deblurring benchmark]
- Potlapalli, V., et al. (2023). “PromptIR: Prompting for All-in-One Blind Image Restoration.” NeurIPS 2023. [Unified restoration]
- Wu, B., et al. (2024). “One-Step Effective Diffusion Network for Real-World Image Super-Resolution.” NeurIPS 2024. [OSEDiff]
- Yu, J., et al. (2024). “Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration in the Wild.” CVPR 2024. [SUPIR]
- Topaz Labs (2024-2026). Topaz Photo AI v3.x and Topaz Video AI v4-5 Technical Whitepapers. https://www.topazlabs.com
- BBC Research and Development (2022-2026). Remaster: AI-Assisted Archive Restoration. BBC R&D White Papers. https://www.bbc.co.uk/rd
- Imperial College London Image & Video Communications Group. https://ivg.doc.ic.ac.uk
- Dabov, K., Foi, A., Katkovnik, V., Egiazarian, K. (2007). “Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering.” IEEE TIP 16(8). [BM3D — classical baseline still used as comparison]
- Lehtinen, J., et al. (2018). “Noise2Noise: Learning Image Restoration without Clean Data.” ICML 2018. [Self-supervised denoising without paired clean images]
- Krull, A., et al. (2019). “Noise2Void — Learning Denoising from Single Noisy Images.” CVPR 2019. [Zero-shot denoising from single image]
- Chan, K.C.K., et al. (2022). “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment.” CVPR 2022. [Video SR with deformable alignment]
- Wang, X., et al. (2019). “EDVR: Video Restoration with Enhanced Deformable Convolutional Networks.” CVPR 2019 Workshops. [Deformable alignment for video restoration]
- Li, B., et al. (2022). “All-In-One Image Restoration for Unknown Corruption.” CVPR 2022. [AirNet — contrastive-learning-based unified restoration]
- Conde, M.V., et al. (2024). “InstructIR: High-Quality Image Restoration Following Human Instructions.” ECCV 2024. [Natural-language instructed restoration]
- Shocher, A., et al. (2018). “Zero-Shot Super-Resolution using Deep Internal Learning.” CVPR 2018. [Internal learning SR without external training data]
- Gu, J., et al. (2022). “VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder.” ECCV 2022. [VQFR — face restoration via VQ codebook]
Metadata
- domain-corrected: infrastructure → artificial-intelligence (image and video restoration is a computer vision and AI domain, not infrastructure; iri, uri, same-as, owl-class, Semantic Classification updated accordingly)
Provenance
- Wang, X., et al. (2021). Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data. ICCV 2021 Workshops. [Landmark real-world blind SR paper, 28K+ GitHub stars]
- Liang, J., et al. (2021). SwinIR: Image Restoration Using Swin Transformer. ICCV 2021 Workshops. [Transformer SR, NTIRE 2022 multi-track winner]
- Chen, L., et al. (2022). Simple Baselines for Image Restoration. ECCV 2022 Best Paper. [NAFNet — SOTA deblurring 40.30 dB GoPro]
- Zhang, K., et al. (2021). Designing a Practical Degradation Model for Deep Blind Image Super-Resolution. ICCV 2021. [BSRGAN — blind SR with randomised degradation]
- Wang, X., et al. (2021). Towards Real-World Blind Face Restoration with Generative Facial Prior. CVPR 2021. [GFPGAN — StyleGAN2-prior face restoration, LPIPS 0.365]
- Zhou, S., et al. (2022). Towards Robust Blind Face Restoration with Codebook Lookup Transformer. NeurIPS 2022. [CodeFormer — codebook face restoration with fidelity control w∈[0,1]]
- Wang, T., et al. (2022). RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs. CVPR 2022. [Cross-attention face restoration]
- Yue, Z., et al. (2023). DifFace: Blind Face Restoration with Diffused Error Contraction. arXiv 2212.06512. [Diffusion blind face restoration, LPIPS 0.2936]
- Suvorov, R., et al. (2022). Resolution-robust Large Mask Inpainting with Fourier Convolutions. WACV 2022. [LaMa — global-receptive-field Fourier convolution inpainting]
- Wang, J., et al. (2023). Exploiting Diffusion Prior for Real-World Image Super-Resolution. arXiv 2305.07015. [StableSR — Stable Diffusion SR]
- Yue, Z., et al. (2023). ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting. NeurIPS 2023. [15-step diffusion SR]
- Wu, B., et al. (2024). One-Step Effective Diffusion Network for Real-World Image Super-Resolution. NeurIPS 2024. [OSEDiff — single-step SR]
- Yu, J., et al. (2024). Scaling Up to Excellence: SUPIR for Photo-Realistic Image Restoration in the Wild. CVPR 2024. [SDXL-based large-scale restoration]
- Chen, X., et al. (2023). HAT: Hybrid Attention Transformer for Image Restoration. CVPR 2023. [Hybrid attention SR, 33.04 dB BSD100 4×]
- Zamir, S.W., et al. (2022). Restormer: Efficient Transformer for High-Resolution Image Restoration. CVPR 2022. [Channel-attention transformer, 40.02 dB GoPro]
- Potlapalli, V., et al. (2023). PromptIR: Prompting for All-in-One Blind Image Restoration. NeurIPS 2023. [Unified prompt-conditioned restoration]
- Zhang, K., et al. (2017). Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising. IEEE TIP 26(7). [DnCNN — foundational residual denoising baseline]
- Wang, X., et al. (2018). ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks. ECCV Workshops 2018. [Perceptual GAN SR with relativistic discriminator]
- Dong, C., et al. (2014). Learning a Deep Convolutional Network for Image Super-Resolution. ECCV 2014. [SRCNN — first deep SR network]
- Blau, Y., Michaeli, T. (2018). The Perception-Distortion Tradeoff. CVPR 2018. [Theoretical proof of PSNR-LPIPS Pareto frontier]
- Zhang, R., et al. (2018). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. CVPR 2018. [LPIPS metric — learned perceptual similarity]
- Wang, Z., et al. (2004). Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE TIP 13(4). [SSIM metric — structural similarity]
- Abdelhamed, A., et al. (2018). A High-Quality Denoising Dataset for Smartphone Cameras. CVPR 2018. [SIDD benchmark — real sRGB noise from 10 smartphones]
- Park, S., et al. (2017). Deep Multi-Scale Convolutional Neural Network for Dynamic Scene Deblurring. CVPR 2017. [GoPro deblurring benchmark — 3214 blurry-sharp pairs]
- Dabov, K., et al. (2007). Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering. IEEE TIP 16(8). [BM3D — classical denoising baseline]
- Lehtinen, J., et al. (2018). Noise2Noise: Learning Image Restoration without Clean Data. ICML 2018. [Self-supervised denoising without clean reference]
- Gu, J., et al. (2022). VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder. ECCV 2022. [VQFR — codebook face restoration]
- Topaz Labs (2024-2026). Topaz Photo AI v3.x and Topaz Video AI v4-5 Technical Whitepapers and Release Notes. https://www.topazlabs.com
- BBC Research and Development (2022-2026). Remaster: AI-Assisted Archive Restoration. BBC R&D White Papers. https://www.bbc.co.uk/rd
- Imperial College London Image & Video Communications Group Publications. https://ivg.doc.ic.ac.uk