Audio Signal Processing is the application of signal processing theory and algorithms to the analysis, transformation, synthesis, and encoding of audio-frequency signals, operating in either the time domain or frequency domain. It encompasses filtering, equalisation, dynamic range control, time-frequency analysis, psychoacoustic coding, and spatial rendering as applied to sound reproduction, communication, and computational audition systems.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:DigitalFilter))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:FastFourierTransform))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:Equalisation))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:DynamicRangeCompression))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:EchoCancellation))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:Beamforming))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:NoiseCancellation))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:hasPart ai:SourceSeparation))

Dependency Relationships

SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:requires ai:RealTimeComputing))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:requires ai:DigitalSignalProcessor))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:requires ai:PulseCodeModulation))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:requires ai:FourierAnalysis))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:dependsOn ai:DigitalSignalProcessing))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:dependsOn ai:AudioParameters))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:dependsOn ai:GPUAcceleration))

Capability Relationships

SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:SpatialAudio))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:SpeechRecognition))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:SpeechSynthesis))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:AudioCompression))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:NoiseSuppression))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:BinauralAudio))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:MusicInformationRetrieval))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:enables ai:AutomaticSpeechRecognition))

Implementation Relationships

SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:implements ai:Convolution))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:implements ai:FourierTransform))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:implements ai:Psychoacoustics))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:uses ai:MFCC))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:uses ai:MelSpectrogram))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:uses ai:DeepLearning))

Reduction Relationships

SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:reducesTo ai:SignalProcessing))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:reducesTo ai:AudioProcessing))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:supports ai:Telecommunications))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:supports ai:AudioEngine))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:supports ai:HearingAids))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:supports ai:WebRTC))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:relatedTo ai:AudioSynthesis))
SubClassOf(ai:AudioSignalProcessing
  ObjectSomeValuesFrom(ai:bridgesTo ai:MachineLearning))

About

  • Audio Signal Processing is the engineering and scientific discipline concerned with representing, transforming, and transmitting acoustic information in digital form. Its mathematical foundations were established across the first half of the twentieth century: Harry Nyquist’s 1928 sampling theorem demonstrated that a bandlimited signal can be perfectly reconstructed from discrete samples taken at a rate exceeding twice the highest frequency component; Claude Shannon’s information theory (1948) formalised the relationship between bandwidth, noise, and channel capacity; and the invention of Pulse-Code Modulation (PCM) by Alec Reeves in 1938 provided the practical encoding scheme that underpins all digital audio. The Cooley-Tukey algorithm for the Fast Fourier Transform (1965) then made frequency-domain analysis computationally tractable, enabling the broad deployment of spectral methods that had previously been prohibitively expensive. Subsequent decades saw rapid industrialisation: Bell Labs’ work on linear prediction coding (LPC) in the 1960s-70s produced the first compressed speech codecs deployed in trans-Atlantic cable telephony; the introduction of the Compact Disc in 1982 brought 16-bit, 44.1 kHz PCM audio to consumers and cemented digital audio as the dominant paradigm for music reproduction; and the development of MPEG audio coding standards from 1989 onward created the global infrastructure for compressed digital audio that reached 4.5 billion music streaming subscribers by 2025.
  • The discipline operates across two complementary representational domains. Time-domain processing manipulates the raw discrete sample sequence directly: FIR (Finite Impulse Response) and IIR (Infinite Impulse Response) Digital Filters implement frequency-selective gain functions with precisely specifiable magnitude and phase responses; dynamic range processors implement side-chain-controlled gain reduction — compressors attenuate peaks beyond a threshold ratio, limiters impose hard ceilings with infinite ratio, gates attenuate signals below a noise floor — and delay-based effects including chorus, flanger, reverb, and echo arise from recirculation of delayed signal copies through feedback topologies. The mathematics of linear time-invariant (LTI) systems, characterised entirely by their impulse response h[n], governs classical filter design: the convolution theorem establishes that time-domain convolution with h[n] is equivalent to multiplication by H(e^{jω}) in the frequency domain, and the z-transform provides the algebraic tool for designing pole-zero configurations that realise target frequency responses. Frequency-domain processing, accessed via the Short-Time Fourier Transform (STFT) — a windowed Fast Fourier Transform applied to overlapping frames of the signal at typically 50% overlap — enables magnitude and phase manipulation that is impractical in the time domain: parametric and graphic Equalisation shapes the spectral balance with surgical precision, allowing independent gain modification of arbitrarily narrow frequency bands; spectral subtraction removes stationary noise by estimating the noise power spectrum during silence periods and subtracting it from subsequent frames; and psychoacoustic codecs such as MP3 (ISO 11172-3, 1993), AAC (ISO 14496-3, 1997), and Opus (IETF RFC 6716, 2012) exploit the Psychoacoustics of simultaneous and temporal auditory masking to discard frequency content that the human auditory system cannot detect, achieving 10:1 to 20:1 compression ratios with acceptable perceptual quality measured by objective PEAQ (Perceptual Evaluation of Audio Quality) and subjective MUSHRA listening tests.
  • The formal treatment of Psychoacoustics — the relationship between physical acoustic parameters and subjective perceptual experience — is essential to understanding audio signal processing design constraints. The basilar membrane of the human cochlea performs a mechanical frequency analysis, separating incoming sound into frequency-specific positions; the critical band concept (Bark scale, 1961) captures the frequency resolution limit of this analysis, approximately 100 Hz wide below 500 Hz and widening proportionally above. Simultaneous masking occurs when a loud tone renders nearby-frequency quieter tones inaudible; temporal masking extends both pre-masking (~5 ms) and post-masking (~100-200 ms) in time around masker events. Perceptual audio codecs exploit these mechanisms by computing a time-varying masking threshold from the signal’s STFT and allocating quantisation bits only where the quantisation noise lies below the threshold, rendering it inaudible. The MFCC (Mel-Frequency Cepstral Coefficients) feature, introduced by Davis and Mermelstein (1980), compresses the mel-scale filterbank envelope into 12-20 cepstral coefficients that capture spectral shape in a perceptually motivated, compact representation aligned with cochlear frequency resolution, making it the dominant hand-crafted feature for Speech Recognition from the early 1980s through to the mid-2010s when deep neural network front-ends began displacing it.
  • The convergence of Audio Signal Processing with Machine Learning from approximately 2012 onward has produced transformative advances across every sub-field. The inflection point was the application of Deep Neural Networks to acoustic modelling in Automatic Speech Recognition by Hinton et al. (2012), which demonstrated dramatic word error rate reductions relative to Gaussian mixture model (GMM) based approaches. Convolutional Neural Networks applied to Mel-Spectrogram representations subsequently provided accurate audio classification (sound event detection, music genre classification, speaker identification) and Feature Extraction that outperformed hand-crafted MFCC features by large margins on standard benchmarks. Recurrent Neural Networks and Transformer architectures applied to temporal audio sequences drive modern Speech Recognition systems such as Whisper (OpenAI, 2022), which achieves near-human word error rates on the LibriSpeech benchmark using a 1.5B parameter encoder-decoder Transformer trained on 680,000 hours of transcribed speech. Neural Audio Codecs (EnCodec, SoundStream, Descript Audio Codec) replace hand-designed perceptual coding with learned encoder-decoder architectures using Residual Vector Quantization (RVQ), achieving transparent quality at bitrates of 6-12 kbps where classical Opus at equivalent rates produces clearly audible degradation. Diffusion Models drive state-of-the-art Source Separation (Demucs v4, MDX23-C achieving SDR improvements of 10+ dB on Musdb18) and speech enhancement. This deep integration has established Audio Signal Processing as a primary application domain for Deep Learning research and the algorithmic substrate for the Generative AI revolution in audio production.
  • The signal chain architecture for a modern Audio Signal Processing system typically follows the path: acoustic transducer (microphone or loudspeaker) → analogue conditioning (pre-amplification, anti-aliasing low-pass filter) → analogue-to-digital conversion (ADC, typically delta-sigma modulator at 64× oversampling followed by decimation filtering) → Digital Signal Processor or general-purpose CPU/GPU executing the processing algorithms → digital-to-analogue conversion (DAC) → reconstruction filter → transducer. The ADC sample rate establishes the Nyquist frequency (half the sample rate) as the upper limit of representable signal bandwidth: at 44.1 kHz (CD standard), the maximum representable frequency is 22.05 kHz, slightly above the nominal 20 kHz upper limit of human hearing. Bit depth determines dynamic range: 16-bit PCM provides 96 dB theoretical dynamic range (6 dB per bit); 24-bit provides 144 dB, sufficient to capture inaudibly quiet signals below the noise floor of virtually any recording environment. In Real-Time Computing contexts — live sound reinforcement, hearing aids, telecommunications — processing latency must remain below perceptual thresholds: under 10 ms for hearing aids (where longer delays cause audibility through bone conduction of the wearer’s own voice), and typically under 20 ms for live sound applications before participants notice delay.

Components / Architecture

  • Analogue-to-Digital / Digital-to-Analogue Conversion: Anti-aliasing low-pass filters before ADC and reconstruction filters after DAC bookend the digital signal chain. Consumer audio uses delta-sigma (ΔΣ) converters operating at 64× or 128× oversampling, then decimated to 44.1 kHz or 48 kHz; the noise-shaping properties of the ΔΣ architecture push quantisation noise to ultrasonic frequencies where it can be filtered without audible impact. Professional audio extends to 96 kHz or 192 kHz with 24-bit resolution, providing approximately 144 dB theoretical dynamic range before dither considerations. Dither — low-level noise intentionally added before bit-depth reduction — linearises quantisation by breaking the statistical correlation between quantisation error and signal level, converting deterministic distortion products into a flat noise floor that is perceptually preferable. Requantisation with noise-shaped dither (MBIT+, UV22HR) is standard practice in mastering to avoid audible truncation distortion when reducing from 24-bit to 16-bit for CD delivery.
  • Digital Filters (IIR and FIR): IIR filters (Butterworth, Chebyshev, elliptic topologies) achieve steep frequency roll-off with small coefficient sets (typically 4-12 coefficients for a 4th-6th order filter) but introduce non-linear phase response below the passband edge; this non-linearity manifests as group delay variation that smears transients and is objectionable in some applications. FIR filters provide exactly linear phase (constant group delay) at the cost of much higher computational order (hundreds to thousands of taps) for equivalent selectivity, required in applications sensitive to group delay such as hearing aid processing and digital crossovers. The windowed sinc design method (Hamming, Kaiser-Bessel windows) and Parks-McClellan equiripple optimisation are standard FIR design techniques. In Real-Time Computing contexts, Biquad IIR cascades (Direct Form II Transposed, the numerically stable topology) are the standard implementation for real-time parametric Equalisation, with each biquad implementing a 2nd-order section and complex filter topologies assembled from cascades of 8-32 biquads. SIMD (AVX-512, NEON) vectorised implementations process four to eight biquads in parallel, achieving sub-microsecond per-sample execution on modern CPUs.
  • Short-Time Fourier Transform (STFT) and Feature Extraction: The fundamental analysis-modification-synthesis tool for time-frequency audio processing. A window function (Hann, Kaiser, Blackman-Harris, chosen for sidelobe suppression requirements) is applied to successive overlapping frames; FFT size N determines frequency resolution Δf = fs/N (at 44.1 kHz / 2048 samples = 21.5 Hz/bin); hop size H determines temporal resolution Δt = H/fs. Magnitude STFT spectrograms visualise signal energy distribution in time and frequency. MFCC features wrap the magnitude spectrum through a mel-scale filterbank (typically 26-128 triangular filters spaced on the mel scale, approximately log-frequency above 1 kHz, matching cochlear frequency resolution) then apply logarithm compression and the Discrete Cosine Transform (DCT) to produce compact feature vectors capturing spectral envelope shape for Speech Recognition and Music Information Retrieval. Mel-Spectrogram — the log magnitude of the mel filterbank output without DCT — has become the preferred intermediate representation for neural audio processing, traded between acoustic models and neural vocoders. Chroma features (12-bin chromagram, pitch class energy) are used for harmonic analysis and chord recognition.
  • Convolution Reverb and Room Acoustics: Applies recorded or synthesised impulse responses (IRs) of acoustic spaces via FFT-based fast convolution (Overlap-Add and Overlap-Save algorithms, both O(N log N) per output block vs O(N²) for direct convolution). Stereo impulse responses of concert halls measured with swept sine excitation at 96 kHz with 4-second decay require convolution with approximately 384,000-sample IRs, achievable in real time on modern hardware via partitioned convolution (which divides the IR into segments and processes them in parallel across multiple threads). Ambisonics B-format IRs measured with a 32-microphone spherical array capture the full 3D spatial character of a room for subsequent binaural or loudspeaker rendering. Room acoustic simulation (finite element method, image source model, ray tracing hybrid) generates synthetic IRs for virtual acoustic environments in Spatial Computing and game audio.
  • Dynamic Range Processing: Compressors, limiters, gates, expanders, and de-essers implement gain functions keyed from side-chain RMS or peak level detectors. Attack times as short as 20 μs protect transducers from clip-inducing transients; release constants of 50-500 ms shape musical dynamics. The gain reduction law G(x) = min(1, (θ/x)^{1-1/R}) (threshold θ, ratio R) captures the compressor characteristic: at 4:1 ratio, input levels 12 dB above threshold are reduced to 3 dB above threshold in the output. Multiband compressors split the signal into 4-7 frequency bands and apply independent gain stages per band, used in broadcast loudness normalisation to EBU R128 / ITU-R BS.1770 standards (Integrated Loudness -23 LUFS for broadcast, -16 LUFS for streaming services). Limiters applied at -1 dBFS true peak prevent inter-sample clipping in lossy codecs. Sidechain-triggered techniques include ducking (lowering background music when speech is detected) and de-essing (compressing sibilant frequencies 5-10 kHz in vocal recordings).
  • Beamforming and Acoustic Echo Cancellation: Multi-microphone Beamforming algorithms enhance signals from target spatial directions while suppressing interference from other directions. Delay-and-sum beamforming applies fractional-sample delays to align wavefronts; the Minimum Variance Distortionless Response (MVDR) beamformer optimises weights to minimise output variance subject to preserving the target direction response; superdirective beamformers with closely-spaced microphones (inter-element spacing << wavelength) achieve narrow beams at low frequencies at the cost of high self-noise amplification. Beamforming is fundamental to voice assistant microphone arrays (Amazon Echo, Google Nest) and conferencing system microphone pods, achieving 8-15 dB directional gain in typical near-field pickup applications. Acoustic Echo Cancellation (AEC) applies adaptive filtering with LMS (Least Mean Squares), NLMS (Normalised LMS), or RLS (Recursive Least Squares) algorithms to estimate and remove the acoustic path from loudspeaker to microphone, cancelling typically 20-40 dB of echo energy. Residual echo suppression using spectral subtraction or neural network post-filtering follows the adaptive filter stage.
  • Neural Audio Codecs: The encoder — a causal 1D Convolutional Neural Network — compresses the waveform to approximately 75 Hz latent frames (SoundStream / EnCodec at 24 kHz produces 75 feature vectors per second regardless of bitrate); Residual Vector Quantization (RVQ) maps each latent vector to a cascade of Q codebook entries (8-32 codebooks of 1024 entries each), with each quantisation stage correcting the residual error of the previous, encoding total information at selectable bitrates from 1.5 kbps (Q=2) to 12 kbps (Q=16). The decoder CNN reconstructs the waveform from the quantised latent. Multi-scale discriminator (operating at multiple sample rates) and multi-period discriminator losses, plus feature matching loss and reconstruction loss, drive perceptual quality via adversarial training. EnCodec (Meta, 2022) and SoundStream (Google, 2021) established the RVQ-GAN paradigm; Descript Audio Codec (2023) improved quality at sub-3 kbps by using improved RVQGAN discriminator architectures. The LRAC 2025 challenge is benchmarking next-generation architectures targeting sub-1 kbps.
  • Source Separation and Audio Unmixing: Source Separation algorithms decompose mixed audio signals into constituent sources (vocals, drums, bass, other instruments; or individual speakers). Classical approaches include Independent Component Analysis (ICA) requiring at least as many microphones as sources, and Non-negative Matrix Factorization (NMF) for single-channel spectrogram decomposition. Deep Learning-based approaches — Conv-TasNet (Luo & Mesgarani, 2019), Demucs (Défossez et al., 2021), MDX-Net — achieve state-of-the-art signal-to-distortion ratios of 9-12 dB improvement on the Musdb18 benchmark for 4-stem music separation. Demucs v4 (a hybrid spectrogram and time-domain model) achieves 9.2 dB average SDR on Musdb18-HQ. Production tools including iZotope Music Rebalance, Stems (Native Instruments), and Splice Stems deploy these models for commercial music production workflows.

Use Cases / Major Families

  • Broadcasting and Streaming: Loudness normalisation to EBU R128 / ITU-R BS.1770 standards (Integrated Loudness -23 LUFS for broadcast, -16 LUFS for streaming) is now mandated across all major streaming platforms, requiring real-time loudness metering (LUFS gating, true-peak limiting) in the broadcast chain. Real-time Audio Compression for delivery — AAC-LC at 128 kbps for stereo music streaming, 192-320 kbps for high-quality tiers, Opus at 64-128 kbps for live event streaming — operates across CDN infrastructure serving billions of streams daily. Platforms such as Spotify, Apple Music, and YouTube apply multi-stage Audio Signal Processing chains including loudness normalisation, dynamic range metering, and adaptive bitrate codec selection before delivery. Object-based audio formats (Dolby Atmos, MPEG-H Audio, Sony 360 Reality Audio) add metadata-described audio objects to scene-based representations, requiring Audio Signal Processing for object-to-channel rendering at the listener’s playback device. Mastering and stem processing at streaming ingestion platforms use automated loudness control and format conversion processing applied to hundreds of millions of tracks.
  • Telecommunications and Conferencing: WebRTC (Web Real-Time Communication) defines a standardised audio processing pipeline incorporated in Chrome, Firefox, Edge, Safari, and all WebRTC-based conferencing applications: 48 kHz fullband audio capture → acoustic Echo Cancellation (frequency-domain LMS, 20-40 dB AEC) → noise suppression (Wiener filter or RNN-based) → automatic gain control (AGC, targeting -18 dBFS RMS) → voice activity detection (VAD, silence gating to reduce upload bandwidth) → Opus codec encoding (variable bitrate 6-510 kbps, targeting 32-64 kbps for speech) → jitter buffer playout with packet loss concealment. Microsoft Teams and Zoom deploy additional AI-based Noise Cancellation layers (deep recurrent network denoising, processing 10 ms frames) on top of classical AEC, marketed as Krisp, RTX Voice (NVIDIA), and proprietary noise suppression SDKs. SILK (Skype codec, 6-40 kbps, 16 kHz wideband) and Opus (IETF RFC 6716, 6-510 kbps, 48 kHz fullband) serve the full range of speech and music Audio Compression requirements in telecommunications. SIP-based telephony infrastructure continues to operate G.711 (64 kbps, 8 kHz, PCM) and G.722 (64 kbps, 16 kHz wideband) narrowband and wideband codecs across PSTN and enterprise telephony networks serving hundreds of millions of users.
  • Hearing Aids and Assistive Listening: Modern hearing aids perform real-time sub-10 ms latency Audio Signal Processing on proprietary Bluetooth-enabled DSP SoCs (Oticon Intent, ReSound OMNIA, Phonak Lumity chipsets) at milliwatt power budgets within the 1 cm³ volume constraint of a behind-the-ear (BTE) device. The processing chain includes: multi-channel filterbank (16-64 channels) for frequency-specific gain application matching the user’s audiogram; directional Beamforming (2-4 microphone array, MVDR superdirective or machine-learned beam patterns) achieving 8-12 dB directional SNR gain; feedback cancellation (adaptive filter suppressing acoustic feedback through the receiver-to-microphone path, which causes the characteristic squeal without suppression); deep neural network denoising (3-7 layer RNN, 1-3 million parameters, compressed to 8-bit integer for on-device inference) running at 7-10 ms latency within the total 10-12 ms end-to-end budget. The Clarity Enhancement Challenge (University of Sheffield, annual since 2021) benchmarks speech intelligibility enhancement algorithms specifically for the hearing aid use case, providing standardised evaluation on the Clarity training set (7,500 hours of conversational speech in noise). Multi-channel binaural processing across bilaterally fitted hearing aids (Bluetooth-linked left and right devices) achieves 3-5 dB additional SNR improvement by enabling spatial noise suppression exploiting the binaural difference signals unavailable to monaural processing.
  • Music Production and Professional Audio: Digital Audio Workstation (DAW) environments — Pro Tools, Logic Pro X, Ableton Live, Cubase, Reaper — host complex signal processing chains involving 50-200+ plugin instances per session on modern hardware. The VST3 (Steinberg), Audio Unit (Apple), and CLAP (open-source, 2022) plugin format APIs define standardised interfaces for third-party processor integration, enabling a multi-billion dollar audio software market of equalisation, compression, reverb, saturation, modulation, and Source Separation plugins. Convolution reverb (Altiverb, Waves IR-1, Eventide Spaces2) applies impulse responses of world-renowned acoustic spaces measured in cathedrals, concert halls, and aircraft hangars. Deep learning-based mastering tools (LANDR, eMastered, Adobe Enhance Speech) automate reference-matching, Dynamic Range Compression, and Equalisation for distribution-ready masters. Stem Source Separation (Demucs v4, MDX23-C, RipX DeepAudio) enables producers to isolate vocals, drums, bass, and harmonic stems from previously mixed commercial recordings for remix, sample clearance dispute, and karaoke generation. Pitch correction processors (Auto-Tune, Melodyne) use short-time Fourier analysis for pitch detection and phase vocoder techniques for pitch shifting while maintaining timbre.
  • Spatial Audio for Extended Reality (XR): Binaural Audio rendering with personalised Head-Related Transfer Functions (HRTFs) enables plausible 3D sound localisation over headphones without physical acoustic crossfeed. HRTFs characterise the filtering applied to sound waves by the listener’s pinnae, head, and torso across frequency (20 Hz-20 kHz) and direction (elevation, azimuth, distance), typically measured via blocked-ear canal microphones with loudspeaker arrays in an anechoic chamber. Generic HRTF databases (CIPIC, SADIE II, HRTF-IEM) enable immediate deployment at the cost of localisation accuracy degradation for individual listeners; automated personalisation from ear photographs or 3D scans is an active research frontier. Ambisonics B-format (first-order: 4 channels W, X, Y, Z; higher-order: 9, 16, 25 channels for 2nd, 3rd, 4th order) captures and encodes spatial sound fields in a format-independent representation that can be decoded to any loudspeaker layout or rendered binaurally. Audio Spatialization and Adaptive Music synthesis are critical components of XR experience design: the Dolby Atmos object-based format (up to 128 audio objects + 10 fixed beds) enables per-object spatial placement that renderer engines (Meta Spatial Audio, Sony 360RA, Apple Spatial Audio via Dolby Atmos) decode to binaural or loudspeaker arrays adaptively. The Microsoft HoloLens and Meta Quest spatial audio stacks process 50-100 audio objects per frame at head-tracking update rates of 90 Hz, applying HRTF convolution and room acoustics simulation per-object using GPU Acceleration.
  • Clinical and Research Audiology: Real-time speech intelligibility enhancement for neurodivergent users (those with autism spectrum disorder, ADHD, or auditory processing disorder) and individuals with sensorineural hearing loss represents a growing application domain where Audio Signal Processing intersects with assistive technology. Adaptive filtering models trained on audiologically defined hearing profiles (from pure-tone audiograms) compensate for frequency-specific hearing threshold elevations and recruitment (abnormal loudness growth) to enhance speech intelligibility without distorting naturalness. Research systems (Newcastle University, ReSound) combining Beamforming, dynamic range compression, and neural network-based speech enhancement have demonstrated statistically significant improvements in speech intelligibility in noise (SIN) for hearing aid users in cocktail-party conditions. Audio diagnostic tools (otoacoustic emission measurement, auditory brainstem response audiometry) use signal-averaging and digital filtering of physiological signals at sub-nanovolt amplitudes from electrode recordings, requiring extremely high dynamic range ADCs (24-bit, 100 kHz) and precision Wiener filtering to extract brainstem response waveforms from electroencephalogram noise. Virtual acoustic reality platforms simulate the acoustic experience of hearing loss for audiologist training and for developing empathy in normally-hearing listeners, by applying audiogram-matched filterbanks and recruitment simulation models to real-time audio streams.

Academic Context

  • The theoretical basis for digital audio was established by Nyquist (1928), Shannon (1948), and the foundational DSP textbook by Oppenheim and Schafer (1989, “Discrete-Time Signal Processing”) which remains a canonical reference widely used in graduate signal processing courses. Julius O. Smith III at Stanford CCRMA (Centre for Computer Research in Music and Acoustics) developed the digital waveguide synthesis framework for physical modelling of stringed and wind instruments and the Spectral Audio Signal Processing textbook (freely available online since 2010, continuously updated). Gold, Morgan, and Ellis at Columbia contributed the “Speech and Audio Signal Processing” textbook covering statistical approaches to speech coding and enhancement. Alan Oppenheim and Ronald Schafer’s partnership through MIT and Georgia Tech anchored the theoretical training of DSP engineers globally. The ICASSP (IEEE International Conference on Acoustics, Speech and Signal Processing) conference series, established in 1976 and typically receiving 5,000-7,000 submissions annually, is the primary venue for audio signal processing research publication, alongside INTERSPEECH (1,000+ speech processing papers annually) and ISMIR (International Society for Music Information Retrieval). The IEEE Signal Processing Society’s Transactions on Audio, Speech, and Language Processing (TASLP) and Transactions on Signal Processing (TSP) serve as the primary archival journals.
  • Specialised research communities include the Computational Auditory Scene Analysis (CASA) community (influenced by the work of AI researcher Guy Brown and psychologist Albert Bregman), which models the auditory system’s ability to segregate simultaneous sound sources; the Music and Audio Research Laboratory (MARL) at New York University; and the Centre for Digital Music (C4DM) at Queen Mary University of London. Challenge-based evaluation has been instrumental: the DCASE Challenge (Detection and Classification of Acoustic Scenes and Events) runs annually since 2013 with tasks on sound event detection, acoustic scene classification, and audio captioning; the Source Separation Challenge (SiSEC) benchmarked Source Separation algorithms systematically from 2008; the ASVspoof Challenge has driven Audio Deepfake detection research since 2015. The CHiME Challenge series benchmarks Speech Recognition in challenging noisy conditions. Academic toolkits including librosa (Python audio analysis), ESPnet (end-to-end speech processing), speechbrain, and Kaldi have democratised access to state-of-the-art algorithms and enabled reproducible research at scale.

Current Landscape (2026)

  • As of mid-2026, the field is characterised by deep integration of classical DSP primitives with neural network components. Real-time neural Noise Suppression has become standard in video conferencing infrastructure: Microsoft Teams, Zoom, and Google Meet all deploy RNN-based denoising models processing audio at 10 ms frame intervals. NVIDIA RTX Voice and Apple Voice Isolation demonstrate hardware-accelerated neural audio at consumer scale. iZotope’s RX 11 suite incorporates diffusion-model-based repair and Source Separation tools that professional audio engineers use in post-production.
  • Neural audio codecs have achieved product deployment: Meta’s Encodec underpins audio generation pipelines; Descript Audio Codec is integrated into developer APIs; and SoundStream derivatives are used in Google’s communication products. The Low-Resource Audio Codec Challenge 2025 (LRAC 2025) evaluated next-generation codecs operating below 1 kbps for highly constrained bandwidth scenarios.
  • Self-supervised audio representation learning — models pre-trained on large unlabelled audio corpora (wav2vec 2.0, HuBERT, data2vec, AudioMAE) — now dominate the front end of Automatic Speech Recognition and audio classification pipelines. ICASSP 2025 featured extensive research on masked spectrogram pre-training, neural architectures for state-space audio modelling, and audio deepfake detection using frequency-time reinforcement learning. Loss functions incorporating auditory spatial perception (binaural ILD, ITD cues) for training spatial audio processing models are an emerging research direction as of 2025-2026.
  • The Hearing Aids sector is undergoing substantial DSP innovation: deep-learning denoisers operating under 12 ms total latency (STFT window: 12 ms, model: 7 ms, ISTFT: 0.5 ms) on hearing aid chipsets are being standardised. Multi-channel binaural processing in bilateral fitting offers 8-12 dB SNR improvement over monaural approaches in cocktail-party scenarios. The ICASSP-adjacent Clarity Challenge series has specifically benchmarked Automatic Speech Recognition for hearing-impaired listeners.

UK Context

  • The UK has a distinguished academic presence in Audio Signal Processing. Queen Mary University of London (QMUL) hosts the Centre for Digital Music (C4DM), one of the world’s leading research groups in music and audio technology, encompassing work on Music Information Retrieval, Source Separation, automatic music transcription, and differentiable DSP. The Machine Listening Lab at QMUL accepted PhD applications for 2025 entry across self-supervised audio learning and music processing topics. The Digital Music Research Network (DMRN+20), hosted at QMUL in December 2025, inaugurated the London Interdisciplinary Music Research Initiative (LIMRI) co-led with Goldsmiths, King’s College London, and Kingston University.
  • Imperial College London has research activity in spatial audio and hearing aid processing, with work on subspace hybrid MVDR Beamforming for augmented hearing. The University of Edinburgh contributes through the Centre for Speech Technology Research (CSTR), home to the Festvox speech synthesis project and research on neural vocoders and speaker adaptation. Cambridge University Engineering Department has historically been prominent in speech signal processing and statistical language models for ASR.
  • Northern English context: The University of Leeds has active audio and music technology programmes, with industry partnerships tied to the city’s electronic music scene. Sheffield’s Music department has strong links to Audio Signal Processing through Music Technology; The University of Manchester houses audio research aligned with its computer science strengths. Newcastle University has contributed to hearing aid and assistive listening technology research. Culturally, Manchester’s electronic music heritage (Factory Records, the Haçienda) and Sheffield’s industrial electronic music scene (Human League, Cabaret Voltaire, Warp Records) represent significant creative contexts in which DSP-based music production tools were pioneered by practitioners in Northern England.
  • The UK government’s UKRI and EPSRC continue to fund audio AI research through programmes such as the AI for Science and Government (ASG) initiative and specific grants in speech and hearing technology. BBC Research and Development remains an important applied research body, particularly in immersive audio, object-based audio (Dolby Atmos, MPEG-H), and broadcast loudness standardisation.

Future Directions (2026-2030)

  • End-to-end learned audio pipelines: The distinction between classical DSP modules and neural network components will continue to dissolve. Differentiable DSP frameworks (DDSP, RAVE, differentiable filters) allow training data-driven models that incorporate interpretable DSP priors, expected to improve data efficiency and generalisation.
  • Real-time neural codec compression below 1 kbps: LRAC 2025 results point toward deployable codecs for extreme bandwidth constraint environments (satellite communications, body-area sensor networks, hearing aid wireless links). Vector quantisation innovations (hierarchical codebooks, finite scalar quantisation) will drive this compression frontier.
  • Personalised spatial audio: Automated HRTF personalisation from ear geometry (photogrammetry, laser scanning, image-based estimation) will enable truly personalised Binaural Audio without acoustical measurement, important for consumer XR and hearing aid beamforming.
  • Neuromorphic audio processing: Research presented at ICASSP 2025 and in preprints through 2026 is exploring reservoir computing and spiking neural networks for end-to-end time-domain audio processing, motivated by the energy efficiency of neuromorphic hardware architectures for Embedded Systems and Hearing Aids.
  • Spatial audio for generative AI: Combining Generative AI with spatial audio synthesis — generating 3D audio environments from text or semantic descriptions — is an emerging research frontier bridging Audio Signal Processing and Audio Generation. Models such as AudioX (2025) already demonstrate unified any-to-audio generation across modalities.
  • AI-assisted audio restoration and archival: Diffusion model-based audio repair (declicking, denoising, bandwidth extension) enables restoration of historical recordings from the pre-digital era, an important cultural heritage application with particular relevance to the UK’s extensive broadcast archive at the BBC.

Research & Literature

    1. Nyquist, H. (1928). “Certain topics in telegraph transmission theory.” Transactions of the American Institute of Electrical Engineers, 47(2), 617-644.
    1. Shannon, C. E. (1948). “A mathematical theory of communication.” Bell System Technical Journal, 27(3), 379-423.
    1. Cooley, J. W., & Tukey, J. W. (1965). “An algorithm for the machine calculation of complex Fourier series.” Mathematics of Computation, 19(90), 297-301.
    1. Oppenheim, A. V., & Schafer, R. W. (1989). Discrete-Time Signal Processing. Prentice Hall. (3rd ed. 2010)
    1. Smith, J. O. (2011). Spectral Audio Signal Processing. W3K Publishing / CCRMA Stanford. (Online edition)
    1. Zwicker, E., & Fastl, H. (1990). Psychoacoustics: Facts and Models. Springer.
    1. Painter, T., & Spanias, A. (2000). “Perceptual coding of digital audio.” Proceedings of the IEEE, 88(4), 451-515.
    1. Brandenburg, K. (1999). “MP3 and AAC explained.” Proceedings of the 17th AES International Conference, San Francisco.
    1. Valin, J.-M., et al. (2012). “Definition of the Opus Audio Codec.” IETF RFC 6716.
    1. Benesty, J., Sondhi, M. M., & Huang, Y. (Eds.) (2008). Springer Handbook of Speech Processing. Springer.
    1. Sainath, T. N., et al. (2015). “Learning the speech front-end with raw waveform CLDNNs.” INTERSPEECH 2015, 1-5.
    1. Davis, S. B., & Mermelstein, P. (1980). “Comparison of parametric representations for monosyllabic word recognition.” IEEE Trans. ASSP, 28(4), 357-366. [MFCC origins]
    1. Hinton, G., et al. (2012). “Deep neural networks for acoustic modeling in speech recognition.” IEEE Signal Processing Magazine, 29(6), 82-97.
    1. Défossez, A., et al. (2022). “High Fidelity Neural Audio Compression.” TMLR, 2023. [EnCodec]
    1. Zeghidour, N., et al. (2021). “SoundStream: An end-to-end neural audio codec.” IEEE/ACM TASLP, 30, 495-507.
    1. Kumar, R., et al. (2023). “High-Fidelity Audio Compression with Improved RVQGAN.” NeurIPS 2023. [Descript Audio Codec]
    1. Luo, Y., & Mesgarani, N. (2019). “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation.” IEEE/ACM TASLP, 27(8), 1256-1266.
    1. Défossez, A. (2021). “Hybrid Spectrogram and Waveform Source Separation.” ISMIR 2021 Workshop. [Demucs]
    1. Baevski, A., et al. (2020). “wav2vec 2.0: A framework for self-supervised learning of speech representations.” NeurIPS 2020, 12449-12460.
    1. Hsu, W.-N., et al. (2021). “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.” IEEE/ACM TASLP, 29, 3451-3460.
    1. Hayes, B., et al. (2023). “A review of differentiable digital signal processing for music and speech synthesis.” Frontiers in Signal Processing, 3, 1284100. [QMUL C4DM]
    1. Hafezi, S., et al. (2023). “Subspace Hybrid MVDR Beamforming for Augmented Hearing.” arXiv:2311.18689. [Imperial College London]
    1. Yamamoto, R., Song, E., & Kim, J.-M. (2020). “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram.” ICASSP 2020.
    1. Purwins, H., et al. (2019). “Deep Learning for Audio Signal Processing.” IEEE J. Selected Topics in Signal Processing, 13(2), 206-219. [Overview survey]
    1. Rafii, Z., et al. (2017). “An overview of lead and accompaniment separation in music.” IEEE/ACM TASLP, 26(8), 1307-1335.
    1. Bengio, Y., et al. (2023). “Advances in Intelligent Hearing Aids: Deep Learning Approaches.” arXiv:2507.07043. (Preprint)
    1. Bogaert, T., Doclo, S., & Wouters, J. (2009). “Speech enhancement with multichannel Wiener filter techniques in multimicrophone binaural hearing aids.” J. Acoust. Soc. Am., 125(1), 360-371.
    1. European Broadcasting Union (2011). “Loudness Normalisation and Permitted Maximum Level of Audio Signals.” EBU R128.

Provenance