An Audio System is an integrated hardware and software architecture responsible for the capture, processing, transmission, and reproduction of sound signals within computing environments, including spatial audio rendering, acoustic signal processing, voice input/output, and environmental sound simulation. In spatial computing and extended reality contexts it delivers positional audio cues that reinforce presence and depth perception, coordinating microphone arrays, digital signal processors, codecs, and loudspeaker or headphone transducers. Audio systems implement psychoacoustic models — including head-related transfer functions (HRTFs) and room acoustics simulation — to produce convincing three-dimensional soundscapes. They underpin voice communication, speech interaction, accessibility features, and immersive media across consumer electronics, professional audio, telecommunication, and mixed-reality platforms.
Overview
- Audio Systems coordinate the full signal chain from acoustic capture to perceptual reproduction, operating in real time across diverse hardware platforms.
- At their core, audio systems convert analogue pressure waves to digital representations (via Pulse-Code Modulation) and back, with processing stages for filtering, mixing, spatialisation, compression, and synthesis interleaved along the path.
- In spatial computing, the audio system must synchronise with the head-tracking and rendering pipeline so that virtual sound sources remain perceptually anchored to objects in the scene, even as the user moves — a property called world-locked audio.
- Low latency is critical: end-to-end audio latency exceeding roughly 20 ms is perceptible and degrades the sense of presence in Augmented Reality and Virtual Reality experiences.
- Modern audio systems are software-defined stacks running on general-purpose CPUs and dedicated DSP hardware, exposing high-level APIs (e.g. Web Audio API, OpenAL, Oboe on Android) while abstracting platform-specific driver details.
- Audio systems are a mature technology in consumer electronics but continue to evolve rapidly in spatial-computing contexts where spatial accuracy, personalisation, and AI-driven enhancement are active research areas.
Key Components
- Input Stage
- Microphone Array — captures sound from the environment; multi-element arrays enable beamforming and noise suppression.
- Analogue-to-Digital Converter (ADC) — converts continuous acoustic signals to discrete digital samples at a fixed sample rate (typically 44.1 kHz or 48 kHz).
- Echo Cancellation — removes the system’s own loudspeaker output from the captured microphone signal to prevent feedback loops in voice communication.
- Processing Stage
- Digital Signal Processor (DSP) — dedicated programmable hardware accelerator for real-time audio computations including filtering, convolution, and mixing.
- Audio Spatialization — positions sound sources in three-dimensional space using Head-Related Transfer Function filtering, inter-aural time differences, and distance attenuation models.
- Ambisonics — a full-sphere surround sound format used for scene-based audio capture and reproduction that decodes gracefully to any loudspeaker layout or binaural headphone rendering.
- Binaural Audio — two-channel audio rendered to simulate three-dimensional sound sources for headphone listening, closely related to HRTF-based spatialisation.
- Room Acoustics Simulation — models early reflections, late reverberation, and surface absorption characteristics to reproduce the acoustic signature of virtual spaces.
- Audio Codec — lossy or lossless compression algorithm (e.g. AAC, Opus, FLAC) that reduces bitrate for storage and transmission while preserving perceptual quality.
- Audio Compression — encompasses both dynamic-range compression (limiting amplitude variation) and perceptual coding (bitrate reduction); the two senses are distinct but both relevant in audio systems.
- Acoustic Model — mathematical representation of sound propagation, including reflections and diffraction, used in simulation and enhancement pipelines.
- Output Stage
- Digital-to-Analogue Converter (DAC) — converts processed digital samples back to analogue signals for transducer driving.
- Headphone Rendering — tailors the output for near-field transducers, applying individualised HRTF corrections and crossfeed.
- Loudspeaker arrays — physical transducers for room-scale playback in professional audio, arcade, and venue contexts.
- Software Stack
- Audio Driver — kernel-mode component interfacing between the OS audio subsystem and hardware; ASIO, ALSA, Core Audio, and WASAPI are major examples.
- Web Audio API — W3C-standardised JavaScript API enabling browser-based audio routing graphs, spatialisation, and synthesis.
- OpenAL — cross-platform 3D audio API widely used in games and spatial-computing applications.
- Real-Time Computing — scheduling discipline ensuring audio buffers are filled within strict deadlines to prevent underruns (clicks and dropouts).
Applications and Use Cases
- Extended Reality (XR)
- Virtual Reality headsets (Meta Quest, Apple Vision Pro, PlayStation VR2) integrate audio systems with head tracking to deliver world-locked spatial audio.
- Augmented Reality glasses layer virtual sounds onto the real acoustic environment, requiring careful loudness matching and occlusion modelling.
- Mixed Reality platforms blend captured room acoustics with synthesised virtual sources, demanding real-time acoustic scene analysis.
- Telepresence and Collaboration
- Telepresence conferencing systems use spatial audio to assign distinct listener directions to remote participants, reducing crosstalk confusion in multi-party calls.
- Telecollaboration platforms for distributed teams use audio systems to simulate shared physical meeting rooms, improving intelligibility and naturalness.
- Voice-over-IP (VoIP) stacks (WebRTC) rely on audio systems for echo cancellation, noise suppression, and adaptive jitter buffering.
- Voice Interfaces
- Speech Recognition engines consume audio system output; the quality of the captured signal directly determines word-error rates.
- Voice User Interface design requires audio systems that can reliably isolate speech from background noise and competing talkers (cocktail-party problem).
- Natural Language Processing pipelines depend on clean audio input from the audio system as their upstream data source.
- Professional and Creative Audio
- Digital audio workstations (DAWs) expose the full audio system stack for music production, post-production, and sound design.
- Broadcast and live-event systems manage dozens of simultaneous channels with precise routing, metering, and format conversion.
- Accessibility
- Audio description (AD) systems narrate on-screen visual content for users with visual impairments, delivered via the platform audio system.
- Hearing-loop and assistive listening systems interface with audio systems to transmit audio directly to hearing aids.
- Accessibility overlays use text-to-speech synthesis embedded in the audio system to convey UI notifications.
Standards and Context
- W3C Web Audio API — browser-side audio processing graph API standardised by the W3C Audio Working Group; enables in-browser spatialisation, synthesis, and analysis without native plugins.
- OpenAL / OpenAL Soft — cross-platform 3D audio API maintained by Creative Technology and subsequently as an open-source implementation; widely used in games engines (Unreal, Unity) and OpenXR runtimes.
- Khronos OpenXR — the XR runtime standard includes spatial audio positioning as part of the scene graph, aligning audio system output with pose data from the XR compositor.
- ISO/IEC 23008-3 (MPEG-H Audio) — scene-based audio standard for broadcast and streaming supporting object audio, channel audio, and Ambisonics in a unified container; used in ATSC 3.0 broadcast.
- IETF RFC 7587 (Opus) — the Opus codec standard, widely deployed in WebRTC and VoIP stacks as the preferred low-latency speech and music codec.
- AES67 — AES/EBU interoperability standard for audio-over-IP (AoIP) transport using RTP, enabling professional-grade networked audio between heterogeneous devices.
- ITU-R BS.1770 — loudness measurement and normalisation recommendation used by streaming platforms and broadcast chains to ensure consistent perceived loudness across content.
- Major platform audio APIs: Core Audio (macOS/iOS), ALSA / PipeWire (Linux), WASAPI / DirectSound (Windows), Oboe (Android) — all provide the hardware abstraction layer that higher-level audio systems build on.
Current Landscape (2026)
- At WWDC 2025 (June) Apple publicly disclosed the Apple Spatial Audio Format (ASAF) and the Apple Positional Audio Codec (APAC, standardised mid-2024), a metadata-driven system that fuses object-based audio with fifth-order Higher Order Ambisonics and renders acoustic cues in real time from listener and object position; it is purpose-built for Vision Pro and mandated for all Apple Immersive Video titles.
- ASAF does not replace Dolby Atmos but wraps and extends it, adding head-tracked binaural rendering plus environment-aware reverb, volume and echo adaptation; APAC streams between 64 kbps and 768 kbps (matching current Atmos streaming ceilings) and falls back to Atmos or stereo on unsupported devices.
- The MPEG-I Immersive Audio standard (ISO/IEC 23090-4) was technically completed in January 2025, reached FDIS in March 2025, and delivers full six-degrees-of-freedom (6DoF) rendering (x/y/z translation plus yaw/pitch/roll) with modelled reverberation, occlusion, diffraction and Doppler; the reference software (ISO/IEC 23090-34) hit FDIS at MPEG 150 in April 2025, with Ericsson-contributed volumetric source rendering.
- An open, royalty-free rival has consolidated around Eclipsa Audio, the Google and Samsung format built on the Alliance for Open Media’s IAMF specification; announced 2024, it shipped across Samsung’s entire 2025 TV and soundbar lineup, gained YouTube upload support in 2025, and by 2025-2026 is being tested for live concerts and sports via Samsung TV Plus, with THX joining AOM for certification.
- Personalised HRTF has moved from novelty to baseline: 2025 work spans spatially-aware transformers reconstructing high-resolution HRTFs from sparse measurements (arXiv, Oct 2025) and camera-based head-width estimation for HRTF selection on commodity Android phones, reducing reliance on Apple’s TrueDepth ear-scan flow.
- Key players now split into three camps: Apple (ASAF/APAC, closed, Vision Pro-centric), the Google/Samsung/AOM open bloc (Eclipsa/IAMF, TV and YouTube-centric), and the MPEG/ISO standards line (Fraunhofer, Ericsson, MPEG-I) targeting cross-platform XR interoperability.
- Open challenges as of 2026 include fragmentation across three incompatible immersive formats, the absence of a shared authoring pipeline (each requires format-specific DAW plugins), correct real-world signal-chain configuration (Eclipsa requires HDMI eARC and non-PCM passthrough to work at all), and generalising personalised HRTFs without per-user calibration.
References
-
- audioXpress (2025). Blackmagic Design Announces DaVinci Resolve 20.1 Adding Support for Apple Spatial Audio and Apple Vision Pro. https://audioxpress.com/news/blackmagic-design-announces-davinci-resolve-20-1-adding-support-for-apple-spatial-audio-and-apple-vision-pro
-
- Apple Developer (2025). Learn about Apple Immersive Video technologies — WWDC25 (Session 403). https://developer.apple.com/videos/play/wwdc2025/403/
-
- Fraunhofer IIS (2025). MPEG-I Immersive Audio (AES Convention Paper 10234). https://www.iis.fraunhofer.de/content/dam/iis/en/doc/ame/Whitepaper/mpeg-i/MPEG-I_Immersive_Audio_AES_Convention_Paper.pdf
-
- Ericsson (2026). Heterogeneous volumetric sound rendering for XR. https://www.ericsson.com/en/blog/2026/3/heterogeneous-sound-sources
-
- Google Open Source Blog (2025). Introducing Eclipsa Audio: immersive audio for everyone. https://opensource.googleblog.com/2025/01/introducing-eclipsa-audio-immersive-audio-for-everyone.html
-
- SamMobile (2026). Samsung reveals new details about Eclipsa Audio, its Dolby Atmos rival. https://www.sammobile.com/news/samsung-eclipsa-audio-dolby-atmos-rival-details-revealed/