Multimodal interaction is a style of human-computer interaction in which users communicate with a system through two or more input or output modalities — such as speech, gesture, gaze, touch, and haptics — used in combination or interchangeably. By fusing complementary signals, multimodal systems interpret intent more robustly and naturally than any single channel, and they are central to spatial and immersive computing where physical input is rich. It is a key paradigm for designing interfaces in augmented and virtual reality.
Overview
- Instead of relying on a single device, multimodal systems let users speak, point, look, and touch, choosing the most natural channel for each task.
- Fusing complementary signals resolves ambiguity — for example, “put that there” disambiguated by a pointing gesture — and increases robustness.
- The paradigm is especially important in spatial computing, where the body and environment become the interface.
Mechanisms
- Early, late, and hybrid fusion combine modality streams at the feature, decision, or intermediate level.
- Temporal alignment synchronises asynchronous inputs like speech and gesture.
- Complementary and redundant use of modalities improves accessibility and error recovery.
Applications
- Voice-plus-gesture control in Augmented Reality and Virtual Reality headsets.
- Conversational assistants that combine speech with on-screen touch.
- Accessible interfaces that adapt to user ability via Universal Design.