Voice input is an interaction modality in which spoken language is captured, recognised, and interpreted as commands or content, allowing users to control systems and enter data hands-free. It combines microphone capture, speech recognition, and natural-language understanding to map utterances onto actions or text. In spatial computing, voice input is a primary modality for immersive and accessible interfaces where conventional keyboards and pointers are impractical.

Overview

  • Voice input lowers the barrier between intent and action: users simply speak, and the system transcribes and interprets the utterance. This is especially valuable in spatial computing, where users’ hands and eyes are occupied with the environment.
  • A complete pipeline captures audio, suppresses noise, recognises the words, and understands their meaning in context, then executes the corresponding action or inserts the dictated text. Robust voice input must handle accents, ambient noise, and ambiguous phrasing.

Mechanisms

  • Microphone capture and noise suppression isolate the speaker’s voice.
  • Wake-word and endpointing detect when a command begins and ends.
  • Speech recognition transcribes audio into text.
  • Natural-language understanding maps transcribed text onto intents and parameters.

Applications

Provenance