The W3C Web Speech API is a browser interface specification that exposes speech recognition (speech-to-text) and speech synthesis (text-to-speech) to web applications through standardised JavaScript objects. It lets pages capture spoken input and produce spoken output without bespoke plugins, underpinning voice-driven and accessibility features on the web. Implementation depth and recognition backends vary across browsers, with some delegating recognition to cloud services.

Content

  • The specification defines two interfaces: SpeechRecognition for capturing and transcribing audio with interim and final results, and SpeechSynthesis with SpeechSynthesisUtterance for configurable voices, pitch, and rate. Cross-browser inconsistency, dependence on remote recognition servers, and privacy implications of streaming audio are the main adoption considerations.