Whisper is an automatic speech recognition model from OpenAI trained on a large multilingual dataset. It transcribes and translates speech across many languages and is released as open source.

Semantic Classification

Content

  • Whisper is an encoder-decoder transformer trained on a large corpus of audio paired with transcripts, which gives it strong performance across languages, accents and noisy conditions. It performs transcription and direct speech translation into English.
  • Because the weights and code are open, Whisper is widely used as a building block in transcription pipelines and as a base for fine-tuning. Several optimised reimplementations exist for faster inference.

Provenance