Whisper is an automatic speech recognition model from OpenAI trained on a large multilingual dataset. It transcribes and translates speech across many languages and is released as open source.
Semantic Classification
Content
- Whisper is an encoder-decoder transformer trained on a large corpus of audio paired with transcripts, which gives it strong performance across languages, accents and noisy conditions. It performs transcription and direct speech translation into English.
- Because the weights and code are open, Whisper is widely used as a building block in transcription pipelines and as a base for fine-tuning. Several optimised reimplementations exist for faster inference.