The transformer architecture was proposed by Vaswani et al. in the paper “Attention is All You Need” published in 2017..
It introduced the concept of self-attention mechanism to capture dependencies between different words in a sequence.
This approach outperformed traditional recurrent neural networks (RNNs) on various natural language processing (NLP) tasks.
Building upon the transformer, Radford et al. introduced “Generative Pre-trained Transformer” or GPT, in their paper “Improving Language Understanding by Generative Pre-training” in 2018. GPT demonstrated impressive performance on language generation and understanding tasks.
It utilized a large unsupervised neural network trained on a massive amount of text data.
BERT achieved state-of-the-art results on a wide range of NLP tasks such as question answering and sentiment analysis. It introduced a new pre-training objective called Masked Language Modeling (MLM) to train a deep bidirectional representation of language.
In 2019, Dai et al. proposed Transformer-XL in the paper “Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context”. Transformer-XL addressed the limitation of the standard transformer regarding its inability to handle long-range dependencies. It introduced the Relative Positional Encoding and Segment-Level Recurrence mechanisms, which improved the model’s ability to capture long-term context.
Facebook AI introduced RoBERTa (A Robustly Optimized BERT Pretraining Approach) in 2019.
RoBERTa was developed by Liu et al. and detailed in their paper “RoBERTa: A Robustly Optimized BERT Pretraining Approach”. It achieved better performance than BERT by optimizing the training process and using larger batch sizes.
The model was trained with significantly more data and for a longer duration, and this approach dominated the field until recently.al results.