OER·harvester

← Back to the library
arXiv HTML resource

Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives

The emergence of Generative Artificial Intelligence (AI) and Large Language Models (LLMs) has marked a new era of Natural Language Processing (NLP), introducing unprecedented capabilities that are revolutionizing various domains. This paper explores the current state of these cutting-edge technologies, demonstrating their remarkable advancements and wide-ranging applications. Our paper contributes to providing a hol…

Licence
OPEN CC-BY-4.0
Authors
Desta Haileselassie Hagos, Rick Battle, Danda B. Rawat
Published
2024-07-20 · arXiv
Language
en
Length
21348 words
Type
narrative text

Cites 52 works

inferred
Open original ↗

III Language Modeling

The use of language models is pervasive in various modern NLP applications. In these models, the probability of different sequences of words is often modeled as the product of local probabilities, as expressed in Equation 13, where $w_{i}$ represents the $i^{th}$ word in the sequence, and $h_{i}$ represents the word history preceding $w_{i}$. The formulation in Equation 13 summarizes the conditional dependencies between words in a sequence, allowing language models to capture complex linguistic patterns. Leveraging such models has proven instrumental in tasks ranging from machine translation and speech recognition to text generation and sentiment analysis [1, 2].

$$ P\left(w_{1},w_{2},\ldots,w_{n}\right)=\prod_{i=1}^{n}P\left(w_{i}\mid h_{i}\right) $$

The following are some of the main approaches to traditional and modern approaches to language modeling.

III-A Statistical Language Models

Statistical language models are based on the idea that the probability of a word appearing in a sentence is related to the probability of the words that came before it [89]. These models are trained on large corpora of text, and they use statistical methods to learn the probabilities of different sequences of words. Such models, including n-gram models and models based on maximum entropy, often use conditional probability to estimate the likelihood of a word given its context [90, 91]. Equation 14 is derived from the maximum likelihood estimation, where the probability of a word given its context is estimated by the ratio of the count of the specific context-word pair to the count of the context alone. In Equation 14, $P\left(w_{1},w_{2},\ldots,w_{n}\right)$ denotes the conditional probability of the word, given the preceding word $w_{n-1}$, $C\left(w_{n-1},w_{n}\right)$ represents is the count of occurrences of the bigram (word $w_{n-1}$, word $w_{n}$) in the training data, and the $C\left(w_{n-1}\right)$ represents the count of occurrences of the word $w_{n-1}$ in the training data. For higher-order n-gram models, the equation is extended to consider a longer history of words as shown in Equation 15.

$$ P\left(w_{n}\mid w_{n-1}\right)=\frac{C\left(w_{n-1},w_{n}\right)}{C\left(w_{n-1}\right)} $$

$$ P\left(w_{n}\mid w_{n-1},w_{n-2},\ldots,w_{1}\right)=\\ \frac{C\left(w_{n-1},w_{n-2},\ldots,w_{1},w_{n}\right)}{C\left(w_{n-1},w_{n-2},\ldots,w_{1}\right)} $$

III-B Neural Network Language Models

Neural network language models, particularly those based on RNNs or transformer architectures, model the probability of a word given its context using a neural network. Actual neural network language models can have variations based on the specific architecture used (e.g., recurrent or transformer-based). However, the simplified representation of such models can be broken down into the hidden state calculation and softmax calculation as shown in Equations 16 and 17 respectively. Equation 16 shows the hidden state calculation where $\mathbf{h}_{n-1}$ denotes the hidden state of the neural network at time step $n-1$, $\mathbf{W}_{h}$ denotes the weight matrix for the hidden state transition, $\mathbf{U}_{h}$ shows the weight matrix for the word embedding transition, $\mathbf{E}_{n-2}$ denotes the embedding vector of the word $w_{n-2}$, and $tanh$ is the hyperbolic tangent activation function. Equation 17 shows the softmax output calculation which computes the conditional probability distribution over the vocabulary for the next word $w_{n}$ where $P\left(w_{n}\mid w_{n-1},w_{n-2},\ldots,w_{1}\right)$ denotes the Conditional probability of the word $w_{n}$ given the history $w_{n-1},w_{n-2},\ldots,w_{1}$, $\mathbf{W}_{o}$ shows the weight matrix for the output layer, $\mathbf{h}_{n_{1}}$ is the hidden state of the neural network at time step $n-1$, where the $softmax$ is the softmax function, converting the network’s output into probabilities.

$$ \mathbf{h}_{n-1}=\tanh\left(\mathbf{W}_{h}\cdot\mathbf{h}_{n-2}+\mathbf{U}_{h}\cdot\mathbf{E}_{n-2}\right) $$

$$ P\left(w_{n}\mid w_{n-1},w_{n-2},\ldots,w_{1}\right)=\\ \operatorname{softmax}\left(\mathbf{W}_{o}\cdot\tanh\left(\mathbf{h}_{n-1}\right)\right) $$

III-C Transformer Language Models

Transformer language models are based on the idea of attention, which allows the model to focus on the most relevant parts of the input sequence when making predictions [18, 20, 4]. Such models leverage pre-training to achieve strong performance across various NLP tasks. According to [18], the transformer architecture offers several advantages over traditional recurrent or convolutional neural networks. It enables significantly more parallelization for faster training, achieves state-of-the-art results in machine translation with shorter training times, reduces the complexity of relating distant input positions, and effectively models long-range dependencies while handling variable-length sequences [18]. The transformer model achieves state-of-the-art results in machine translation by employing attention mechanisms, enabling it to capture long-range dependencies and process variable-length sequences without padding or truncation [18]. Moreover, it simplifies the computation of relationships between distant positions, leading to enhanced parallelization, faster training, and superior performance compared to traditional neural networks.

Self-Attention Mechanism. The Transformer architecture revolutionized sequence modeling by introducing a self-attention mechanism, eliminating the need for recurrent or convolutional structures. The self-attention mechanism essentially computes a weighted sum of input representations, where each position in the input sequence is allowed to attend to all other positions with different weights. This mechanism allows the model to capture long-range dependencies between distant words in a sentence, which is important for tasks such as machine translation and text summarization. Given an input sequence $X=\left\{x_{1},x_{2},\ldots,x_{n}\right\}$, the self-attention mechanism computes the output vector $Y=\left\{y_{1},y_{2},\ldots,y_{n}\right\}$. As shown in Equation 18, the attention mechanism computes a set of attention scores, which are then used to calculate a weighted sum of the input vectors. Here, $Q_{i}$, $K_{j}$, $v_{j}$ are the query, key, and value vectors for the $i^{t}h$ output element and $j^{t}h$ input element, respectively, and $d_{k}$ is the dimension of the key vectors [18]. The attention score $a_{ij}$ for the $i^{th}$ element in the output sequence and the $j^{th}$ element in the input sequence is computed as shown in Equation 19. Here, $e_{ij}$, commonly represented as $Q^{T}_{i}\cdot K_{j}$, is the attention energy or compatibility function between the $i^{th}$ element in the output sequence and the $j^{th}$ element in the input sequence. Once the attention scores are computed, the weighted sum of the input vectors is calculated to obtain the context vector for each output element as shown in Equation 20 where $V_{j}$ is the value vector for the $j^{th}$ input element.

$$ y_{i}=\sum_{j=1}^{n}\frac{\exp\left(e_{ij}\right)}{\sum_{k=1}^{n}\exp\left(e_{ik}\right)}\cdot v_{j},\text{~where}~e_{ij}=\frac{\left(Q_{i}\cdot K_{j}\right)}{\sqrt{d_{k}}} $$

$$ a_{ij}=\frac{\exp\left(e_{ij}\right)}{\sum_{k=1}^{n}\exp\left(e_{ik}\right)} $$

$$ c_{i}=\sum_{j=1}^{n}a_{ij}\cdot V_{j} $$

Multi-Head Self-Attention. The multi-head self-attention mechanism is a variant of the self-attention mechanism that introduces multiple attention heads to capture different aspects of the relationships in the input sequence [18]. The transformer model uses multiple self-attention heads in parallel across multiple heads to capture different aspects of the relationships within the input sequence instead of performing a single attention function with $d_{model}$-dimensional keys, value vectors, and queries. This allows the model to learn more complex representations of the input, which can improve performance on a variety of NLP tasks. As shown in Equation 21, the outputs of these heads are concatenated and linearly transformed [18] where the transformations are parameter matrices $W_{i}^{Q}\in\mathbb{R}^{d_{\text{model }}\times d_{k}},W_{i}^{K}\in\mathbb{R}^{d_{\text{model }}\times d_{k}},W_{i}^{V}\in\mathbb{R}^{d_{\text{model }}\times d_{v}}$ and $W^{O}\in\mathbb{R}^{hd_{v}\times d_{\text{model }}}$. Here, $W_{i}^{Q}$, $W_{i}^{K}$, $W_{i}^{V}$, and $W^{O}$ are learned weight matrices. This allows the model to learn a wider range of relationships between words in the input sequence.

$$ \text{MultiHead}(Q,K,V)=\text{Concat}\left({head}_{1},\ldots,{head}_{\mathrm{h}}\right).W^{O}\\ \text{Where}~\\ \text{head}{{}_{i}}=\text{SelfAttention}\left(QW_{i}^{Q},KW_{i}^{K},VW_{i}^{V}\right) $$

FFN (FFN). The FFN is an important component of the transformer architecture. It is responsible for processing information from the self-attention mechanism across all positions in the input sequence [18]. The FFN consists of two fully connected linear transformations with a ReLU (ReLU) activation function in between them. This structure allows the FFN to learn complex non-linear relationships between the input features [18]. The FFN is applied independently to each position in the input sequence, ensuring that each position can interact with all other positions [18]. This parallelized approach makes the FFN computationally efficient and scalable to long input sequences. The output of the self-attention mechanism is then passed through a position-wise feed-forward network as shown in Equation 22, where the learned parameters $W_{1}$ and $W_{2}$ are learned weight matrices while $b_{1}$ and $b_{2}$ are the learned bias vectors. As shown in Equation 23, other works have proposed replacing the ReLU activation function with other nonlinear activation functions such as GELU (GELU) $(x)=x\Phi(x)$ [92] where $\Phi(x)$ is the standard Gaussian cumulative distribution function, and $\operatorname{Swish}_{\beta}(x)=x\sigma(\beta x)$ [93].

$$ \text{FFN(}x\text{)}=max(0,x.W_{1}+b_{1}).W_{2}+b_{2} $$

$$ FFN_{\text{GELU }}\left(x,W_{1},W_{2}\right)={\text{GELU}}\left(xW_{1}\right)W_{2} \tag{23} $$

$$ FFN_{\text{Swish}}\left(x,W_{1},W_{2}\right)={\text{Swish}}_{1}\left(xW_{1}\right)W_{2} $$

Year of Release LLMs Number of Parameters Number of Training Tokens Learning Rate (Default) Developer
2017 Transformer [18] 530 Million Not explicitly stated 1x$10^{-3}$ Google AI
2018 BERT [20] 340 Million 250 Billion 5x$10^{-5}$ Google AI
2019 GPT-2 [94] 1.5 Billion 40 Billion 1x$10^{-5}$ OpenAI
2020 T5 [19] 11 Billion 1 Trillion 5x$10^{-5}$ Google AI
2020 GPT-3 [4] 175 Billion 300 Billion 6x$10^{-5}$ OpenAI
2020 Gopher [95] 280 Billion 300 Billion 4x$10^{-5}$ Google AI
2021 Jurassic-1 Jumbo [96] 178 Billion 300 Billion 6x$10^{-5}$ AI21 Labs
2021 Megatron-Turing NLG [97] 530 Billion 270 Billion 5x$10^{-5}$ NVIDIA
2022 Chinchilla [98] 70 Billion 1.4 Trillion 1.25x$10^{-4}$ Deep Mind
2022 LaMDA [26] 137 Billion 768 Billion Not explicitly stated Google AI
2022 GPT-3.5 (InstructGPT) [99] 175 Billion Not explicitly stated[^3] Not explicitly stated OpenAI
2022 GPT-3.5 (ChatGPT) 175 Billion[^4] Not explicitly stated[^5] 5x$10^{-5}$ OpenAI
2022 PaLM [100, 101] 540 Billion 780 Billion Not explicitly stated Google AI
2023 LLaMA [102] 65 Billion 1.4 Trillion 1.5x$10^{-4}$ Meta AI
2023 Llama 2 [102] 70 Billion 2 Trillion 1.5x$10^{-4}$ Meta AI
2023 PaLM 2 [103] 340 Billion[^6] 3.6 Trillion Not explicitly stated Google AI
2023 GPT-4 [104] 1-1.76 Trillion[^7] Not explicitly stated[^8] Not explicitly stated OpenAI
2023 Gemini [105] Not explicitly stated[^9] Not explicitly stated[^10] Not explicitly stated Google AI

TABLE II: A list of some of the state-of-the-art LLMs well-suited for a wide range of NLP tasks.

Fig. 1: Timeline and model size of LLMs (M=millions, B=billions).

Model MMLU GSM8K ARC (Challenge)
Gemini Pro 79.1 86.5 –
Gemini Ultra 90.0 94.4 –
GPT-3 53.9 55.0 53.2
GPT-3.5 (ChatGPT) 70.0 57.1 85.2
GPT-4 86.5 92.0 96.3
LLaMA (65B) 63.4 50.9 56.0
Llama 2 (70B) 68.9 56.8 67.3
PaLM (540B) 69.3 82.1 87.1
PaLM 2 81.2 91.0 95.1

TABLE III: A performance comparison of some of the state-of-the-art LLMs well-suited for a wide range of NLP tasks, as reported on PapersWithCode. (MMLU = Massive Multitask Language Understanding, GSM8K = Grade School Math, ARC = Abstraction and Reasoning Corpus)

In the context of language models, the transformer architecture facilitates the training of LLMs, such as GPT [28]. LLMs are a type of generative AI model that is specifically trained on large corpora of text data. In recent years, LLMs have emerged as transformative breakthroughs in the field of AI, NLG (NLG), and NLU (NLU) [106] due to their remarkable capabilities in understanding and generating human-like text and other forms of content [21]. LLMs are trained on massive datasets comprising text and code, and they exhibit the ability to learn and perform a wide range of language tasks, including text generation, language translation [107], text summarization [99], sentiment analysis [108], and question answering [109]. These models are more powerful and versatile than traditional language models. LLMs have revolutionized the way we interact with and leverage natural language data, and they are now used in a wide variety of applications, including chatbots [88], machine translation systems [7, 1], and search engines. These models have experienced significant growth in terms of scale, complexity, and performance. Recently, several LLMs have been introduced, with some of the largest dense language models that have scaled to billions of model sizes [97, 95, 96, 4, 26]. These powerful models demonstrate the capability to perform a wide range of innovative NLP tasks, including machine translation, text summarization, question answering, and code completion. To provide a comprehensive comparison of some well-known state-of-the-art LLMs, we have presented a list in Table II. Figure 1 shows a trend of some of the LLMs and their corresponding number of parameters (model sizes). Some of the well-known state-of-the-art LLMs include GPT [28] T5 [19], Gopher [95], LaMDA [102], etc. These models have demonstrated the power of pre-trained, massive neural networks for NLP tasks. For example, GPT can be used to generate realistic and coherent text, while BERT can be used to extract complex meaning from text. Table III shows a performance comparison of some of the state-of-the-art LLMs well-suited for a wide range of NLP tasks, as reported on PapersWithCode[^11].

III-D Architecture of Transformer Models

Transformer architectures have revolutionized NLP tasks, such as sequence modeling, by effectively capturing long-range dependencies and modeling relationships between words. The advantages of the transformer architecture include enhanced parallelization, faster training, and the ability to model long-range dependencies efficiently. The attention mechanism allows the model to focus on relevant parts of the input sequence, contributing to its success in handling variable-length sequences without sacrificing performance. Recognizing the shift from encoder-decoder to decoder-only architectures, understanding pre-training strategies, and the advantages of transformer models provide a more nuanced perspective on their capabilities in Generative AI and various NLP tasks. Here, we will distinguish between the original encoder-decoder architecture and the decoder-only architecture and the pre-training strategies of transformer models.

Encoder-Decoder Architecture. The encoder-decoder architecture serves as a fundamental structure in Transformer models, employed for sequence-to-sequence tasks such as machine translation, where an input sequence (source language) is transformed into an output sequence [18, 110]. In an encoder-decoder architecture, the model consists of two main components featuring multiple layers of self-attention and feedforward layers: an encoder and a decoder network. The encoder network processes the input sequence, capturing relevant information and creating a contextualized representation that encompasses semantic and syntactic details of the input. Subsequently, the decoder network, in turn, utilizes this contextualized representation from the encoder to generate the output sequence step by step. At each step, the decoder attends to various parts of the encoder’s output, facilitating the alignment of source and target language information. Both the encoder and decoder components typically employ the self-attention mechanism [18]. This mechanism enables the model to weigh the importance of different positions in the input sequence during the generation of the output sequence, thereby allowing for the capture of long-range dependencies. Encoder-decoder architectures are commonly trained in a supervised and unsupervised manner, where the model is first pre-trained on a large corpus, then fine-tuned on provided with pairs of input sequences and corresponding target output sequences. The model learns to map input sequences to output sequences by minimizing a suitable loss function [18]. However, the field has witnessed a significant shift with the emergence of decoder-only architectures, indicating a transition towards more flexible and potentially more powerful models.

Decoder-Only Architecture. The decoder-only architecture utilizes only the decoder component of the Transformer model [110, 21]. In this architecture, the model generates output sequences autoregressively, predicting one token at a time based on the preceding tokens without relying on an explicit encoder [110]. The absence of an encoder implies that the model does not receive direct information about the input sequence but instead uses its autoregressive nature to capture dependencies within the generated sequence itself. Decoder-only architectures leverage a specific variant of the self-attention mechanism. This mechanism allows the model to attend to different positions within the already generated sequence while predicting each new token, effectively capturing the necessary contextual information for generating coherent output [110]. These models are typically pre-trained on massive text corpora in an unsupervised manner [21]. During this pre-training phase, the model learns general language representations, capturing both syntactic and semantic information [21]. Subsequently, fine-tuning on specific tasks with labeled data allows the model to adapt to various downstream applications. One well-known example of the decoder-only architecture is the GPT [4]. GPT employs a stack of transformer decoder layers for autoregressive sequence generation [4].

III-E Pre-training Strategies in Transformer Language Models

One of the key factors behind the success of transformer-based language models is their pre-training on massive amounts of text data using self-supervised learning techniques [18]. This pre-training stage equips the models with a robust understanding of language structure and semantics, enabling exceptional performance on various downstream NLP tasks [20, 21]. Transformer language models, leveraging pre-training, have demonstrated outstanding performance across diverse NLP tasks. In machine translation, the transformer’s attention mechanism allows it to capture long-range dependencies, yielding state-of-the-art results without the need for excessive padding or truncation [18]. Beyond translation, decoder-only architectures like GPT have proven effective in tasks such as sentiment analysis, named entity recognition, and text completion [21].

Self-Supervised Learning for Pre-training. Unlike traditional supervised learning methods that demand extensive labeled data, self-supervised learning leverages the unlabeled nature of textual data. Common pre-training objectives in transformers involve tasks like predicting the next word in a sequence, also known as MLM (MLM) [20, 111], or reconstructing a sentence where certain words are replaced with special tokens (masked tokens) [20]. By tackling these tasks, the model learns contextual relationships between words and develops a strong understanding of grammatical structures. The pre-training phase serves as a critical foundation for downstream NLP tasks. The learned representations from vast amounts of text data can be fine-tuned for specific tasks like sentiment analysis, question answering, or machine translation. This approach requires significantly less labeled data compared to training a model from scratch [112]. Consequently, self-supervised learning not only improves the efficiency of NLP model training but also enables them to perform effectively on tasks where obtaining large amounts of labeled data might be challenging.

III-F Long Sequence Language Models

Long sequence language models are neural network architectures specifically designed to effectively handle long textual input sequences by leveraging the Transformer architecture [113]. While various architectures can handle longer sequences, Transformers are dominant due to their self-attention mechanisms, enabling parallel processing and capturing long-range dependencies, overcoming the sequential limitations of RNNs. This, unlike traditional language models, enables long-sequence language models to efficiently capture long-range dependencies and relationships between words [113]. Several long sequence language models address the limitations of standard Transformers by introducing modifications and additional features to their architectures.

Transformer-XL. Transformer-XL is an extension of the standard Transformer model designed to overcome the limitations of fixed-length contexts in traditional models [114]. It addresses the inherent limitation of the standard Transformer model, which employs a fixed-length context window, by employing two advanced mechanisms. These mechanisms enable the model to learn dependencies beyond a fixed length in language modeling and retain information from previous segments of the input sequence, thus enhancing its ability to process longer sequences more effectively [114]. The first mechanism, segment-level recurrence, allows the model to reuse hidden states from previous segments by propagating them through recurrent connections. This enables information flow across segments, facilitating the retention of context from previous segments and extending the context beyond a fixed length. Incorporating recurrence at the segment level empowers Transformer-XL to capture longer-term dependencies in the data [114]. In addition to the segment-level recurrence mechanism, Transformer-XL employs a novel relative positional encoding scheme [114]. This encoding scheme is crucial for enabling state reuse without causing temporal confusion, thereby allowing the model to effectively capture dependencies across longer sequences. By utilizing relative positional encodings instead of absolute ones, Transformer-XL ensures that information can be propagated across longer sequences without sacrificing temporal coherence. This encoding scheme plays a vital role in allowing the model to learn dependencies that extend beyond the fixed context length [114]. Furthermore, Transformer-XL incorporates a state reuse mechanism by caching a sequence of hidden states from previous segments, which can be reused during evaluation. As demonstrated in [114], this state reuse mechanism significantly accelerates evaluation and enables the model to maintain context from earlier segments, contributing to its ability to capture long-term dependencies in sequences.

XLNet Architecture. XLNet represents a pre-training method for NLU tasks [115]. Building upon BERT’s bidirectional context modeling [20], XLNet addresses its limitations, such as the fixed-length context constraint. Unlike BERT, which relies on masked language modeling, XLNet achieves bidirectional context learning by maximizing the expected likelihood over all permutations of the factorization order [115]. By utilizing an autoregressive formulation, XLNet ensures consistency between pretraining and fine-tuning stages, a limitation observed in BERT. As shown in Equation 24, instead of predicting the next word in a sentence given all previous words, XLNet predicts based on randomly chosen permutations of the input sequence. This approach encourages the model to consider all possible input permutations, effectively capturing the dependencies within the sequence. This involves randomly shuffling the order of elements in the sequence and then predicting each element based on its permuted context. In Equation 24, $x_{1},x_{2},\ldots,x_{n}$ represents the input sequence, $\pi$ is a random permutation of indices, and $P\left(x_{\pi(i)}\mid x_{1},x_{2},\ldots,x_{\pi(i-1)}\right)$ is the conditional probability of predicting the $i^{th}$ token given the previously predicted tokens and the current input sequence. During training, XLNet receives permuted sequences as input and predicts each element based on the surrounding elements in the shuffled order [115]. This forces the model to learn contextual representations that are not dependent on the order of elements. Additionally, as shown in Equation 25, $x_{1},x_{2},\ldots,x_{n}$ denotes the the input sequence, and $P\left(x_{i}\mid x_{1},x_{2},\ldots,x_{i-1}\right)$ is the conditional probability of predicting the $i^{th}$ token given the previously predicted tokens $x_{1},x_{2},\ldots,x_{n}$. Equation 25 represents the probability of generating the entire sequence $x_{1},x_{2},\ldots,x_{n}$ by factorizing it into conditional probabilities conditioned on the previously generated tokens. XLNet incorporates a generalized autoregressive objective, similar to the one used in GPT models, allowing diverse and coherent text generation. In Equation 25, Integrating ideas from Transformer-XL enhances XLNet’s ability to handle long-range dependencies and capture contextual information efficiently [115]. XLNet can be applied across a wide range of NLP tasks including question answering, natural language inference, sentiment analysis, and document ranking [115]. Its generalized autoregressive pretraining method enables effective handling of bidirectional contexts, long-range dependencies, and ensures consistency across pretraining and fine-tuning stages. Furthermore, XLNet’s integration of Transformer-XL and advanced architectural designs improves performance on tasks involving longer text sequences and explicit reasoning.

$$ P\left(x_{1},x_{2},\ldots,x_{n}\right)=\prod_{i=1}^{n}P\left(x_{\pi(i)}\mid x_{1},x_{2},\ldots,x_{\pi(i-1)}\right) $$

$$ P\left(x_{1},x_{2},\ldots,x_{n}\right)=\prod_{i=1}^{n}P\left(x_{i}\mid x_{1},x_{2},\ldots,x_{i-1}\right) $$

Longformer. In the context of long sequence language models, the Longformer is a specialized architecture designed to improve the processing of long textual inputs [116]. It shares the transformer architecture’s foundation but introduces modifications to the attention mechanism to accommodate the challenges posed by long sequences. It uses a locality-sensitive attention mechanism where each token attends only to its relevant local context and a few globally important tokens [113, 116]. This attention only considers relevant subsequences around each token, improving efficiency for long sequences and it is adjusted as shown in Equation 26 where $Q_{i}$, $K_{j}$, and $V_{j}$ are the query, key, and value vectors for positions $i$ and $j$ in the input sequence, respectively. The $mask\_matrix$ is used to mask certain positions, such as preventing attending to future positions during training or ignoring padding positions. In Equation 26, the division by $\sqrt{d_{k}}$ is a scaling factor that helps stabilize the gradients during training, where $d_{k}$ is the dimensionality of the key vectors. As described in [116, 113], Longformer’s attention mechanism scales linearly with the sequence length, making it feasible to process long documents efficiently. It combines local windowed attention with task-motivated global attention. Local attention is primarily used to build contextual representations, while global attention allows Longformer to create full sequence representations for prediction [116]. In standard transformers, the self-attention mechanism considers interactions between all pairs of positions in the input sequence, leading to quadratic complexity.

$$ \operatorname{Attention}\left(Q_{i},K_{j},V_{j},\text{mask\_matrix}\right)= \tag{26} $$

$$ \operatorname{softmax}\left(\frac{Q_{i}K_{j}^{T}}{\sqrt{d_{k}}}+\text{mask\_matrix}_{ij}\right)\cdot V_{j} $$

Sparse Transformers. The standard transformer’s attention mechanism calculates attention scores for all pairs of positions in a sequence, leading to quadratic time complexity [18]. Sparse Transformers address this issue by considering only a subset of positions during attention computation [117]. This introduces sparsity, significantly reducing memory requirements and computational load, making them suitable for longer sequences [117]. As shown in Equation 27, the sparse Transformer is a modified version of the standard attention mechanism used in transformers [117]. Sparse Transformer’s attention, given a sequence of input embeddings $X$ with dimensions $N\times d$, where $N$ is the sequence length and $d$ the embedding dimension, the attention scores for position $i$ attending to position $j$ can be computed as shown in Equation 27, where $Sp$ represents the sparse attention, $Q_{i}$ represents the query vector for position $i$, $K_{j}$ denotes the key vector for position $j$, ${Q_{i}K_{j}^{T}}$ represents the dot product of query and the key vectors, capturing the pairwise interactions between positions in the input sequence, $V_{j}$ represents the value vector for position $j$, and $M_{ij}$ denotes a binary mask element indicating whether vector position $i$ attends to vector position $j$. This demonstrates that the attention mechanism is computed for each pair of positions $i$ and $j$ based on their corresponding query, key, and value vectors. In global sparse attention, the mask $M_{ij}$ is generated by randomly selecting a fixed number of positions for each position $i$ to attend to. This introduces sparsity by limiting the attention to a small subset of positions in the sequence [117]. However, for local sparse attention, the mask $M_{ij}$ ensures that each position attends to a nearby local neighborhood. This reduces the computational complexity associated with attending to all positions and helps capture short-range dependencies efficiently [117]. The division by $\sqrt{d_{k}}$ serves as a scaling factor for numerical stability, where $d_{k}$ represents the dimensionality of the key vectors. Additionally, the binary mask $M_{ij}$ controls the sparsity pattern by allowing only certain positions to contribute to the attention scores. The $softmax$ function is applied to the masked and scaled dot product to normalize the scores and finally, the result is multiplied element-wise with the value matrix $V_{j}$. Some variations of Sparse Transformers incorporate adaptively determined sparsity based on the input sequence, task, or training phase [118], enhancing the model’s flexibility and performance in handling diverse sequences [118].

$$ \operatorname{Sp}\left(Q_{i},K_{j},V_{j},M_{ij}\right)=\operatorname{softmax}\left(\frac{Q_{i}K_{j}^{T}\cdot M_{ij}}{\sqrt{d_{k}}}\right)\cdot V_{j} $$

III-G Applications of LLMs

LLMs are a specific type of Generative AI designed primarily for generating and understanding human language. In addition to the applications of Generative AI explained in Section II, LLMs can be employed for various other important tasks, such as the following.

Language Understanding. In the context of NLU, LLMs are employed to extract meaning from human language. LLMs are being used for a variety of NLU [106] and other language-related tasks, including sentiment analysis and named entity recognition [4]. These models can analyze and comprehend the context of a given text, making them valuable for a wide range of applications.

Machine Translation. In the context of machine translation, LLMs are used to automatically translate text between different languages [119]. For example, Google Translate utilizes LLMs to seamlessly translate text, documents, and websites from one language into another. This capability demonstrates the practical utility of LLMs in bridging language barriers and enhancing communication, achieved through training on extensive multilingual datasets [7]. The quality of translation relies on the underlying capabilities of LLMs for natural language understanding and generation. The work in [107] introduced the concept of attention to neural machine translation architecture, leading to significant advancements in language translation quality.

Question Answering. LLMs are effectively employed in question-answering tasks across a variety of topics, enabling them to provide relevant answers to user queries [4, 3]. This capability has applications in virtual assistants, information retrieval systems, and educational platforms. For example, the AI assistant from Google can answer questions about a variety of topics, such as current events, history, and science.

Chatbots. The NLP capabilities of LLMs contribute significantly to the development of intelligent chatbots [99]. This adaptability enhances the overall user experience, making interactions with virtual assistants more intuitive and effective. LLMs are widely employed in creating chatbots for customer support and other interactive applications, enabling these intelligent virtual assistants to engage with humans, answer queries, and help in a natural and informative way [26]. The ability of LLMs to understand and respond to natural languages has opened up new possibilities in customer service, education, entertainment, and healthcare [120]. For example, companies like Facebook and Microsoft have successfully integrated LLMs in their chatbot systems, such as Facebook’s Messenger platform and Microsoft’s Azure Bot Service. These platforms utilize the power of LLMs to provide users with personalized and context-aware responses, demonstrating the practical applications of these models in real-world interactive environments.

Speech Recognition. Older speech recognition systems often relied on RNNs or hybrid models combining HMMs (HMMs) with DNNs (DNNs) [121, 122]. However, these approaches faced limitations. RNNs process input sequences one element at a time, leading to slow processing and difficulties handling long-range dependencies in audio signals [123]. Additionally, hybrid models were complex and required careful integration of separate components. To address these limitations, researchers have explored and applied LLMs to speech recognition tasks, yielding promising results [124]. The core technology for speech recognition remains ASR (ASR) models specifically trained on vast amounts of speech data. These models excel at converting audio features into text but lack the broader language understanding capabilities of LLMs. While traditional ASR systems often rely on specialized architectures, the use of LLMs, particularly transformer-based models, has gained attention for end-to-end speech recognition [18]. LLMs can analyze the output of ASR models and suggest corrections based on their understanding of language and context, improving the accuracy of transcriptions, especially in noisy environments or with unclear pronunciations [125]. Additionally, LLMs can be leveraged to provide context to the speech recognition process. By considering surrounding text or information about the speaker and situation, LLMs can assist ASR models in making better decisions about what is being said.

Text Summarization. LLMs have demonstrated successful applications in various text summarization tasks, such as summarizing documents and news articles [24, 126]. For example, the work presented in [3] introduces a sequence-to-sequence pre-training model, which has proven highly effective in abstractive summarization tasks. Modern LLMs, empowered with powerful NLP capabilities, can understand the context of a document, and generate concise and coherent summaries quickly while preserving the overall meaning of the original text.

Code Completion. In addition to the capabilities of LLMs to generate human-like text and perform various NLP tasks, LLMs have also demonstrated the ability to understand the context of code and generate relevant and accurate code suggestions [127]. Code completion with LLMs involves predicting the next set of characters in a code snippet based on the provided context [128]. These models leverage their extensive pre-trained knowledge of programming languages and coding patterns to generate pertinent code suggestions [129]. This approach has been shown to improve developer productivity [130].