IV Challenges of Generative AI and LLMs
Despite their wide range of immense potential for society, Generative AI and LLMs also pose several critical challenges that need to be carefully considered and addressed. These challenges include:
Bias and Fairness. One of the main challenges associated with Generative AI and LLMs is the inheriting biases from the training data, which can lead to biased, unfair, and discriminatory outputs. Biased outputs from Generative AI and LLMs can have significant real-world consequences. For example, biased hiring algorithms may discriminate against certain job applicants. Potential bias problems like these can be mitigated by developing algorithms that are explicitly designed to be fair and unbiased by using approaches such as fairness-aware training [131], counterfactual analysis [132, 133, 134], and adversarial debiasing [135].
Interpretability. Understanding and interpreting the decision-making process of LLMs presents a significant challenge. The inherent lack of interpretability in these models raises serious concerns, especially in critical applications that require explainable decision-making. Addressing interpretability challenges in Generative AI and LLMs involves several approaches. One solution is to design LLMs with inherent explainability features, such as employing interpretable model architectures and incorporating constraints that promote understandable decision-making. Another approach is to develop advanced techniques that provide insights into the inner workings of LLMs, such as saliency maps [136], attention mechanisms [18], and feature attribution methods. Additionally, implementing post-hoc interpretability methods [137, 138] including feature importance analysis and model-agnostic interpretation techniques, can offer valuable insights into the factors influencing the outputs of the model.
Fine-tuning and Adaptability. Fine-tuning large LLMs for specific domains is challenging due to their inherent limitations in generalization. In addition to their limited ability to generalize, LLMs may face difficulty in understanding and reasoning complex concepts, hindering their ability to adapt to new tasks. Addressing the challenges associated with fine-tuning and adaptability in Generative AI and LLMs involves exploring various approaches. One approach involves employing transfer learning techniques that leverage knowledge from pre-trained models on diverse datasets, allowing the model to capture a broader range of knowledge by accelerating learning and improving generalization [139, 140]. Additionally, incorporating domain-specific data during fine-tuning can enhance the model’s adaptability to particular tasks ensuring it learns domain-specific patterns and relationships. Incorporating symbolic reasoning capabilities into LLMs can also enhance their ability to understand and manipulate abstract concepts [141]. Leveraging meta-learning techniques to enable LLMs to learn how to quickly learn also improves their ability to adapt to new tasks and data distributions [142].
Domain Adaptation. Most of the high-performing models being released are already fine-tuned for instruction-following. However, adapting these pre-trained LLMs, which have been specifically fine-tuned for a specific domain (such as chat), to a new task (such as generating text formats or answering your questions) not formatted for instruction-following without compromising its performance in the original domain is challenging. The challenge lies in preserving the model’s ability to understand and follow instructions while also enabling it to generate coherent and informative text in the new domain. This requires careful consideration of the training data, the model architecture, and the fine-tuning process. However, fine-tuning LLMs for an entirely new domain introduces the risk of negative transfer [143]. This occurs when the model’s new knowledge conflicts with its existing knowledge. Additionally, domain adaptation often requires access to a large amount of high-quality data from the new domain. This can be challenging to obtain, especially for specialized domains. Potential strategies for addressing this challenge include leveraging the weights of the pre-trained LLMs as a starting point for the fine-tuning process, synthesizing additional data from the new domain to supplement the existing data, and simultaneous multi-task training involving both the original and new tasks.
Data Privacy and Security. LLMs are trained on massive and diverse datasets that may contain sensitive personal information. The potential for unintentional disclosure of private or sensitive information during text generation is a significant concern. For instance, when applied in healthcare, the use of LLMs raises concerns regarding patient privacy and the potential for misdiagnosis. There is also a risk of AI systems being exploited for malicious purposes, such as generating fake identities, which raises privacy concerns. This, for example, has caused ChatGPT to be temporarily outlawed in Italy[^12][^13]. Addressing privacy concerns in Generative AI and LLMs requires a multifaceted approach that includes enhancing model training with privacy-preserving techniques, such as federated learning, homomorphic encryption, or differential privacy [144, 145]. Additionally, fine-tuning models on curated datasets that exclude sensitive information can help minimize the risk of unintentional disclosures. Ethical guidelines and regulations specific to AI applications, such as in healthcare, can provide further safeguards against privacy breaches [146, 147]. LLMs should be able to handle adversarial attacks, noisy data, and out-of-distribution inputs. In addition to this, it is worth mentioning that beyond model privacy, addressing concerns related to the privacy and security of the training and deployment data itself is important.
Computational Cost. Training and deploying LLMs demand significant computational resources. This poses infrastructure challenges, energy consumption particularly for large-scale deployments, and accessibility of high-performance computing resources. As shown in Figure 1, the increase in model sizes comes with challenges related to computational requirements and resource accessibility. Reducing the computational cost of LLMs involves several approaches. Firstly, optimizing model architectures and algorithms can enhance efficiency, reducing the computational burden without compromising performance. Secondly, leveraging distributed computing frameworks and specialized hardware accelerators, such as GPUs or TPUs (TPUs), can significantly improve training speed and resource utilization [148]. In addition to this, employing quantization techniques [149] to models that have already been trained[^14] is also important.
Deepfake Generation. Generative AI models are widely used for deepfake generation [150]. Deepfakes utilize various generative models, including GANs, to manipulate or generate realistic-looking content, primarily for image or video production [151, 152]. Despite their potential applications in various domains including education and entertainment, deepfakes also pose several potential risks due to their potential misuse, including the spread of misinformation and identity theft [153]. The deepfakes technology can be exploited to create fake videos or audio recordings of individuals leading to the spread of misinformation or disinformation, which can have devastating consequences for individuals and society. It is, therefore, important to develop advanced techniques to mitigate the risks associated with deepfakes.
Human-AI Collaboration. LLMs should be designed to enable seamless human-AI collaboration, enabling them to effectively understand and respond to human instructions and provide clear explanations for their outputs [154]. To achieve effective human-AI collaboration, it is important to integrate humans into the design process of LLMs to ensure that they are aligned with human needs and expectations [155, 156]. To incorporate human feedback into the training process, we can utilize techniques such as RLHF (RLHF) [157, 158] and DPO (DPO) [159] for training RL (RL) agents using human feedback. Additionally, employing XAI (XAI) techniques for LLMs can enhance the transparency and understandability of their decision-making processes [160]. Developing natural language interfaces that facilitate natural human-LLM interactions is another key aspect of enhancing human-AI collaboration [161]. Conversational AI, intelligent chatbots, and voice assistants are examples of technologies that enable intuitive human-AI interactions.
Long-Term Planning. Generative models, particularly autoregressive models that generate text one token at a time, face challenges in long-term planning [162]. These models tend to focus on the immediate local context, making it difficult to maintain consistency over longer text passages. This limitation comes from the model’s lack of a global view of the entire sequence it generates. Additionally, autoregressive models struggle to plan for situations with future uncertainties. To address the long-term planning challenge with LLMs, we can employ several approaches including hierarchical attention, which allows LLMs to focus on different parts of the input at different times that can help the models capture long-range dependencies [163]. Equipping LLMs with memory that allows them to store information about the past, which can be used to inform future decisions, is another approach to address this challenge [164].
Limited Context Window. Having a limited context window is a fundamental challenge for LLMs since they can only process a limited amount of text at a time. This limitation comes from their reliance on attention mechanisms [18], which allow them to focus on the most relevant parts of the text when generating content. The context window defines the number of tokens considered by the model during prediction, and a smaller context window can limit the model’s ability to understand and generate a contextually relevant text, especially in long passages or documents. Several techniques can be employed to address the challenge of a limited context window. A common approach involves using hierarchical attention, which enables models to focus on different levels of context [163]. Additionally, the parallel context window approach allows for parallel processing of multiple context windows [165]. This feature allows the models to store information beyond the immediate context window, enabling better handling of long-term dependencies [166].
Long-Term Memory. LLMs are trained on a massive corpus of text and code, but their completely stateless nature limits their ability to store and retrieve information from past experiences [167]. This inherent lack of explicit memory restricts their ability to maintain context and engage in natural conversations, leading to less coherent responses, especially across multi-turn dialogues or tasks requiring information retention. Without the ability to remember past interactions, LLMs cannot personalize their responses to specific users. This means they cannot adapt their communication style based on the user’s preferences, interests, or previous interactions. Challenges associated with this limitation include issues of consistency and task continuity. To address these challenges, various approaches and techniques can be considered. Beyond context window techniques, integrating external memory mechanisms like memory networks or attention mechanisms with an external memory matrix can enhance the model’s ability to access and update information across different turns [168]. Alternatively, designing applications that externally maintain session-based context allows the model to reference past interactions within a session. Additionally, retrieval-based techniques enable LLMs to access relevant information from past conversations or external sources during inference, enhancing their ability to maintain context and deliver more consistent responses [169].
Measuring Capability and Quality. Traditional statistical quality measures, such as Accuracy and F-score do not easily translate to generative tasks [170], especially long-form generative tasks. Furthermore, the accessibility of test sets in numerous benchmark datasets provides an avenue for the potential manipulation of leaderboards by unethical practitioners. This involves the inappropriate training of models on the test set, a practice likely employed by researchers seeking funding through achieving top positions on public leaderboards, such as Hugging Face’s Open LLM Leaderboard[^15]. At the time of writing this paper, a 7 billion parameter model is outperforming numerous 70 billion parameter models. A prospective and pragmatic approach to appraising model outputs is to utilize an auxiliary model for evaluating the generated content from the original model [171]. However, this methodology may prove ineffective if the judgment model lacks training within the specific domain it is employed to assess.
A Concerning Trend Towards “Closed” Science. As models transition from experimental endeavors to commercially viable products, there is a diminishing inclination to openly share the progress achieved within research laboratories[^16]. This shift poses a significant obstacle to the collaborative advancement of knowledge, hindering the ability to build upon established foundations when essential details are withheld. Furthermore, replicating published results becomes arduous when the prompts employed in the experimentation are not disclosed, since subtle alterations to prompts can, in some cases, significantly affect the performance of the model. Compounding these concerns, accessing the necessary resources to reproduce results often entails financial obligations to the publishers of the models, creating yet another barrier to entry into the scientific landscape for low-resource researchers. This situation prompts reflection on the current situation and the potential impediments it imposes on the pursuit of knowledge and innovation.