CHAPTER 3
Artificial Intelligence in Drug Discovery and Development
39
3.1 AI-BASED DRUG DESIGN AND TARGET IDENTIFICATION
Artificial intelligence (AI) has emerged as an important technology for improving the efficiency and precision of modern drug discovery. Traditional drug discovery is a lengthy and expensive process involving disease understanding, target identification, hit discovery, lead optimization, preclinical evaluation, and clinical development. AI-based approaches can assist researchers by analyzing large and heterogeneous datasets, identifying relationships that may be difficult to detect manually, predicting molecular properties, and prioritizing promising drug candidates. Machine learning (ML), deep learning (DL), natural language processing, graph neural networks, and generative AI are increasingly being integrated into computer-aided drug discovery workflows (Singh et al., 2024). (Figure 3.1)
Figure 3.1: AI-Based Drug Design and Target Identification
3.1.1 AI-Based Target Identification
Target identification is one of the earliest and most critical stages of drug discovery. A drug target is generally a biological molecule, such as a protein, enzyme, receptor, nucleic acid, or signaling component, whose modulation can produce a desired therapeutic effect. Selecting an inappropriate target can
40
result in significant failure during later stages of development. AI can improve target identification by integrating information from genomics, transcriptomics, proteomics, metabolomics, electronic health records, biomedical literature, molecular interaction databases, and clinical datasets.
Machine learning models can identify associations between genes, proteins, diseases, and phenotypes. Network-based approaches are particularly useful because diseases are often caused by complex interactions among multiple biological pathways rather than by a single molecular component. AI can therefore analyze biological networks and identify highly connected or functionally important molecules that may represent potential therapeutic targets. Recent approaches increasingly combine network biology, multimodal data integration, and perturbation-aware modelling to identify context-specific and functionally actionable targets (Patne et al., 2024).
Natural language processing (NLP) also contributes to target discovery by extracting information from scientific publications, patents, clinical reports, and biomedical databases. Millions of scientific articles contain information about gene–disease relationships, protein interactions, mechanisms of disease, and experimental findings. AI-based text-mining systems can process this information rapidly and identify previously unrecognized relationships. This allows researchers to construct knowledge graphs connecting diseases, genes, proteins, pathways, drugs, and phenotypic outcomes. Such approaches can help prioritize targets for experimental validation.
Another important contribution of AI is the prediction of protein structure. The development of AlphaFold demonstrated the ability of deep-learning methods to predict highly accurate three-dimensional protein structures from amino-acid sequences (Jumper et al., 2021). Structural information is highly valuable for target assessment because understanding the three-dimensional architecture of a protein can reveal binding pockets, active sites, conformational features, and opportunities for interaction with drug molecules. However, predicted structures do not automatically establish that a protein is therapeutically relevant; biological validation and experimental evidence remain essential.
41
3.1.2 AI-Based Drug Design
Once a promising target has been identified, AI can support the design and prioritization of molecules capable of interacting with that target. Conventional computer-aided drug design commonly relies on molecular docking, quantitative structure–activity relationships (QSAR), molecular dynamics, and pharmacophore modelling. AI extends these approaches by learning relationships between molecular structures and biological properties from large datasets.
One important application is virtual screening. AI models can predict the probability that a compound will interact with a particular target, allowing researchers to prioritize a smaller number of molecules for laboratory testing. This can reduce the number of compounds requiring expensive experimental screening. Deep-learning architectures, including graph neural networks and attention-based models, can represent molecules as graphs or sequences and predict properties such as binding affinity, activity, toxicity, solubility, and pharmacokinetic characteristics (Lavecchia, 2024).
Generative AI has further expanded the possibilities of drug design. Instead of simply selecting compounds from an existing chemical library, generative models can create new molecular structures according to predefined objectives. Variational autoencoders, generative adversarial networks, recurrent neural networks, normalizing flows, and Transformer-based models have been investigated for de novo molecular design. These systems can generate molecules optimized for multiple characteristics, such as target affinity, selectivity, molecular stability, and drug-like properties (Lavecchia,
2024). Structure-based generative design is particularly promising because AI models can incorporate information about the three-dimensional structure of a target protein while generating candidate molecules. Recent research has focused on protein-structure-based molecular generation, enabling models to explore chemical structures compatible with specific binding sites (Xu et al., 2024).
3.1.2 Advantages and Challenges
AI-based drug design can potentially reduce the time required for hit identification, improve molecular prioritization, explore larger chemical
42
spaces, and support more systematic decision-making. It can also facilitate drug repurposing by identifying new relationships between existing drugs, molecular targets, and diseases. Nevertheless, AI does not eliminate the fundamental uncertainties of drug discovery.
Model performance depends heavily on the quality, quantity, and representativeness of training data. Data scarcity, experimental bias, inconsistent measurements, model interpretability, and difficulties in translating computational predictions into biological effects remain important challenges (Gangwal et al., 2024).
Therefore, AI should be viewed as an augmentation of experimental drug discovery rather than a complete replacement for laboratory research. The most effective workflow combines computational prediction with biochemical assays, cellular experiments, structural studies, animal models, and ultimately clinical evaluation. AI can narrow the search space and prioritize hypotheses, while experimental science determines whether those predictions are biologically and clinically meaningful.
Overall, AI-based target identification and drug design represent a major evolution in pharmaceutical research. By integrating biological knowledge with advanced computational models, AI enables researchers to move from conventional trial-and-error approaches toward more data-driven and hypothesis-guided discovery. Continued improvements in multimodal data integration, protein structure prediction, generative modelling, explainable AI, and experimental validation are expected to further strengthen the role of AI in the development of safer, more selective, and potentially more effective therapeutics.
3.2 VIRTUAL SCREENING AND MOLECULAR MODELING
Virtual screening and molecular modeling have become important components of modern computer-aided drug discovery. Instead of experimentally testing thousands or millions of chemical compounds individually, computational methods can prioritize molecules that are more likely to interact with a selected biological target. This approach helps researchers reduce the size of experimental libraries, lower screening costs, and focus laboratory resources on the most promising candidates. Virtual screening is broadly divided into
43
ligand-based virtual screening (LBVS) and structure-based virtual screening (SBVS). LBVS relies primarily on information from known active compounds, whereas SBVS uses the three-dimensional structure of a biological target to identify molecules capable of fitting into its binding site (Cheng et al., 2012; Kučera, 2016). (Figure 3.2)
Figure 3.2: Virtual Screening and Molecular Modeling
3.2.1 Principles of Virtual Screening
The central objective of virtual screening is to identify potential “hit” compounds from a large chemical library. In a typical structure-based workflow, researchers first identify an appropriate protein target and determine or predict its three-dimensional structure. Structural information may be obtained using experimental techniques such as X-ray crystallography, nuclear magnetic resonance, or cryo-electron microscopy, or through computational protein-structure prediction and molecular modeling. The target is then prepared by assigning appropriate protonation states, removing or retaining water molecules as appropriate, adding hydrogen atoms, and defining the binding pocket (Cheng et al., 2012).
The compound library is similarly prepared. Molecules may be filtered according to molecular weight, physicochemical properties, chemical stability, or undesirable structural features. Pharmacophore filters, similarity searches,
44
and drug-likeness rules can further reduce the number of compounds that require computationally intensive analysis. This preprocessing is important because the quality of the input library can substantially influence the performance of a virtual-screening campaign.
3.2.2 Molecular Docking and Modeling
Molecular docking is one of the most widely used approaches in structure-based virtual screening. Docking attempts to predict how a small molecule, or ligand, can fit within the binding site of a protein and estimates the strength or favorability of the resulting interaction using a scoring function. The docking process generally involves two major components: pose generation and scoring. Pose generation explores possible orientations and conformations of the ligand, while scoring functions rank the resulting complexes according to predicted interaction quality (Pinzi & Rastelli, 2019).
Molecular modeling extends beyond docking by providing a broader computational representation of molecular structure, dynamics, and interactions. Molecular mechanics, molecular dynamics (MD), pharmacophore modeling, homology modeling, and free-energy calculations can all contribute to drug-design workflows. MD simulations, for example, can provide information about protein flexibility and ligand–protein interactions over time. This is particularly relevant because proteins are not rigid structures; they undergo conformational changes that can influence ligand binding. Consequently, ensemble docking, in which multiple protein conformations are used during screening, can sometimes improve the representation of receptor flexibility (Gorgulla et al., 2023).
3.2.3 Integration of Artificial Intelligence
Artificial intelligence (AI) is increasingly being integrated into virtual screening and molecular modeling. Machine-learning and deep-learning algorithms can learn relationships between molecular structures, biological activities, and target interactions from large datasets. Molecules can be represented using molecular fingerprints, descriptors, graphs, three-dimensional coordinates, or learned embeddings. Deep-learning approaches can then predict properties such as binding affinity, activity, toxicity, or drug– target interactions and use these predictions to prioritize compounds (Jiménez-Luna et al., 2021).
45
One important advantage of AI-assisted screening is its ability to process chemical and biological datasets at a scale that would be difficult to handle manually. AI models can be used as an initial filtering stage before computationally expensive docking or molecular simulation. Alternatively, docking results can be combined with machine-learning models to improve compound ranking. Recent research has increasingly explored hybrid workflows combining molecular docking, pharmacophore modeling, machine learning, and interpretable AI techniques for lead identification and optimization (Taha & Daoud, 2026).
3.2.3 Workflow of AI-Assisted Virtual Screening
A modern AI-assisted virtual-screening workflow can be summarized as:
Target identification → Protein-structure preparation → Chemical-library preparation → Molecular/property filtering → AI-based prioritization → Molecular docking → Rescoring → Molecular dynamics/free-energy analysis → Selection of candidate hits → Experimental validation
This multi-stage approach is generally more effective than relying on a single computational method. Recent developments have also enabled ultra-large virtual screening, in which computational workflows can examine hundreds of millions or even billions of compounds.
Such approaches combine highly efficient docking, machine learning, distributed computing, and intelligent compound prioritization to explore chemical space on an unprecedented scale (Chee et al., 2026).
3.2.4 Advantages and Limitations
Virtual screening offers several advantages. It can examine very large chemical libraries rapidly, prioritize compounds before laboratory testing, support the identification of structurally diverse molecules, and provide hypotheses about ligand–target interactions. It is particularly useful during the early stages of drug discovery, where selecting appropriate compounds for experimental testing can substantially influence the efficiency of subsequent development.
However, virtual screening is not a substitute for experimental validation. Docking scores do not necessarily correspond directly to experimentally
46
measured binding affinities, and scoring functions can produce false positives and false negatives. Protein flexibility, solvent effects, protonation states, metal ions, ligand entropy, and induced-fit effects can complicate accurate prediction of molecular recognition (Cheng et al., 2012). Moreover, AI models can inherit biases from their training datasets and may perform poorly when applied to chemical spaces or biological targets that differ substantially from those represented in the training data. Recent evaluations have therefore emphasized problems such as dataset bias, data leakage, limited generalizability, and the complex geometric nature of molecular recognition (Suay-García & Falcó, 2026).
3.2.5Future Perspectives
The future of virtual screening is likely to involve deeper integration between AI, molecular simulation, structural biology, and experimental chemistry. Deep-learning-based docking and scoring functions, protein-structure prediction, generative molecular design, ultra-large-scale screening, and physics-based simulations are increasingly being combined into integrated computational pipelines. Recent reviews indicate that docking-based virtual screening continues to evolve from conventional library screening toward large-scale and AI-enhanced workflows (Zou et al., 2026).
Overall, virtual screening and molecular modeling provide a computational bridge between biological target identification and experimental drug discovery. Their greatest value lies not in replacing laboratory experiments, but in making those experiments more selective and information-rich. When combined with high-quality structural data, appropriate molecular modeling, AI-based prioritization, and rigorous experimental validation, these approaches can accelerate the identification and optimization of promising drug candidates.
3.3 AI IN CLINICAL TRIALS AND DRUG REPURPOSING
Artificial intelligence (AI) is increasingly becoming an important component of modern drug development, particularly in clinical trials and drug repurposing. Traditional clinical research is often associated with lengthy timelines, high costs, complex protocols, difficulties in patient recruitment, and substantial rates of trial failure. AI and machine learning (ML) can address several of these limitations by processing large and heterogeneous datasets,
47
identifying patterns that may not be readily visible to researchers, and supporting evidence-based decisions throughout the clinical development process (Zhang et al., 2025). (Figure 3.3)
Figure 3.3: AI in Clinical Trials and Drug Repurposing
3.3.1 AI in Clinical Trials
Clinical trials represent one of the most important and resource-intensive stages of drug development. AI can be applied across almost the entire clinical-trial lifecycle, beginning with protocol development and continuing through patient recruitment, trial monitoring, data analysis, and post-trial evaluation. During trial design, machine-learning algorithms can analyze historical clinical-trial data to identify factors associated with successful or unsuccessful studies. Such analysis can support the selection of appropriate endpoints, sample sizes, eligibility criteria, treatment arms, and trial sites. AI-based approaches may also facilitate adaptive trial designs in which aspects of a study can be modified in response to accumulating evidence while maintaining appropriate statistical and regulatory controls.
48
One of the most significant applications of AI is patient recruitment and selection. Clinical trials frequently experience delays because eligible participants are difficult to identify and enroll. Natural language processing (NLP) can analyze electronic health records (EHRs), medical notes, laboratory results, and other clinical information to identify individuals who potentially satisfy trial eligibility criteria. Machine-learning models can additionally stratify participants according to demographic, clinical, genetic, or disease characteristics. This can improve the identification of suitable participants and potentially reduce recruitment time and screening costs.
AI can also support patient stratification and personalized clinical trials. By integrating genomic, proteomic, imaging, clinical, and real-world data, AI systems can identify subgroups that are more likely to respond to a particular intervention or experience adverse effects. This approach is particularly valuable in diseases with substantial biological heterogeneity, such as cancer, neurological disorders, and autoimmune diseases. More precise patient selection can improve the probability of detecting treatment effects and may contribute to more efficient clinical development.
Another important application is trial monitoring and data management. AI can continuously analyze information generated through EHRs, laboratory systems, wearable devices, imaging platforms, and electronic patient-reported outcomes. Algorithms can identify missing or inconsistent data, detect unusual patterns, and support the early identification of safety signals. AI can also assist in endpoint assessment and automate certain repetitive data-processing tasks. Nevertheless, these systems require rigorous validation because errors, biased datasets, or inappropriate algorithmic assumptions can affect clinical conclusions.
AI also has potential in predicting clinical-trial outcomes. By learning from previous trials, disease characteristics, drug properties, and patient-level information, predictive models can estimate the probability of trial success or failure. Such predictions can help pharmaceutical companies prioritize promising candidates and allocate research resources more efficiently. However, prediction should be considered a decision-support function rather than a substitute for clinical judgment, experimental evidence, or regulatory assessment.
49
3.3.2 AI in Drug Repurposing
Drug repurposing, also called drug repositioning, involves identifying a new therapeutic application for an existing drug that was originally developed or approved for another indication. Because many repurposed drugs already have information regarding pharmacokinetics, toxicity, manufacturing, dosing, or human safety, repurposing can potentially reduce some of the time, cost, and uncertainty associated with developing an entirely new medicine (Fu et al.,
2026). AI provides powerful methods for identifying potential repurposing opportunities. Conventional approaches may depend heavily on individual discoveries or experimental screening, whereas AI can systematically integrate information from drug databases, disease databases, molecular networks, genomic datasets, scientific publications, EHRs, and real-world evidence. Machine-learning and deep-learning models can examine relationships among drugs, proteins, genes, diseases, pathways, and phenotypes to predict previously unrecognized drug–disease associations. Knowledge graphs and graph neural networks (GNNs) are particularly promising for drug repurposing. These approaches represent biological and medical entities as interconnected networks, allowing algorithms to investigate complex relationships between drugs and diseases. For example, Huang et al. (2024) developed TxGNN, a graph-based foundation model designed to predict therapeutic candidates for diseases, including conditions with limited available treatments. The model illustrates how AI can extend drug-repurposing analysis beyond diseases for which well-established treatment options already exist. Generative AI is another emerging approach. Large language models can analyze and synthesize information from extensive scientific and biomedical literature to generate hypotheses about potential therapeutic relationships. In one study, generative AI was used to prioritize potential drugs for Alzheimer’s disease, after which selected candidates were examined using large clinical datasets. The findings demonstrated how generative AI can contribute to hypothesis generation and candidate prioritization, although computational predictions still require rigorous experimental and clinical validation. 50
AI-supported repurposing can therefore establish a computational-to-clinical pipeline: first, AI identifies candidate drugs; second, molecular and laboratory experiments evaluate biological plausibility; third, observational or real-world data can provide supporting evidence; and finally, appropriately designed clinical trials determine efficacy and safety for the new indication. This integration is important because an AI prediction alone does not establish therapeutic effectiveness.
Despite its considerable potential, AI-based clinical trials and drug repurposing face several challenges. Data quality, missing information, population bias, privacy, algorithmic transparency, reproducibility, cybersecurity, and regulatory acceptance remain important concerns. AI models trained on limited or nonrepresentative datasets may generate misleading predictions. Furthermore, complex models can be difficult to interpret, creating challenges when researchers or regulators need to understand why a particular recommendation was produced. Current discussions involving regulators therefore emphasize validation, human oversight, transparency, and appropriate governance when AI is incorporated into drug development.
Overall, AI is not expected to replace clinical researchers, physicians, pharmacologists, or regulatory authorities. Instead, its greatest value lies in augmenting human expertise by rapidly processing complex information and identifying patterns or hypotheses that can guide further investigation. When combined with high-quality biomedical data, rigorous validation, and appropriate regulatory oversight, AI can make clinical trials more efficient and expand the search for new therapeutic applications of existing medicines. Consequently, AI-driven clinical research and drug repurposing are likely to become increasingly important components of precision medicine and next-generation drug development.
3.4 CHALLENGES AND FUTURE SCOPE IN PHARMACEUTICAL AI
Artificial intelligence (AI) is increasingly being integrated into pharmaceutical research, covering target identification, virtual screening, molecular design, lead optimization, toxicity prediction, clinical trial design, drug repurposing, manufacturing, and post-market surveillance. The technology offers the potential to reduce the time and cost associated with conventional drug
51
development while enabling researchers to explore chemical and biological spaces that are difficult to investigate experimentally. However, the practical impact of pharmaceutical AI remains constrained by limitations in data quality, model reliability, biological complexity, validation, regulation, and ethical governance. Recent research emphasizes that improvements in computational performance alone are insufficient; successful pharmaceutical AI requires integration with experimental science and clinical decision-making (Ghislat et al., 2024; Su et al., 2026).
3.4.1 Challenges in Pharmaceutical AI
1. Data quality, availability, and heterogeneity: The effectiveness of AI models depends heavily on the quality and representativeness of training data. Pharmaceutical datasets are frequently heterogeneous, incomplete, biased, noisy, or relatively small. Experimental measurements generated by different laboratories may use different protocols, instruments, assay conditions, and reporting standards, making direct comparison difficult. In addition, negative experimental results are often underreported, producing datasets that may not accurately represent the complete chemical or biological landscape. Such problems can lead to biased models and unreliable predictions (Ghislat et al., 2024). Data scarcity is particularly important for rare diseases and novel biological targets, where only limited experimental observations may be available (Gangwal et al., 2024). 2. Limited interpretability and explainability: Many advanced AI systems, particularly deep-learning and generative models, operate as complex "black boxes." Although a model may accurately predict molecular activity, toxicity, or treatment response, researchers may not always understand why a particular prediction was generated. Lack of interpretability can reduce confidence among medicinal chemists, clinicians, regulators, and patients. Explainable AI (XAI) approaches are therefore becoming increasingly important for identifying the biological or chemical factors responsible for predictions and for supporting scientifically defensible decisions (Alam et al., 2024). 3. Translational and experimental validation: Strong performance on retrospective datasets does not necessarily translate into successful drug candidates. AI models may identify promising molecules computationally, but 52
these candidates must subsequently undergo synthesis, biochemical testing, pharmacokinetic studies, toxicological assessment, and clinical evaluation. Consequently, the gap between in silico prediction and real-world biological performance remains a major obstacle. Recent analyses highlight unrealistic benchmarks, inadequate prospective testing, and uncertainty estimation as important reasons why impressive computational results do not always produce clinically meaningful outcomes (Ghislat et al., 2024).
4. Regulatory uncertainty: Pharmaceutical products are subject to stringent regulatory requirements, and AI introduces additional questions concerning model credibility, validation, data provenance, reproducibility, change management, and accountability. Regulatory agencies must determine how AI-generated evidence should be assessed without unnecessarily restricting innovation. In January 2025, the U.S. Food and Drug Administration (FDA) issued draft guidance proposing a risk-based credibility assessment framework for AI models used to support regulatory decisions concerning drug and biological product safety, effectiveness, or quality. This represents an important step toward establishing clearer expectations for pharmaceutical AI, although regulatory frameworks will continue to evolve as technologies mature (FDA, 2025). 5. Privacy, security, and ethical concerns: Pharmaceutical AI increasingly uses clinical records, genomic information, real-world evidence, and other sensitive datasets. Improper data handling may create risks related to privacy, cybersecurity, consent, and unauthorized secondary use. AI systems may also reproduce existing demographic or socioeconomic biases if their training data are not representative. The World Health Organization emphasizes that AI in health should be developed around principles including transparency, accountability, inclusiveness, equity, and protection of human autonomy (WHO, 2021). 6. Lack of interdisciplinary expertise: Effective pharmaceutical AI requires collaboration among pharmacologists, medicinal chemists, biologists, clinicians, data scientists, computational scientists, and regulatory experts. A technically sophisticated model may still be unsuitable if it does not address a meaningful pharmaceutical problem. Therefore, future development must 53
focus not simply on building increasingly complex algorithms but on designing AI systems around clearly defined scientific and clinical needs.
3.4.2 Future Scope of Pharmaceutical AI
The future of pharmaceutical AI is likely to move from isolated prediction tools toward integrated, multimodal, and experimentally connected platforms. Generative AI is expected to support the design of novel molecules with desired potency, selectivity, physicochemical characteristics, and safety profiles. Instead of screening millions of existing compounds alone, generative systems can propose previously unexplored molecular structures for experimental evaluation (Gangwal & Lavecchia, 2024).
Another important direction is the integration of multimodal biological data, including genomics, transcriptomics, proteomics, imaging, clinical records, chemical structures, and real-world evidence. Combining these data sources may improve understanding of disease mechanisms and enable more accurate target identification and patient stratification. AI may consequently contribute to precision medicine by identifying which patient subgroups are most likely to respond to particular therapies.
Foundation models and large language models (LLMs) also have considerable future potential. These systems can assist researchers in extracting knowledge from scientific literature, generating hypotheses, interpreting experimental results, designing molecules, and integrating information from diverse biological databases. Recent developments indicate that AI is expanding beyond individual stages of drug discovery toward broader workflows spanning target identification, molecular design, lead optimization, clinical research, and drug repurposing (Su et al., 2026).
Future pharmaceutical AI will also increasingly incorporate human-in-the-loop systems, in which AI provides predictions and recommendations while scientists retain decision-making authority. This approach can combine computational scalability with human scientific judgment and experimental expertise. Similarly, AI-enabled automated laboratories may create closed-loop systems in which algorithms design experiments, robotic platforms conduct them, experimental data are automatically analyzed, and subsequent experiments are selected based on the results.
54
Regulatory science is another major area of future development. Risk-based validation frameworks, standardized reporting practices, model monitoring, and international regulatory harmonization will be essential for trustworthy adoption. The FDA has already recognized increasing AI use throughout nonclinical, clinical, manufacturing, and post-marketing stages of the drug product life cycle and is working toward responsible integration of AI into drug development (FDA, 2026).
Overall, the future of pharmaceutical AI is not likely to involve the complete replacement of researchers. Rather, AI is expected to function as an advanced scientific partner capable of processing large datasets, generating hypotheses, prioritizing experiments, and identifying patterns beyond conventional human analysis. The greatest value will emerge when high-quality data, explainable algorithms, experimental validation, interdisciplinary expertise, and responsible regulation are developed together. Thus, the long-term success of pharmaceutical AI will depend less on algorithmic sophistication alone and more on its ability to produce reproducible, biologically meaningful, clinically relevant, and regulatory acceptable outcomes.
55