1 Introduction
Artificial Intelligence, and more specifically, Generative AI, is emerging as an important technology. Over the past few years a number of prominent generative AI technologies have been developed and have received widespread attention; ranging from text generation via large language models (ChatGPT, Claude, Llama, Gemini, DeepSeek, Qwen, etc), image generation via diffusion models (Dall-E, Stable Diffusion, etc), to scientific generative AI techniques used for protein generation (e.g., Watson et al., 2023, etc), DNA sequence editing (e.g., Ruffolo et al., 2025, etc), among others.
Such methods have been quickly adopted by end users and institutions, both via direct usage, as well as integrated in other tools such as code assistants and web search agents. The scientific community has shown significant interest in using generative AI models, achieving a number of breakthrough results (see e.g., Davies et al., 2021, Hayes et al., 2025, etc), culminating in a 2024 Nobel Prize in Chemistry awarded in part for work with a significant component in protein structure design and generation (The Royal Swedish Academy of Sciences, 2024).
Yet, the adoption of generative AI (GenAI) methods more generally is hindered by their lack of reliability (see e.g., Farquhar et al., 2024, Strauss et al., 2025, Manduchi et al., 2025, etc). At their core, these methods rely on sampling from probability distributions over complex spaces that are learned from huge datasets. At the outset, GenAI does not provide any guarantees about correctness, safety, or any other desired criteria. While the performance and reliability of GenAI models is increasing steadily, so far, issues around reliability have not been successfully eliminated.
Generative AIThe construction of probabilistic models over large semantic spaces (text, images, etc.) that allows sampling from these models, given certain inputs.
Statistical methods offer potential opportunities to improve the reliability of GenAI systems. In this paper, we review several examples, highlighting statistical methods with proven or potential applications in generative AI. We focus on four topics: improving and changing the behavior of systems, diagnostics and uncertainty quantification, AI evaluation, as well as interventions and experiment design. We highlight here that the approaches we discuss are as of now mainly in the research phase, and they are usually not yet deployed in mainstream generative AI products. Their eventual usefulness remains to be determined.
1.1 About This Review
Generative AI models are commonly studied separately, for each specific modality that they pertain to (text, images, video, etc), or based on the underlying technology (diffusion models, large language models or LLMs, etc). There are already a few reviews with significant coverage of statistics related to these topics individually, including Chen et al. (2024), Zhang et al. (2025) for diffusion models, Suh & Cheng (2024) for deep learning more generally, but also touching on generative models, and Ji et al. (2025) for language models.
Our focus is different, rendering our work largely non-overlapping with the above works. We focus on techniques that are applicable to all generative AI models, regardless of their modality. Moreover, with a few exceptions, do not focus on statistical methods that are applicable to generative AI only when the tasks of interest are essentially simple classification/regression tasks (e.g., multiple-choice question-answering with LLMs). For these, there are already numerous useful references.
We also focus specifically on statistical methodology for AI and omit discussion about statistical theory of generative AI, as well as statistics-adjacent methods that primarily leverage optimization or other techniques. Further, we omit certain topics, such as watermarking and ranking from human preferences, which have already been discussed in detail in the above works. Due to space limitations, we mainly consider simple methods that have clear theoretical motivation and often provable guarantees. Moreover, we also omit discussion of how generative AI models can be used to improve statistical analysis, see e.g., Bashari et al. (2025) for a representative example.
Target audience. Our target audience includes statisticians eager to see how their expertise can drive impact in generative AI, AI researchers interested in how statistical methods can strengthen their tools, and scientists looking to better understand this emerging area. For this reason, our paper aims to be largely self-contained, with prerequisites that include knowledge of introductory undergraduate-level probability and statistics, and a basic familiarity with AI at the advanced undergraduate level.
1.2 What is Generative AI?
| Generative AI Model $\hat{p}$ | Input Space $\mathcal{X}$ | Output Space $\mathcal{Y}$ |
|---|---|---|
| Language Models | Text | Text |
| Diffusion Models | Text | Images |
| Multimodal Language Models | Images, text | Images, text, sound, video |
| Protein Structure Generation | Amino acid sequence | 3D structure |
Table 1: Representative types of generative AI models and their input and output spaces.
Generative AI usually refers to the use of generative models, which are learned probability distributions one can sample from. Concretely, consider an input space $\mathcal{X}$ (e.g., images, text, documents, their combinations, etc., represented in an appropriate way) and an output space $\mathcal{Y}$ (similarly, this could be images, text, audio, video, etc). See Table 1 for some examples.
Formally speaking, this includes as a special case standard statistical machine learning problems such as classification (when $\mathcal{Y}$ consists of the classes) and regression (when $\mathcal{Y}=\mathbb{R}$, for instance). However, the cases of interest in generative AI are usually high-dimensional spaces $\mathcal{Y}$ representing objects that are semantically meaningful to humans, such as text—viewed as a sequence of symbols $(x_{1},\ldots,x_{k})\in V^{k}$ for a finite set $V$ of symbols—or images, viewed as tensors representing pixels. Generative ModelA generative model $\hat{p}$ provides a way to sample an output $Y\sim\hat{p}(\cdot\mid x)$ from the conditional distribution of $\hat{p}$ given any input $x\in\mathcal{X}$.

Figure 1: General workflow of a generative model: inputs (e.g., text prompts, images) are processed through a black-box model to produce outputs.
Generative AI models are often designed for interaction with humans. A simple protocol is as follows: The user inputs a specific $x\in\mathcal{X}$, for instance, a text prompt such as “How can I fix a broken lamp?”. Then, the generative model $\hat{p}$ provides a way to draw a sample $Y\sim\hat{p}(\cdot\mid x)$ from the conditional distribution $\hat{p}$ given $x$; for instance, a textual response by the language model such as “To fix a broken lamp, you need to […]”. This is then returned to the user. See Figure 1 for an illustration. The ability to provide a user input corresponds to being able to sample conditionally from the generative model. This crucially unlocks a huge range of applications, by being able to be responsive to the specific needs of the user.[^1] The interaction can also continue. For simplicity, we will mostly restrict our discussion to one round of interaction.
1.3 How is a Generative Model Learned?
The GenAI model $\hat{p}$ is usually obtained by empirical loss minimization, in a manner that is conceptually similar to that used in most standard statistical modeling and machine learning. This is performed by running an algorithm—often a stochastic gradient descent-based method or a variant—aiming to minimize a loss function over a large function class using a massive data set.
For instance, for language models, the training data consists of text represented as a collection of sequences $x=(x_{1},x_{2},\ldots,x_{k})$, where for a finite set $V$ usually referred to as a vocabulary, each $x_{j}\in V$, $j\leqslant k$. The length $k$ of the strings can vary, up to a so-called context length $L$. Instead of viewing text as a sequence of letters, usually, text is encoded in tokens which are adjacent groups of letters that can offer more efficiency in the modeling process. For instance, “encoded” might consist of the tokens “en+code+d”.
The loss used is often the negative log-likelihood $\theta\mapsto-\sum_{x\in\mathcal{D}}\log p_{\theta}(x)$. The function class $\theta\mapsto p_{\theta}$ usually consists of huge neural nets parametrized in very special ways, with up to hundreds of billions of parameters. The dataset used for training consists of text data crawled from the internet, enriched with high information content (Wikipedia, arXiv), and other sources such as books. Typical costs for training powerful Generative AI models can start from millions of US dollars, which means that only organizations with significant financial resources can perform the initial training.
1.4 Access Mode to the Generative Model
An important consideration is the mode of access that we have to the generative model of interest. At the time of writing, the most powerful GenAI models are closed-source and run by commercial providers on their own cluster infrastructure, accessible only through querying. This leads to a black box mode of access, meaning that for any given input $x$, we can only observe the output $Y$, but not any internal components of the generation process $\hat{p}$. Sometimes some additional information is provided in a gray box access mode; for instance, the probability $\hat{p}(Y|x)$ may also be returned. Black Box AccessAn access model where we can only observe the output of a GenAI model, and not its internal workings.
Open-source or open-weights GenAI models may be run on local machines depending on the available hardware.[^2] In such cases, it is possible to inspect the internal workings of the models. However, since generative models tend to be highly complicated neural networks, using the internal information is challenging. Therefore, to maintain generality, we will usually focus on methods applicable to black box GenAI models. In a few cases, we will also discuss methods that require gray or white box access.