2 Statistical Methods in Generative AI
Our goal is to discuss a few emerging areas of research where statistical methods or ideas can be used in generative AI. A key starting point is that AI systems can be wrong. They can make any type of mistake, and they have no guarantees by default about correctness, content, logical consistency, safety, etc.[^3] This stems intrinsically from their structure as sampling methods.[^4]
While there are a variety of engineering approaches to improve reliability, such as endowing the AI models with external tools, such as calculators, web search, or access to a computer where they can run programs, the use of these tools is in turn orchestrated by a sampling-based generative AI model, which can still have reliability problems. Moreover, while there are constrained sampling methods that aim to ensure certain basic formatting and correctness criteria, their current scope is limited; for example, at the moment they cannot ensure logical correctness.
For these reasons, statistical methods that aim to improve the behavior of generative models—sometimes with provable guarantees—are particularly significant; we begin our discussion with this topic. Crucially, to have an impact in this area, statistical methods must directly align with AI practice and goals; endowing practically useful AI-enhancement methods with desirable guarantees. To put it another way, statistical methods act as simple tunable wrappers that can be calibrated to meet explicit error budgets with finite sample guarantees.
2.1 Improving and Changing Behavior
To improve the performance of a generative AI model, there are numerous of standard approaches relying on variants of standard training (e.g., supervised fine-tuning for LLMs). Once these have been exhausted, there is room for alternative techniques that change the behavior of the generative model in a non-standard way that can conceivably improve certain accuracy metrics, for instance, by returning a trimmed version of the input from which false claims have been deleted (see e.g., Mohri & Hashimoto, 2024, etc). These techniques can be roughly categorized into changing (a) the output $y$, (b) the input $x$, or (c) the internal workings on the generative model $\hat{p}$.
Moreover, many of these techniques require a degree of hyperparameter tuning; for instance, determining how much to trim the outputs. This process of tuning can sometimes be endowed with statistical correctness guarantees (see Table 2), and so this is the first topic we review in this work.
| Technique | Type | Examples |
|---|---|---|
| Change output | Additional output type | Highlight parts of output (Sun et al., 2022, Vasconcelos et al., 2025) |
| Abstain from generation when a risk score is high (Farquhar et al., 2024, Yadkori et al., 2024) | ||
| Add “Everything Else” as a possible answer (Noorani et al., 2025) | ||
| Set of outputs | Construct prediction interval for each output coordinate (Horwitz & Hoshen, 2022, Teneggi et al., 2023) | |
| Generate set of outputs (Quach et al., 2024, Gui et al., 2024, Nag et al., 2025) | ||
| Trimmed output | Delete parts of output until correctness is achieved (Khakhar et al., 2023, Mohri & Hashimoto, 2024) | |
| Find small parent set of possible outputs in a directed acyclic graph (Zhang et al., 2024) | ||
| Regenerated output | Reformulate output until it is appropriately correct and specific (Jiang et al., 2025) | |
| Task-specific output | Train model to improve performance in downstream task (Band et al., 2024) | |
| Construct prediction intervals for latent variables of a generated output (Sankaranarayanan et al., 2022) | ||
| Interactively ask questions that maximize the informativeness of the answers (Chan et al., 2025) | ||
| Change input | Set of inputs | Retrieve sets of documents in RAG (Li et al., 2024) |
| Select prompts that control risk (Zollo et al., 2024) | ||
| Change other algorithm settings | — | Accelerate generation by early exit (Schuster et al., 2021, Schuster et al., 2022, Jazbec et al., 2024) |
| Reduce ambiguity by seeking additional input (Ren et al., 2023, Ren et al., 2024) | ||
| Control a “size” component of the sampling mechanism (Ravfogel et al., 2023, Deutschmann et al., 2024, Ulmer et al., 2024) | ||
| Switch between models when risk score is high (Overman & Bayati, 2025) |
Table 2: Types of methods that change the behavior of generative AI systems; most of them endowed with statistical guarantees. Some methods belong to multiple categories.
2.1.1 An example: Controlling the probability of refusal/abstention
To get a sense of the types of problems that can be solved, as well as the types of statistical methods that are used, we will explain one specific example in some detail. We will consider the example of abstaining from generation when a risk score is high (see e.g., Farquhar et al., 2024, Yadkori et al., 2024, etc).
Consider a given loss function[^5] $\ell:\mathcal{X}\times\mathcal{Y}\to\mathbb{R}.$ This could measure the quality or safety of an input-output pair. There are many examples, including the negative log likelihood $\ell(x,y)=-\log\hat{p}(y|x)$ specified by the generative model itself, or the negative of a pre-trained reward function (measuring for instance safety), etc. The loss could depend on both $x$ and $y$, or only on one of the two. If the loss only depends on the input $x$, it can capture either input ambiguity, or the dispersion in outputs generated by the model (Lin et al., 2024); or some combination thereof.
Refusal/AbstentionWhen a generative AI model does not return an output. Can be useful for improving safety. To improve user experience, a strategy is to refuse/refrain/abstain from answering when the loss is high. Specifically, we want to find a threshold $\tau$ such that when $\ell(x,Y)>\tau$ we should instead return a special message like ‘‘Sorry I cannot answer.’’, where $Y\sim\hat{p}(\cdot|x)$ is generated by the model $\hat{p}$. There is a trade-off: decreasing the threshold will ensure that only higher quality—lower loss—generations/answers are returned, but higher refusal also hampers utility to users.
The threshold $\tau$ can be set by standard hyperparameter tuning, by checking the loss values and abstention rates on a dataset. However, there is also a statistical approach, which can provide provable guarantees on the behavior of the system under certain conditions. This approach is based on predictive inference/conformal prediction (Vovk et al., 2005), and the ideas date back to work on tolerance regions (e.g., Wilks, 1941, Wald, 1943, etc). Predictive InferenceThe goal of endowing the outputs of predictive models—including GenAI—with statistical guarantees.
The statistical approach aims to guarantee generalization to a distribution $D$ of prompts. The goal is then to control the abstention probability over the distribution $D$, which can be written as $\Pr_{X\sim D,\,Y\sim\hat{p}(\cdot|X)}\!\big(\ell(X,Y)>\tau\big).$ We do not fully know the distribution $D$, because it represents the behavior of future users. However, we assume that we have a calibration dataset—also referred to as a validation or hold-out dataset—$D_{n}=\{X_{1},\dots,X_{n}\}$ of prompts which we view as an i.i.d. sample from $D$. This is collected based on user interactions that are representative of the distribution, and we assume that they have not been used for model training. Calibration datasetGiven a trained GenAI model, a separate dataset used to endow the model with various statistical properties.
Then, we aim to construct an estimated threshold $\hat{\tau}=\hat{\tau}(D_{n})$ using the calibration dataset $D_{n}$ such that the abstention probability[^6] is controlled at a user-specified level $\alpha>0$, i.e., $\Pr_{X\sim D,\,Y\sim\hat{p}(\cdot|X),D_{n}}\!\big(\ell(X,Y)>\hat{\tau}(D_{n})\big)\leqslant\alpha.$
The key observation is the following: suppose we generate responses $Y_{i}\sim\hat{p}(\cdot|X_{i})$ for each of our inputs $i=1,\dots,n$ from the calibration dataset. Then the values $\ell_{i}:=\ell(X_{i},Y_{i})$, $i=1,\ldots,n$ are i.i.d. random variables with the same distribution as the test loss $\ell(X,Y)$ where $X\sim D$ is a test data point and $Y\sim\hat{p}(\cdot|X)$ is a corresponding outcome sampled from the generative model. Of course, the distribution of these loss values is in general still unknown, because it depends on the unknown target distribution $D$.
Exchangeability. However, since the $\ell_{i}$ are i.i.d., conditional on the set (or multiset) of their values $S_{n+1}=\{\ell_{1},\ldots,\ell_{n},\ell(X,Y)\}$, their ordering is uniform given $S_{n+1}$. This corresponds to exchangeability, and it is their only property used here. ExchangeabilityInformally, the property that a sequence of random variables is equally likely to be presented in any order.
Therefore, assuming for simplicity of exposition that there are no ties,[^7] the rank of $\ell(X,Y)$ among $\ell_{1},\ldots,\ell_{n},\ell(X,Y)$ is distributed uniformly over $\{1,\ldots,n+1\}$, conditional on $S_{n+1}$. Now, for any $\beta\in[0,1]$, let $Q_{\beta}$ be the $\beta$-th quantile of $\{\ell_{1},\ldots,\ell_{n}\}$, namely $Q_{\beta}=\inf\{t:\#\{i:\ell_{i}\leqslant t\}\geqslant\beta n\}.$ We have that $\ell(X,Y)\geqslant Q_{(1-\alpha)(1+1/n)}$ if and only if the rank of $\ell(X,Y)$ among $\{\ell_{1},\ldots,\ell_{n},\ell(X,Y)\}$ is at most $\lfloor\alpha(n+1)\rfloor$.
Consequences of exchangeability. By exchangeability, this occurs with probability at most $\lfloor\alpha(n+1)\rfloor/(n+1)\leqslant\alpha$, conditional on $S_{n+1}$. Hence, if we choose $\hat{\tau}(D_{n}):=Q_{(1-\alpha)(1+1/n)},$ we find $\Pr\big(\ell(X,Y)\geqslant\hat{\tau}(D_{n})\big|S_{n+1})\leqslant\alpha,$ for any $S_{n+1}$. Since this holds for any set $S_{n+1}$, it also holds unconditionally, i.e., $\Pr\big(\ell(X,Y)\geqslant\hat{\tau}(D_{n})\big)\leqslant\alpha,$ as desired.
The above argument explains how, by choosing a threshold for abstention equal to a particular quantile of the calibration losses, we can control the abstention rate. All that is needed is that the test loss is exchangeable with the calibration losses.
2.1.2 An overview of applications and techniques
The above discussion is quite representative of a variety of methods designed to improve the behavior of generative AI models. Some of the common elements include: (1) introduction of a loss function; (2) introduction of a small number of tunable hyperparameters; (3) formulation of a desired goal in terms of expectations of the losses and probabilistic properties, and (4) using distribution-free or only weakly distributionally dependent probabilistic tools—such as the distribution of order statistics or concentration inequalities—to ensure the desired goal. We refer to the works cited in Table 2 for details.
For instance, one approach proposes to delete claims from the output $Y$ of a large language model until correctness is reached (Khakhar et al., 2023, Mohri & Hashimoto, 2024). This approach defines a deletion operator $\Delta$, typically implemented by another large language model, and a sequence $Y^{(k)}=\Delta(Y^{(k-1)})$, $k\geqslant 1$, starting with $Y^{(0)}=Y$. The loss function $\ell$ is defined based on whether $Y^{(k)}$ has any claim that contradicts a ground truth answer $y^{*}$ for the prompt $x$. This is also evaluated by another large language model. The tunable hyperparameter is the number $k$ of deletions; and the goal is to ensure correctness with probability at least $1-\alpha$. Then the required number of deletions can be determined based on a calibration dataset, similarly to above.
Many of the methods discussed above rely on some form of non-parametric statistics, distribution-free predictive inference, conformal prediction, and variants. The idea of distribution-free prediction sets dates back at least to the pioneering works of Wilks (1941), Wald (1943), etc. Distribution-free inference has been extensively studied in recent works (see, e.g., Saunders et al., 1999, Vovk et al., 1999, Papadopoulos et al., 2002, Vovk et al., 2005, Vovk, 2012, Lei et al., 2013, Lei & Wasserman, 2014, Lei et al., 2018, Guan, 2023, Romano et al., 2020, Dobriban & Yu, 2025, etc). Predictive inference methods have been developed under various assumptions (Geisser, 2017, Bates et al., 2021, Park et al., 2022b, Park et al., 2022a, Sesia et al., 2023, Qiu et al., 2023, Li et al., 2022, Kaur et al., 2022, Si et al., 2024, Lee et al., 2024, see, e.g.,). Overviews of the field are provided by Vovk et al. (2005), Shafer & Vovk (2008), and Angelopoulos et al. (2023).
2.2 Diagnostics and Uncertainty Quantification
When AI systems encounter problems, one should of course aim to improve the behavior of the AI system. A crucial step toward this is to precisely diagnose the problem. A variety of approaches exist for this task, ranging from constructing unit tests to fine-tuning the model. There are also a number of methods based on computing certain specific diagnostic scores (e.g., Farquhar et al., 2024, Yadkori et al., 2024, Lin et al., 2024, etc). Such diagnostics are already used in many of the methods discussed in Section 2.1 to change or improve the generative model; for instance, if a safety score for the input is low, the model can refrain from generating an output.
In this section, we are specifically interested in diagnostics that aim to quantify uncertainty, as these have close connections to probability and statistics. There are a variety of interpretations of uncertainty quantification, see e.g., Baan et al. (2023), Shorinwa et al. (2024), Liu et al. (2025), Abbasli et al. (2025), Xia et al. (2025), He et al. (2025), Campos et al. (2024), Trivedi & Nord (2025). Due to space reasons, here we can only discuss a few specific approaches, see Table 3.
| Approach | Type | Examples |
|---|---|---|
| Defining Uncertainty | Epistemic & Aleatoric Unc. | Define and estimate epistemic and aleatoric uncertainty through input clarification ensembling (Hou et al., 2024) |
| Semantic Uncertainty | Cluster outputs to capture semantic uncertainty (Kuhn et al., 2023) | |
| Soft-cluster outputs with partially overlapping meaning from black-box models (Lin et al., 2024) | ||
| Other measures | Estimate pseudo-entropy of a prompting-induced sampling distribution (Abbasi Yadkori et al., 2024) | |
| Approximate Bayesian posterior uncertainty by updating model (Yang et al., 2024, Wang et al., 2024) | ||
| Calibration | Re-calibrate probabilities in multiple choice/classification problems (Jiang et al., 2021) | |
| Calibrate uncertainty to predict performance (Huang et al., 2024, Liu et al., 2024) |
Table 3: Types and examples of uncertainty quantification methods for generative AI.
2.2.1 Epistemic and aleatoric uncertainty
We start by introducing the notions of epistemic and aleatoric uncertainty. To set the stage, we observe that given an input $x$, the output $y$ is not always uniquely determined. For instance, the query $x=$ “Write a paragraph about an economist” has ambiguity, since it does not specify a particular economist. This is sometimes referred to as epistemic uncertainty (Der Kiureghian & Ditlevsen, 2009, Hüllermeier & Waegeman, 2021). It can be reduced by collecting more information. In particular, the AI system could query “Which economist?”, to which the answer, e.g., “Adam Smith”, could greatly reduce the uncertainty of the answer to be generated.
In practice, there are usually many such sources of epistemic uncertainty for any given query. For instance, even after knowing which economist to consider, we still do not know the desired number of sentences, the target audience (children, general public, scientists, or some other group), etc. Some of these might be more important than others to the user, but either way they contribute to the uncertainty of the possible answers.
We can contrast epistemic uncertainty with aleatoric uncertainty. For instance, in the query “Choose between A and B uniformly at random.”, all information is perfectly well specified (so the epistemic uncertainty vanishes), yet there is still irreducible random uncertainty in the desired output, which is sometimes referred to as aleatoric uncertainty (Der Kiureghian & Ditlevsen, 2009). Epistemic and aleatoric uncertaintyRoughly speaking, uncertainty due to lack of knowledge, and due to random chance, respectively. Can be hard to define precisely.
While multiple definitions exist (see, e.g., Schweighofer et al. (2025)), including approaches tailored to estimating them in generative AI models (see, e.g., Hou et al. (2024)), in many cases the definition of—say—aleatoric uncertainty reduces to specifying what we choose not to predict, rather than to something intrinsically fixed.
2.2.2 Uncertainty in model generations
While the discussion in Section 2.2.1 refers to uncertainty in ideal “ground truth” answers, in practice we need to take into account that we only have an empirical model $\hat{p}$, not the ground truth; and need to handle the uncertainty in the answers generated by $\hat{p}$. Equivalently, we should quantify to what extent the model is certain. There have been several approaches aimed to extract this form of uncertainty from generative AI models.
For language models, a special ability is that they might potentially be able to express uncertainty in words. However, this capability is not guaranteed to work well by default, and special fine-tuning techniques have been developed to induce this behavior in certain special cases (Lin et al., 2022).
An approach that applies more generally to all generative models, regardless of their modality, is to compute some uncertainty or confidence score[^8] based on the input $x$, the output $y$, and/or the model $\hat{p}$ (e.g., Farquhar et al., 2024, Yadkori et al., 2024, Lin et al., 2024, etc). For instance, one can consider the probability $\hat{p}(y|x)$, Which reflects how likely the generated output is according to the model; and thus can be viewed as a very basic form of a confidence score. Alternatively, for generations whose length can vary, such as for standard language models, one can consider a length-normalized version $\hat{p}(y|x)^{1/|y|},$ where $|y|$ is the length of $y$; aiming to correct for the effect that longer generations tend to have smaller probabilities. Uncertainty and confidence scoresNumerical values computed based on the input, output, or other characteristics of the GenAI model, aiming to capture the level of uncertainty.
However, it is not always straightforward to use and interpret such scores. There are multiple key challenges:
Challenge 1: Inability to recover “true” probabilities and lack of calibration. The probabilities $\hat{p}(y|x)$ represent only the model’s internal beliefs about the likelihood of output $y$ given input $x$; by default, they do not correspond to any notion of “true” probabilities. Because the input and output spaces are extremely high-dimensional, the probabilities produced by a generative AI model should not be expected to be consistent for any “ground truth”. However, we might hope to achieve weaker forms of correctness.
One such relaxation is calibration, which is a general property associated with probabilistic forecasts (Gneiting & Katzfuss, 2014) that only asks for a restricted set of probabilities to reflect real probabilities. For instance, for a calibrated weather forecaster, if we predict “50% chance of rain tomorrow”, then over all such days, we expect that it rains half the time (Lichtenstein et al., 1977, Van Calster & Vickers, 2015, Van Calster et al., 2019). CalibrationThe property of a predicted probability that it reflects the empirical frequencies of a specific class of events. There are a variety of notions of calibration relevant to GenAI, and empirical work has found that model calibration is not guaranteed by default. Instead, it can depend strongly on model training, model size, etc; see e.g., Kadavath et al. (2022), Achiam et al. (2023).
A direct way to apply calibration to answers generated by an LM $\hat{p}$ is to construct an additional probability predictor $\hat{q}$ for the claim “The chance that my answer is right is $\hat{q}$.” Such a probability predictor can be obtained via re-calibration on separate calibration data (Mincer & Zarnowitz, 1969, Guo et al., 2017), but it might require a lot of calibration data.
If less calibration data is available, one may still be able to approximately satisfy a weaker form of calibration, e.g., that the average accuracy increases with the predicted probability of success, a behavior termed rank-calibration (Huang et al., 2024). Rank-calibrationThe property that the average accuracy increases with the predicted probability of success.
Challenge 2: Semantic multiplicity. Another key challenge is that there are often many equivalent answers. For instance, in text generation, answers such as “15 pages” and “fifteen pages” are semantically equivalent. We usually want to pool them together when determining the model’s confidence. Semantic uncertaintyUncertainty after semantic equivalences have been accounted for.
An approach to this problem—termed semantic uncertainty—was proposed by Kuhn et al. (2023), who suggested generating multiple outputs $Y_{1},\ldots,Y_{K}\sim\hat{p}(\cdot|x)$ i.i.d., clustering them based on their semantics (via another LLM), and then estimating uncertainty based on the resulting distribution induced over the clusters.
2.3 AI Evaluation
Evaluating generative AI models is important in order to properly understand the capabilities that these models possess. However, model evaluation can be surprisingly challenging, and in particular, it can bring novel challenges compared to the evaluation of more standard machine learning models, see e.g., Burden et al. (2025) for a review.
A typical current workflow for evaluating a GenAI model—in particular, a large language model—is as follows. Suppose we want to measure reasoning ability in mathematical problems. To evaluate this ability, we collect test data consisting of such problems. Then we evaluate the accuracy of the model on these problems and report the results.
This simple workflow is mired with a number of challenges. First of all, the specific data required for evaluation (say mathematical problems), can be quite complex, and finding genuinely new test problems that the model has not seen during training is hard. Indeed, information leakage from standard public test datasets into the model training sets is a genuine concern (see e.g., Matton et al., 2024, etc). This leads to potential biases in model performance evaluation, where the models score higher because they have already seen the problems during training.
A potential approach is to have private test data sets that are not released to the public. Another potential approach is to use dynamically generated AI evaluation environments, such as based on debates (Moniri et al., 2025). Due to these reasons, and as collecting large, high quality, and genuinely new evaluation datasets can be expensive, high quality test datasets sometimes have relatively small sample sizes.[^9]
Second, checking correctness can be non-trivial and ambiguous outside of simple problems with clear, well-defined answers. For instance, in a mathematical problem, it can be straightforward to evaluate the correctness of a numerical answer, but it can be much harder to evaluate the correctness of a reasoning process. For this reason, often heuristics such as other LLMs are used for checking answers, which in turn raises questions about reliability.
Third, evaluating the largest models can be expensive, which further poses a limit on the sample sizes that we can collect for evaluation depending on the available budget.
Due to these reasons, evaluation can involve dealing with small sample sizes and various biases. Thus, statistical methods and thinking can be valuable for reliable and efficient evaluation.
2.3.1 A basic statistical formulation of model evaluation
We consider a basic setting of model evaluation, in which we have some inputs $x$ for which we wish to evaluate the performance of a GenAI model. For mathematical reasoning, this could correspond to the problem statement, and may also include instructions to the model. The problem has a ground truth answer $y^{*}$, which can be an entire solution/reasoning path or just the final result. Then, we sample a candidate answer $Y\sim\hat{p}(\cdot|x)$. Again, this may include intermediate steps, and a final answer is extracted at the end.
As in Section 2.1.1, the quality of the answer is evaluated via a loss function $\ell$, such that $\ell(x,y^{*},y)$ measures the (negative) utility of answer $y$ for input $x$ with ground truth $y^{*}$. In some cases, designing loss functions is straightforward. For instance, for an integer answer $y^{*}$, we may use the binary loss $I(y\neq y^{*}).$ However, for more elaborate problems, designing a loss function can be non-trivial. For instance, for a reasoning problem, we want to make sure that all valid and concise reasoning paths receive low loss, not just the reference path.
Tasks in AI evaluation. Given these components, there are several possible tasks of interest. For a distribution $D$ of inputs, we may want to estimate or perform statistical inference (confidence intervals, tests) for the task performance $\theta=\mathbb{E}_{(X,Y^{*})\sim D,\;Y\sim\hat{p}(\cdot\mid X)}\,\ell(X,Y^{*},Y).$
Given a dataset $D_{n}=\{(X_{i},Y_{i}^{*}):i=1,\dots,n\}$ of question–answer pairs sampled i.i.d. from $D$, we can generate outputs $Y_{i}\sim\hat{p}(\cdot\mid X_{i})$ independently, and compute the loss values $\ell_{i}=\ell(X_{i},Y_{i}^{*},Y_{i}),$ $i=1,\dots,n$; as in Section 2.1.1.
Then, these loss values $\ell_{1},\dots,\ell_{n}$ are sampled i.i.d. from a distribution whose population mean is the unknown true task performance $\theta$. Thus, this problem becomes that of inference for a population mean, for which many statistical methods exist (Casella & Berger, 2024, Lehmann & Romano, 2005).
Notably, in many important examples we are interested in (A) binary losses, leading to inference for a Binomial parameter; (B) or bounded losses (for which concentration inequalities such as Hoeffding’s inequality can be used); (C) or given a large sample size (so that an asymptotic normal approximation works well). [h]
2.3.2 AI Evaluation and Statistical Inference
AI evaluation with limited data has a very close link to statistical inference.
An important observation here is that AI evaluation with limited data has a very close link to statistical inference. Beyond this core setting, there are a variety of important additional scenarios. For instance, we may be interested in comparing the performance of two models $\hat{p}_{1},\hat{p}_{2}$. If we can query both models on the same inputs $X_{i},i=1,\dots,n$, this can be formulated as statistical inference for the parameter $\Delta=\mathbb{E}_{(X,Y^{*})\sim D,\;Y_{1}\sim\hat{p}_{1}(\cdot\mid X),\;Y_{2}\sim\hat{p}_{2}(\cdot\mid X)}\big[\ell(X,Y^{*},Y_{1})-\ell(X,Y^{*},Y_{2})\big].$ Considerations and methods similar to the ones above apply.
Standard methods for the above two problems have been reviewed in Miller (2024); where other considerations, such as power analysis and clustered data arising from repeated generations for the same input, are also considered. However, this work focuses on a signal-plus-noise model for the observed losses, which may need to be relaxed.
| Technique | Type | Examples |
|---|---|---|
| Inference on Performance | Confidence intervals | Review of standard large-sample methods (Miller, 2024) |
| Construct CIs with improved finite-sample coverage on model accuracy under i.i.d. and clustered data settings (Bowyer et al., 2025) | ||
| Develop asymptotically valid CIs for comparing the KL divergence to the true distribution of two models (Gao & Sun, 2025) | ||
| Construct uniform upper bound on the CDF of a performance metric (Vincent et al., 2024) | ||
| Construct confidence interval for probability of biased answers on counterfactual prompts (Chaudhary et al., 2025) | ||
| Hypothesis testing | Test hypothesis about which policy achieves higher reward, choosing number of trials adaptively (Snyder et al., 2025) | |
| Small-data Evaluation | Small-sample performance | Estimate model accuracy on multiple questions and models leveraging item response theory (Polo et al., 2024) |
| Synthetic + human labels | Combine synthetic and human labels for unbiased performance estimates and CIs (Boyeau et al., 2024, Fisch et al., 2024, Oosterhuis et al., 2024) | |
| Rank models with hybrid label sets (Chatzi et al., 2024) | ||
| Multi-task Evaluation | Active testing | Actively sample and evaluate in multitask settings (Anwar et al., 2025) |
Table 4: Types and examples of statistical evaluations of generative AI models.
2.3.3 Additional methods
There are a variety of works addressing other settings in AI evaluation, see Table 4 for examples. A few of them are discussed in more detail below. However, a comprehensive and unified statistical methodology that addresses most of the common evaluation problems with a unified terminology and set of methods remains to be developed.
- Bowyer et al. (2025) study methods for producing confidence intervals on model performance, focusing on inference for Bernoulli parameters of model accuracy. They include single-model performance (for i.i.d. and clustered data), two-model comparison (both independent data and paired samples). They conclude that the most straightforward asymptotic normality-based confidence intervals can be inaccurate for small datasets at most $n=100$ datapoints. They argue for using Bayesian credible intervals, which they argue have adequate frequentist coverage when one can specify appropriate prior distributions.
- Gao & Sun (2025) develop methods for comparing the Kullback-Leibler (KL) divergence of two generative methods for which the probabilities $\hat{p}$ can be computed. They show how to construct an asymptotically valid confidence interval for the difference of KL divergences.
- Polo et al. (2024) develop methods for estimating accuracy using a small number of datapoints, leveraging methods item response theory. They consider settings where the performance of a model $\hat{p}$ on an example $x$ is captured by (unknown) model-specific and example-specific latent variables $\theta_{\hat{p}}$ and $\gamma_{x}$. For instance, we may model the probability $Q(\hat{p},x)$ of a correct answer by $\hat{p}$ on the input $x$ via a logistic model $\mathrm{logit}(Q(\hat{p},x))=\theta_{\hat{p}}^{\top}\gamma_{x}+\beta_{x}$. Then, these parameters are estimated on a small dataset, and the correctness probability predictions they induced are used on new test examples to extrapolate correctness; leading to significant savings in the number of test examples needed. See also Zhou et al. (2025), Gignac & Ilić (2025), Kipnis et al. (2025) for other uses of item response theory and related methods.
- Boyeau et al. (2024), Fisch et al. (2024), Oosterhuis et al. (2024) develop methods to use a large set of synthetically generated labels along with a small set of human labels for unbiased model evaluation, including confidence intervals for model performance. See Chatzi et al. (2024) for ranking.
- Anwar et al. (2025) develop methods for multi-task evaluation of (robot) policies with active testing, where they pool information on performance of several policies across several tasks, prioritizing tasks with high information gain leveraging Bayesian active learning (Houlsby et al., 2011).
- Chowdhury et al. (2025) develop a variational lower bound on the expected loss incurred by a language model, and use it to find prompts that elicit problematic behavior. Concretely, let $\ell$ be a loss, $\hat{p}$ be the target LLM. Our goal is to find prompts $x$ to make $\ell(x,Y)$ large when $Y\sim\hat{p}(\cdot|x)$. Formally, we aim to make $S(x)=\log\mathbb{E}_{Y\sim\hat{p}(\cdot|x)}\exp(S(x,Y))$ large. To find such $x$, we rely on an auxiliary LLM $\hat{q}$ for which the loss tends to be larger for all $x$. Due to Jensen’s inequality, we have the variational lower bound $S(x)\geqslant\mathbb{E}_{Y\sim\hat{q}(\cdot|x)}[\log\hat{q}(Y|x)-\log\hat{p}(Y|x)+S(x,Y)]$. This lower bound is estimated by sampling $Y\sim\hat{q}(\cdot|x)$ repeatedly, which can be more efficient than estimating $S(x)$ directly.
2.4 Interventions and Experiment Design
Interventions refer to systematically modifying or perturbing the inputs of an AI system, to gain understanding or control of its behavior. This approach has become one of the most widely used and most powerful tools in a variety of AI research directions, including interpretability, robustness, and fairness (e.g., Zhao et al., 2018, Rudinger et al., 2018, Belinkov, 2022, Kotek et al., 2023, etc). The ideas underlying interventions are closely connected to statistical causality and experiment design; see also Pearl (2001), Soumm (2024).
2.4.1 Basic setting for interventions
In a basic setting for interventions, we have a generative model $\hat{p}$ to which we can provide an input $x$ (e.g., a query to an LLM). In contrast to the other parts covered in this review, for interventions, it is often the case that the intermediate computations are of crucial importance. The reason is that, empirically, certain internal mechanisms can sometimes be responsible for specific behaviors, such as biases and harmful outputs (see e.g., Mikolov et al., 2013b, Turner et al., 2023, Rimsky et al., 2024, Zou et al., 2023, etc).
Therefore, in this section, we will sometimes also assume that we have access to intermediate computations $e(x)$ (e.g., representations, intermediate/chain of thought tokens) of the model. Most often, vector-valued intermediates $e(x)$ are considered. Finally, we also consider the output layer $o(x)$ of the model (e.g., last-layer predicted probabilities or log-probabilities), as well as the final model output $y$. These quantities can be either deterministic or random.
We want to understand or control a certain components of the behavior of the AI system. We consider components measured through the input, intermediate computation, or output. For instance, which components of an LLM (activations, neurons) contribute to gender bias? How can we intervene to reduce such biases? How does an LLM behave internally when it is non-truthful, and does this differ from truthful behavior? Are there specific components that are activated when the LLM generates harmful output, and can we intervene to suppress this behavior?
InterventionsPerturbing components of the model (input or intermediate computations) to achieve a desired effect, such as reducing biases. To do this, we find a way to intervene by perturbing the input $x$ to induce the condition of interest. For example, to understand how harmfulness is propagated, we can change part of a harmful input to a harmless concept: e.g., $x=$ “how to build a bomb?” $\rightarrow$ $x^{\prime}=$ “how to build a chair?”. We can also intervene on an intermediate computation in the AI system. Then, we track the change in either the intermediate stage or the final output, depending on what we are interested in.
Example 2.1.****
Contextual concept vectors measure the difference in embeddings that a change in a concept leads to, in the form $C_{x\rightarrow x^{\prime}}:=e(x^{\prime})-e(x)$, where $x$ is an input and $x^{\prime}$ is the corresponding input with the concept changed, e.g., for the concept of gender, $x$=“king”, $x^{\prime}$=“queen”; $x$=“actor”, $x^{\prime}$=“actress”, etc. Early work investigating related questions dates back at least to Mikolov et al. (2013a), Mikolov et al. (2013b), Pennington et al. (2014) for word embeddings, and more recently has studied human biases (Bolukbasi et al., 2016), developed steering vectors (Turner et al., 2023, Rimsky et al., 2024) and introduced representation engineering (Zou et al., 2023).
Contextual concept vectorThe effect of changing a concept in the input on some vector in the intermediate computation of the genAI model.
To obtain a more stable and generalizable picture about the effect of the intervention, it is common to consider a distribution $D$ of interest, and the associated mean $\mathbb{E}_{X\sim D}[C_{X\rightarrow X^{\prime}}]$ or top principal component of the covariance matrix $\text{Cov}_{X\sim D}(C_{X\rightarrow X^{\prime}})$ (Zou et al., 2023). These are typically estimated using the standard plug-in estimators. Let $\hat{c}$ be such an estimated concept vector.
Steering vectorA quantity that used in—typically added to—the intermediate representations of a model to make a desired behavior more likely.
Steering vectors. These estimates can be used as steering vectors (Turner et al., 2023, Rimsky et al., 2024) to make certain behaviors more likely. A common approach is to take any input $x$, compute its intermediate representation $e(x)$, and add a scaled version $\lambda\cdot\hat{c}$ for some $\lambda>0$ to obtain a new intermediate representation $e^{\prime}=e(x)+\lambda\cdot\hat{c}$. The computation then continues identically to obtain the final output. Here $\lambda$ is a hyperparameter that requires careful tuning. This operation approximates a shift of the representation of the original input towards the representation of a changed input $e(x^{\prime})$. For instance, in the above example, the goal would to approximately remove the harmful concept. Empirically, it has been observed that the resulting final output can sometimes indeed correspond to the desired concept change (Turner et al., 2023, Rimsky et al., 2024); which however comes with caveats (Tan et al., 2024).
Assessing biases. Analogously, to assess biases[^10] (e.g., gender bias), one can choose a representative output variable $o(x)$, such as the probability of a gendered word, and then repeat the above analysis. For instance, to study gender bias, Kotek et al. (2023) intervene to modify gender in an input such as $x=\text{ ``The doctor called the nurse because \lx@text@underline{he} was late. Who was late?"}$ They change this to $x^{\prime}=\text{ ``The doctor called the nurse because \lx@text@underline{she} was late. Who was late?"}$
Then, they evaluate its effect on an output $o$ which they choose as a measure of the probability of the output “nurse”. Specifically, they compute $O_{x\rightarrow x^{\prime}}=o(x^{\prime})-o(x)$, which measures how much more likely the model is to output “nurse” solely due to the change $\text{``he"}\rightarrow\text{``she"}$, and thus it can be interpreted as a form of gender bias. Kotek et al. (2023) also design an improved version that also permutes “doctor” and “nurse”, aiming to control for the effect of syntactic position.
Probing. A related concept is that of probing (see e.g., Alain & Bengio, 2016, Belinkov, 2022, etc). To understand if a feature $e$ captures a concept $x\mapsto x^{\prime}$, in probing one trains a classifier of datapoints $X\sim D$ versus their transformed counterparts $X^{\prime}$, using a simple function—often linear—of the features $e$. If this classifier has a high accuracy, then it is concluded that the feature captures the concept. This approach has been leveraged in generative AI, e.g., to understand where models store spatial information about the input (Gurnee & Tegmark, 2024). ProbingTraining models based on intermediate features to see if they contain information about a specific concept.
| Technique | Type | Examples |
|---|---|---|
| Understand Behavior via Intervention | Learn bias or association | Learn gender bias in output by modifying input (Bolukbasi et al., 2016, Zhao et al., 2018, Rudinger et al., 2018) |
| Identify internal/intermediate component associated with bias or factual association via causal mediation analysis (Vig et al., 2020, Meng et al., 2022, Dai et al., 2022) | ||
| Learn effect of circuits (sub-networks) by pruning to the circuit and observing behavior (Nanda et al., 2023) | ||
| Learn effect of thoughts (intermediate outputs) by modifying them (Bogdan et al., 2025) | ||
| Learn concept | Learn concept or steering vector by inducing concept modifying input (Mikolov et al., 2013a, Mikolov et al., 2013b, Pennington et al., 2014, Turner et al., 2023, Rimsky et al., 2024, Zou et al., 2023) | |
| Evaluate performance | Perform ablation study: change algorithm setting and test behavior | |
| Design perturbed dataset to evaluate LLM reasoning robustness (Wu et al., 2024, Shi et al., 2023, Mirzadeh et al., 2025) | ||
| Evaluate alignment | Design prompt eliciting behavior that would modify AI system and observe behavior (Greenblatt et al., 2023) | |
| Understand Behavior via Probing | Identify neurons associated with sentiment (Radford et al., 2017) or neurons that represent world state (Li et al., 2023a) | |
| Identify sparse linear combinations of neurons that represent features (Gurnee et al., 2023) | ||
| Change Behavior via Intervention | Add gradient of concept classifier (Dathathri et al., 2020) or steering vector (Subramani et al., 2022, Turner et al., 2023, Zou et al., 2023, Li et al., 2023b) to elicit behavior | |
| Patch activations from one input into the activations of another input (Meng et al., 2022, Zhang & Nanda, 2024) |
Table 5: Types and examples of interventions and experiment design in generative AI.
See Table 5 for some examples of related methods. A few examples are discussed below:
- There is work aiming to identify sub-networks (not just representations) responsible for specific tasks, by pruning to the networks and checking if they can still perform the computation (Nanda et al., 2023). Further, Zhang & Nanda (2024) systematized activation patching methods to localize causal computations in LLMs, providing best practices for intermediate-stage interventions.
- Greenblatt et al. (2023) used intervention-based prompts to elicit deceptive behavior from an LLM, finding that LMs may internally simulate misaligned objectives while faking alignment.
- There has been work to design perturbations of standard mathematical datasets to evaluate LLM reasoning robustness (Shi et al., 2023, Mirzadeh et al., 2025).
2.4.2 Causal mediation analysis
Figure 2: Diagram to represent computational flow and interventions, for use with causal mediation analysis Solid arrows denote standard computational flows; dashed arrows denote interventions or their effects.
Causal mediation analysis (Pearl, 2001) is a more advanced technique from statistical causality, which can be used to identify the precise effects of intermediate components of generative AI models (e.g., Vig et al., 2020, etc.). In a basic setting for causal mediation analysis, we consider an input $x$, and a changed input $x^{\prime}$, Where we intervene via an intervention that we would like to study, for instance changing the sentiment of a review $x$ from positive to negative.
We aim to study a generative model of interest. We consider an intermediate representation/activation $e$ whose effect we aim to study; in causal mediation analysis $e$ is known as the mediator. The final output representation $o$ of the generative model depends on the intermediate representation $e$, as well as on other model components, which together we denote by $e^{\perp}$. Algebraically, we write the output representation in the functional form $o(x)=g(e(x),e^{\perp}(x))$ for all $x\in\mathcal{X}$, for some set of computations denoted by $g$. See Figure 2 for a diagram representing this setting.
Then, $o(x^{\prime})-o(x)$ represents the overall effect of the intervention $x\to x^{\prime}$. Typically, we are interested not just in the particular query $x$, but rather about the average behavior over a distribution of interest. The total average effect of $x\to x^{\prime}$ is $\mathbb{E}\left[o(X^{\prime})-o(X)\right]$. This can be decomposed into a the sum of natural direct and indirect effects.
Natural direct effect. The natural direct effect of $x\to x^{\prime}$ on $o$ is the effect that happens through pathways other than the mediator $e$. This expression keeps $e$ fixed:
$$ \mathbb{E}\left[o\left(e(X),e^{\perp}(X^{\prime})\right)-o(X)\right]=\mathbb{E}\left[o\left(e(X),e^{\perp}(X^{\prime})\right)-o\left(e(X),e^{\perp}(X)\right)\right] $$
If the direct effect is small, this can be interpreted as the mediator $e$ capturing most of the effect of $x^{\prime}$ on $o$. When the direct effect is small, we can view the mediator as having an important role in enacting the effect $x\mapsto x^{\prime}$, making it a promising target for interventions if we aim to mitigate this effect. Natural Direct EffectThe effect of an input on an output that happens through pathways other than the mediator under study.
Natural indirect effect. To complement this, the natural indirect effect of $x\to x^{\prime}$ on $o$ captures the remaining part of the total effect, which goes through the mediator $e(x)\to e(x^{\prime})$:
$$ \mathbb{E}\left[o(X^{\prime})-o\left(e(X),e^{\perp}(X^{\prime})\right)\right]=\mathbb{E}\left[o\left(e(X^{\prime}),e^{\perp}(X^{\prime})\right)-o\left(e(X),e^{\perp}(X^{\prime})\right)\right]. $$
Natural Indirect EffectThe effect of an input on an output that happens through the mediator under study. This decomposition of effects into direct and indirect ones has been used, among others, to identify components responsible for gender bias (Vig et al., 2020) as well as other factual associations (Meng et al., 2022, Dai et al., 2022) in LLMs. In some cases, $x^{\prime}$ can correspond to “adding noise” to tokens that contain specific information, e.g., “The Space Needle is in” $\to$“[i.i.d. Gaussian activations] is in”; by acting at the levels of token embeddings of $x$. This allows capturing the effect of deleting information from the input. However, fully rigorous and well-justified methods for interventions on the identified mediators have not yet been developed.