OER·harvester

← Back to the library
arXiv HTML resource

Statistical Methods in Generative AI

Generative Artificial Intelligence is emerging as an important technology, promising to be transformative in many areas. At the same time, generative AI techniques are based on sampling from probabilistic models, and by default, they come with no guarantees about correctness, safety, fairness, or other properties. Statistical methods offer a promising potential approach to improve the reliability of generative AI te…

Licence
OPEN CC-BY-4.0
Authors
Edgar Dobriban
Published
2025-09-08 · arXiv
Language
en
Length
14443 words
Type
narrative text

Cites 36 works

inferred
Open original ↗

3 Discussion

We have presented overviews of some applications of statistical ideas to generative AI, focusing on topics such as improving and changing the behavior of GenAI models, diagnostics and uncertainty quantification, evaluation, as well as interventions and experiment design. These leverage ideas from classical statistical inference, distribution-free predictive inference, forecasting and calibration, as well as causality.

At the moment, generative AI models are exceedingly complex, and are usually best viewed as black boxes. To ensure usefulness in GenAI, one needs to develop methods that are light on assumptions. Moreover, in order to to maximize impact, the methods need to be illustrated on current GenAI models, which requires both a familiarity with ongoing developments in AI, and adequately large computational resources. For statisticians, collaboration with AI researchers can help ensure that these requirements are met.

[SUMMARY POINTS]

  1. GenAI lacks guarantees. Generative AI models are stochastic black boxes: as probability distributions over large semantic spaces (text, images) from which we can sample. While showing promising performance in a variety of areas, they do not have any guarantees about correctness, safety, etc., by default.
  2. Statistical methods for GenAI need to handle black box models. In order to be applicable to black-box generative AI models, statistical methods need to be light on assumptions and able to handle structured semantic input and output spaces.
  3. The flexibility of statistical “wrappers”. There are a variety of approaches to change the behavior of AI models, both in terms of their inputs and their outputs. Statistical “wrappers” can be used in order to precisely control the performance of these approaches.
  4. Uncertainty quantification must be calibrated and handle semantics. Quantifying the uncertainty of a GenAI model could be a promising way to make it more reliable; however, the issues of semantic multiplicity and lack of calibration need to be handled.
  5. Evaluation is statistical inference. AI evaluation, especially with small datasets, presents opportunities for leveraging statistical inference methods.
  6. The power of interventions. Interventions on generative AI systems, building on ideas from causal inference, have the potential to identify components responsible for specific capabilities and to induce desired behaviors.
  7. The promise of dataset and experiment design. Calibration, evaluation, and intervention all hinge on carefully collected, held-out calibration sets and targeted perturbations, which offer opportunities for statistical thinking.

[FUTURE ISSUES]

  1. Statistical methods aimed at improving AI models need to be developed by taking into account the black-box nature of AI, where often only the inputs and outputs of the models are available, and the intermediate computations are unknown.
  2. A comprehensive statistical framework for the evaluation of generative AI systems is yet to be developed.
  3. Well-justified methods for interventions on mediators identified in generative AI models remain to be introduced.