OER·harvester

← Back to the library
arXiv HTML resource

Statistical Methods in Generative AI

Generative Artificial Intelligence is emerging as an important technology, promising to be transformative in many areas. At the same time, generative AI techniques are based on sampling from probabilistic models, and by default, they come with no guarantees about correctness, safety, fairness, or other properties. Statistical methods offer a promising potential approach to improve the reliability of generative AI te…

Licence
OPEN CC-BY-4.0
Authors
Edgar Dobriban
Published
2025-09-08 · arXiv
Language
en
Length
14443 words
Type
narrative text

Cites 36 works

inferred
Open original ↗

References

  • Abbasi Yadkori Y, Kuzborskij I, György A, Szepesvari C. 2024. To believe or not to believe your llm: Iterative prompting for estimating epistemic uncertainty. Advances in Neural Information Processing Systems 37:58077–58117
  • Abbasli T, Toyoda K, Wang Y, Witt L, Ali MA, et al. 2025. Comparing uncertainty measurement and mitigation methods for large language models: A systematic review. arXiv preprint arXiv:2504.18346
  • Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
  • Alain G, Bengio Y. 2016. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644
  • Angelopoulos AN, Bates S, et al. 2023. Conformal prediction: A gentle introduction. Foundations and Trends® in Machine Learning 16(4):494–591
  • Anwar A, Gupta R, Merchant Z, Ghosh S, Neiswanger W, Thomason J. 2025. Efficient evaluation of multi-task robot policies with active experiment selection. arXiv preprint arXiv:2502.09829
  • Baan J, Daheim N, Ilia E, Ulmer D, Li HS, et al. 2023. Uncertainty in natural language generation: From theory to applications. arXiv preprint arXiv:2307.15703
  • Band N, Li X, Ma T, Hashimoto T. 2024. Linguistic calibration of long-form generations, In Proceedings of the 41st International Conference on Machine Learning, pp. 2732–2778
  • Bashari M, Lotan RM, Lee Y, Dobriban E, Romano Y. 2025. Synthetic-powered predictive inference. arXiv preprint arXiv:2505.13432
  • Bates S, Angelopoulos A, Lei L, Malik J, Jordan M. 2021. Distribution-free, risk-controlling prediction sets. Journal of the ACM (JACM) 68(6):1–34
  • Belinkov Y. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics 48(1):207–219
  • Bogdan PC, Macar U, Nanda N, Conmy A. 2025. Thought anchors: Which llm reasoning steps matter? arXiv preprint arXiv:2506.19143
  • Bolukbasi T, Chang KW, Zou JY, Saligrama V, Kalai AT. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings, In Advances in Neural Information Processing Systems (NeurIPS), vol. 29
  • Bowyer S, Aitchison L, Ivanova DR. 2025. Position: Don’t use the clt in llm evals with fewer than a few hundred datapoints. ICML (Spotlight Position Paper)
  • Boyeau P, Angelopoulos AN, Yosef N, Malik J, Jordan MI. 2024. Autoeval done right: Using synthetic data for model evaluation. arXiv preprint arXiv:2403.07008
  • Burden J, Tešić M, Pacchiardi L, Hernández-Orallo J. 2025. Paradigms of ai evaluation: Mapping goals, methodologies and culture. arXiv preprint arXiv:2502.15620
  • Campos M, Farinhas A, Zerva C, Figueiredo MA, Martins AF. 2024. Conformal prediction for natural language processing: A survey. Transactions of the Association for Computational Linguistics 12:1497–1516
  • Casella G, Berger R. 2024. Statistical inference. Chapman and Hall/CRC
  • Chan KHR, Ge Y, Dobriban E, Hassani H, Vidal R. 2025. Conformal information pursuit for interactively guiding large language models. arXiv preprint arXiv:2507.03279
  • Chatzi I, Straitouri E, Thejaswi S, Rodriguez M. 2024. Prediction-powered ranking of large language models. Advances in Neural Information Processing Systems 37:113096–113133
  • Chaudhary I, Hu Q, Kumar M, Ziyadi M, Gupta R, Singh G. 2025. Certifying counterfactual bias in LLMs, In The Thirteenth International Conference on Learning Representations
  • Chen M, Mei S, Fan J, Wang M. 2024. An overview of diffusion models: Applications, guided generation, statistical rates and optimization
  • Chowdhury N, Schwettmann S, Steinhardt J, Johnson DD. 2025. Surfacing pathological behaviors in language models. https://transluce.org/pathological-behaviors
  • Dai D, Dong L, Hao Y, Sui Z, Chang B, Wei F. 2022. Knowledge neurons in pretrained transformers, In Proc. of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 8493–8502
  • Dathathri S, Madotto A, Lan J, Hung J, Frank E, et al. 2020. Plug and play language models: A simple approach to controlled text generation, In Proc. of the International Conference on Learning Representations (ICLR)
  • Davies A, Veličković P, Buesing L, Blackwell S, Zheng D, et al. 2021. Advancing mathematics by guiding human intuition with ai. Nature 600(7887):70–74
  • Der Kiureghian A, Ditlevsen O. 2009. Aleatory or epistemic? does it matter? Structural safety 31(2):105–112
  • Deutschmann N, Alberts M, Martinez MR. 2024. Conformal autoregressive generation: Beam search with coverage guarantees, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 11775–11783
  • Dobriban E, Yu M. 2025. Symmpi: predictive inference for data with group symmetries. Journal of the Royal Statistical Society Series B: Statistical Methodology :qkaf022
  • Farquhar S, Kossen J, Kuhn L, Gal Y. 2024. Detecting hallucinations in large language models using semantic entropy. Nature 630(8017):625–630
  • Fisch A, Maynez J, Hofer RA, Dhingra B, Globerson A, Cohen WW. 2024. Stratified prediction-powered inference for hybrid language model evaluation. arXiv preprint arXiv:2406.04291
  • Gao Z, Sun Y. 2025. Statistical inference for generative model comparison
  • Geisser S. 2017. Predictive inference: an introduction. Chapman and Hall/CRC
  • Gignac GE, Ilić D. 2025. Psychometrically derived 60-question benchmarks: Substantial efficiencies and the possibility of human-ai comparisons. Intelligence 110:101922
  • Gneiting T, Katzfuss M. 2014. Probabilistic forecasting. Annual Review of Statistics and Its Application 1(1):125–151
  • Greenblatt R, Denison C, Wright B, Roger F, MacDiarmid M, et al. 2023. Alignment faking in large language models. arXiv preprint arXiv:2412.14093
  • Guan L. 2023. Localized conformal prediction: A generalized inference framework for conformal prediction. Biometrika 110(1):33–50
  • Gui Y, Jin Y, Ren Z. 2024. Conformal alignment: Knowing when to trust foundation models with guarantees, In The Thirty-eighth Annual Conference on Neural Information Processing Systems
  • Guo C, Pleiss G, Sun Y, Weinberger KQ. 2017. On calibration of modern neural networks, In International conference on machine learning, pp. 1321–1330, PMLR
  • Gurnee W, Nanda N, Pauly M, Harvey K, Troitskii D, Bertsimas D. 2023. Finding neurons in a haystack: Case studies with sparse probing. Transactions on Machine Learning Research
  • Gurnee W, Tegmark M. 2024. Language models represent space and time, In The Twelfth International Conference on Learning Representations
  • Hayes T, Rao R, Akin H, Sofroniew NJ, Oktay D, et al. 2025. Simulating 500 million years of evolution with a language model. Science 387(6736):850–858
  • He J, Yu L, Li C, Yang R, Chen F, et al. 2025. Survey of Uncertainty Estimation in Large Language Models -Sources, Methods, Applications, and Challenge. Working paper or preprint
  • Horwitz E, Hoshen Y. 2022. Conffusion: Confidence intervals for diffusion models. arXiv preprint arXiv:2211.09795
  • Hou B, Liu Y, Qian K, Andreas J, Chang S, Zhang Y. 2024. Decomposing uncertainty for large language models through input clarification ensembling, In Proceedings of the 41st International Conference on Machine Learning, pp. 19023–19042
  • Houlsby N, Huszár F, Ghahramani Z, Lengyel M. 2011. Bayesian active learning for classification and preference learning. arXiv preprint arXiv:1112.5745
  • Huang X, Li S, Yu M, Sesia M, Hassani H, et al. 2024. Uncertainty in language models: Assessment through rank-calibration, In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, eds. Y Al-Onaizan, M Bansal, YN Chen
  • Hüllermeier E, Waegeman W. 2021. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine learning 110(3):457–506
  • Jazbec M, Timans A, Hadži Veljković T, Sakmann K, Zhang D, et al. 2024. Fast yet safe: Early-exiting with risk control. Advances in Neural Information Processing Systems 37:129825–129854
  • Ji W, Yuan W, Getzen E, Cho K, Jordan MI, et al. 2025. An overview of large language models for statisticians. arXiv preprint arXiv:2502.17814
  • Jiang Z, Araki J, Ding H, Neubig G. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9:962–977
  • Jiang Z, Liu A, Van Durme B. 2025. Conformal linguistic calibration: Trading-off between factuality and specificity. arXiv preprint arXiv:2502.19110
  • Kadavath S, Conerly T, Askell A, Henighan T, Drain D, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221
  • Kaur R, Jha S, Roy A, Park S, Dobriban E, et al. 2022. idecode: In-distribution equivariance for conformal out-of-distribution detection, In Proceedings of the AAAI Conference on Artificial Intelligence
  • Khakhar A, Mell S, Bastani O. 2023. Pac prediction sets for large language models of code, In International Conference on Machine Learning, pp. 16237–16249, PMLR
  • Kipnis A, Voudouris K, Buschoff LMS, Schulz E. 2025. metabench - a sparse benchmark of reasoning and knowledge in large language models, In The Thirteenth International Conference on Learning Representations
  • Kotek H, Dockum R, Sun D. 2023. Gender bias and stereotypes in large language models, In Proceedings of the ACM collective intelligence conference, pp. 12–24
  • Kuhn L, Gal Y, Farquhar S. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, In The Eleventh International Conference on Learning Representations
  • Lee Y, Dobriban E, Tchetgen ET. 2024. Conditional predictive inference for missing outcomes. arXiv preprint arXiv:2403.04613
  • Lehmann EL, Romano JP. 2005. Testing statistical hypotheses. Springer Science & Business Media
  • Lei J, G’Sell M, Rinaldo A, Tibshirani R, Wasserman L. 2018. Distribution-free predictive inference for regression. Journal of the American Statistical Association 113(523):1094–1111
  • Lei J, Robins J, Wasserman L. 2013. Distribution-free prediction sets. Journal of the American Statistical Association 108(501):278–287
  • Lei J, Wasserman L. 2014. Distribution-free prediction bands for non-parametric regression. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 76(1):71–96
  • Li K, Hopkins AK, Bau D, Viégas F, Pfister H, Wattenberg M. 2023a. Emergent world representations: Exploring a sequence model trained on a synthetic task, In The Eleventh International Conference on Learning Representations
  • Li K, Patel O, Viégas F, Pfister H, Wattenberg M. 2023b. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems 36:41451–41530
  • Li S, Ji X, Dobriban E, Sokolsky O, Lee I. 2022. Pac-wrap: Semi-supervised pac anomaly detection, In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
  • Li S, Park S, Lee I, Bastani O. 2024. Traq: Trustworthy retrieval augmented question answering via conformal prediction, In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 3799–3821
  • Lichtenstein S, Fischhoff B, Phillips LD. 1977. Calibration of probabilities: The state of the art, In Decision Making and Change in Human Affairs: Proceedings of the Fifth Research Conference on Subjective Probability, Utility, and Decision Making, Darmstadt, 1–4 September, 1975, pp. 275–324, Springer
  • Lin S, Hilton J, Evans O. 2022. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research
  • Lin Z, Trivedi S, Sun J. 2024. Generating with confidence: Uncertainty quantification for black-box large language models. Transactions on Machine Learning Research
  • Liu H, Dou ZY, Wang Y, Peng N, Yue Y. 2024. Uncertainty calibration for tool-using language agents, In Findings of the Association for Computational Linguistics: EMNLP 2024, eds. Y Al-Onaizan, M Bansal, YN Chen. Miami, Florida, USA: Association for Computational Linguistics
  • Liu X, Chen T, Da L, Chen C, Lin Z, Wei H. 2025. Uncertainty quantification and confidence calibration in large language models: A survey. arXiv preprint arXiv:2503.15850
  • Manduchi L, Meister C, Pandey K, Bamler R, Cotterell R, et al. 2025. On the challenges and opportunities in generative AI. Transactions on Machine Learning Research Survey Certification
  • Matton A, Sherborne T, Aumiller D, Tommasone E, Alizadeh M, et al. 2024. On leakage of code generation evaluation datasets. arXiv preprint arXiv:2407.07565
  • Meng K, Bau D, Andonian A, Belinkov Y. 2022. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems 35:17359–17372
  • Mikolov T, Chen K, Corrado G, Dean J. 2013a. Efficient estimation of word representations in vector space
  • Mikolov T, Yih Wt, Zweig G. 2013b. Linguistic regularities in continuous space word representations, In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies, pp. 746–751
  • Miller E. 2024. Adding error bars to evals: A statistical approach to language model evaluations. arXiv preprint arXiv:2411.00640
  • Mincer JA, Zarnowitz V. 1969. The evaluation of economic forecasts. In Economic forecasts and expectations: Analysis of forecasting behavior and performance. NBER, 3–46
  • Mirzadeh SI, Alizadeh K, Shahrokhi H, Tuzel O, Bengio S, Farajtabar M. 2025. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models, In The Thirteenth International Conference on Learning Representations
  • Mohri C, Hashimoto T. 2024. Language models with conformal factuality guarantees, In Forty-first International Conference on Machine Learning
  • Moniri B, Hassani H, Dobriban E. 2025. Evaluating the performance of large language models via debates, In Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)
  • Nag S, Ghosh U, Ta CK, Bose S, Li J, Roy-Chowdhury AK. 2025. Conformal prediction and mllm aided uncertainty quantification in scene graph generation, In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 11676–11686
  • Nanda N, Rajamanoharan S, Kramár J, Shah R. 2023. Fact-finding: Attempting to reverse-engineer factual recall
  • Noarov G, Mallick S, Wang T, Joshi S, Sun Y, et al. 2025. Foundations of top-$k$ decoding for language models. arXiv preprint arXiv:2505.19371
  • Noorani S, Kiyani S, Pappas G, Hassani H. 2025. Conformal prediction beyond the seen: A missing mass perspective for uncertainty quantification in generative models. arXiv preprint arXiv:2506.05497
  • Oosterhuis H, Jagerman R, Qin Z, Wang X, Bendersky M. 2024. Reliable confidence intervals for information retrieval evaluation using generative ai, In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2307–2317
  • Overman W, Bayati M. 2025. Conformal arbitrage: Risk-controlled balancing of competing objectives in language models. arXiv preprint arXiv:2506.00911
  • Papadopoulos H, Proedrou K, Vovk V, Gammerman A. 2002. Inductive confidence machines for regression, In European Conference on Machine Learning, pp. 345–356, Springer
  • Park S, Dobriban E, Lee I, Bastani O. 2022a. PAC prediction sets for meta-learning, In Advances in Neural Information Processing Systems
  • Park S, Dobriban E, Lee I, Bastani O. 2022b. PAC prediction sets under covariate shift, In International Conference on Learning Representations
  • Parmar G, Kumar Singh K, Zhang R, Li Y, Lu J, Zhu JY. 2023. Zero-shot image-to-image translation, In ACM SIGGRAPH 2023 conference proceedings, pp. 1–11
  • Pearl J. 2001. Direct and indirect effects. Probabilistic and Causal Inference: The Works of Judea Pearl :373
  • Pennington J, Socher R, Manning CD. 2014. Glove: Global vectors for word representation, In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543
  • Polo FM, Weber L, Choshen L, Sun Y, Xu G, Yurochkin M. 2024. tinybenchmarks: evaluating llms with fewer examples, In International Conference on Machine Learning, pp. 34303–34326, PMLR
  • Qiu H, Dobriban E, Tchetgen Tchetgen E. 2023. Prediction sets adaptive to unknown covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology 85(5):1680–1705
  • Quach V, Fisch A, Schuster T, Yala A, Sohn JH, et al. 2024. Conformal language modeling, In The Twelfth International Conference on Learning Representations
  • Radford A, Jozefowicz R, Sutskever I. 2017. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444
  • Ravfogel S, Goldberg Y, Goldberger J. 2023. Conformal nucleus sampling, In Findings of the Association for Computational Linguistics: ACL 2023, eds. A Rogers, J Boyd-Graber, N Okazaki. Toronto, Canada: Association for Computational Linguistics
  • Ren AZ, Clark J, Dixit A, Itkina M, Majumdar A, Sadigh D. 2024. Explore until confident: Efficient exploration for embodied question answering, In First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024
  • Ren AZ, Dixit A, Bodrova A, Singh S, Tu S, et al. 2023. Robots that ask for help: Uncertainty alignment for large language model planners, In Conference on Robot Learning, pp. 661–682, PMLR
  • Rimsky N, Gabrieli N, Schulz J, Tong M, Hubinger E, Turner A. 2024. Steering llama 2 via contrastive activation addition, In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15504–15522
  • Romano Y, Sesia M, Candes E. 2020. Classification with valid and adaptive coverage. Advances in Neural Information Processing Systems 33:3581–3591
  • Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. 2022. High-resolution image synthesis with latent diffusion models, In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695
  • Rudinger R, Naradowsky J, Leonard B, Van Durme B. 2018. Gender bias in coreference resolution, In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pp. 8–14
  • Ruffolo JA, Nayfach S, Gallagher J, Bhatnagar A, Beazer J, et al. 2025. Design of highly functional genome editors by modelling crispr–cas sequences. Nature :1–8
  • Sankaranarayanan S, Angelopoulos A, Bates S, Romano Y, Isola P. 2022. Semantic uncertainty intervals for disentangled latent spaces., In NeurIPS
  • Saunders C, Gammerman A, Vovk V. 1999. Transduction with confidence and credibility, In IJCAI
  • Schuster T, Fisch A, Gupta J, Dehghani M, Bahri D, et al. 2022. Confident adaptive language modeling, In Advances in neural information processing systems, eds. S Koyejo, S Mohamed, A Agarwal, D Belgrave, K Cho, A Oh, vol. 35, p. 17456–17472, Curran Associates, Inc.
  • Schuster T, Fisch A, Jaakkola T, Barzilay R. 2021. Consistent accelerated inference via confident adaptive transformers, In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, eds. MF Moens, X Huang, L Specia, SWt Yih. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics
  • Schweighofer K, Aichberger L, Ielanskyi M, Hochreiter S. 2025. On information-theoretic measures of predictive uncertainty, In The 41st Conference on Uncertainty in Artificial Intelligence
  • Sesia M, Favaro S, Dobriban E. 2023. Conformal frequency estimation using discrete sketched data with coverage for distinct queries. Journal of Machine Learning Research 24(348):1–80
  • Shafer G, Vovk V. 2008. A tutorial on conformal prediction. Journal of Machine Learning Research 9(Mar):371–421
  • Shi F, Chen X, Misra K, Scales N, Dohan D, et al. 2023. Large language models can be easily distracted by irrelevant context, In International Conference on Machine Learning, pp. 31210–31227, PMLR
  • Shorinwa O, Mei Z, Lidard J, Ren AZ, Majumdar A. 2024. A survey on uncertainty quantification of large language models: Taxonomy, open research challenges, and future directions. arXiv preprint arXiv:2412.05563
  • Si W, Park S, Lee I, Dobriban E, Bastani O. 2024. PAC prediction sets under label shift. International Conference on Learning Representations
  • Snyder D, Hancock AJ, Badithela A, Dixon E, Miller P, et al. 2025. Is your imitation learning policy better than mine? policy comparison with near-optimal stopping. arXiv preprint arXiv:2503.10966
  • Soumm M. 2024. Causal inference tools for a better evaluation of machine learning. arXiv preprint arXiv:2410.01392
  • Strauss I, Moure I, O’Reilly T, Rosenblat S. 2025. Real-world gaps in ai governance research. arXiv preprint arXiv:2505.00174
  • Subramani N, Suresh N, Peters ME. 2022. Extracting latent steering vectors from pretrained language models, In Findings of the Association for Computational Linguistics (ACL Findings)
  • Suh N, Cheng G. 2024. A survey on statistical theory of deep learning: Approximation, training dynamics, and generative models. Annual Review of Statistics and Its Application 12
  • Sun J, Liao QV, Muller M, Agarwal M, Houde S, et al. 2022. Investigating explainability of generative ai for code through scenario-based design, In Proceedings of the 27th International Conference on Intelligent User Interfaces, pp. 212–228
  • Tan D, Chanin D, Lynch A, Paige B, Kanoulas D, et al. 2024. Analysing the generalisation and reliability of steering vectors. Advances in Neural Information Processing Systems 37:139179–139212
  • Teneggi J, Tivnan M, Stayman W, Sulam J. 2023. How to trust your diffusion model: A convex optimization approach to conformal risk control, In International Conference on Machine Learning, pp. 33940–33960, PMLR
  • The Royal Swedish Academy of Sciences. 2024. Press release: The nobel prize in chemistry 2024. https://www.nobelprize.org/prizes/chemistry/2024/press-release/. Accessed: 2 September 2025
  • Trivedi S, Nord BD. 2025. On the need to align intent and implementation in uncertainty quantification for machine learning. arXiv preprint arXiv:2506.03037
  • Turner AM, Thiergart L, Leech G, Udell D, Vazquez JJ, et al. 2023. Steering language models with activation engineering. arXiv preprint arXiv:2308.10248
  • Ulmer D, Zerva C, Martins A. 2024. Non-exchangeable conformal language generation with nearest neighbors, In Findings of the Association for Computational Linguistics: EACL 2024, eds. Y Graham, M Purver. St. Julian’s, Malta: Association for Computational Linguistics
  • Van Calster B, McLernon DJ, Van Smeden M, Wynants L, Steyerberg EW, et al. 2019. Calibration: the achilles heel of predictive analytics. BMC medicine 17(1):230
  • Van Calster B, Vickers AJ. 2015. Calibration of risk prediction models: impact on decision-analytic performance. Medical decision making 35(2):162–169
  • Vasconcelos H, Bansal G, Fourney A, Liao QV, Wortman Vaughan J. 2025. Generation probabilities are not enough: Uncertainty highlighting in ai code completions. ACM Trans. Comput.-Hum. Interact. 32(1)
  • Vig J, Gehrmann S, Belinkov Y, Qian S, Nevo D, et al. 2020. Investigating gender bias in language models using causal mediation analysis. Advances in neural information processing systems 33:12388–12401
  • Vincent JA, Nishimura H, Itkina M, Shah P, Schwager M, Kollar T. 2024. How generalizable is my behavior cloning policy? a statistical approach to trustworthy performance evaluation. IEEE Robotics and Automation Letters
  • Vovk V. 2012. Conditional validity of inductive conformal predictors, In Asian conference on machine learning, pp. 475–490, PMLR
  • Vovk V, Gammerman A, Saunders C. 1999. Machine-learning applications of algorithmic randomness, In International Conference on Machine Learning
  • Vovk V, Gammerman A, Shafer G. 2005. Algorithmic learning in a random world. Springer Science & Business Media
  • Wald A. 1943. An extension of wilks’ method for setting tolerance limits. The Annals of Mathematical Statistics 14(1):45–55
  • Wang Y, Shi H, Han L, Metaxas D, Wang H. 2024. Blob: Bayesian low-rank adaptation by backpropagation for large language models. Advances in Neural Information Processing Systems 37:67758–67794
  • Watson JL, Juergens D, Bennett NR, Trippe BL, Yim J, et al. 2023. De novo design of protein structure and function with rfdiffusion. Nature 620(7976):1089–1100
  • Wilks SS. 1941. Determination of sample sizes for setting tolerance limits. The Annals of Mathematical Statistics 12(1):91–96
  • Wu Z, Qiu L, Ross A, Akyürek E, Chen B, et al. 2024. Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks, In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), eds. K Duh, H Gomez, S Bethard. Mexico City, Mexico: Association for Computational Linguistics
  • Xia Z, Xu J, Zhang Y, Liu H. 2025. A survey of uncertainty estimation methods on large language models. arXiv preprint arXiv:2503.00172
  • Yadkori YA, Kuzborskij I, Stutz D, György A, Fisch A, et al. 2024. Mitigating llm hallucinations via conformal abstention. arXiv preprint arXiv:2405.01563
  • Yang AX, Robeyns M, Wang X, Aitchison L. 2024. Bayesian low-rank adaptation for large language models, In The Twelfth International Conference on Learning Representations
  • Zhang B, Li S, Bastani O. 2024. Conformal structured prediction. arXiv preprint arXiv:2410.06296
  • Zhang F, Nanda N. 2024. Towards best practices of activation patching in language models: Metrics and methods, In Proc. of the International Conference on Learning Representations (ICLR)
  • Zhang H, Wang P, Chen S, Zhang Z, Qu Q. 2025. Generalization of diffusion models: Principles, theory, and implications. SIAM News https://www.siam.org/publications/siam-news/articles/generalization-of-diffusion-models-principles-theory-and-implications/
  • Zhang L, Rao A, Agrawala M. 2023. Adding conditional control to text-to-image diffusion models, In Proceedings of the IEEE/CVF international conference on computer vision, pp. 3836–3847
  • Zhao J, Wang T, Yatskar M, Ordonez V, Chang KW. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods, In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
  • Zhou H, Huang H, Zhao Z, Han L, Wang H, et al. 2025. Lost in benchmarks? rethinking large language model benchmarking with item response theory. arXiv preprint arXiv:2505.15055
  • Zollo TP, Morrill T, Deng Z, Snell J, Pitassi T, Zemel R. 2024. Prompt risk control: A rigorous framework for responsible deployment of large language models, In The Twelfth International Conference on Learning Representations
  • Zou A, Phan L, Chen S, Campbell J, Guo P, et al. 2023. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405