9 References
Adhya, Raahi. “The Appearance of Interpretation: Teaching Humanities in the Age of AI.” Economic and Political Weekly (EPW Engage) 61, no. 8 (2026). DOI: 10.71279/epw.v61i8.48455.
Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. 2021. DOI: 10.1145/3442188.3445922.
Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning 2, no. 1 (1979): 67–90.
Du, Mingxuan, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. “DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.” arXiv:2506.11763 (2025).
Espeland, Wendy Nelson, and Mitchell L. Stevens. “A Sociology of Quantification.” European Journal of Sociology 49, no. 3 (2008): 313–343.
Gadamer, Hans-Georg. Truth and Method. 2nd rev. ed. Translated by Joel Weinsheimer and Donald G. Marshall. New York: Continuum, 2004.
Gao, Tianyu, Howard Yen, Jiatong Yu, and Danqi Chen. “Enabling Large Language Models to Generate Text with Citations.” arXiv:2305.14627 (2023).
Goodhart, Charles A. E. “Problems of Monetary Management: The U.K. Experience.” In Papers in Monetary Economics, vol. 1. Sydney: Reserve Bank of Australia, 1975.
Grafton, Anthony. The Footnote: A Curious History. Cambridge, MA: Harvard University Press, 1997.
Liu, Nelson F., Tianyi Zhang, and Percy Liang. “Evaluating Verifiability in Generative Search Engines.” In Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. 7001–7025. DOI: 10.18653/v1/2023.findings-emnlp.467.
Matthias, Andreas. “The Responsibility Gap: Ascribing Responsibility for the Actions of Learning Automata.” Ethics and Information Technology 6, no. 3 (2004): 175–183. DOI: 10.1007/s10676-004-3422-1.
Kane, Michael T. “Validating the Interpretations and Uses of Test Scores.” Journal of Educational Measurement 50, no. 1 (2013): 1–73. DOI: 10.1111/jedm.12000.
Li, Ruizhe, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. “DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report.” arXiv:2601.08536 (2026).
MacIntyre, Alasdair. After Virtue: A Study in Moral Theory. 3rd ed. Notre Dame, IN: University of Notre Dame Press, 2007.
Messeri, Lisa, and M. J. Crockett. “Artificial Intelligence and Illusions of Understanding in Scientific Research.” Nature 627, no. 8002 (2024): 49–58. DOI: 10.1038/s41586-024-07146-0.
Messick, Samuel. “Validity.” In Educational Measurement, 3rd ed., edited by Robert L. Linn, 13–103. New York: Macmillan, 1989.
Muller, Jerry Z. The Tyranny of Metrics. Princeton, NJ: Princeton University Press, 2018.
Power, Michael. The Audit Society: Rituals of Verification. Oxford: Oxford University Press, 1997.
Raji, Inioluwa Deborah, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. “AI and the Everything in the Whole Wide World Benchmark.” arXiv:2111.15366 (2021).
Schlangen, David. “Dialogue Games for Benchmarking Language Understanding: Motivation, Taxonomy, Strategy.” arXiv:2304.07007 (2023).
Rashkin, Hannah, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. “Measuring Attribution in Natural Language Generation Models.” Computational Linguistics 49, no. 4 (2023): 777–840. DOI: 10.1162/coli_a_00486.
Ru, Dongyu, Lin Qiu, Xiangkun Hu, Tianhang Zhang, Peng Shi, Shuaichen Chang, Jiayang Cheng, Cunxiang Wang, Shichao Sun, Huanyu Li, Zizhao Zhang, Binjie Wang, Jiarong Jiang, Tong He, Zhiguo Wang, Pengfei Liu, Yue Zhang, and Zheng Zhang. “RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation.” arXiv:2408.08067 (2024).
Samarinas, Chris, Alexander Krubner, Alireza Salemi, Youngwoo Kim, and Hamed Zamani. “Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation.” In Findings of the Association for Computational Linguistics: ACL 2025, 13468–13482. 2025. DOI: 10.18653/v1/2025.findings-acl.693.
Sharma, Aman. “Proxy-Based Evaluation and the Limits of Assurance in AI Governance.” ACM AI Letters (2026). DOI: 10.1145/3828672.
Sharma, Manasi, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang, Ankit Aich, Huy Nghiem, Tahseen Rabbani, Ye Htet, Brian Jang, Sumana Basu, Aishwarya Balwani, Denis Peskoff, Marcos Ayestaran, Sean M. Hendryx, Brad Kenstler, and Bing Liu. “ResearchRubrics: A Benchmark of Prompts and Rubrics for Evaluating Deep Research Agents.” arXiv:2511.07685 (2025).
Skinner, Quentin. “Meaning and Understanding in the History of Ideas.” History and Theory 8, no. 1 (1969): 3–53. https://www.jstor.org/stable/2504188.
Song, Yixiao, Yekyung Kim, and Mohit Iyyer. “VeriScore: Evaluating the Factuality of Verifiable Claims in Long-form Text Generation.” In Findings of the Association for Computational Linguistics: EMNLP 2024, 9447–9474. 2024. DOI: 10.18653/v1/2024.findings-emnlp.552.
Wei, Jerry, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le. “Long-form Factuality in Large Language Models.” arXiv:2403.18802 (2024).
Zhao, Jujia, Zhaoxin Huan, Zihan Wang, Xiaolu Zhang, Jun Zhou, Suzan Verberne, and Zhaochun Ren. “ReportLogic: Evaluating Logical Quality in Deep Research Reports.” In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8470–8502. 2026. DOI: 10.18653/v1/2026.acl-long.384.