OER·harvester

← Back to the library
arXiv PDF resource

Future of Information Retrieval Research in the Age of Generative AI

In the fast-evolving field of information retrieval (IR), the integration of generative AI technologies such as large language models (LLMs) is transforming how users search for and interact with information. Recognizing this paradigm shift at the intersection of IR and generative AI (IR-GenAI), a visioning workshop supported by the Computing Community Consortium (CCC) was held in July 2024 to discuss the future of …

Licence
OPEN CC-BY-4.0
Authors
James Allan, Eunsol Choi, Daniel P. Lopresti, Hamed Zamani
Published
2024-12-03 · arXiv
Language
en
Length
16966 words
Type
narrative text

Cites 6 works

inferred
Open ↗ Download Open original ↗

A. APPENDICES

A.1 Glossary

A.1.1 Acronyms and abbreviations

  • AI: Artificial Intelligence

  • API: Application Programming Interface

  • CCC: The Computing Community Consortium whose goal is to catalyze and empower the U.S. computing research community to pursue audacious, high-impact research. https://cra.org/ccc/

  • CIIR: Center for Intelligent Information Retrieval. https://ciir.cs.umass.edu/

  • CLEF: Conference and Labs of the Evaluation Forum whose goal is to promote research, innovation, and development of information access systems with an emphasis on multilingual and multimodal information with various levels of structure. https://www.clef-initiative.eu/

  • CRA: Computing Research Association, a non-profit association of North American academic, governmental, and industry institutions related to computer science and engineering. https://cra.org/

  • DPO: Direct Preference Optimization, an algorithm for large language model alignment.

  • FIRE: Forum for Information Retrieval Evaluation. https://dl.acm.org/conference/fire

  • GenAI: Generative Artificial Intelligence, models and systems that learn to generate new content, including but not limited to text, audio, image, and video.

  • GPU: Graphics Processing Unit

  • HCI: Human-Computer Interaction

  • LLM: Large Language Model

  • IR: Information Retrieval

  • NLP: Natural Language Processing

  • NTCIR: Japan’s NII (National Institute of Informatics) Test Collection for Information Resources. https://research.nii.ac.jp/ntcir/

  • RAG: Retrieval-Augmented Generation

  • RLHF: Reinforcement Learning from Human Feedback

  • SIGIR: ACM’s Special Interest Group for Information Retrieval. Also the premiere research conference in the field. https://sigir.org

  • TREC: The Text REtrieval Conference organized annually by NIST, the U.S. National Institute of Standards and Technology. https://trec.nist.gov A.1.2. Vocabulary terms

  • Digital Twin/Digital Shadow: The terms “digital twin” and “digital shadow” have their origins in work to create close digital representations of physical systems for design and testing purposes. The use of simulators by NASA in the 1960’s to study and model the Apollo moon missions are considered an early example of digital twins. These concepts have been adapted and extended to a wide range of applications involving infrastructure, manufacturing and healthcare. Here, we use “digital twin” with reference to fine-grained modeling of human behavior when using IR-AI systems. While “digital twin” is currently the more common terminology, “digital shadow” is likely to be more correct in the situations contemplated herein. The former refers to a digital model that interacts with the real-world entity it is intended to represent, whereas the latter is a stand-in for a real-world entity that may not actually exist and where there is no feedback loop connection.

  • IR-GenAI: The intersection of information retrieval and generative artificial intelligence research.

  • Multimodal data: data with different formats, such as text, images, videos, and audios.

A.2 CCC Workshop Participants and Report Contributors

First Name Last Name Affiliation
Eugene Agichstein Emory University
Radhika Agrawal Computing Research Association
James Allan University of Massachusetts Amherst
Michael Bendersky Google DeepMind
Paul Bennett Spotify
Jonathan Berant Tel Aviv University / Google DeepMind
Nene Bundu Computing Research Association
Jamie Callan Carnegie Mellon University
Haw-Shiuan Chang UMass Amherst
Eunsol Choi UT Austin
Charles Clarke University of Waterloo
Arman Cohan Yale University
Nick Craswell Microsoft
Jeff Dalton University of Edinburgh
Maarten de Rijke University of Amsterdam
Fernando Diaz Carnegie Mellon University
Andrew Drozdov Databricks
Greg Durrett UT Austin
Nicola Ferro University of Padua
Grace Hui Yang Georgetown University
Petruce Jean-Charles Computing Research Association
Jean Joyce UMass Amherst Center for Intelligent Information Retrieval
Dawn Lawrie HLTCOE at Johns Hopkins University
Michael Littman National Science Foundation
Daniel Lopresti Lehigh University
Mary Lou Maher Computing Research Association
Julian McAuley UC San Diego
Timothy McKinnon IARPA
Qiaozhu Mei School of Information, University of Michigan
Bhaskar Mitra Microsoft Research
Brian Mosley Computing Research Association
Vanessa Murdock AWS AI/ML
Jian-Yun Nie University of Montreal
Negin Rahimi University of Massachusetts Amherst
Siva Reddy Mila / McGill
Mark Sanderson RMIT University
Ian Soborrof National Institute of Standards and Technology
Johanne Trippas RMIT University
Dan Weld Allen Institute for AI
Yiming Yang Carnegie Mellon University
Scott Yih FAIR, Meta
Hamed Zamani University of Massachusetts Amherst
ChengXiang Zhai University of Illinois at Urbana-Champaign
Yongfeng Zhang Rutgers University
Guido Zuccon The University of Queensland