2. HOW THIS DOCUMENT CAME ABOUT
2.1. Pre-workshop activities: How we assembled
Given the importance of the topic, the UMass Amherst Center for Intelligent Information Retrieval (CIIR) organized a relatively small, regional, invitation-only Brainstorming Session on Information Retrieval Research in the Age of Generative AI on December 1-2, 2023, in Amherst, Massachusetts (CIIR hosts Two-Day Brainstorming Session on IR Research in the Age of Generative AI, n.d.). The brainstorming session included seven faculty members from University of Massachusetts Amherst (UMass), six faculty members or senior researchers from five other institutions, and 15 doctoral students from UMass and Carnegie Mellon University. In this two-day event, the future challenges in the intersection of IR and generative AI research were discussed and there was a unanimous consensus that the organization of a follow-up workshop was needed.
The organizers thus reached out to the CCC with a proposal to run this visioning workshop. Upon CCC’s approval, the organizers worked together to expand the organization team and compile a list of potential participants from diverse research communities, backgrounds, demographics, and institutions. A questionnaire was sent to the potential participants to collect feedback on their interest in joining the visioning workshop and preference on timing and location. Given the feedback from over 30 potential participants, the organizers and the CCC decided to hold the visioning workshop in Washington DC on July 19-20, 2024, immediately following the ACM SIGIR International Conference on Research and Development in Information Retrieval, the premier conference for information retrieval research.
A few weeks before the workshop, the participants were asked via a second survey to think about the core topics of the workshop and provide their individual ideas and indicate which of those particularly interested them. The responses were used by the organizers to frame initial ideas for the workshop.
2.2. Workshop activities: What we discussed
In all, 44 experts from various computing disciplines, including 40 invited researchers and 4 organizers, in addition to 5 members of the CRA staff and 1 member of the CIIR staff joined the visioning workshop in person at the Planet Word Museum in Washington, D.C. The participating experts comprised 30 academics, 5 government employees, and 9 industrial researchers; 34 US-based attendees and 10 from
outside the US; 24 primarily affiliated with information retrieval, 12 with NLP, and 8 with other AI topics. A list of participants and their affiliations is provided in the appendix.
The workshop was structured as follows:
- Opening: The organizers kicked off the workshop by describing the visioning process, defining the scope of the IR-GenAI workshop, highlighting participation rules for breakout sessions, and presenting the workshop agenda. There also was a quick introduction of each participant.
- Kicking off the discussions: The organizers presented an overview of the discussions that happened during the CIIR Brainstorming Session on IR Research in the Age of Generative AI and the MSR Workshop on Task Focused IR in the Era of Generate AI. Seven participants were selected by the organizers based on their responses to the pre-workshop questionnaire, aiming for a diverse set of topics of broad interest. They each presented a two-minute visioning idea (a “flash prompt”) to catalyze discussions on those and related topics. These talks were followed by substantial open-floor discussions suggesting alternative topics, expanding on those previously offered, and offering combinations and divisions. Workshop organizers, staff, and participants synchronously took notes on the discussions in a shared document and added additional possible topics or thoughts within the document.
- Forming breakout sessions: The participants then compiled a list of potential breakout sessions. (We also used existing commercial generative AI technologies to produce a list of breakout sessions from the notes; unfortunately without useful output.) The list of potential breakout sessions was iteratively refined based on the feedback from the participants in an open-floor discussion and periodic votes of interest, resulting in eight breakout sessions within IR-GenAI: a. Evaluation: This breakout group focused on (1) developing datasets, methods, and platforms for evaluating the next generation of (LLM-powered) information access systems, and (2) using generative AI technologies to develop more reliable and cheaper (personalized) evaluation methodologies for existing information access tasks. b. Training, Feedback, and Reasoning: This breakout group discussed (1) effective methods to collect and use explicit or implicit feedback from users for training the next generation of retrieval-enhanced generative AI systems, and (2) the obstacles and potential solutions in unlocking higher-level reasoning capabilities in the current generative AI technologies. c. Understanding and Modeling Users: This group discussed major research questions related to what users need and expect from new generative AI-powered information access systems, how to build effective user models for these new technologies, and how to mitigate issues related to ethics and privacy. d. Social Ramifications: This breakout group focused on socio-technical challenges in developing the new generation of information access systems and potential solutions to mitigate or address societal and ethical issues these technologies may create.
e. Personalization: This breakout group explored challenges, opportunities, and potential solutions in developing “digital twins” or “digital shadows” for users to develop effective personal assistants for information discovery and access. These terms are defined in Appendix A.1.2. f. Scaling Across Compute, Data, and Human Efforts: This group built on the idea that current success in deep learning and generative AI research is largely due to various scaling efforts, to discuss potential challenges and opportunities in scaling efforts and how to continue this progress in the scope of IR and generative AI research efficiently. g. AI Agents and Information Retrieval: This group discussed the challenges and opportunities in developing intelligent agents that are ubiquitous, effective, inexpensive and able to provide information, or accomplish tasks on behalf of users. This group also explored how different agents can interactively communicate and solve complex problems, such as planning and decision making. h. Foundation Models for Information Access and Discovery: This breakout group focused on the needs for and potential next steps in developing foundation models, specifically designed for information access and discovery.
- Breakout sessions: Each participant chose one breakout session and participated in detailed discussions throughout Friday afternoon and Saturday. Each breakout group had chosen a leader and a scribe, though writing was typically distributed throughout the group. Each breakout session presented a brief oral summary of their work-so-far to all participants twice, with responses by the full set of participants entered directly into that group’s notes. This process ensured that all voices were heard and feedback was received. The last part of the workshop was devoted to writing up descriptions of sessions and compiling the challenges and opportunities discussed in the breakout session (and presented below).
- Closing: Each breakout session presented two things to the entire set of participants: (1) a major and compelling recommendation that arose in their discussions and (2) a burning question whose answer could improve the group’s report. The workshop ended with an open-floor discussion of the questions, challenges, and any other topic that someone felt needed to be raised.
2.3. Post-workshop activities: How we produced this report
The workshop organizers gathered immediately after the workshop, read the reports from each breakout session, and discussed the outline and potential content of the report. The first draft of the report was produced solely by the organizers using all the material produced by each breakout session (i.e., notes, summaries, comments, challenges, questions, recommendations) and extending or modifying them as needed. The draft of the report was shared with CCC and workshop participants for further feedback and it was revised to shape the report at hand. This report has been reviewed internally by a member of the CCC council and externally by 2 members of the IR/AI research community.