OER·harvester

← Back to the library
arXiv PDF resource

Future of Information Retrieval Research in the Age of Generative AI

In the fast-evolving field of information retrieval (IR), the integration of generative AI technologies such as large language models (LLMs) is transforming how users search for and interact with information. Recognizing this paradigm shift at the intersection of IR and generative AI (IR-GenAI), a visioning workshop supported by the Computing Community Consortium (CCC) was held in July 2024 to discuss the future of …

Licence
OPEN CC-BY-4.0
Authors
James Allan, Eunsol Choi, Daniel P. Lopresti, Hamed Zamani
Published
2024-12-03 · arXiv
Language
en
Length
16966 words
Type
narrative text

Cites 6 works

inferred
Open ↗ Download Open original ↗

5. ADDITIONAL RECOMMENDATIONS FOR FUNDING AGENCIES AND THE RESEARCH COMMUNITIES

5.1. Recommendations for evaluation campaigns

Funding evaluation pays off in three distinct ways: (1) there are evaluation artifacts such as test collections that can be used to measure systems; (2) those artifacts are built on a task concept, an understanding of what the user is trying to do in the large, that itself can be studied; and (3) the artifacts persist over a long period of time, supporting research well beyond the term of the funded program. Refer to the TREC economic impact study from 2010 (Rowe et al., 2010).

One area which information retrieval in particular holds in high regard (and in which the information retrieval community has made important advances) is evaluation. Pushing the state-of-the-art on shared reproducible benchmarks and test collections has always been important for advancing IR – and all other research areas touched on by this report. Much of the progress toward shifting the state-of-the-art across the information access community has been enabled by large-scale evaluation campaigns that challenge the community to work on a common problem and provide broadly agreed upon measures for reliably comparing system effectiveness.

These benchmarks are now common in several styles:

  1. Challenges often affiliated with a conference or workshop. For example, the RecSysChallenge at the ACM RecSys conference (Recommender Systems Challenge, 2021), and the prototype LLMJudge Challenge run by the LLM4Eval workshop at SIGIR 2024 (LLMJudge Challenge, n.d.).
  2. Challenges sponsored by commercial organizations. For example, the Netflix Prize and the Alexa Prize competitions.
  3. Campaigns run by organizations such as NIST’s TREC (in the US), CLEF (EU), NTCIR (Japan), and FIRE (India), all of which have provided venues for numerous community-wide evaluations. We note that those campaigns are well-known and trusted in the IR community, though it is not clear that they are similarly viewed outside of that community. As information access moves more toward interactive, GenAI-based approaches, we anticipate the need for new evaluation campaigns and new demands upon the evaluations. We discuss some overarching changes that we believe should impact most of these campaigns.

Evaluation campaigns within the information access community predominantly focus on measuring how well systems can match “gold standard” output. Some campaigns are more user-focused, attempting to measure how well systems support a task, either directly or through side-by-side preference comparisons or what is commonly called A/B testing. Those sorts of evaluations have been incredibly successful in understanding and improving systems. However, in the context of societal challenges posed by GenAI, which impact a broad range of people across various professions, backgrounds, and abilities, a narrow perspective on evaluation is no longer adequate.

5.1.a. Support for human-centered campaigns Experimental evaluation frameworks that do not directly leverage human input – e.g., leveraging LLMs to determine whether an answer is biased or using digital twins as synthetic users (E.* and P.*) – continue to be valuable and will likely continue to provide some gains. However, they ultimately risk a type of confirmation bias, where the LLMs agree with LLMs, ignoring human input. Unfortunately, the cost of providing human assessments is prohibitive at the scale needed for training and measuring the impacts of the mechanisms we list. For that reason, we strongly encourage efforts to support the type of human-based evaluations necessary to reduce those risks.

5.1.b. Engage with stakeholders Evaluation – indeed underlying research – must engage with all feasible stakeholders to understand the impact of GenAI and IR (at least) across those groups and the individuals within them (see Societal Ramifications). We thus strongly encourage a rethinking of evaluation campaigns to be stakeholder-engaged. It is critical that specific task-based evaluations be designed and carried out – from start to finish and even beyond – in consultation with the people who are involved with that task and with anyone who is impacted by the results. 5.1.c. Multidisciplinary campaigns To better understand the social issues underlying a challenge and the societal impact of potential solutions, it is necessary to involve research communities outside of computing (in addition to the stakeholders). For example, social and behavioral scientists are skilled in studying human motivations and reactions and are often much better attuned to the impact of technology on marginalized groups. As outlined in Societal Ramifications, we recommend multi-pronged campaigns where studies of users surface challenges, technology innovations are developed, their use by people is explored to identify new challenges and opportunities, technology advances, more exploration, and so on. 5.1.d. Non-proprietary campaigns We strongly encourage community-wide evaluation campaigns where resources can be made available broadly, amortizing the cost of creating judgments across the entire research community. 5.1.e. Hardware (efficiency) aware evaluations Evaluation practices should be geared towards modeling a heterogeneous set of controlled but realistic hardware environments that capture different operating points and constraints (e.g., environment with no GPU resources, embedded systems with specialized GPUs, small-scale, mid-scale and large-scale GPU environments). We envision these taking the form of shared tasks where such hardware environments are standardized and controlled – mimicking the familiar concept of a test-collection in TREC, but from a hardware environment standpoint. We recognize initial inroads towards addressing this goal have been recently made. For example the ReNeuIR shared task at SIGIR 2024 featured a fixed hardware environment and a common set of efficiency measurements (Fröbe et al., 2024) – but it still lacked a holistic evaluation approach, and the modeling of a diversity of constrained hardware environments.

5.2. Recommendations for shared computing infrastructure and resources

There is a need for greater national and international computing facilities specifically that support research on the development and use of generative AI in information retrieval systems. The funding agencies should invest in infrastructure suitable for IR-GenAI. Traditional supercomputers are often not well suited for these tasks. It may require significant storage and GPU capacity at scale to process large-scale multimodal datasets. This requires careful engineering and collocation of compute with storage (petabyte scale), high memory (nodes with multi-terabyte of memory for in-memory dense retrieval), and GPUs optimized for a mixture of both training and inference.

5.2.a. Open-source software for IR-GenAI We call for significant investment in open-source tooling (such as frameworks like DSPy (DSPy: The framework for programming—not prompting—foundation models, n.d.) to maintain and evolve both core agentic tools as well as evaluation methods. 5.2.b. Computing research infrastructure support We recommend computing research infrastructure grants in this area. Funding agencies should support creating such shared resources to ensure everyone can run experiments on a common ground and improve reproducibility. 5.2.c. A non-profit for open science for IR-GenAI. Related to this, we recommend the creation of a non-profit organization for developing open models, open code, and open data for foundation models with applications to information access.

5.3. Funding programs supporting collaborative research

IR-GenAI is an interdisciplinary research topic that involves researchers from various communities, such as information retrieval, natural language processing, human-computer interaction, machine learning, and broadly artificial intelligence. We recommend funding agencies to provide support for collaborative research as major progress will not happen without the cooperation of various communities within and outside the computing field.

In addition to AI-related areas, we recommend programs that address the joint development of IR-GenAI and specialized hardware to support it. GenAI and IR systems' efficiency is constrained by the hardware capabilities upon which they are executed. Direct interaction between researchers developing new hardware and those optimizing GenAI and IR systems would ensure that future AI & IR technologies can fully leverage the next generation of hardware, ensuring short-circuiting the path from hardware ideation to usage.