OER·harvester

← Back to the library
arXiv PDF resource

Future of Information Retrieval Research in the Age of Generative AI

In the fast-evolving field of information retrieval (IR), the integration of generative AI technologies such as large language models (LLMs) is transforming how users search for and interact with information. Recognizing this paradigm shift at the intersection of IR and generative AI (IR-GenAI), a visioning workshop supported by the Computing Community Consortium (CCC) was held in July 2024 to discuss the future of …

Licence
OPEN CC-BY-4.0
Authors
James Allan, Eunsol Choi, Daniel P. Lopresti, Hamed Zamani
Published
2024-12-03 · arXiv
Language
en
Length
16966 words
Type
narrative text

Cites 6 works

inferred
Open ↗ Download Open original ↗

4. SHORT-AND LONG-TERM RESEARCH TOPICS AND RECOMMENDATIONS

In this section we provide more context, additional details, and extra challenges and opportunities from each breakout group, building on the summaries provided above. Each includes summary information followed by a list of short-term (within five years) and long-term (within ten years) research challenges.

4.1. Evaluation

We identified two broad classes of evaluation challenges: (1) the application of genAI to IR evaluation and

(2) the evaluation challenges of generative IR. Generative AI offers two key opportunities to the evaluation of classic document or passage retrieval systems. First, a classic approach to offline evaluation in IR is the use of a test collection. The construction of these collections is resource intensive, in particular the creation of the lists of the documents that are relevant or not relevant to queries. There is growing evidence that LLMs may be useful in labeling documents. Research to explore and refine this labeling process holds great promise. A second key aspect of offline evaluation is the so-called user evaluation: the observation of people in how they initiate or react to the actions of a search engine. There is the prospect of exploiting GenAI to simulate the actions of people. One might even speculate that human “digital twins” could be created for the purposes of testing an information retrieval system. For the purposes of evaluating a generative IR system, there is much to consider. Retrieval Augmented Generation represents a coalition of IR and generative systems to accomplish a task. It is necessary to consider evaluation of these component systems in concert as well as individually. It is also necessary to establish the domain of competency of generative IR systems. A challenge with the output of a generative system is that it always expresses an answer with apparent confidence. There is great value in knowing the knowledge domain that a system is competent in. How do we establish and communicate that competency? In a related topic, we don’t know how an LLM is achieving its answers. A research challenge is to establish a methodology for building confidence in the output of LLMs. Finally the reproducibility of systems needs to be considered. Note that recommendations for carrying out community-wide evaluations – some of which might be impacted by the research suggested here – are presented later in this report.

4.1.1. Recommendations for short-term (the next five years) research programs

4.1.1.a. Using LLMs to support the relevance assessment/labeling process

To support the evaluation of classic search, researchers rely on a fixed set of information needs and the identification of relevant documents. There is currently an explosion of interest in having LLMs do the labeling of documents to control costs. There is a spectrum of ways to use LLMs to that end, from replacing human judgments with LLM labels to utilizing them as an assistant. Using LLMs as labelers directly suffers from various issues. For instance, LLMs make mistakes, create misinformation through hallucination, and suffer from various types of biases. Alternatively one could use an LLM to assist an assessor who is the one responsible for providing the labels. In particular an LLM could help assessors be more consistent by having LLMs partner with assessors to identify documents that may need to be rejudged because relevance judgments are inconsistent. Here continual fine-tuning of a personal, topic-specific LLM could support label suggestion.

4.1.1.b. Proposing relevance rationales (one-line comments) and highlighting

For many evaluation data sets, relevance labels are applied at the document level, but there is no recording of why a document was labeled as relevant or – when not all of the document is relevant to the information need – which part of the document contains the relevant information. A research direction could investigate whether LLMs could be used to generate written descriptions of why a particular document is relevant. Alternatively, LLMs could be used to identify which particular parts of a document provide relevant information.

4.1.1.c. Evaluation protocols able to assess the RAG pipeline as a whole, digging into the

interaction between retrieval and generative components

The current evaluation methodologies for Retrieval-Augmented Generation (RAG) pipelines assess retrieval and generation separately, assuming that the better the retrieval (evaluated on its own), the better will the generation (evaluated on its own) based on the top-ranked documents. However, there is not an a priori guarantee that the text of the most relevant document used as a prompt will lead to the generation of the best output; it could be that the second or the third relevant document would provide a better priming of the LLM, inducing it to produce a better response. Current evaluation is not designed to investigate these aspects and this kind of interaction and there is an opportunity to develop more comprehensive approaches able to assess the whole RAG pipeline and clarify the effects of this interaction. The benefit of these new evaluation methods would be to not only lead to improving RAG itself but also to shed some light on the inner workings of LLMs.

4.1.1.d. Domain of competency for LLMs and Prompt Performance Prediction

It is currently thought that LLMs embed knowledge distilled through their learning process, but it is not fully clear what this knowledge actually is and what domains it covers. There is a broad understanding that general purpose models need to be specialized or fine-tuned to specific domains and this led to a proliferation of models for specific languages or areas. There is also the idea that bigger and bigger models should be able to perform better and better in more and more domains, even if it could still be the case that they lack enough training data in a given domain or that the relative size of the training data for a domain becomes smaller and smaller as the other domains grow bigger, making the model less

effective in that domain. In general, there is a lack of an evaluation methodology able to investigate and determine which is the actual domain of competence of an LLM and/or how an LLM could perform in a given domain (or outside it). This kind of evaluation would be needed for understanding the extent to which an LLM is suitable for one or more domains. Moreover, it would open the possibility for a new area of investigation, i.e. Prompt Performance Prediction (PPP), aimed at estimating whether a given prompt could be effectively dealt with by an LLM, given its domain(s) of competency.

4.1.1.e. Guarantee reproducibility and generalizability of the experiments, in order to improve the

validity of the conclusions drawn, e.g. by controlling for the effect of (small) variations in prompts Reproducibility is a primary concern in almost any area of science and it has been discussed in IR for many years as well. LLMs, and particularly proprietary LLMs, bring the reproducibility concerns to a completely new and different level. There is an urgent need for much more “open source” and shared models and datasets, in terms of code, training data, training regimes, intermediate snapshots, and so on. This is needed just to be able to conduct experiments on a common ground and to strive for reproducibility. Another angle of reproducibility concerns the extreme variability in the functioning of LLMs themselves where even small changes in the prompt may lead to substantially different answers. Therefore, there is a need for a more systematic investigation of the behavior of such systems in order to understand how to control the effect of such small variations or how to account for them when assessing and comparing systems, e.g. by suitably computed “confidence intervals”. The ultimate goal is to improve the validity and generalizability of the conclusions drawn from the experiments.

4.1.2. Recommendations for long-term (the next ten years) research programs

4.1.2.a. Simulating user studies and A/B tests / Digital Twins

Simulation has been pursued for a long time in IR, and LLMs provide a concrete opportunity to push its boundaries and open up to the evaluating system by credibly simulating the user behavior or even by creating whole synthetic experimental collections (e.g., topics, documents, relevance judgements, clicks). Ultimately this kind of simulation would correspond to the creation of digital twins of people/users (or better of groups of people/users) which can be exploited to assess systems. Such digital twins could be useful when testing new systems for which there are no existing interactions and one wants to be proactive. Such a simulation would clearly offer the possibility to scale-up and make evaluation much more fine-grained, allowing for more systematic testing before moving to the user-side. Simulation naturally fits in the continuum of evaluation which ranges from lab-based evaluation based on experimental collections, i.e. a kind of static simulation of user information needs and answers to them, to interactive evaluation, e.g. A/B testing with real users in an operational setting. One open question is whether current LLMs are enough to create such digital twins or whether we should imagine more advanced models which, besides the language generation capability, include behavioral models of users, browsing or interaction components, and more. In some sense, these enhanced models could parallel what multimedia models are today, instead of being constituted by multiple media, they would be constituted by multiple components of such a digital twin. The perspective of such models raises several questions: how do we validate them? How stable would they be? Would their training data be rich and diverse enough to (even partially) capture what a real person is? It is important to recall that this will also be an approximation of users and thus in the end cannot substitute for actual user studies. Note that

developing “digital twins” is also discussed in Section 4.5 for a different purpose, which is related to user simulation but has some fundamental differences.

4.1.2.b. Designing evaluation that is capable of accumulating evidence about an LLM's ability to

provide the correct answer because of the correct processing/inference

LLMs are systems initially designed for a specific task, i.e. language generation, that are now used for many different purposes and downstream tasks. The current evaluation methodology is designed to assess systems that were specifically designed for one single purpose, e.g., sentiment classification. As a consequence, it is based on verifying that for a given input, the expected output is produced, e.g., positive or negative sentiment. However, LLMs are not designed for a specific purpose and there is not an a priori guarantee that the correct answer is generated because of the correct process/inference/reasoning. There is a need and an opportunity for developing deeper evaluation methodologies to investigate and provide evidence about how and why a given answer has been generated. One can consider this linked to some sort of “explainability”, or can view this as a kind of “debugging”, or if one imagines an LLM as a big matrix or an extremely complex circuit, this kind of evaluation tries to understand which parts of the matrix (or the circuit) deliver classification, which ones ranking, etc. In this sense, this is different from AI explainability which is instead targeted at providing an explanation for a specific answer, e.g. which features mattered most. We note that this is related to the “Domain of Competency” listed as a short-term challenge.

4.1.2.c. Evaluating results of systems that generate responses rather than rank documents

Responses to information needs that are knowledge requests where a generative response is appropriate will need new approaches to evaluation. This is in contrast to item-based retrieval where a ranked list is an appropriate response. The overarching evaluation of generated output will need to verify that the output is responsive to the information need. Different evaluation approaches will be needed depending on whether the generated output includes citations to source documents or not. With source documents, existing approaches that utilize document labels could be integrated into the evaluation; however, whether or not there are document citations, there will need to be approaches to evaluating whether or not the content that appears in the generated text is responsive to the request. LLMs have a role to play in answering such questions. It is likely that in addition to new frameworks for evaluation, new metrics for evaluation will be needed. The reproducibility of the evaluation is an important consideration and how to perform the evaluation in a way that enables consistent comparison of systems (ie. techniques) is also critical.

4.2. Training, feedback, and reasoning

At the heart of retrieval-augmented generation (RAG) systems, and more broadly retrieval-enhanced generative AI (RE-GenAI) systems, is the interaction between the information retrieval / access model and the language model (LM), and eventually the user. Current training methods focus on specific and narrow definition of RE-GenAI and mainly focus on training RAG components separately, but we see promise in tightly coupling the training of the information access and language model, because such training would allow components to adapt to one another for the purpose of doing well on the downstream task. Some near challenges in joint training are to learn when to retrieve (e.g., based on measures of model uncertainty), when to rely on parametric knowledge, and how to deal with conflicting information.

Moreover, the granularity and the representation space of retrieval are also unclear, i.e., retrieving sentences, passages, entire documents, or even latent representations from a memory. When it comes to new sources of human feedback, it is unclear what constitutes meaningful feedback when click data (user clicks on relevant documents) is no longer available. Some sources of implicit feedback include absence of follow-up questions or question reformulation, which provide explicit scalar or natural language feedback, and how to perform credit assignment from target answer to source documents. We also envision that components can self-improve – the information access and LM components can provide training signals to one another. Nevertheless, we realize the importance of human annotations, and advocate for a mix of human and machine annotations and argue research on principled integration of human and machine labeling is required. We also envision some paradigm changes where retrieval is possible even when it requires complex reasoning over document collections, e.g., when queries are not similar to the underlying knowledge but are more abstract and/or require aggregation and synthesis at different granularity.

4.2.1. Recommendations for short-term (the next five years) research programs 4.2.1.a. Joint and end-to-end training of retrieval and language models at both pre-training and fine-tuning phases The information access models in RE-GenAI systems may benefit from being trained with awareness of the GenAI model. Similarly, the GenAI model may benefit from being trained with awareness of the information they are likely to encounter as the output of an information access system. This diverges from current practice where the information access and the GenAI systems are trained separately and only combined at test time. Even though there are some preliminary studies in this space, many open research questions exist, such as learning optimal communication protocols and languages between IR and GenAI systems. See 4.1.1.c for recommendations for evaluating such a pipeline.

4.2.1.b. Learning when to retrieve vs. generate

Current RAG systems are typically implemented as a “hard pipeline”: retrieve, then feed the inputs into the reader for generation of the answer. Some work has shown how to do retrieval more dynamically, either with sequential queries or even based on the output of previous outputs from information access systems. A key question here is how to estimate generation uncertainty: if we know when a language model will produce an inaccurate answer, we have found a point in the generation where retrieval is more likely to improve the reliability of the results. Defining optimal communication language between GenAI and retrieval models is of significant importance to make retrieval versus generation decisions.

4.2.1.c. Uncertainty modeling in alignment

Related to the above point, we expect it to be important to develop techniques for calibration of LLMs in the context of alignment. For language modeling, it is still not understood how different alignment techniques (RLHF, direct preference optimization (DPO)) impact calibration. For retrieval-augmented settings, the uncertainty estimation needs to consider not just the language model’s behavior and output, but also the retrieved documents.

4.2.1.d. Self-improving RAG systems

Systems where the retrieval and generation component provide training signals to one another to progressively and jointly improve (what are good search queries, what do good documents look like).

4.2.1.e. Training better retrievers given feedback from users only on generated answers

Unlike in search engines where users can provide explicit and implicit feedback at document level (e.g., by clicking on documents in the search engine result page), translating user feedback on generated answers to document-level feedback on retrieved information is challenging. Future work can potentially address these issues through: (1) credit assignment from generated answers to source documents. Can we identify where a generated answer referenced a particular source document? This potentially requires reader models that generate attributed answers or interpretability techniques that enable us to recognize and manipulate how closely the generated answer matches the source. (2) coming up with new training signals based on user behavior (e.g., eye tracking, session data, interactions with citations). See 4.1.2.c for ideas about evaluation when there are generated answers.

4.2.1.f. Training from a mix of human annotation and machine labeling

We envision a hybrid paradigm where the retriever is trained with machine labels as above, or humans can directly intervene to indicate where poor retrieval has hampered the performance of the system. Again, this notion requires some notion of epistemic uncertainty to know where human annotation is necessary. This can be further improved by leveraging methods from active inference (i.e., choose examples to evaluate on rather than train on for reliable unbiased estimation of model performance).

4.2.1.g. Training a retrieval model (a search engine) for multiple GenAI Systems

Similar to search engines that provide service to many users and learn from their interactions, we can envision a retrieval model or a search engine that provides service to many GenAI systems, depending on their needs. These systems need to formalize a standard query language that is shared among different GenAI systems. A major benefit in working on this approach is that it can benefit from feedback from a diverse set of GenAI systems, helping the retrieval model to generalize across GenAI models and learn more effective retrieval functions. Calibrating feedback across GenAI systems in order to optimize the search engine is among the important research directions in this area. Personalizing the result list for each GenAI system is of great importance, as each demonstrates different behavior and consumes the result lists differently.

4.2.1.h. Developing information access models that require reasoning

Complex questions that are posed to GenAI systems are often different from the types of queries fed into conventional search engines. Queries may require additional reasoning, which poses problems for search engines. Queries may also require skills like aggregating over opinions, e.g., to summarize or give insights from product reviews. We believe that additional benchmarks and reliable, reusable, and scalable evaluation methodologies should be developed to test this setting. Furthermore, different types of retrieval machinery may be needed.

4.2.1.i. Instruction-following information access models

How can we get IR systems to follow prompts in queries similar to LLMs, which define constraints external to the semantics of queries. To develop such systems, we must (1) move away from symmetric encoders as the query encoders need to incorporate instructions in addition to query text, (2) define constraints on properties of the documents-prefer formal sources, prefer personal opinions, etc. (3) invest in developing generative information access models. There is no doubt that instruction-following retrieval models are necessary to build general-purpose retrieval models. Progress in this area requires large-scale training datasets that include a diverse set of instructions. Generalization to unseen instructions is a significant challenge in this area.

4.2.2. Recommendations for long-term (the next ten years) research programs

4.2.2.a. Fresh generative models

As opposed to using RAG to put information into the context of an LLM, we can imagine future systems where retrieved documents are directly used to update the parametric knowledge of an LLM on the fly. For instance, knowledge editing methods promise to enable updates that LLMs can successfully reproduce to answer new queries. Can we get to a point where a new document can be added to an LLM and reproduced losslessly when the LLM is queried appropriately? Challenges include the difficulty of knowledge editing methods to correctly update the desired knowledge.

4.2.2.b. Continual joint training of parametric and nonparametric memory

When models are trained from scratch, they encode some facts reliably in their parametric knowledge. As we develop RE-GenAI systems, we may wish to establish a boundary between what information needs to be memorized (and memorized completely) vs. what information a language model should ignore and assume it will be provided at test time. A model should learn what to put in each type of memory, and when to take non-parametric information and on-the-fly move it into parametric memory when it proves generally useful.

4.2.2.c. Implicit user feedback modeling

Current feedback channels include explicit supervision and implicit supervision from users like clicking on links in generated responses. However, generative AI systems bring up a host of new challenges and potential new feedback types given the emphasis on long-form text generation. For instance, users may re-query an LLM with a modified version of their query, like in traditional IR, but they also may simply ask the LLM “show me something different” or “elaborate on [aspect] of the [response]”. Additionally, the way a user processes LLM responses could be further studied with mechanisms like gaze tracking. All of these contribute possible methods for additional user feedback.

4.2.2.d. Reasoning over user context and recommendations

The input to the retrieval model from the GenAI system does not have to be an explicit query describing the information need (in either natural language or latent representation). Instead the input could be some contextual information (such as user’s short- or long-term history, user’s location, device, time, etc.) that could initiate a recommendation functionality depending on the context (without an explicit query). This will enable us to develop systems that consume, synthesize and reason over the provided contextual

information and recommended items and proactively interact with the users when needed. Such advancements are necessary to develop truly mixed-initiative intelligent agents.

4.3. Understanding and modeling users

Generative AI provides us with new capabilities that allow us to imagine richer interactions with users. Future information access systems and intelligent agents should offer richer understanding of the user's tasks, goals, context, experience, and cognitive state. To support complex tasks and decisions, these systems should generate plans to decompose the tasks and track task progress and user state jointly. The systems should be more interactive and proactive, offering the opportunities to users to do multi-turn interactions, to integrate multimodal data, and to accomplish more complex tasks. The new interactive interface raises new challenges that should be dealt with in future research: user goal detection, integration of rich context information, task and user modeling, planning and choosing appropriate actions and determining the right type of answer or results for the user. GenAI can help implement these functionalities but we will also need to develop specific models for them. In addition, data and evaluation will be important challenges. What is envisioned for future users is an integrated system assisting users to accomplish various tasks that often require information access.

4.3.1. Recommendations for short-term (the next five years) research programs

4.3.1.a. Multimodal interactions

In the next five years, we expect that the GenAI research community will make advances in foundation models, creating the ability to query and return results from multimodal data. The research community could look to make progress toward intent detection in multimodal context: what is the user’s need in the current context?

4.3.1.b. New information access use cases enabled by generative AI

The current chatbot model focuses on linear question-response multi-turn interactions, with some customization based on prior statements or conversations. However, this ad hoc query approach encompasses only one of many potential use cases for GenAI. A gap that research could address is the creation of a taxonomy of prompts in the context of GenAI. A unified set of terminology could aid the research space in identifying new use cases for LLMs and other GenAI, including push-based models that provide information to users without the need for a stated query. Based on the identified needs, a system should identify the right type of input (multimodal) and response—including problem breakdown and generation of single mode or multimodal responses. Evaluation approaches (as described elsewhere in this report) could focus on the extent to which users’ diverse needs are met.

4.3.1.c. Developing robust user models within IR-GenAI systems

An important prerequisite for achieving the aforementioned future goals is a more robust model of the user. At present, there are two disparate representations of the user: the model that the user would use to describe themselves and their interests, and the representation that a system has for the user. In the future, we would expect to bring these closer together–creating a cognitive state representation of the user, which could be operationalized by LLMs and other GenAI. Given the need to maintain accuracy, the cognitive model should be both explainable to the user, as well as correctable by the user—potentially

supported by natural language conversation with the model of the user. Additionally, there is a need to evaluate the representativeness of the model of the user.

4.3.1.d. Privacy and economic implications of IR-GenAI technology

To what extent is this model an accurate representation of the user? Taking a cue from eDiscovery and posthumous search spaces (i.e., selective disclosure), identifying what information should be provided or withheld from systems for privacy reasons could be treated as an information retrieval problem. Given the potential level of detail provided by and about the user, should they be able to control sharing of this data–for instance, by receiving payment for their personal model?

4.3.1.e. Cultural, social biases, and trustworthiness in user models

Over the next few years, strides should be made toward modeling biases in both foundation models as well as user models. Consistent with the recommendations provided in Section 4.4, this could include cultural and social biases inherent in the training data and the information conveyed by or about the user. Training data is a fundamental area for additional research. Models learn from information that is provided to them, but not everything that represents reality is written down. The user can be influenced (intentionally or unintentionally) by the results that the GenAI system produces. As a result, to make sure the user sees a result that is accurate, additional research could focus on elucidation and definition of data such as common knowledge: things that users generally know that are unlikely to be written down in a format that a model could see. The trustworthiness of the information should also be taken into account and indicated to the user. This is a key area, because of the risk of providing the user results that do not represent reality. For example, in the medical field, researchers may publish negative findings or contraindications, but not positive or neutral results. These types of research gaps or imbalances could lead into biases within data that could lead to weighting toward negative responses, which do not represent the full range of human understanding. To enable models to provide more balanced results, the research community could also consider new publication approaches that emphasize a broader range of results.

4.3.2. Recommendations for long-term (the next ten years) research programs

4.3.2.a. Cognitive models for information access

The research community can investigate at a deep level the ability to implement cognitive models for information access. While there is the potential for significant time savings, there are also significant implications for privacy and security related to a digital twin.

4.3.2.b. Robust modeling of the user's state of knowledge

Technology can focus on creating user models that can trace the user’s state of knowledge in a robust way, identifying preferences and gaps and how to fill those gaps. Future GenAI-driven AI systems could take proactive steps related to filling user needs based on their knowledge of the user and their interests, in some cases preemptively providing information built from diverse sources, without the user having to enter some input. The research community could make strides toward how to plan and solve complex research tasks, driving toward the user’s goal state.

4.3.2.c. Controllable information access

Considerations should be taken to understand the need to maintain an ability for the user to choose their actions. Rather than falling into a scenario where the GenAI capability suggests something to the user and they blindly follow the top course of action or recommendation produced by the system, there should be some capability maintained for user choice. Serendipitous discovery and novelty should be built into the systems.

4.3.2.d. Addressing digital amnesia

As a research community we should develop an accepted methodology to measure the effects of GenAI on people’s ability to find and process and critically evaluate the results. One expected effect is amplification of so called “digital amnesia” - people degrading ability and interest to retain information that can be easily retrieved. As GenAI provides capabilities to digest and summarize and recommend and synthesize knowledge, users will need to develop new skills to evaluate the results, and the GenAI systems need to help users to use human intuition and common sense and provide background knowledge to help evaluate the answers.

4.4. Social ramifications

What are the consequences of developing and deploying information retrieval technology in the real world? What are the mechanisms through which information retrieval and GenAI technology bring about these consequences? And what are the corresponding risks to relevant stakeholders? In this section we provide a framework for thinking about and clarifying these questions.

Identifying risks, challenges, and opportunities of generative AI for information retrieval requires an interdisciplinary approach informed by socio-technical perspectives and co-developed with social science scholars, legal scholars, civil society representatives, and policy makers among others.

We are motivated by previous literature (Mitra et al., 2024) at surveying some of these sociotechnical implications of generative AI for information access. That work identifies several systemic consequences and corresponding mechanisms and risks of generative IR. We focus our recommendations on the “mechanisms” that represent sites of potential and concrete mitigation. That work identifies a list of 16 mechanisms that contribute to different corresponding consequences and risks: content pollution, the “game of telephone” effect, search engine manipulation, degrading retrieval quality, direct model access, the paradox of reuse, compute and data moat, AI persuasion, AI alignment, appropriation of data labor, bias amplification, AI exploitation & doxing, industry capture, pollution of research artifacts, resource demand & waste, and persuasive advertising.

Below, we indicate short-term and long-term challenges for two of these mechanisms, illustrating how we can develop research agendas corresponding to each of these mechanisms. We suspect that some of those mechanisms will not admit to this measure-and-mitigate process, creating a set of preliminary challenges.

We also describe important implications for research into socio-technical challenges in the Evaluation Campaigns section below.

4.4.1. Recommendations for short-term (the next five years) research programs

4.4.1.a. The “game of telephone” effect mechanism

This mechanism addresses the tendency for generative models to produce information that is incorrect, putting the burden of interpretation on the user: “hallucinations” is a well-known instance of this mechanism. There are broadly two complementary research directions that require attention in the short-term. (1) We need to study and develop quantitative measures for how often LLMs misrepresent information from retrieved documents in their responses, either by providing incorrect information or providing information without context that leads to misinterpretations. Correspondingly, we need to also study the impact of the incorrectness on users. (2) We also need HCI research to understand how information should be presented in a conversational context that allows for and encourages searchers to appropriately inspect and verify the presented information. For example, we should study the effectiveness of providing references in the generated responses. For e.g.,

  • Do searchers pay attention to them?
  • Do they actually click-through the references to verify presented information?
  • Can the number of references actually mislead the searcher to trust the presented information when that’s not justified? In the same vein, we need to explore newer ways for information presentation towards empowering searchers to more effectively verify the information they are presented with. Answering these important questions will require substantial effort devoted to user studies, with the attendant funding needs that come with supporting such work.

4.4.1.b. The AI persuasion mechanism

This mechanism considers the possibility that generative systems may inadvertently (or intentionally) persuade users on opinions or perspectives that they do not already hold – for example, by masquerading as a trustworthy source or appealing to learned biases of the user. Correspondingly, we need research to identify and understand the mechanisms and potential mitigations. Research is also necessary to understand how persuasion capabilities interact with the incentives structures that information access systems may optimize towards, such as monetization. Recent experiences with the negative impacts of human-generated misinformation add urgency to such work, since AI will become even more pervasive with the capability of scaling without limits (so-called “Dark LLMs”). Finally, research is necessary to identify possible safeguard mechanisms and develop participatory processes and mechanisms for civil society and external scholars to investigate and audit these phenomena.

4.4.1.c. Cost and footprint

The computational demands of generative AI models pose challenges in community involvement in research. Historically, AI models have been small and efficient enough to run on commodity hardware, while today’s large AI are effectively too inefficient and computationally demanding to do so. This has implications in the availability of these models for use by certain groups, including small organizations, students, and many underdeveloped and developing countries, which only exacerbates inequality in society given the growing role AI now plays. Advances in the efficiency of these models have the potential

to democratize the technology, and more work in this space is needed. The environmental implications of the high computational of both generative AI and Neural IR are significant. Improvements to training and inference efficiency of these models – as well as the intersection between them – can help mitigate these concerns.

4.4.2. Recommendations for long-term (the next ten years) research programs

4.4.2.a. The “game of telephone” effect

Future research questions for the “game of telephone” effect should explore how conversational agents (and LLMs in general) can be employed to intervene and encourage users to reflect on and critique the information they are presented with by posing appropriate questions, e.g., serving as a “devil’s advocate”. These “critical literacy” agents may also find application in mediating dialog between searchers for knowledge production and consensus building.

4.4.2.b. AI persuasion mechanism

For the AI persuasion mechanism, we need to develop frameworks to reflect on questions on how conversational agents should ethically interact with people, who they represent and how to be transparent about that, and what higher level objectives they are optimized for in these contexts.

4.5. Personalization

Improving current GenAI technology through personalization is undoubtedly an important research area with significant impact on society. Here we discuss new problems and challenges in the intersection of personalization, IR, and GenAI, which can have major applications in developing the next generation of recommender systems, question answering, writing assistants, and intelligent agents. We iterate over challenges and opportunities in developing digital twins (or “digital shadows”), which are capable of representing new models per-user, based on users’ historical interactions. Such models can be used for personalized retrieval or to generate writing in a user’s style. (They can also be used for evaluation as outlined earlier at 4.1.2.a.) We argue that systems for recommendation / retrieval will gradually push the boundaries of retrieving existing content to synthesizing new content. While traditional IR can return what exists, future IR systems will generate what you want even if it doesn’t exist (or a process for making it). While LLMs already synthesize content to some extent, we are specifically interested in content synthesis that is (a) more personalized to individual users than what is currently possible; and (b) able to synthesize types of content that go well beyond what is currently possible with LLMs, ranging from mixed media to (designs of) physical objects. Future personalized recommender systems will not just recommend items but might persuade a user that a specific recommendation is the best (e.g. via a conversational interface). We also highlight that personalized interfaces need to become proactive and adaptable to user needs, expectations, and preferences. We acknowledge that researchers in this area must be aware of privacy considerations associated with using user’s data. Collecting reusable personalized datasets that also preserves user privacy is quite important in advancing personalization in generative AI models.

4.5.1. Recommendations for short-term (the next five years) research programs

4.5.1.a. Personalized intelligent assistants in operating systems

As we build stronger generative AI systems that can leverage user’s personal data, preferences, and past history, we envision the development of personalized intelligent systems within operating systems that can not only execute actions, but can also retrieve and recommend files, directories, and applications. These systems should gain user trust given the information they need access to.

4.5.1.b. Personalized dashboard for controllable digital twins

We acknowledge the importance of providing control to the user to decide what information to be used in the development of their digital twin. This calls for the development of customizable dashboards that allow users to select what information is most relevant to the retrieval task or problem solving. Can we generate interfaces that are adaptable to users in their context and allow the user to be intentional about what information is included in the retrieval and/or generative task? Ultimately, there is a large range of possible interfaces that combine elements from “traditional” IR systems with generative components (e.g. mixed retrieval and conversational interfaces).

4.5.1.c. Online learning for personalized intelligent assistants

The twin must “move with the person” as they conduct their digital life. Since recent experience is often necessary for full contextualization, updates should be immediate. This calls for the development of new online learning and machine teaching algorithms for keeping personalized intelligent assistants up-to-date.

4.5.1.d. Personalized question answering systems

Despite their limitations in hallucination and lack of up-to-date knowledge, generative AI systems have shown great promise in answering user’s questions. However, users have different background knowledge: for the same question, a user may need extended background information, while another user may need a brief and to-the-point answer. Developing resources that enable personalized question answering research and developing models that preserve the privacy of users together with providing personalized experience are among major research directions in this space.

4.5.1.e. Modeling the trade-off between privacy and personalization

Developing effective personalized generative AI systems requires training on and/or consuming user’s personal data, which can potentially harm user’s privacy. We acknowledge that this is an important and sensitive topic and should be considered seriously when designing personalized systems. Modeling the trade-off between the personalization effectiveness and privacy preservation would smooth the path towards building trustworthy personalized systems.

4.5.2. Recommendations for long-term (the next ten years) research programs

4.5.2.a. Exploring various approaches for developing digital twins

Even though immediate actions can be imagined in this space, developing effective, robust, and controllable digital twins requires significant long-term research investments. We must explore in what situations and tasks we need to use a non-personalized intelligent assistant and in what situations we

must use the user’s digital twin. Exploring the most efficient way to create different models tailored to specific contexts and integrate them seamlessly for holistic personalization.

4.5.2.b. Personalized content synthesis

Using generative AI technologies we sometimes move from retrieving information items to generating content. Such content generation approaches can better address user’s information needs through synthesizing new content in a personalized way, building upon existing work on personalized search, recommendation, intelligent agents, and text generation. Preserving the privacy of users should be considered in this research.

4.5.2.c. Persuasive information access

Can generative AI be used to build persuasive recommenders that (e.g., through a conversational interface) guide users toward certain items or towards achieving personal goals? To achieve this, recommender systems should provide explanations and reasoning. These explanations are not just why the system generated a recommendation list, but why and how the user can benefit from the recommendation.

4.5.2.d. Personalized result generation and presentation

How can we enable users to interact with and manipulate their information in a visceral environment? Ultimately we envision IR systems in which information is presented in a way that is more personalized / contextualized to individual users, moving away from more “rigid” IR interfaces and closer to more “human-to-human” interactions.

4.5.2.e. On-device digital twins

As research in developing effective generative AI models on a small scale that can be handled and processed using devices with limited resources, such as smartphones, progresses, we can envision development of privacy-preserving digital twins on personal devices, such as smartphones. Here, developing in-device online updating of generative AI systems using personal data is a major challenge.

4.6. Scalability and efficiency

Current generative AI systems and the applications built upon them require great expertise and massive data, and computational resources. Focusing research effort and funding on more efficient systems that require fewer computational, data, and human resources lowers costs, increases access, and enables application development for smaller devices and a wider range of human activities.

Better exploiting emerging zero-shot capabilities would reduce or eliminate training and data costs in IR models. The paradigm of prompting LLMs with instruction or in-context examples has enabled the generalization to new tasks or domains without needing to collect and use training data (“zero-shot”). This paradigm can deliver substantial savings in terms of computational and data costs, and by extension human effort (i.e., task-specific training data need not be collected). Explorations on using LLMs for retrieval and ranking tasks in a zero-shot setting are fast emerging. However, leveraging these capabilities efficiently remains a key challenge. Zero-shot effectiveness appears to emerge at large model sizes, thereby requiring high computational costs to leverage. Meanwhile, these models can be distilled into

smaller or more efficient architectures, though this comes with computational costs due to the additional distillation step. Therefore, the challenge remains on how to efficiently adapt LLMs to new corpora, contexts, and tasks.

Moreover, search engines will continue to be an important method of accessing information in standalone applications and as components in other systems due to their low storage, computational, and expertise requirements. LLMs will be used to generate new ways of representing queries, documents, and users that greatly improve accuracy in these efficient search engine architectures. However, existing search algorithms and data structures are based on patterns of human language usage that are well-understood. The use of new representations generated or motivated by LLMs is likely to disrupt the assumptions on which existing data structures and algorithms for efficient retrieval are based.

4.6.1. Recommendations for short-term (the next five years) research programs

4.6.1.a. Continual learning

One emerging opportunity in this domain is developing robust methods for continual learning of IR models and LLMs. Emergence of new knowledge and newer techniques will both motivate continual learning. As developing models from scratch is getting prohibitively expensive, we should develop techniques to re-use existing models and indexed corpus to build the next generation of models. This could come in different forms: (1) Updating existing models with new knowledge by updating their parameters, (2) Developing new architectures which adapt representations or parameters from existing models.

4.6.1.b. Scaling laws for training IR models

As GenAI models are getting more expensive to train, the community started to study how to train compute-optimal models: training over more tokens and training bigger models both will bring performance gain at the cost of compute, so how should we balance these? Studying these relationships enable researchers to determine optimal allocation of a fixed compute budget. While IR models are also getting much bigger, we lack study on such trade-offs in training compute-optimal IR models.

4.6.1.c. Efficient training of LLMs through IR-based data curation

Current generative AI models are trained on massive amounts of data. There is currently evidence that being selective in the process of training data curation can lead to stronger models than those that are less selective. This data selection process relates strongly to core ideas in information retrieval, including the efficient organization and clustering of data. Therefore, there is potential for IR techniques to benefit the training process of GenAI models. The benefits could take many potential forms. For instance, diversification techniques, which are well-studied in IR, could reduce the redundancy between training samples, thereby enabling the training procedure to be more sample-efficient. Alternatively, curricular learning methods could help gradually increase the estimated difficulty of the training samples, which may help GenAI models learn from appropriately challenging training samples as its training progresses.

4.6.1.d. Model compression, distillation, quantization

As the scale of future IR and RAG systems keeps increasing, methods for model compression, distillation, and quantization become more and more critical. Methods such as model pruning, low-rank

approximations, memory and parameter efficient finetuning are going to be essential in managing the demands of large large-scale models. Similarly, improved quantization methods can be explored in query and document representations, corpus indexing, as well as the parameter space of RAG systems.

4.6.1.e. Smaller LLMs with larger search engines

LLMs capture knowledge about the world, knowledge about language usage, and instruction-following behavior that emerges at large scale. The most reliable way to increase accuracy has been to scale models ever larger, which increases the expense and obstacles to creating and deploying large language models in varied settings. We seek to separate these components into a smaller LLM that retains the language understanding, instruction-following, and other novel components of large language models while relying on a large and efficient search engine as the source of most information. This architecture helps with hallucination problems, is easy to update, and is far less computationally complex. It also focuses research attention on what produces desirable behaviors at scale and how to reproduce those behaviors at a smaller scale.

4.6.1.f. Improving how the outputs from IR systems are integrated into LLM

When augmenting LLMs with the documents retrieved from the search engine, the current practice is passing the top k documents from the IR system to the input to LLMs. Decreasing the amount of tokens prepended to LLM will improve the efficiency. Thus, future research should be directed towards investigating methods for compressing the outputs from IR systems, and selecting the minimal subset of documents needed to enable LLMs to generate valid outputs. We should also consider how optimizing IR outputs will interact with the development of long-context LLMs. See Section 4.2 for other ideas about integrating search and LLMs.

4.6.2. Recommendations for long-term (the next ten years) research programs

4.6.2.a. New hardware paradigms