We recognize emerging efforts in the long-term development of hardware and computational paradigms alternative to the current one based on the Von Neumann architecture. For example, quantum computing represents a groundbreaking shift in computational technology, utilizing principles of quantum mechanics to perform complex calculations at speeds unattainable by traditional computers, promising enhanced performance for computationally intensive tasks and superior energy efficiency over existing GPU systems. The reduction of energy consumption and environmental impact is essential to make generative AI and IR sustainable and scalable. Yet, it is largely unclear how current generative AI technology based on Transformers could be executed in the current experimental quantum computer. Similarly, biocomputing and neuromorphic computing promise reductions in energy consumption and high degree of parallelization. While these technologies are still underdeveloped and highly speculative, we believe the IR field should explore how techniques in the field can benefit from new and emerging computing paradigms.
4.6.2.b. Reasoning engines are coming
So far, we have largely considered efficiency from the compute and data perspectives. IR and GenAI have the potential to improve efficiency from the perspective of human efforts too. Current LLMs have basic
reasoning and instruction-following capabilities that will only improve over time. These more powerful reasoning engines will support a wider range of information-seeking and information-usage tasks that enable new modes of accessing and analyzing information to facilitate more efficient human effort. For example, information systems will help people generate hypotheses, seek support for components of a hypothesis, and investigate information provenance and reasoning chains. The results will be faster and more powerful scientific discovery. For instance, LLMs and search techniques could improve mining patterns about protein interactions to help inform new hypotheses to test about new interactions – a very laborious process when done manually. The IR community thought about these issues in its early days, but largely drifted away from them in the web era. The community will now have the tools to investigate them more fully.
4.6.2.c. Constrained resources
Computer science tends to assume that computational power and storage grow continuously over time. Research on more advanced capabilities within the same hardware footprint would provide environmental and societal benefits, as well as support future long-term human endeavors such as traveling to Mars, where the timeframe between new hardware deployments will be less rapid.
4.7. AI agents and information retrieval
We envision intelligent agents that are ubiquitous, effective, inexpensive and able to provide information, or accomplish tasks on behalf of users. Technologies like augmented reality enable multimodal real-time contextual input and provide multimodal output (voice, in-ear; vision as augmentations to images; information overlays on the scene). Physically-embodied agents will be able to accomplish physical tasks in the real world, acting as a bridge between the digital and analog world.
One core “tool” for an agentic IR system will be an LLM, which allows generation from IR outputs. If the user’s query has multiple subtopics, the agent can issue multiple queries or a single query that covers all subtopics. The agent may have access to multiple tools, in which case it is not only generating appropriate queries but also choosing where to send them. To complete a task with these tools, an agent executes a series of actions, often called a “plan”. Today these plans are often hard-coded, but autonomous planning is an essential component for future systems that interact with one or more machines, or agents. Once we have built such IR agents, we need to continue improving them. IR agents acting on behalf of people need feedback in order to adapt to individuals, their contexts, and their tasks. For example, an agent running on a device with a camera or a microphone might collect real-time implicit signals such as speech, facial expressions, gestures, or other actions the user is taking in the moment.
4.7.1. Recommendations for short-term (the next five years) research programs 4.7.1.a. Planning Truly capable agents will need to be able to formulate novel plans to answer challenging questions. Several LLM-powered methods have been proposed for generating these plans and recovering from execution failures, However, the generated plans are often incorrect and new reliable methods must be developed. We also consider the efficiency of planning. We envision proactive agents that gather additional information that may lead to even more dynamic plans. To test the planning capabilities of
agents we need benchmarks that are sufficiently challenging and require multiple actions to complete, as well as appropriate evaluation metrics that emphasize both the correctness of the final outputs and value of the intermediate steps. There is much to do in building evaluation platforms, and models should learn to use new tools as they are added as well.
4.7.1.b. Interpreting feedback for Agentic IR system
Related to the Feedback section above, moving beyond simple implicit feedback such as clicks will require more deeply understanding use behavior both for new sensors (e.g., cameras, microphones) and new presentation methods (e.g., free text, video, audio). This research can be grounded in familiar properties such as relevance and utility as well as new properties based on interaction with agents such as social norms. Second, as IR agents satisfy user needs through multi-step plans, determining the quality of a specific step traditionally requires attributing a final quality signal to a given intermediate step. In lieu of this, agents can benefit from instantaneous or partial feedback for specific steps in the plan. These partial signals can be specified by the system designer (e.g., “avoid tools that use too much time”), learned by the model (through attribution or real-time signals), or negotiated (when the tool itself is an agent). Lastly, as agentic systems become increasingly interactive, they may ask for clarification questions, or allow users—human or machine—to specify constraints or edit part of their plans. We will need to develop techniques to leverage these interventions to improve the effectiveness of future interactions.
4.7.1.c. Verification of Agentic IR system outputs
AI agents can generate texts containing hallucinations. How can agents expose or explain their behavior in a way that demonstrably enables human oversight? Future work should attempt to develop and evaluate (see 4.1.2.b and 4.1.2.c) the effectiveness of methods such as
- Attribution: What excerpts from primary sources are best surfaced to provide evidence for generative summaries? Do these passages truly entail the summary? How should the authoritativeness of the primary sources be signaled?
- Process: For some types of information gathering queries, the process (or plan) followed by the agent is itself evidence of the answer’s correctness. For example, given the question “Which system performs best on the XYZ benchmark” in the absence of a leaderboard, the only way to answer the question is for the agent to read each paper which cites the benchmark, extract the systems’ scores, and report the max. How and when should agents expose this type of reasoning?
- Critical Thinking: Naive approaches to explanation generations often tend towards persuasive text that may mislead; perhaps, directing the LLM to emphasize critical thinking may lead to explanations that actually enable verification.
- Graceful Degradation: In the process of generating a response, an agent may fail for various reasons: some APIs it tried to access were unavailable, or returned unexpected results, or the answer was simply not available anywhere in the consulted sources. In these cases, the agent should (a) try to recover, e.g., by trying alternate ways to query the API, (b) consider alternative
ways to fulfill the request, (c) if a+b fail, be able to reliably report its failure and not hallucinate an answer.
Since intuitions are often wrong, it’s important to evaluate with real people in order to measure whether people can correctly recognize when the agent is correct and when returned answers are flawed.
4.7.1.d. Long-horizon agents
Instead of instant answers or fast generation that mimic previous generations of web search, the aim of long-horizon agents is to leverage agentic capabilities that can spend additional time planning, interacting with external tools, and iterating. It naturally includes complex multi-agent coordination with asynchronous collaboration on shared information tasks involving iterative and complex research tasks. The actions for long-term agents allow a much wider range of information actions versus existing RAG agentic models that are often a small number of calls and interactions. Long-term agents enable a much broader set of strategies leveraging interaction, external systems and tools. Such strategies could entail creation or specialization of new models as part of the task.
4.7.2. Recommendations for long-term (the next ten years) research programs
4.7.2.a. Mixed-initiative reasoning
IR agents are mixed-initiative systems that gather information proactively on behalf of the user, or accomplish tasks in the background proactively, without the direct initiation by the user. While mixed initiative systems exist today, there are open questions about how to integrate multimodal results, when and how to take an action on behalf of the user, how to provide the user the ability to examine or verify an action that was taken, and how to manage the dialog with the user, given the recent language capabilities of LLMs.
4.7.2.b. Multi-agent coordination
Everyday objects (such as documents and calendars) will be imbued with intelligence to enhance the user experience. For example, a document might detect that the people collaboratively writing need a reference to cite, and it might identify the references and add the citations to the text proactively.
4.7.2.c. Modular architecture
Future intelligent systems will be composite systems of multiple smaller bespoke models, with less reliance on LLMs (see 4.6.1.e). IR agents can be composed of multiple smaller models, and any architecture (knowledge bases, APIs, tools, other agents, etc.). This will democratize research and development in this area, as universities and smaller companies will not need vast resources to develop agents. The challenge will be developing smaller agents that communicate with each other, that produce transparent, explainable and trusted results. Evaluating modular systems presents a challenge as errors upstream compound downstream, and in modular collections of agents it may be difficult to find the source of mistakes.
4.7.2.d. Evolving capabilities
IR agents will need the ability to improve their effectiveness through additional interaction. They should be able to generate code that designs new IR models as well as new data processing techniques. They should be able to self-refine and continually improve through internal and external feedback.
4.8. Foundation models for information access and discovery
The recent technological advances in generative AI are poised to extend human capabilities further by changing how people are able to rapidly capture knowledge digitally, communicate intent with an AI-based device, and rely on the AI-device to perform numerical and semantic computation. We believe these new capabilities of generative AI as well as anticipated ongoing innovations will spur a new range of innovations where information access and generative AI techniques primarily serve as a way of assisting and augmenting the intelligence and abilities of every person – something we envision in an embodied form as a knowledge extender.
Inspired by the great success of general foundation LLMs as a general way to augment Intelligence and extend human capabilities, we envision the possibility of developing foundation models for supporting all kinds of information access applications. What do we mean precisely by a foundation model? A foundation model is meant to support many different kinds of applications (on top of it), thus generality is the first requirement. It is meant to be "task-agnostic" where these methods, tools, or systems can be applied to a wide range of tasks without being specifically designed for any single task.
4.8.1. Recommendations for short-term (the next five years) research programs
4.8.1.a. Efficient, robust, and scalable generative retrieval and re-ranking
Large transformer networks are shown to be able to produce a ranking of identifiers associated with information items in a given corpus. Large-scale efficiency and reliability of such methods, sometimes called differentiable search index or generative retrieval, are yet to be addressed or improved. Research in this area should focus on better optimization methods, effective identifier assignment, and end-to-end training as well as generalization to unseen information items.
4.8.1.b. Instruction-following foundation models for information access
We feel that in the domain of information access, the main inputs to the model should include: (1) information about a user (including the user’s profile, specification, historical behavior, user intent), (2) information sources and collections (including APIs for information that can be accessed, document collections, the Web, (3) the information access functionality (i.e., information access instruction), which defines the objectives and a space of actions, and (4) the context (such as the state of the outer world and the environment). The output should be a sequence of system actions and predicted user decisions, or behavior in general that best match to provide information for the particular user with the proposed content or decisions that are characterized by the context. Ideally, the system should also provide justifications and explanations of the reasoning behind the output it generates. This helps establish user trust and is a step toward making these systems more steerable. We also highlight the importance of robustness in this space, meaning that foundation models should be robust to various reformulations or paraphrasing of the same information access request.
4.8.1.c. Generating a long sequence of actions for complex information seeking
Instead of only targeting labels as output, the focus should be on multi-step chains of decisions and behaviors – where progress on this challenge will be able to simulate a user’s behavior over longer future sequences. Building high-quality pre-training data through algorithmic synthetic data generation to make model training more efficient and effective is also an important direction to consider. In addition, the key to learning a foundational information access model will be to rely on techniques of learning by self-supervision, instruction, and demonstration.
4.8.1.d. Complex multimodal information access
To enable effective customization of such models, we also need to address challenges such as integrating multimodal input/output, determining how to best incorporate and represent a user’s past behavior and decisions in the foundational layer, incorporating audio input, generating video on the fly, and providing accessible outputs customized for user’s preference or current device type.
4.8.1.e. User simulation for training foundation models
Training these models requires diverse data sources (especially user behaviors in diverse contexts), significant computing power, and the ability to learn from small demonstrations. One could anticipate that the new foundation model will start with the behavior data from a mixture of limited real users and a large number of simulated users, and simulating user behaviors through user simulation agents is essential. Sections 4.3 and 4.5 present other ideas for how to simulate users with digital twins.
4.8.2. Recommendations for long-term (the next ten years) research programs
4.8.2.a. A real-time learner knowledge extender
Challenges will evolve to building foundational task models and user models that reflect the diversity of behaviors and objectives and that support real-time interactions in the real world. We will be able to use these models to augment human abilities – e.g. instantaneous recall of similar past patients while collaboratively listening to a patient intake, augmented reality projection and correction of physical tasks, suggestion and synthesis of information relevant to a person’s current context. In a ten-year vision, the information access foundation model is anticipated to continuously learn, self-improve, mitigate and reduce data biases, and correct inherited biases from input backbone models.