OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

1 Introduction

Two main pursuits of embodied intelligence and robotics are achieving physical grounding [37, 83], the capability of robots to sense, comprehend, and interact efficiently with the physical environment, and human-level reasoning on abstract concepts [37, 24, 5, 72], which remain a central goal of artificial intelligence (AI) research in general. Recently, the convergence of two powerful paradigms, generative AI and Reinforcement Learning (RL), has shown significant promise in advancing this objective [163, 94, 151, 101, 103, 184, 71]. On the one hand, tools from generative AI, including Large Language Models (LLMs), multi-modal foundation models [14] such as Vision-Language Models (VLMs), and world or video prediction models, have excelled in processing and generating diverse data types such as text, code, and images [10, 33, 128, 60]. Trained on vast multi-modal datasets, these models encapsulate rich, generalizable knowledge representations [28, 57, 123]. Moreover, diffusion models excel in generative capabilities and robust training processes for downstream tasks such as state representation and policy generation [61, 35]. RL, on the other hand, offers a framework for agents to learn optimal behaviors through interaction with their environment [139, 111].

The potential synergy between generative AI and RL is particularly compelling in the context of robotics. Generative models can serve as robust priors for information fusion in RL agents—i.e., fusing information flows from different sources to produce a robot control policy—providing extensive world knowledge, grounding language understanding, and generating rich behaviors [51, 101, 151]. Conversely, RL can “embody” generative AI models used as generative policies for robotics, usually trained through pure Imitation Learning (IL) [119], allowing them to interact with and learn from dynamic physical environments and suboptimal data [140]. This integration could lead to robotic systems with enhanced adaptability, generalization, and overall intelligence through feature alignment with experience [141], similarly to RL from human feedback training for LLMs [186].

Foundation models, in particular, have demonstrated exceptional proficiency in processing information for downstream robotics tasks [73, 134, 12]. Additionally, specialized foundation models for robotics exist [51]. Foundation models have also been adopted in the field of learning and control of dynamical systems, where robotics represents a key area of application. Transformer-based pre-trained models have been proposed [52, 46] for zero-shot prediction of outputs from a class of dynamical systems in response to any query input sequence, and for in-context state estimation [22]. These advancements highlight the potential of foundation models to introduce innovative approaches and paradigms to traditional dynamical control systems, facilitating a shift towards data-driven estimation and control synthesis for classes of dynamical systems, rather than for single, specific systems.

However, despite significant progress in foundation models, their application with RL in robotics remains underexplored. Recent developments suggest this promising union to enhance robotic learning and generalization, yet challenges persist, especially for real-world robotic applications [99].

(a) The figure illustrates the number of papers published each year integrating both generative AI and RL in robotics, categorized by the type of model employed.

(a) The figure illustrates the number of papers published each year integrating both generative AI and RL in robotics, categorized by the type of model employed.

(b) The heatmap highlights the notable rise in the use of diffusion models as tools for RL in robotics from 2023 up to February 2025 (dark purple block). Notice the recent introduction of few seminal works in RL-based fine-tuning and distillation of generative policies.

(b) The heatmap highlights the notable rise in the use of diffusion models as tools for RL in robotics from 2023 up to February 2025 (dark purple block). Notice the recent introduction of few seminal works in RL-based fine-tuning and distillation of generative policies.

(a) The figure illustrates the number of papers published each year integrating both generative AI and RL in robotics, categorized by the type of model employed.

1.1 Scope and Contribution

Despite rapid advancements in both generative AI and RL [47, 2, 111, 126], there is a noticeable lack of comprehensive studies systematically exploring their integration in robotics (see Table 1). Few research either focuses on these areas in isolation [69, 162, 51, 181], or examines narrow aspects of their combination [23, 184]. Our review addresses this gap by providing an in-depth analysis of current research trends at the intersection of generative AI and RL in robotics. Specifically, it examines Transformer- and Diffusion-based generative AI models (LLMs, VLMs, diffusion models (DMs), world models (WMs) and video prediction models (VPMs)) currently used to enhance RL, considering the diverse data modalities, roles, and applications (see Sections 3, 4, and 5). Figure 1(a) illustrates the growing adoption of generative AI models for RL, particularly diffusion models as tools for RL in 2023–2024, as shown in Figure 1(b). Additionally, particular attention is given to the unique relationship between generative policies, which are a specific and relatively narrow type of generative models for robotic action generation, and their integration with RL (see Sections 6, 7, and 8), with insights on their good fitting with learning-based control techniques for grounding into downstream robotic tasks (see Section 10). To the best of our knowledge, RL fine-tuning of generalist generative policies has also not been classified in previous surveys [69, 162, 51, 181].

Survey RFM Tools for RL RL Policies
LLM VLM DM WM/VPM Train Fine-Tune
Hu et al. [69] ✓ $\star$ $\star$ ✓
Xiao et al. [162] ✓ $\star$ $\star$
Firoozi et al. [51] ✓ $\star$ $\star$ ✓
Zhou et al. [181] ✓ $\star$ $\star$
Wang et al. [148] ✓ $\star$
Cao et al. [23] $\star$ ✓
Zhu et al. [184] $\star$ ✓
Ours $\star$ ✓ ✓ ✓ ✓ ✓ ✓

Table 1: Comparison of survey papers across robotics learning categories and methods: Foundation Models for Robotics (RFM), Generative Tools for RL (Tools for RL) and RL for Generative Policies (RL Policies). $\star$ means that the survey falls into that category but it does not broadly focus on the RL analysis, instead focusing on a specific subtopic.

The main contributions of our work are:

  1. A comprehensive review of the intersection between Transformer- and Diffusion-based generative models and RL for robotics.
  2. The first, to the best of our knowledge, dual analysis of how generative AI tools improve RL and RL improves generative policy models for robotics.
  3. The identification of best practices and challenges when using generative AI models as tools for RL.
  4. A detailed classification of RL-based training, fine-tuning and distillation methods for generalist generative policies.
  5. The identification of three new research directions integrating generative AI and RL to enhance robotics.
  6. A unified new taxonomy and continuously updated repository [^1] for tracking progress in this field.

The remainder of our paper is structured as follows: Section 2 presents our taxonomy based on the duality between generative AI transformer- and diffusion-based tools and RL. Section 3 (Base Model), Section 4 (Modality) and Section 5 (Task) investigate deeper the dimension of generative AI models used as modular tools for the RL training loop, analyzing Generative Tools for RL. On the other hand, Section 6 explores RL-Based Training of generative policies with consequent RL-Based Fine-Tuning for downstream tasks (Section 7) and Model Distillation (Section 8), looking into RL for Generative Policies. Section 9 discusses challenges. We conclude with Section 10, suggesting our perspective on possible future research directions based on our findings.

1.2 Methodology

We conducted this comprehensive review on the integration of transformer- and diffusion-based models with RL because these are two of the most rapidly emerging research areas in robotics learning. As noted in the previous section, we have observed a significant increase in the number of published papers on these topics in recent years, but especially, at their intersection. Given the exponential growth of research in generative AI for robotics, we chose to focus our review specifically on transformer- and diffusion-based models. To compile our dataset, we searched for relevant papers published between 2019, when we identified the first preliminary works integrating transformers with RL for robotics applications, and February 2025. We used keyword-based searches to filter papers aligned with the scope of our review, excluding those that did not meet our criteria, following the widely recognized PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) methodology[^2]. In the identification phase, we selected 15 relevant keywords to identify papers relevant to our review. We used combinations of terms such as RL with foundation models, LLM, VLM, diffusion models, world models, and video prediction models for filtering works on Generative Tools for RL. For RL for Generative Policies, we combined RL with transformer, diffusion, flow matching, generative policy, pre-training, fine-tuning, and distillation. Our search spanned multiple platforms, including Scopus, DBLP, IEEE Xplore, Google Scholar, and ArXiv, as well as references found in other related surveys listed in Table 1. In the screening phase, two researchers manually reviewed the abstracts of the identified papers to ensure their relevance to the robotics domain. During the selection stage, 169 papers that strictly met the inclusion criteria were chosen for analysis. To structure our taxonomy, we identified relevant categories based on two major dimensions of our study. Papers included in our review were required to be directly related to robotics and RL, published in English, and ideally peer-reviewed. However, given the fast-paced nature of this field, many relevant works have been published on preprint servers such as ArXiv and have yet to undergo formal peer review for top-tier journals or conferences.