OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

7 RL Fine-Tuning

Generative models, such as transformer-based and diffusion-based policies, are often trained with IL and they demand flexible RL fine-tuning strategies, if expert data for supervised fine-tuning (SFT) are limited. Recent works effectively integrate these expressive models into RL pipelines without instability or loss of efficiency [66, 107]. In this section, we explore two major classes of generative policies where we identified RL fine-tuning is particularly relevant, (i) large transformer policies (mostly VLA models), and (ii) diffusion policies (usually smaller in size).

Recent advances in Transformer-based policies have transformed robotic control, with robotics foundation models excelling through large-scale pre-training. A major step toward generative policy architectures and datasets for robotics is the OpenX Embodiment project [119], which introduced the largest open-source real-robot dataset—the Open X-Embodiment Dataset. This dataset supports the development of pivotal VLA models like RT-X models [18, 19], Octo [118], and OpenVLA [80]. However, state-of-the-art generalist VLA policies are typically trained via IL and often require additional training to adapt to out-of-distribution scenarios [89, 99, 162, 69]. RL has become an alternative to improve policy performance, particularly when it comes to managing distribution shifts or leading generalist models toward specialization. This has motivated the development of new RL-based fine-tuning algorithms, particularly actor-critic methods [56, 93]. Actor-Critic approaches enable very reliable and effective policy updates by utilizing two networks: the critic to assess actions and the actor to decide on them. This method aids in RL fine-tuning by continuously enhancing action selections in response to critic feedback [139]. Through the use of on-policy algorithms such as PPO, cautious learning rates, and distinct actor-critic networks, the FLaRe framework [66] demonstrates RL-based fine-tuning by matching pre-trained transformer policies with new experience. FLaRe reduces the sim-to-real gap by using domain randomization and extensive simulations, increasing success rates by roughly +30 % on both real robots and in simulation. PPO is the recommended RL technique for VLA fine-tuning, according to Liu et al. [93]. Its advantages include shared actor-critic backbones, short training epochs, and great generalization on unseen objects and dynamic situations. However, the low-frequency action generation (reported in panel a in Figure 5(a)) of VLA models and the requirement of large amounts of interaction data even for minor adjustments, make online fine-tuning difficult and limited in rel world settings. Standard fine-tuning methods still need further exploration in case of VLA models, and current research focuses on large-scale RL fine-tuning for transformer policies [66].

An alternative area where RL fine-tuning shows potential is with Diffusion-based policies, which offer many possibilities with RL due to their ability to predict smooth trajectories non-autoregressively at each inference step [152, 34]. However, key problems in applying RL to diffusion policies are include the high-dimensional nature of the action space, the slow iterative denoising process involved in action generation, and training instability. Standard policy gradient methods have been viewed as inefficient for such models, due to the increased number of steps over which rewards must be predicted, introduced by iterative denoising [168]. Recent advances, such as Diffusion Policy Policy Optimization (DPPO), have demonstrated that RL can be effectively integrated with diffusion policies by formulating the denoising process as a Markov Decision Process [127]. This approach enables policy gradients to propagate through the diffusion steps, leveraging the structured noise removal process to facilitate more stable training. A major advantage of RL-fine-tuned diffusion policies is their ability to engage in on-manifold exploration, meaning that the policy remains close to the expert data distribution while still improving performance through RL. This structured exploration contrasts with traditional RL methods, which often struggle with off-manifold exploration, leading to unstable training and suboptimal policies [121]. Panel b of Figure 5(a) shows the working principle of diffusion policies, while panel c depicts the state-of-the-art methods for RL-based fine-tuning generative policies.

In general, most research focuses on transformer-based or diffusion-based policy architectures separately, as they align with different goals and fine-tuning methods, although Actor-Critic RL is commonly used. While some researchers are exploring these pathways, others aim to develop a generalized RL-based fine-tuning framework for any generative policy, regardless of size and backbone model [107].

(a) This figure illustrates the transformer policy (panel a shows VLA inference time) and the diffusion policy (panel b) principles. In panel c we show recent RL methods that can be used to fine-tune a base generative policy.

(a) This figure illustrates the transformer policy (panel a shows VLA inference time) and the diffusion policy (panel b) principles. In panel c we show recent RL methods that can be used to fine-tune a base generative policy.

(b) This figure illustrates the challenge of deploying a black-box generative policy in real-world dynamic environments without access to model-based information.

(b) This figure illustrates the challenge of deploying a black-box generative policy in real-world dynamic environments without access to model-based information.

(a) This figure illustrates the transformer policy (panel a shows VLA inference time) and the diffusion policy (panel b) principles. In panel c we show recent RL methods that can be used to fine-tune a base generative policy.