OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

8 Policy Distillation

Policy distillation is a well-known concept in the RL literature [131]. Recently, with the emergence of large generalist generative policies, RL-based methods for policy distillation have been applied to VLA models. In particular, recent works have explored policy distillation in OpenVLA [80] and Octo [118]. The goal of policy distillation is to transfer pre-trained knowledge from a teacher policy to a student policy. The recent literature can be categorized into two opposing approaches:

  • From Generalist to Expert: Jülg et al. [78] developed a method to create task-specific RL agents distilled from a pre-trained VLA model. While VLA models are known for their strong generalization capabilities, they often struggle to achieve highly precise results compared to task-specific RL policies. The key idea is that the internal knowledge of VLA models can be useful in guiding RL exploration for specific agents. However, this work primarily presents preliminary simulation results, emphasizing that sim-to-real transfer remains the main challenge when training RL policies in simulation. This is particularly relevant given that OpenVLA and Octo typically perform better in real-world conditions [119, 89].
  • From Experts to Generalist: Xu et al. [164] propose a method to improve OpenVLA and Octo using demonstrations from expert RL policies. Their approach focuses on generating high-quality training data to fine-tune VLA models, as these models performance is highly dependent on the quality of their training data. Their experiments demonstrate that this method outperforms models trained solely on human demonstrations, achieving superior results in real-world precision manipulation tasks.