OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

References

  • Abdelkareem, Y., et al., 2022. Advances in preference-based reinforcement learning: A review, in: 2022 IEEE SMC, pp. 2527–2532.
  • Achiam, J., et al., 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 .
  • Adeniji, A., et al., 2023. Language reward modulation for pretraining reinforcement learning. arXiv preprint arXiv:2308.12270 .
  • Agia, C., et al., 2025. Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress, in: CoRL, PMLR. pp. 689–723.
  • Ahn, M., et al., 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 .
  • Ajay, A., et al., 2022. Is conditional generative modeling all you need for decision-making? arXiv preprint arXiv:2211.15657 .
  • Alakuijala, M., et al., 2023. Learning reward functions for robotic manipulation by observing humans, in: 2023 IEEE ICRA, IEEE. pp. 5006–5012.
  • Alayrac, J.B., et al., 2022. Flamingo: a visual language model for few-shot learning. NeurIPS 35, 23716–23736.
  • Authors, G., 2024. Genesis: A universal and generative physics engine for robotics and beyond.
  • Baumli, K., et al., 2023. Vision-language models as a source of rewards. arXiv preprint arXiv:2312.09187 .
  • Beck, J., et al., 2023. A survey of meta-reinforcement learning. arXiv preprint arXiv:2301.08028 .
  • Bhat, V., et al., 2024. Grounding llms for robot task planning using closed-loop state feedback. arXiv preprint arXiv:2402.08546 .
  • Bhateja, C., et al., 2023. Robotic offline rl from internet videos via value-function pre-training. arXiv preprint arXiv:2309.13041 .
  • Bommasani, R., et al., 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 .
  • Bonatti, R., et al., 2023. Pact: Perception-action causal transformer for autoregressive robotics pre-training, in: 2023 IEEE/RSJ IROS, IEEE. pp. 3621–3627.
  • Bordes, F., et al., 2024. An introduction to vision-language modeling. CoRR .
  • Brehmer, J., et al., 2024. Edgi: Equivariant diffusion for planning with embodied agents. NeurIPS 36.
  • Brohan, A., et al., 2022. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817 .
  • Brohan, A., et al., 2023. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818 .
  • Bruce, J., et al., 2024. Genie: Generative interactive environments, in: ICML.
  • Bucker, A., et al., 2023. Latte: Language trajectory transformer, in: 2023 IEEE ICRA, IEEE. pp. 7287–7294.
  • Busetto, R., et al., 2024. In-context learning of state estimators. IFAC 58, 145–150.
  • Cao, Y., et al., 2024. Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods. arXiv preprint arXiv:2404.00282 .
  • Carta, T., et al., 2023. Grounding large language models in interactive environments with online reinforcement learning, in: ICML, PMLR. pp. 3676–3713.
  • Chan, B., et al., 2024. Offline-to-online reinforcement learning for image-based grasping with scarce demonstrations, in: CoRL Workshop on Mastering Robot Manipulation in a World of Abundant Data.
  • Chebotar, Y., et al., 2023. Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions, in: CoRL, PMLR. pp. 3909–3928.
  • Chen, A.S., et al., 2021a. Learning generalizable robotic reward functions from" in-the-wild" human videos. arXiv preprint arXiv:2103.16817 .
  • Chen, B., et al., 2023a. Open-vocabulary queryable scene representations for real world planning, in: 2023 IEEE ICRA, pp. 11509–11522.
  • Chen, C., et al., 2024a. Simple hierarchical planning with diffusion. arXiv preprint arXiv:2401.02644 .
  • Chen, H., et al., 2022. Offline reinforcement learning via high-fidelity generative behavior modeling. arXiv preprint arXiv:2209.14548 .
  • Chen, H., et al., 2023b. Score regularized policy optimization through diffusion behavior. arXiv preprint arXiv:2310.07297 .
  • Chen, L., et al., 2021b. Decision transformer: Reinforcement learning via sequence modeling. NeurIPS 34, 15084–15097.
  • Chen, W., et al., 2024b. Vision-language models provide promptable representations for reinforcement learning. arXiv preprint arXiv:2402.02651 .
  • Chen, Y., et al., 2025. Fdpp: Fine-tune diffusion policy with human preference. arXiv preprint arXiv:2501.08259 .
  • Chi, C., et al., 2023. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137 .
  • Chu, K., et al., 2023. Accelerating reinforcement learning of robotic manipulations via feedback from large language models. Arxiv preprint arXiv:2311.02379 .
  • Cohen, V., et al., 2024. A survey of robotic language grounding. arXiv preprint arXiv:2405.13245 .
  • Colas, C., et al., 2020. Language as a cognitive tool to imagine goals in curiosity driven exploration. NeurIPS 33, 3761–3774.
  • Colas, C., et al., 2023. Augmenting autotelic agents with large language models, in: CLLA, PMLR. pp. 205–226.
  • Cui, Y., et al., 2022. Can foundation models perform zero-shot task specification for robot manipulation?, in: LDCC, PMLR. pp. 893–905.
  • Dalal, M., et al., 2024. Plan-seq-learn: Language model guided rl for solving long horizon robotics tasks. arXiv preprint arXiv:2405.01534 .
  • Di Palo, N., et al., 2023. Towards a unified agent with foundation models. arXiv preprint arXiv:2307.09668 .
  • Ding, Z., Jin, C., 2023. Consistency models as a rich and efficient policy class for reinforcement learning. arXiv preprint arXiv:2309.16984 .
  • Du, Y., et al., 2023a. Guiding pretraining in reinforcement learning with large language models, in: ICML, PMLR. pp. 8657–8677.
  • Du, Y., et al., 2023b. Learning universal policies via text-guided video generation. arXiv e-prints , arXiv–2302.
  • Du, Z., et al., 2023c. Can transformers learn optimal filtering for unknown systems? IEEE Control Systems Letters 7, 3525–3530.
  • Dubey, A., et al., 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 .
  • Duchoň, F., et al., 2014. Path planning with modified a star algorithm for a mobile robot. Procedia engineering 96, 59–69.
  • Escontrela, A., et al., 2024. Video prediction models as rewards for reinforcement learning. NeurIPS 36.
  • Esser, P., et al., 2024. Scaling rectified flow transformers for high-resolution image synthesis, in: ICML.
  • Firoozi, R., et al., 2023. Foundation models in robotics: Applications, challenges, and the future. arXiv preprint arXiv:2312.07843 .
  • Forgione, M., et al., 2023. From system models to class models: An in-context learning paradigm. IEEE Control Systems Letters .
  • Fu, J., et al., 2020. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219 .
  • Gu, S., et al., 2022. A review of safe reinforcement learning: Methods, theory and applications. arXiv preprint arXiv:2205.10330 .
  • Gui, Y., et al., 2024. Conformal alignment: Knowing when to trust foundation models with guarantees. arXiv preprint arXiv:2405.10301 .
  • Guo, Y., et al., 2025. Improving vision-language-action model with online reinforcement learning. arXiv preprint arXiv:2501.16664 .
  • Ha, D., Schmidhuber, J., 2018. Recurrent world models facilitate policy evolution. NeurIPS 31.
  • Hansen, N., et al., 2024. Td-mpc2: Scalable, robust world models for continuous control.
  • Hansen-Estruch, P., et al., 2023. Implicit q-learning as an actor-critic method with diffusion policies. arXiv preprint arXiv:2304.10573 .
  • Hassan, M., et al., 2024. Gem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition control. arXiv preprint arXiv:2412.11198 .
  • He, H., et al., 2024. Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning. NeurIPS 36.
  • Hegde, S., et al., 2023. Generating behaviorally diverse policies with latent diffusion models. NeurIPS 36, 7541–7554.
  • Ho, J., Salimans, T., 2021. Classifier-free diffusion guidance, in: NeurIPS Workshop on Deep Generative Models and Downstream Applications.
  • Hu, H., Sadigh, D., 2023. Language instructed reinforcement learning for human-ai coordination, in: ICML, PMLR. pp. 13584–13598.
  • Hu, J., et al., 2023a. Instructed diffuser with temporal condition guidance for offline rl. arXiv preprint arXiv:2306.04875 .
  • Hu, J., et al., 2024a. Flare: Achieving masterful and adaptive robot policies with large-scale reinforcement learning fine-tuning. CoRR abs/2409.16578.
  • Hu, S., et al., 2023b. Prompt-tuning decision transformer with preference ranking. arXiv preprint arXiv:2305.09648 .
  • Hu, S., et al., 2024b. HarmoDT: Harmony multi-task decision transformer for offline reinforcement learning, in: ICML, pp. 19182–19197.
  • Hu, Y., et al., 2023c. Toward general-purpose robots via foundation models: A survey & meta-analysis. arXiv preprint arXiv:2312.08782 .
  • Huang, S., et al., 2024a. Understanding and mitigating noise in vlm rewards. arXiv preprint arXiv:2409.15922 .
  • Huang, T., et al., 2024b. Diffusion reward: Learning rewards via conditional video diffusion. ECCV .
  • Huang, W., et al., 2023a. Grounded decoding: Guiding text generation with grounded models for embodied agents, in: NeurIPS.
  • Huang, W., et al., 2023b. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973 .
  • Huang, X., et al., 2024c. Diffuseloco: Real-time legged locomotion control with diffusion from offline datasets. arXiv preprint arXiv:2404.19264 .
  • Jain, V., Ravanbakhsh, S., 2023. Learning to reach goals via diffusion. arXiv preprint arXiv:2310.02505 .
  • Janner, M., et al., 2022. Planning with diffusion for flexible behavior synthesis, in: ICML.
  • Janson, L., et al., 2018. Deterministic sampling-based motion planning: Optimality, complexity, and performance. IJRR 37, 46–61.
  • Jülg, T., et al., 2025. Refined policy distillation: From vla generalists to rl experts. arXiv preprint arXiv:2503.05833 .
  • Kang, B., et al., 2024. Efficient diffusion policies for offline reinforcement learning. NeurIPS 36.
  • Kim, M.J., et al., 2024a. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246 .
  • Kim, S., et al., 2024b. Stitching sub-trajectories with conditional diffusion model for goal-conditioned offline rl. arXiv preprint arXiv:2402.07226 .
  • Kim, W.K., et al., 2024c. Robust policy learning via offline skill diffusion. arXiv preprint arXiv:2403.00225 .
  • Kollar, T., et al., 2017. Generalized grounding graphs: A probabilistic framework for understanding grounded commands. arXiv preprint arXiv:1712.01097 .
  • Kumar, A., et al., 2022. Pre-training for robots: Offline rl enables learning new tasks in a handful of trials. RSS .
  • Lee, K., et al., 2024a. Refining diffusion planner for reliable behavior synthesis by automatic detection of infeasible plans. NeurIPS 36.
  • Lee, O.Y., et al., 2024b. Affordance-guided reinforcement learning via visual prompting. arXiv preprint arXiv:2407.10341 .
  • Li, P.h., et al., 2024a. Online foundation model selection in robotics. arXiv preprint arXiv:2402.08570 .
  • Li, W., et al., 2023a. Hierarchical diffusion for offline decision making, in: ICML, PMLR. pp. 20035–20064.
  • Li, X., et al., 2024b. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941 .
  • Li, Z., et al., 2023b. Beyond conservatism: Diffusion policies in offline multi-agent reinforcement learning. arXiv preprint arXiv:2307.01472 .
  • Liang, Z., et al., 2023. Adaptdiffuser: Diffusion models as adaptive self-evolving planners. arXiv preprint arXiv:2302.01877 .
  • Liu, J., et al., 2023. Dipper: Diffusion-based 2d path planner applied on legged robots. arXiv preprint arXiv:2310.07842 .
  • Liu, J., et al., 2025. What can rl bring to vla generalization? an empirical study. arXiv preprint arXiv:2505.19789 .
  • Liu, S., et al., 2024. Rl-gpt: Integrating reinforcement learning and code-as-policy. arXiv preprint arXiv:2402.19299 .
  • Lu, C., et al., 2023. Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning, in: ICML, PMLR. pp. 22825–22855.
  • Lubana, E.S., et al., 2023. Fomo rewards: Can we cast foundation models as reward functions? arXiv preprint arXiv:2312.03881 .
  • Luo, J., et al., 2024. Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning. arXiv preprint arXiv:2410.21845 .
  • Ma, R., et al., 2024a. Guiding exploration in reinforcement learning with large language models. arXiv preprint arXiv:2403.09583 .
  • Ma, Y., et al., 2024b. A survey on vision-language-action models for embodied ai. arXiv preprint arXiv:2405.14093 .
  • Ma, Y.J., et al., 2022. Vip: Towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030 .
  • Ma, Y.J., et al., 2023a. Eureka: Human-level reward design via coding large language models. arXiv preprint arXiv: Arxiv-2310.12931 .
  • Ma, Y.J., et al., 2023b. Liv: Language-image representations and rewards for robotic control, in: ICML, PMLR. pp. 23301–23320.
  • Ma, Y.J., et al., 2024c. Dreureka: Language model guided sim-to-real transfer. arXiv preprint arXiv:2406.01967 .
  • Mahmoudieh, P., et al., 2022. Zero-shot reward specification via grounded natural language, in: ICML, PMLR. pp. 14743–14752.
  • Majumdar, A., et al., 2024. Where are we in the search for an artificial visual cortex for embodied intelligence? NeurIPS 36.
  • Mao, Z., et al., 2024. Zero-shot safety prediction for autonomous robots with foundation world models. arXiv preprint arXiv:2404.00462 .
  • Mark, M.S., et al., 2024. Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone. arXiv preprint arXiv:2412.06685 .
  • Mazoure, B., et al., 2023. Value function estimation using conditional diffusion models for control. arXiv preprint arXiv:2306.07290 .
  • Mazzaglia, P., et al., 2024. Multimodal foundation world models for generalist embodied agents. arXiv preprint arXiv:2406.18043 .
  • Mezghani, L., et al., 2023. Think before you act: Unified policy for interleaving language reasoning with actions. arXiv preprint arXiv:2304.11063 .
  • Mnih, V., et al., 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 .
  • Mnih, V., et al., 2015. Human-level control through deep reinforcement learning. nature 518, 529–533.
  • Nair, S., et al., 2022. Learning language-conditioned robot behavior from offline data and crowd-sourced annotation, in: CoRL, PMLR. pp. 1303–1315.
  • Nakamura, K., et al., 2025. Generalizing safety beyond collision-avoidance via latent-space reachability analysis. arXiv preprint arXiv:2502.00935 .
  • Ni, F., et al., 2023. Metadiffuser: Diffusion model as conditional planner for offline meta-rl, in: ICML, PMLR. pp. 26087–26105.
  • Nottingham, K., et al., 2023. Do embodied agents dream of pixelated sheep: Embodied decision making using language guided world modelling, in: ICML, PMLR. pp. 26311–26325.
  • Nuti, F., et al., 2024. Extracting reward functions from diffusion models. NeurIPS 36.
  • Oier, M., et al., 2024. Octo: An open-source generalist robot policy, in: Proceedings of RSS, Delft, Netherlands.
  • Padalkar, A., et al., 2023. Open X-Embodiment: Robotic learning datasets and RT-X models.
  • Pérez-Dattari, R., et al., 2024. Puma: deep metric imitation learning for stable motion primitives. Advanced Intelligent Systems 6, 2400144.
  • Permenter, F., Yuan, C., 2024. Interpreting and improving diffusion models from an optimization perspective, in: ICML, JMLR.org.
  • Psenka, M., et al., 2023. Learning a diffusion model policy from rewards via q-score matching. arXiv preprint arXiv:2312.11752 .
  • Qi, H., et al., 2025. Strengthening generative robot policies through predictive world modeling. arXiv preprint arXiv:2502.00622 .
  • Radford, A., et al., 2021. Learning transferable visual models from natural language supervision, in: ICML, Pmlr. pp. 8748–8763.
  • Ramesh, A., et al., 2021. Zero-shot text-to-image generation, in: ICML, Pmlr. pp. 8821–8831.
  • Reed, S., et al., 2022. A generalist agent. arXiv preprint arXiv:2205.06175 .
  • Ren, A.Z., et al., 2024. Diffusion policy policy optimization, in: arXiv preprint arXiv:2409.00588.
  • Rocamonde, J., et al., 2023. Vision-language models are zero-shot reward models for reinforcement learning. arXiv preprint arXiv:2310.12921 .
  • Rombach, R., et al., 2022. High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF, pp. 10684–10695.
  • Ross, S., Bagnell, D., 2010. Efficient reductions for imitation learning, in: Teh, Y.W., Titterington, M. (Eds.), ICAIS, PMLR, Chia Laguna Resort, Sardinia, Italy. pp. 661–668.
  • Rusu, A.A., et al., 2015. Policy distillation. arXiv preprint arXiv:1511.06295 .
  • Seo, Y., et al., 2023a. Masked world models for visual control, in: CoRL, PMLR. pp. 1332–1344.
  • Seo, Y., et al., 2023b. Multi-view masked world models for visual robotic manipulation, in: ICML, PMLR. pp. 30613–30632.
  • Shridhar, M., et al., 2022. Cliport: What and where pathways for robotic manipulation, in: CoRL, PMLR. pp. 894–906.
  • Song, J., et al., 2023. Self-refined large language model as automated reward function designer for deep reinforcement learning in robotics. arXiv preprint arXiv:2309.06687 .
  • Sontakke, S., et al., 2024. Roboclip: One demonstration is enough to learn robot policies. NeurIPS 36.
  • Suh, H.T., et al., 2023. Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching, in: CoRL, PMLR. pp. 2878–2904.
  • Sun, J., et al., 2024. Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning, in: 2024 IEEE ICRA, IEEE. pp. 16236–16242.
  • Sutton, R., Barto, A., 1998. Reinforcement learning: An introduction. IEEE Transactions on Neural Networks 9, 1054–1054.
  • Sutton, R.S., 1988. Learning to predict by the methods of temporal differences. Machine learning 3, 9–44.
  • Tian, R., et al., 2024. Maximizing alignment with minimal feedback: Efficiently learning rewards for visuomotor robot policy alignment. arXiv preprint arXiv:2412.04835 .
  • Trabucco, B., et al., 2022. Anymorph: Learning transferable polices by inferring agent morphology, in: ICML, PMLR. pp. 21677–21691.
  • Triantafyllidis, E., et al., 2024. Intrinsic language-guided exploration for complex long-horizon robotic manipulation tasks, in: 2024 IEEE ICRA, IEEE. pp. 7493–7500.
  • Van de Ven, G.M., et al., 2024. Continual learning and catastrophic forgetting. arXiv preprint arXiv:2403.05175 .
  • Venkatraman, S., et al., 2023. Reasoning with latent diffusion in offline reinforcement learning. arXiv preprint arXiv:2309.06599 .
  • Venuto, D., et al., 2024. Code as reward: Empowering reinforcement learning with vlms. arXiv preprint arXiv:2402.04764 .
  • Wabersich, K.P., Zeilinger, M.N., 2021. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems. Automatica 129, 109597.
  • Wang, J., et al., 2024a. Large language models for robotics: Opportunities, challenges, perspectives. arXiv preprint arXiv:2401.04334 .
  • Wang, L., et al., 2023a. Gensim: Generating robotic simulation tasks via large language models. ArXiv abs/2310.01361.
  • Wang, Y., et al., 2023b. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455 .
  • Wang, Y., et al., 2024b. Rl-vlm-f: Reinforcement learning from vision language foundation model feedback, in: ICML.
  • Wang, Z., et al., 2022. Diffusion policies as an expressive policy class for offline reinforcement learning. arXiv preprint arXiv:2208.06193 .
  • Wang, Z., et al., 2023c. Cold diffusion on the replay buffer: Learning to plan from known good states, in: CoRL, PMLR. pp. 3277–3291.
  • Wen, M., et al., 2022. Multi-agent reinforcement learning is a sequence modeling problem. NeurIPS 35, 16509–16521.
  • Wen, M., et al., 2023. Large sequence models for sequential decision-making: a survey. Frontiers of Computer Science 17, 176349.
  • Wolczyk, M., et al., 2023. On the role of forgetting in fine-tuning reinforcement learning models, in: ICLR Workshop on Reincarnating Reinforcement Learning.
  • Wołczyk, M., et al., 2024. Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem. arXiv preprint arXiv:2402.02868 .
  • Wu, J., et al., 2024. ivideogpt: Interactive videogpts are scalable world models. arXiv preprint arXiv:2405.15223 .
  • Wu, P., et al., 2023. Daydreamer: World models for physical robot learning, in: CoRL, PMLR. pp. 2226–2240.
  • Xiao, W., et al., 2023a. Safediffuser: Safe planning with diffusion probabilistic models. arXiv preprint arXiv:2306.00148 .
  • Xiao, W., et al., 2025. Safediffuser: Safe planning with diffusion probabilistic models, in: ICLR.
  • Xiao, X., et al., 2023b. Robot learning in the era of foundation models: A survey. arXiv preprint arXiv:2311.14379 .
  • Xie, T., et al., 2023. Text2reward: Reward shaping with language models for reinforcement learning, in: ICLR.
  • Xu, C., et al., 2024. Rldg: Robotic generalist policy distillation via reinforcement learning. arXiv preprint arXiv:2412.09858 .
  • Xu, M., et al., 2023. Hyper-decision transformer for efficient online policy adaptation. arXiv preprint arXiv:2304.08487 .
  • Xue, H., et al., 2024. Full-order sampling-based mpc for torque-level locomotion control via diffusion-style annealing. arXiv:2409.15610.
  • Yang, J., et al., 2024. Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning, in: 2024 IEEE ICRA, IEEE. pp. 4804–4811.
  • Yang, L., et al., 2023a. Policy representation via diffusion probability model for reinforcement learning. arXiv preprint arXiv:2305.13122 .
  • Yang, M., et al., 2023b. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114 .
  • Ye, W., et al., 2023. Towards embodied generalist agents with foundation prior assistance. arXiv preprint arXiv:2310.02635 .
  • Yu, W., et al., 2023. Language to rewards for robotic skill synthesis. Arxiv preprint arXiv:2306.08647 .
  • Zala, A., et al., 2024. Envgen: Generating and adapting environments via llms for training embodied agents, in: COLM.
  • Zeng, A., et al., 2022. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598 .
  • Zhang, E., et al., 2024a. Language control diffusion: Efficiently scaling through space, time, and tasks, in: ICLR.
  • Zhang, J., et al., 2023. Bootstrap your own skills: Learning to solve new tasks with large language model guidance. arXiv preprint arXiv:2310.10021 .
  • Zhang, J., et al., 2024b. Vision-language models for vision tasks: A survey. IEEE Transactions on PAMI .
  • Zhang, Y., et al., 2024c. Motiongpt: Finetuned llms are general-purpose motion generators, in: Proceedings of AAAI, pp. 7368–7376.
  • Zhao, W., et al., 2024. Vlmpc: Vision-language model predictive control for robotic manipulation, in: RSS.
  • Zhou, C., et al., 2022. A review of motion planning algorithms for intelligent robots. Journal of Intelligent Manufacturing 33, 387–424.
  • Zhou, G., et al., 2024a. Diffusion model predictive control. arXiv preprint arXiv:2410.05364 .
  • Zhou, H., et al., 2024b. Language-conditioned learning for robotic manipulation: A survey. arXiv preprint arXiv:2312.10807 .
  • Zhou, S., et al., 2024c. Adaptive online replanning with diffusion models. NeurIPS 36.
  • Zhu, D., et al., 2023a. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592 .
  • Zhu, Z., et al., 2023b. Diffusion models for reinforcement learning: A survey. arXiv preprint arXiv:2311.01223 .
  • Zhu, Z., et al., 2023c. Madiff: Offline multi-agent learning with diffusion models. arXiv preprint arXiv:2305.17330 .
  • Ziegler, D.M., et al., 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593 .