The Duality of Generative AI and Reinforcement Learning in Robotics: A Review
Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…
Licence
OPEN
CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Adeniji, A., et al., 2023. Language reward modulation for pretraining reinforcement learning. arXiv preprint arXiv:2308.12270 .
Agia, C., et al., 2025. Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress, in: CoRL, PMLR. pp. 689–723.
Ahn, M., et al., 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 .
Ajay, A., et al., 2022. Is conditional generative modeling all you need for decision-making? arXiv preprint arXiv:2211.15657 .
Alakuijala, M., et al., 2023. Learning reward functions for robotic manipulation by observing humans, in: 2023 IEEE ICRA, IEEE. pp. 5006–5012.
Alayrac, J.B., et al., 2022. Flamingo: a visual language model for few-shot learning. NeurIPS 35, 23716–23736.
Authors, G., 2024. Genesis: A universal and generative physics engine for robotics and beyond.
Baumli, K., et al., 2023. Vision-language models as a source of rewards. arXiv preprint arXiv:2312.09187 .
Beck, J., et al., 2023. A survey of meta-reinforcement learning. arXiv preprint arXiv:2301.08028 .
Bhat, V., et al., 2024. Grounding llms for robot task planning using closed-loop state feedback. arXiv preprint arXiv:2402.08546 .
Bhateja, C., et al., 2023. Robotic offline rl from internet videos via value-function pre-training. arXiv preprint arXiv:2309.13041 .
Bommasani, R., et al., 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 .
Bonatti, R., et al., 2023. Pact: Perception-action causal transformer for autoregressive robotics pre-training, in: 2023 IEEE/RSJ IROS, IEEE. pp. 3621–3627.
Bordes, F., et al., 2024. An introduction to vision-language modeling. CoRR .
Brehmer, J., et al., 2024. Edgi: Equivariant diffusion for planning with embodied agents. NeurIPS 36.
Brohan, A., et al., 2022. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817 .
Brohan, A., et al., 2023. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818 .
Bucker, A., et al., 2023. Latte: Language trajectory transformer, in: 2023 IEEE ICRA, IEEE. pp. 7287–7294.
Busetto, R., et al., 2024. In-context learning of state estimators. IFAC 58, 145–150.
Cao, Y., et al., 2024. Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods. arXiv preprint arXiv:2404.00282 .
Carta, T., et al., 2023. Grounding large language models in interactive environments with online reinforcement learning, in: ICML, PMLR. pp. 3676–3713.
Chan, B., et al., 2024. Offline-to-online reinforcement learning for image-based grasping with scarce demonstrations, in: CoRL Workshop on Mastering Robot Manipulation in a World of Abundant Data.
Chebotar, Y., et al., 2023. Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions, in: CoRL, PMLR. pp. 3909–3928.
Chen, A.S., et al., 2021a. Learning generalizable robotic reward functions from" in-the-wild" human videos. arXiv preprint arXiv:2103.16817 .
Chen, B., et al., 2023a. Open-vocabulary queryable scene representations for real world planning, in: 2023 IEEE ICRA, pp. 11509–11522.
Chen, C., et al., 2024a. Simple hierarchical planning with diffusion. arXiv preprint arXiv:2401.02644 .
Chen, H., et al., 2022. Offline reinforcement learning via high-fidelity generative behavior modeling. arXiv preprint arXiv:2209.14548 .
Chen, H., et al., 2023b. Score regularized policy optimization through diffusion behavior. arXiv preprint arXiv:2310.07297 .
Chen, L., et al., 2021b. Decision transformer: Reinforcement learning via sequence modeling. NeurIPS 34, 15084–15097.
Chen, W., et al., 2024b. Vision-language models provide promptable representations for reinforcement learning. arXiv preprint arXiv:2402.02651 .
Chen, Y., et al., 2025. Fdpp: Fine-tune diffusion policy with human preference. arXiv preprint arXiv:2501.08259 .
Chi, C., et al., 2023. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137 .
Chu, K., et al., 2023. Accelerating reinforcement learning of robotic manipulations via feedback from large language models. Arxiv preprint arXiv:2311.02379 .
Cohen, V., et al., 2024. A survey of robotic language grounding. arXiv preprint arXiv:2405.13245 .
Colas, C., et al., 2020. Language as a cognitive tool to imagine goals in curiosity driven exploration. NeurIPS 33, 3761–3774.
Colas, C., et al., 2023. Augmenting autotelic agents with large language models, in: CLLA, PMLR. pp. 205–226.
Cui, Y., et al., 2022. Can foundation models perform zero-shot task specification for robot manipulation?, in: LDCC, PMLR. pp. 893–905.
Dalal, M., et al., 2024. Plan-seq-learn: Language model guided rl for solving long horizon robotics tasks. arXiv preprint arXiv:2405.01534 .
Di Palo, N., et al., 2023. Towards a unified agent with foundation models. arXiv preprint arXiv:2307.09668 .
Ding, Z., Jin, C., 2023. Consistency models as a rich and efficient policy class for reinforcement learning. arXiv preprint arXiv:2309.16984 .
Du, Y., et al., 2023a. Guiding pretraining in reinforcement learning with large language models, in: ICML, PMLR. pp. 8657–8677.
Du, Y., et al., 2023b. Learning universal policies via text-guided video generation. arXiv e-prints , arXiv–2302.
Du, Z., et al., 2023c. Can transformers learn optimal filtering for unknown systems? IEEE Control Systems Letters 7, 3525–3530.
Dubey, A., et al., 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 .
Duchoň, F., et al., 2014. Path planning with modified a star algorithm for a mobile robot. Procedia engineering 96, 59–69.
Escontrela, A., et al., 2024. Video prediction models as rewards for reinforcement learning. NeurIPS 36.
Esser, P., et al., 2024. Scaling rectified flow transformers for high-resolution image synthesis, in: ICML.
Firoozi, R., et al., 2023. Foundation models in robotics: Applications, challenges, and the future. arXiv preprint arXiv:2312.07843 .
Forgione, M., et al., 2023. From system models to class models: An in-context learning paradigm. IEEE Control Systems Letters .
Fu, J., et al., 2020. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219 .
Gu, S., et al., 2022. A review of safe reinforcement learning: Methods, theory and applications. arXiv preprint arXiv:2205.10330 .
Gui, Y., et al., 2024. Conformal alignment: Knowing when to trust foundation models with guarantees. arXiv preprint arXiv:2405.10301 .
Guo, Y., et al., 2025. Improving vision-language-action model with online reinforcement learning. arXiv preprint arXiv:2501.16664 .
Ha, D., Schmidhuber, J., 2018. Recurrent world models facilitate policy evolution. NeurIPS 31.
Hansen, N., et al., 2024. Td-mpc2: Scalable, robust world models for continuous control.
Hansen-Estruch, P., et al., 2023. Implicit q-learning as an actor-critic method with diffusion policies. arXiv preprint arXiv:2304.10573 .
Hassan, M., et al., 2024. Gem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition control. arXiv preprint arXiv:2412.11198 .
He, H., et al., 2024. Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning. NeurIPS 36.
Hegde, S., et al., 2023. Generating behaviorally diverse policies with latent diffusion models. NeurIPS 36, 7541–7554.
Ho, J., Salimans, T., 2021. Classifier-free diffusion guidance, in: NeurIPS Workshop on Deep Generative Models and Downstream Applications.
Hu, H., Sadigh, D., 2023. Language instructed reinforcement learning for human-ai coordination, in: ICML, PMLR. pp. 13584–13598.
Hu, J., et al., 2023a. Instructed diffuser with temporal condition guidance for offline rl. arXiv preprint arXiv:2306.04875 .
Hu, J., et al., 2024a. Flare: Achieving masterful and adaptive robot policies with large-scale reinforcement learning fine-tuning. CoRR abs/2409.16578.
Hu, S., et al., 2023b. Prompt-tuning decision transformer with preference ranking. arXiv preprint arXiv:2305.09648 .
Hu, S., et al., 2024b. HarmoDT: Harmony multi-task decision transformer for offline reinforcement learning, in: ICML, pp. 19182–19197.
Hu, Y., et al., 2023c. Toward general-purpose robots via foundation models: A survey & meta-analysis. arXiv preprint arXiv:2312.08782 .
Huang, S., et al., 2024a. Understanding and mitigating noise in vlm rewards. arXiv preprint arXiv:2409.15922 .
Huang, T., et al., 2024b. Diffusion reward: Learning rewards via conditional video diffusion. ECCV .
Huang, W., et al., 2023a. Grounded decoding: Guiding text generation with grounded models for embodied agents, in: NeurIPS.
Huang, W., et al., 2023b. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973 .
Huang, X., et al., 2024c. Diffuseloco: Real-time legged locomotion control with diffusion from offline datasets. arXiv preprint arXiv:2404.19264 .
Jain, V., Ravanbakhsh, S., 2023. Learning to reach goals via diffusion. arXiv preprint arXiv:2310.02505 .
Janner, M., et al., 2022. Planning with diffusion for flexible behavior synthesis, in: ICML.
Janson, L., et al., 2018. Deterministic sampling-based motion planning: Optimality, complexity, and performance. IJRR 37, 46–61.
Jülg, T., et al., 2025. Refined policy distillation: From vla generalists to rl experts. arXiv preprint arXiv:2503.05833 .
Kang, B., et al., 2024. Efficient diffusion policies for offline reinforcement learning. NeurIPS 36.
Kim, M.J., et al., 2024a. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246 .
Kim, S., et al., 2024b. Stitching sub-trajectories with conditional diffusion model for goal-conditioned offline rl. arXiv preprint arXiv:2402.07226 .
Kim, W.K., et al., 2024c. Robust policy learning via offline skill diffusion. arXiv preprint arXiv:2403.00225 .
Kollar, T., et al., 2017. Generalized grounding graphs: A probabilistic framework for understanding grounded commands. arXiv preprint arXiv:1712.01097 .
Kumar, A., et al., 2022. Pre-training for robots: Offline rl enables learning new tasks in a handful of trials. RSS .
Lee, K., et al., 2024a. Refining diffusion planner for reliable behavior synthesis by automatic detection of infeasible plans. NeurIPS 36.
Lee, O.Y., et al., 2024b. Affordance-guided reinforcement learning via visual prompting. arXiv preprint arXiv:2407.10341 .
Li, P.h., et al., 2024a. Online foundation model selection in robotics. arXiv preprint arXiv:2402.08570 .
Li, W., et al., 2023a. Hierarchical diffusion for offline decision making, in: ICML, PMLR. pp. 20035–20064.
Li, X., et al., 2024b. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941 .
Li, Z., et al., 2023b. Beyond conservatism: Diffusion policies in offline multi-agent reinforcement learning. arXiv preprint arXiv:2307.01472 .
Liang, Z., et al., 2023. Adaptdiffuser: Diffusion models as adaptive self-evolving planners. arXiv preprint arXiv:2302.01877 .
Liu, J., et al., 2023. Dipper: Diffusion-based 2d path planner applied on legged robots. arXiv preprint arXiv:2310.07842 .
Liu, J., et al., 2025. What can rl bring to vla generalization? an empirical study. arXiv preprint arXiv:2505.19789 .
Liu, S., et al., 2024. Rl-gpt: Integrating reinforcement learning and code-as-policy. arXiv preprint arXiv:2402.19299 .
Lu, C., et al., 2023. Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning, in: ICML, PMLR. pp. 22825–22855.
Lubana, E.S., et al., 2023. Fomo rewards: Can we cast foundation models as reward functions? arXiv preprint arXiv:2312.03881 .
Luo, J., et al., 2024. Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning. arXiv preprint arXiv:2410.21845 .
Ma, R., et al., 2024a. Guiding exploration in reinforcement learning with large language models. arXiv preprint arXiv:2403.09583 .
Ma, Y., et al., 2024b. A survey on vision-language-action models for embodied ai. arXiv preprint arXiv:2405.14093 .
Ma, Y.J., et al., 2022. Vip: Towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030 .
Ma, Y.J., et al., 2023a. Eureka: Human-level reward design via coding large language models. arXiv preprint arXiv: Arxiv-2310.12931 .
Ma, Y.J., et al., 2023b. Liv: Language-image representations and rewards for robotic control, in: ICML, PMLR. pp. 23301–23320.
Ma, Y.J., et al., 2024c. Dreureka: Language model guided sim-to-real transfer. arXiv preprint arXiv:2406.01967 .
Mahmoudieh, P., et al., 2022. Zero-shot reward specification via grounded natural language, in: ICML, PMLR. pp. 14743–14752.
Majumdar, A., et al., 2024. Where are we in the search for an artificial visual cortex for embodied intelligence? NeurIPS 36.
Mao, Z., et al., 2024. Zero-shot safety prediction for autonomous robots with foundation world models. arXiv preprint arXiv:2404.00462 .
Mark, M.S., et al., 2024. Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone. arXiv preprint arXiv:2412.06685 .
Mazoure, B., et al., 2023. Value function estimation using conditional diffusion models for control. arXiv preprint arXiv:2306.07290 .
Mazzaglia, P., et al., 2024. Multimodal foundation world models for generalist embodied agents. arXiv preprint arXiv:2406.18043 .
Mezghani, L., et al., 2023. Think before you act: Unified policy for interleaving language reasoning with actions. arXiv preprint arXiv:2304.11063 .
Mnih, V., et al., 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 .
Mnih, V., et al., 2015. Human-level control through deep reinforcement learning. nature 518, 529–533.
Nair, S., et al., 2022. Learning language-conditioned robot behavior from offline data and crowd-sourced annotation, in: CoRL, PMLR. pp. 1303–1315.
Nakamura, K., et al., 2025. Generalizing safety beyond collision-avoidance via latent-space reachability analysis. arXiv preprint arXiv:2502.00935 .
Ni, F., et al., 2023. Metadiffuser: Diffusion model as conditional planner for offline meta-rl, in: ICML, PMLR. pp. 26087–26105.
Nottingham, K., et al., 2023. Do embodied agents dream of pixelated sheep: Embodied decision making using language guided world modelling, in: ICML, PMLR. pp. 26311–26325.
Nuti, F., et al., 2024. Extracting reward functions from diffusion models. NeurIPS 36.
Oier, M., et al., 2024. Octo: An open-source generalist robot policy, in: Proceedings of RSS, Delft, Netherlands.
Padalkar, A., et al., 2023. Open X-Embodiment: Robotic learning datasets and RT-X models.
Pérez-Dattari, R., et al., 2024. Puma: deep metric imitation learning for stable motion primitives. Advanced Intelligent Systems 6, 2400144.
Permenter, F., Yuan, C., 2024. Interpreting and improving diffusion models from an optimization perspective, in: ICML, JMLR.org.
Psenka, M., et al., 2023. Learning a diffusion model policy from rewards via q-score matching. arXiv preprint arXiv:2312.11752 .
Qi, H., et al., 2025. Strengthening generative robot policies through predictive world modeling. arXiv preprint arXiv:2502.00622 .
Radford, A., et al., 2021. Learning transferable visual models from natural language supervision, in: ICML, Pmlr. pp. 8748–8763.
Ramesh, A., et al., 2021. Zero-shot text-to-image generation, in: ICML, Pmlr. pp. 8821–8831.
Reed, S., et al., 2022. A generalist agent. arXiv preprint arXiv:2205.06175 .
Seo, Y., et al., 2023a. Masked world models for visual control, in: CoRL, PMLR. pp. 1332–1344.
Seo, Y., et al., 2023b. Multi-view masked world models for visual robotic manipulation, in: ICML, PMLR. pp. 30613–30632.
Shridhar, M., et al., 2022. Cliport: What and where pathways for robotic manipulation, in: CoRL, PMLR. pp. 894–906.
Song, J., et al., 2023. Self-refined large language model as automated reward function designer for deep reinforcement learning in robotics. arXiv preprint arXiv:2309.06687 .
Sontakke, S., et al., 2024. Roboclip: One demonstration is enough to learn robot policies. NeurIPS 36.
Suh, H.T., et al., 2023. Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching, in: CoRL, PMLR. pp. 2878–2904.
Sun, J., et al., 2024. Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning, in: 2024 IEEE ICRA, IEEE. pp. 16236–16242.
Sutton, R., Barto, A., 1998. Reinforcement learning: An introduction. IEEE Transactions on Neural Networks 9, 1054–1054.
Sutton, R.S., 1988. Learning to predict by the methods of temporal differences. Machine learning 3, 9–44.
Tian, R., et al., 2024. Maximizing alignment with minimal feedback: Efficiently learning rewards for visuomotor robot policy alignment. arXiv preprint arXiv:2412.04835 .
Trabucco, B., et al., 2022. Anymorph: Learning transferable polices by inferring agent morphology, in: ICML, PMLR. pp. 21677–21691.
Triantafyllidis, E., et al., 2024. Intrinsic language-guided exploration for complex long-horizon robotic manipulation tasks, in: 2024 IEEE ICRA, IEEE. pp. 7493–7500.
Van de Ven, G.M., et al., 2024. Continual learning and catastrophic forgetting. arXiv preprint arXiv:2403.05175 .
Venkatraman, S., et al., 2023. Reasoning with latent diffusion in offline reinforcement learning. arXiv preprint arXiv:2309.06599 .
Venuto, D., et al., 2024. Code as reward: Empowering reinforcement learning with vlms. arXiv preprint arXiv:2402.04764 .
Wabersich, K.P., Zeilinger, M.N., 2021. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems. Automatica 129, 109597.
Wang, J., et al., 2024a. Large language models for robotics: Opportunities, challenges, perspectives. arXiv preprint arXiv:2401.04334 .
Wang, L., et al., 2023a. Gensim: Generating robotic simulation tasks via large language models. ArXiv abs/2310.01361.
Wang, Y., et al., 2023b. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455 .
Wang, Y., et al., 2024b. Rl-vlm-f: Reinforcement learning from vision language foundation model feedback, in: ICML.
Wang, Z., et al., 2022. Diffusion policies as an expressive policy class for offline reinforcement learning. arXiv preprint arXiv:2208.06193 .
Wang, Z., et al., 2023c. Cold diffusion on the replay buffer: Learning to plan from known good states, in: CoRL, PMLR. pp. 3277–3291.
Wen, M., et al., 2022. Multi-agent reinforcement learning is a sequence modeling problem. NeurIPS 35, 16509–16521.
Wen, M., et al., 2023. Large sequence models for sequential decision-making: a survey. Frontiers of Computer Science 17, 176349.
Wolczyk, M., et al., 2023. On the role of forgetting in fine-tuning reinforcement learning models, in: ICLR Workshop on Reincarnating Reinforcement Learning.
Wołczyk, M., et al., 2024. Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem. arXiv preprint arXiv:2402.02868 .
Wu, J., et al., 2024. ivideogpt: Interactive videogpts are scalable world models. arXiv preprint arXiv:2405.15223 .
Wu, P., et al., 2023. Daydreamer: World models for physical robot learning, in: CoRL, PMLR. pp. 2226–2240.
Xiao, W., et al., 2023a. Safediffuser: Safe planning with diffusion probabilistic models. arXiv preprint arXiv:2306.00148 .
Xiao, W., et al., 2025. Safediffuser: Safe planning with diffusion probabilistic models, in: ICLR.
Xiao, X., et al., 2023b. Robot learning in the era of foundation models: A survey. arXiv preprint arXiv:2311.14379 .
Xie, T., et al., 2023. Text2reward: Reward shaping with language models for reinforcement learning, in: ICLR.
Xu, C., et al., 2024. Rldg: Robotic generalist policy distillation via reinforcement learning. arXiv preprint arXiv:2412.09858 .
Xu, M., et al., 2023. Hyper-decision transformer for efficient online policy adaptation. arXiv preprint arXiv:2304.08487 .
Xue, H., et al., 2024. Full-order sampling-based mpc for torque-level locomotion control via diffusion-style annealing. arXiv:2409.15610.
Yang, J., et al., 2024. Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning, in: 2024 IEEE ICRA, IEEE. pp. 4804–4811.
Yang, L., et al., 2023a. Policy representation via diffusion probability model for reinforcement learning. arXiv preprint arXiv:2305.13122 .
Yang, M., et al., 2023b. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114 .
Ye, W., et al., 2023. Towards embodied generalist agents with foundation prior assistance. arXiv preprint arXiv:2310.02635 .
Yu, W., et al., 2023. Language to rewards for robotic skill synthesis. Arxiv preprint arXiv:2306.08647 .
Zala, A., et al., 2024. Envgen: Generating and adapting environments via llms for training embodied agents, in: COLM.
Zeng, A., et al., 2022. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598 .
Zhang, E., et al., 2024a. Language control diffusion: Efficiently scaling through space, time, and tasks, in: ICLR.
Zhang, J., et al., 2023. Bootstrap your own skills: Learning to solve new tasks with large language model guidance. arXiv preprint arXiv:2310.10021 .
Zhang, J., et al., 2024b. Vision-language models for vision tasks: A survey. IEEE Transactions on PAMI .
Zhang, Y., et al., 2024c. Motiongpt: Finetuned llms are general-purpose motion generators, in: Proceedings of AAAI, pp. 7368–7376.
Zhao, W., et al., 2024. Vlmpc: Vision-language model predictive control for robotic manipulation, in: RSS.
Zhou, C., et al., 2022. A review of motion planning algorithms for intelligent robots. Journal of Intelligent Manufacturing 33, 387–424.
Zhou, G., et al., 2024a. Diffusion model predictive control. arXiv preprint arXiv:2410.05364 .
Zhou, H., et al., 2024b. Language-conditioned learning for robotic manipulation: A survey. arXiv preprint arXiv:2312.10807 .
Zhou, S., et al., 2024c. Adaptive online replanning with diffusion models. NeurIPS 36.
Zhu, D., et al., 2023a. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592 .
Zhu, Z., et al., 2023b. Diffusion models for reinforcement learning: A survey. arXiv preprint arXiv:2311.01223 .
Zhu, Z., et al., 2023c. Madiff: Offline multi-agent learning with diffusion models. arXiv preprint arXiv:2305.17330 .
Ziegler, D.M., et al., 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593 .
Metadata record
One description, two standard projections
Built from what the sources declared and what the gates observed.
Nothing absent has been filled in here.
1 value read out of the text by the enrichment rules and 98 links to or from other resources — a lab's files, the pages it links, the works it cites stand under their elements, marked inferred, and are kept apart in the exports.
DCMI Metadata Terms. Dublin Core has no element that separates the original file from the text extracted out of it, and none for LOM's educational characterisation. Both survive here as provenance statements and in the record itself, not in the projection.
the standard ↗
The groups below are this library's, for reading. DCMI Terms itself has no categories; each term keeps its standard name.
Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality between generative AI and RL for robotics downstream tasks. Specifically, we investigate: (1) The role of prominent generative AI tools as modular priors for multi-modal input fusion in RL tasks. (2) How RL can train, fine-tune and distill generative models for policy generation, such as VLA models, similarly to RL applications in large language models. We then propose a new taxonomy based on a considerable amount of selected papers. Lastly, we identify open challenges accounting for model scalability, adaptation and grounding, giving recommendations and insights on future research directions. We reflect on which generative AI models best fit the RL tasks and why. On the other side, we reflect on important issues inherent to RL-enhanced generative policies, such as safety concerns and failure modes, and what are the limitations of current methods. A curated collection of relevant research papers is maintained on our GitHub repository, serving as a resource for ongoing research and development in this field: https://github.com/clmoro/Robotics-RL-FMs-Integration.
The abstract, and the depositor's additional notes after it when the source has a field for them.
Where the work was published, in the source's own words: a journal with its volume and pages, a conference, an imprint.
Relations and custody
1/9
Is part ofdcterms:isPartOf
none found — the source did not declare it
The repository, book or record this resource was found inside.
Has partdcterms:hasPart
none found — the source did not declare it
What this resource is made of, when the source lists its parts.
Is version ofdcterms:isVersionOf
none found — the source did not declare it
Has versiondcterms:hasVersion
none found — the source did not declare it
Referencesdcterms:references
inferred
read its reference list, line 447citation
arXiv:2303.08774
arXiv:2308.12270
arXiv:2204.01691
arXiv:2211.15657
arXiv:2312.09187
arXiv:2301.08028
arXiv:2402.08546
arXiv:2309.13041
arXiv:2108.07258
arXiv:2212.06817
arXiv:2307.15818
arXiv:2404.00282
and 86 more, every one of them in the export
Works this one cites, when the source declares them as relations. What its text links to and its reference list cites is inferred, and stands under it apart.
Is referenced bydcterms:isReferencedBy
none found — the source did not declare it
What links to or cites this one, when a source declares it; inferred from the texts held otherwise, and set apart.
Requiresdcterms:requires
none found — the source did not declare it
What this resource needs in order to be used, when a source declares it. The files a lab works on are inferred, and stand under it apart.
Is required bydcterms:isRequiredBy
none found — the source did not declare it
What needs this resource, when a source declares it. The lab a component belongs to is inferred, and stands under it apart.
Provenancedcterms:provenance
this engine
source
conversion
latexml-html
Retrieved from arXiv on 2026-10-09 in response to the search string “(all:"artificial intelligence" OR all:"machine learning" OR all:"generative AI" OR all:"deep learning" OR all:"reinforcement learning" OR all:"large language model") AND (all:"AI concepts" OR all:"types of AI" OR all:"AI fundamentals" OR all:"recognizing AI" OR all:"recognising AI" OR all:"general versus narrow AI" OR all:"narrow AI" OR all:"general AI" OR all:"machine intelligence" OR all:"AI strengths and weaknesses" OR all:"traditional software" OR all:"rule-based systems" OR all:"introduction to AI" OR all:"introduction to artificial intelligence" OR all:"artificial intelligence introduction" OR all:"AI primer" OR all:"foundations of artificial intelligence" OR all:"overview of AI" OR all:"understanding AI" OR all:"history of AI" OR all:"AI essentials" OR all:"AI terminology" OR all:"metaphors for AI" OR all:"AI fundamental concepts" OR all:"AI key concepts" OR all:"philosophy of AI" OR all:"critical AI literacy")”. arXiv served the resource and is not asserted to be its publisher or author.
Text extracted from latexml-html to Markdown by arxiv-html; the original is retained unchanged beside it.
Where it was collected from, what was converted, and what container it came out of — the custody statements that would otherwise be mistaken for authorship.
IEEE 1484.12.1 Learning Object Metadata. LOM has no element for an SPDX identifier or a licence URI, so both are written into 6.3 Rights.Description. Flattening this record into simple Dublin Core would lose more again, which is why the two projections exist side by side rather than one being generated from the other.
the standard ↗
Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality between generative AI and RL for robotics downstream tasks. Specifically, we investigate: (1) The role of prominent generative AI tools as modular priors for multi-modal input fusion in RL tasks. (2) How RL can train, fine-tune and distill generative models for policy generation, such as VLA models, similarly to RL applications in large language models. We then propose a new taxonomy based on a considerable amount of selected papers. Lastly, we identify open challenges accounting for model scalability, adaptation and grounding, giving recommendations and insights on future research directions. We reflect on which generative AI models best fit the RL tasks and why. On the other side, we reflect on important issues inherent to RL-enhanced generative policies, such as safety concerns and failure modes, and what are the limitations of current methods. A curated collection of relevant research papers is maintained on our GitHub repository, serving as a resource for ongoing research and development in this field: https://github.com/clmoro/Robotics-RL-FMs-Integration.
The abstract, and the depositor's additional notes after it as a second LangString when the source has a field for them.
Role, entity and date per declared contribution. A role outside LOM's vocabulary is reported in the entry's description instead.
3 Meta-metadata
4/4
Identifier3.1
this engine
resource_id
URI: tag:aim-pro.eu,2026:oer/5835121c01b5/record
The identifier of this metadata record — the resource's own, with /record after it, because the record is a description of the resource and not the resource.
Contribute3.2
this engine
source
AIM-PRO WP3 OER harvester (arXiv) — creator
Who generated this record and when — a statement about the record, not about the resource.
Metadata schema3.3
LOMv1.0
aimpro-oer-profile/1
LOMv1.0, and the profile this was built against.
Language3.4
language gate
read the declared field
en
4 Technical
3/7
Format4.1
conversion
latexml-html
text/markdown
One value per form held: the original as the source published it, and the Markdown this engine extracted.
Size4.2
conversion
latexml-html
127545
Bytes. The original's, because the resource is the file and not our conversion of it.
Location4.3
this engine
resource_id
https://arxiv.org/abs/2410.16411
Where the source serves it.
Requirement4.4
not collected — this library does not fill it
Software or hardware needed to use it. No source declares it.
Installation remarks4.5
not collected — this library does not fill it
No source declares it.
Other platform requirements4.6
not collected — this library does not fill it
No source declares it.
Duration4.7
not collected — this library does not fill it
Playing time, for audio and video. The corpus holds neither.
5 Educational
1/11
Interactivity type5.1
not collected — this library does not fill it
Active, expositive or mixed. A judgement about how the resource is used; no source declares it.
Yes unless the licence reserves nothing — attribution is a restriction. The conditions after the dash are the licence gate's reading; the export carries LOM's bare term.
The SPDX id, the licence URI and the conditions. LOM has no element for any of the three, so this is where they survive.
7 Relation
2/2
Kind7.1
ispartof
Resource7.2
inferred
read its reference list, line 447citation
Information Fusion Volume 129, May 2026, 104003
references: arXiv:2303.08774
references: arXiv:2308.12270
references: arXiv:2204.01691
references: arXiv:2211.15657
references: arXiv:2312.09187
references: arXiv:2301.08028
references: arXiv:2402.08546
references: arXiv:2309.13041
references: arXiv:2108.07258
references: arXiv:2212.06817
references: arXiv:2307.15818
references: arXiv:2404.00282
and 86 more, every one of them in the export
The container a file was found inside, and the Markdown extracted from the original. What a lab requires, and the lab a component belongs to, are inferred and stand apart.
8 Annotation
0/3
Entity8.1
not collected — this library does not fill it
Comments on the resource's educational use, by whoever made them. The platform's review grades competencies, which are classification (9), and writes no comment here.
Date8.2
not collected — this library does not fill it
Description8.3
not collected — this library does not fill it
9 Classification
0/4
Purpose9.1
not collected — this library does not fill it
Empty in the record: no source declares a competency. The alignment reads the resource for them and stands beside the record, never in it, and a taxon path derived from the search string that found it would be a claim about the query.
Taxon path9.2
not collected — this library does not fill it
Where the competency framework goes. Empty in the record for the reason above.
Description9.3
not collected — this library does not fill it
Keyword9.4
not collected — this library does not fill it
What could not be established3
Where the source's metadata could not be carried
over as it was — missing, contradictory, with no matching term in the
standard, restructured, or taken from the repository — and what was done
instead. Without these notes, an empty element would look like something
the harvester missed.
Status
Field
Why
Not available
rights_holder
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it
Not available
publisher
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim
Not available
educational
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them