Research Papers research paper arxiv ai artificial-intelligence

Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds

arXivMarch 31, 202610 min read0 views

arXiv:2603.18532v2 Announce Type: replace-cross Abstract: The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of t — Andrew Choi, Xinjie Wang, Zhizhong Su, Wei Xu

View PDF

Abstract:The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of transforming a broadly pretrained model into an overfitted, scene-specific policy. Training in simulation can instead provide access to diverse scenes, but designing those scenes is also costly. In this work, we show that VLAs can be RL fine-tuned without sacrificing generality and with reduced labor by leveraging 3D world generative models. Using these models together with a language-driven scene designer, we generate hundreds of diverse interactive scenes containing unique objects and backgrounds, enabling scalable and highly parallel policy learning. Starting from a pretrained imitation baseline, our approach increases simulation success from 9.7% to 79.8% while achieving a 1.25$\times$ speedup in task completion time. We further demonstrate successful sim-to-real transfer enabled by the quality of the generated digital twins together with domain randomization, improving real-world success from 21.7% to 75% and achieving a 1.13$\times$ speedup. Finally, we further highlight the benefits of leveraging the effectively unlimited data from 3D world generative models through an ablation study showing that increasing scene diversity directly improves zero-shot generalization.

Subjects:

Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Cite as: arXiv:2603.18532 [cs.RO]

(or arXiv:2603.18532v2 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2603.18532

arXiv-issued DOI via DataCite

Submission history

From: Andrew Choi [view email] [v1] Thu, 19 Mar 2026 06:22:11 UTC (44,258 KB) [v2] Sat, 28 Mar 2026 07:03:02 UTC (44,254 KB)

Original source

arXiv

https://arxiv.org/abs/2603.18532

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

CountriesLive

UCLA Researchers Highlight the ‘Body Gap’ in AI: Why Lacking Human Experience Could Impact Safety - Bioengineer.org

<a href="https://news.google.com/rss/articles/CBMiuwFBVV95cUxNekVOelBPcWt4Y2lGcHFvalZ4UWdRLUxLMnN5NndWVURBNUt2T25DYjVsT2F4VVdUUE1XbEFsTW91d3JDaFJsV1NiSjBDenU2Z1ZCaUhrUl81ME5TTlN4OC1lSzhmV1hrUThSc3U5UkRyelRGYWVDdE9fUF84VUZSWEpFYlNBTnU0WHBMVTVHUlFqd1VXZ1hna3NNSjNfeXdTdkdJTHdJQzhvdE52eVZLVkV6R0hETW84a2pR?oc=5" target="_blank">UCLA Researchers Highlight the ‘Body Gap’ in AI: Why Lacking Human Experience Could Impact Safety</a> Bioengineer.org

Google News: AI Safety

1m19 minutes ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNd09pNzVMTG0wTVNDZGgyWlZKNFpURjg4YkZTYVpKNWd2WnUzSUllZXdsYVgxM0tRblhmbXNDQ050cGlmY0gwVVpxVXZxLW5XX1V5U1NUXy1xaXVoVmlLb29RSG1aa0hNSzBaZjh3Q1N0a1pIQnNpc3lKQnJDdG5rUjlZUjE0NUhtUWstUGwxSHdtSWUzelUydXFQbzdZaTB1QnNjYWF2WWQ5RnB3YV9vNk00VkhyQUdKSnFzX1VoZWFzZElkLVh6a2QzMm1pY21EeURBVFhvMUZMNDFZTFpmd0k2OWJ1MFpYd0wydi1BSUFyalJhUGRfeWFHY21UZzVGYm5USU9iV3dCQjdHR2hUVGw1UUk5aU5xVkExX1RBckhQYk1OcTQwRDJsNldSRWdIZ2ljdlg0SVdYQWRkQkx2eG1feWtfdXM2YWFNdEpuLXZEcGVqSVRzNFdid0dwd2QwQUFQWEItUVp2VktzeG1LM09KOHM3bEltekZyRjNSbDhiZnRleGVnU0VBSTI0NDhMZEx3RThvSFFKVFRFMms0cEtVeXpuQ3Z2OWQyeGstaDJOLWJmNFpzWA?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 23 hours ago

ProductsLive

Save the Sun Shrimp!

The supposition that we live in a "goldilocks zone" is frankly just nonsense built up by an anthropocentric need to feel self-important, like Copernicus I am here to rescue us from a self-absorbed disaster of thought. Indeed, what is required for life to form is the ability to create complex structures with causal persistence times above a threshold. With this in mind we are able to find many areas where organisms could persist, if we just had the eyes to see them, namely the Sun! The surface of the Sun is frankly massive, mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space:

LessWrong AI

15m29 minutes ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 197 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersLive

Researchers to use robotics and AI to help sheep producers - University of Nevada, Reno

<a href="https://news.google.com/rss/articles/CBMic0FVX3lxTFB4UmxpREpFODBJN0lKakYwRVVtdlZPNmNiTExRelVFaDYzYW9kX2RCc0pEZjlmX01fT1dWYTlxZE1ET2ZKVVgzSVZIenY3bDlHa3FXS1dUdVBmTEdLa1hUR2x3OWxHbkE2RnROSjl6VHVHQ2c?oc=5" target="_blank">Researchers to use robotics and AI to help sheep producers</a> University of Nevada, Reno

Google News: AI

1mabout 2 hours ago

Research PapersLive

AIRA_2: Breaking Bottlenecks In AI Research Agents - Forbes

<a href="https://news.google.com/rss/articles/CBMiowFBVV95cUxNNmtndHhmQ2lpZGdPdTJwY25xejcyV1c1SWNLdWFOWnNwbjRUQTF0ZWdOZFNaclNBNWVsaUgtU0JUM2xrakhoOXVLMVJzVTNkajdrMmJGeS1lYUpMUG1NMkZNMDJFREZZdXU2ZVdEbkNZSDNBRjJBLVYyZE9XeEY4T0RJY3J5aDVWcEZVQ2lWUjhUYXBsUk16d09NdGdsQ3lxb3gw?oc=5" target="_blank">AIRA_2: Breaking Bottlenecks In AI Research Agents</a> Forbes

Google News: Machine Learning

1mabout 1 hour ago

Research PapersFresh

Can Science Predict When a Study Won’t Hold Up?

Conducting research is hard; confirming the results is, too. And artificial intelligence isn’t yet ready to help, a major new study finds.

NYT Technology

1mabout 2 hours ago

Research PapersFresh

Oracle Layoffs Recast Costs To Back US$50b AI Infrastructure Bet - simplywall.st

<a href="https://news.google.com/rss/articles/CBMivwFBVV95cUxQNWpZb2ZQVDBIOGVZTTBtLThzaGwxS3NkMnJBSS1wek5pQlJXRWdTOEh5aTdPTE9Cd3JHdjZDeWRtVzdMUUdESHJOQXZDdGNVdGZtTTBhanpfb3UxQnRobVlzNGdVUXJLZWptV2V6NXlNSWllX3FxOU5XYTF0RkM2TnJIaFJkcVBFOGc2alBSLTZEeU85QU1oTjBrMVZSTl84dm9GeFl5OGtUMjc3LVd1dS1fcHZ1RG9HcV82T2JFWdIBxAFBVV95cUxOSE5XVXh0QkM4Yi1WbXNhWkJ2Z2dLRlBGNjAwaTcyNFJWMWRPdXo5WjRQQkRGTG9IamxxbmdhMHpsaEJ6RDQwZl9ENGl5WDc5a2lrTXZ1bVpFbGdsdndHYjFINnZPSnNKX1dZamszUXByR1BlRXF6d1pKOHpBU3M5UFhUSldlUWtIMlRNQzdvTk9haEJKeDI1ZEg0WWQ1SXYzLUZCWElQc3pzR19ucGExdVpnc2hBQXlQNVpOZFVBVzRkLXFE?oc=5" target="_blank">Oracle Layoffs Recast Costs To Back US$50b AI Infrastructure Bet</a> simplywall.st

GNews AI USA

1mabout 5 hours ago