Live
Black Hat USAAI BusinessBlack Hat AsiaAI Business"Final Year Student? Here's Exactly What You Need to Get a Dev Job in 2026"DEV CommunityHow I Launched 14 SaaS Products in 6 Months as a Solo Founder Using LovableDEV CommunityFDB Just Launched the First MCP Server for Medication DecisionsDEV CommunityClaude Code Unpacked: A Visual GuideDEV CommunityGoogle Deepmind study exposes six "traps" that can easily hijack autonomous AI agents in the wild - the-decoder.comGoogle News: DeepMind3 Lines of Code Saved Anthropic 250K API Calls Per DayDEV CommunityClaude Knows When You're Mad — And Uses Regex, Not AIDEV CommunityInside Claude Code: 12 Hidden Features Anthropic Didn't Want You to SeeDEV CommunityCameo partners with TikTok to boost popularityTechCrunch AI🔐 AES-256 Finally Makes Sense (And It’s Way Simpler Than You Think)DEV CommunityI Built an OPA Plugin That Turns It Into an AuthZEN-Compatible PDPDEV CommunityMonorepo Architecture with pnpm Workspace, Turborepo & Changesets 📦DEV CommunityBlack Hat USAAI BusinessBlack Hat AsiaAI Business"Final Year Student? Here's Exactly What You Need to Get a Dev Job in 2026"DEV CommunityHow I Launched 14 SaaS Products in 6 Months as a Solo Founder Using LovableDEV CommunityFDB Just Launched the First MCP Server for Medication DecisionsDEV CommunityClaude Code Unpacked: A Visual GuideDEV CommunityGoogle Deepmind study exposes six "traps" that can easily hijack autonomous AI agents in the wild - the-decoder.comGoogle News: DeepMind3 Lines of Code Saved Anthropic 250K API Calls Per DayDEV CommunityClaude Knows When You're Mad — And Uses Regex, Not AIDEV CommunityInside Claude Code: 12 Hidden Features Anthropic Didn't Want You to SeeDEV CommunityCameo partners with TikTok to boost popularityTechCrunch AI🔐 AES-256 Finally Makes Sense (And It’s Way Simpler Than You Think)DEV CommunityI Built an OPA Plugin That Turns It Into an AuthZEN-Compatible PDPDEV CommunityMonorepo Architecture with pnpm Workspace, Turborepo & Changesets 📦DEV Community

Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds

arXivMarch 31, 202610 min read0 views
Source Quiz

arXiv:2603.18532v2 Announce Type: replace-cross Abstract: The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of t — Andrew Choi, Xinjie Wang, Zhizhong Su, Wei Xu

View PDF

Abstract:The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of transforming a broadly pretrained model into an overfitted, scene-specific policy. Training in simulation can instead provide access to diverse scenes, but designing those scenes is also costly. In this work, we show that VLAs can be RL fine-tuned without sacrificing generality and with reduced labor by leveraging 3D world generative models. Using these models together with a language-driven scene designer, we generate hundreds of diverse interactive scenes containing unique objects and backgrounds, enabling scalable and highly parallel policy learning. Starting from a pretrained imitation baseline, our approach increases simulation success from 9.7% to 79.8% while achieving a 1.25$\times$ speedup in task completion time. We further demonstrate successful sim-to-real transfer enabled by the quality of the generated digital twins together with domain randomization, improving real-world success from 21.7% to 75% and achieving a 1.13$\times$ speedup. Finally, we further highlight the benefits of leveraging the effectively unlimited data from 3D world generative models through an ablation study showing that increasing scene diversity directly improves zero-shot generalization.

Subjects:

Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Cite as: arXiv:2603.18532 [cs.RO]

(or arXiv:2603.18532v2 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2603.18532

arXiv-issued DOI via DataCite

Submission history

From: Andrew Choi [view email] [v1] Thu, 19 Mar 2026 06:22:11 UTC (44,258 KB) [v2] Sat, 28 Mar 2026 07:03:02 UTC (44,254 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by AI News Hub · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
Scaling Sim…researchpaperarxivaiartificial-…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 197 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers