Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning
arXiv:2603.27400v1 Announce Type: cross Abstract: Several approaches have been proposed to improve the sample efficiency of online reinforcement learning (RL) by leveraging demonstrations collected offline. The offline data can be used directly as transitions to optimize RL objectives, or offline policy and value functions can first be learned from the data and then used for online finetuning or to provide reference actions. While each of these strategies has shown compelling results, it is unclear which method has the most impact on sample efficiency, whether these approaches can be combined, — Dwait Bhatt, Shih-Chieh Chou, Nikolay Atanasov
View PDF HTML (experimental)
Abstract:Several approaches have been proposed to improve the sample efficiency of online reinforcement learning (RL) by leveraging demonstrations collected offline. The offline data can be used directly as transitions to optimize RL objectives, or offline policy and value functions can first be learned from the data and then used for online finetuning or to provide reference actions. While each of these strategies has shown compelling results, it is unclear which method has the most impact on sample efficiency, whether these approaches can be combined, and if there are cumulative benefits. We classify existing demonstration-augmented RL approaches into three categories and perform an extensive empirical study of their strengths, weaknesses, and combinations to isolate the contribution of each strategy and determine effective hybrid combinations for sample-efficient online RL. Our analysis reveals that directly reusing offline data and initializing with behavior cloning consistently outperform more complex offline RL pretraining methods for improving online sample efficiency.
Comments: Accepted to ICRA 2026
Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
Cite as: arXiv:2603.27400 [cs.RO]
(or arXiv:2603.27400v1 [cs.RO] for this version)
https://doi.org/10.48550/arXiv.2603.27400
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Dwait Bhatt [view email] [v1] Sat, 28 Mar 2026 20:34:20 UTC (397 KB)
Sign in to highlight and annotate this article

Conversation starters
Daily AI Digest
Get the top 5 AI stories delivered to your inbox every morning.
More about
researchpaperarxiv‘It’s all very possible’: Michael Patrick King on The Comeback return’s shocking AI twist – and why And Just Like That will age well
<p>Could AI write an entire sitcom series? That’s the premise of the new season of comedy drama The Comeback. Its co-creator talks about being shocked by his research – and why the world needs to catch up with AJLT</p><p>TV veteran Michael Patrick King has had a long, lively career, writing, directing and producing on shows including Murphy Brown, Will & Grace and 2 Broke Girls. He’s best-known, though, for his work on the Sex and the City franchise, serving as its showrunner for the bulk of its run, writing and directing its two films, and masterminding its controversial <a href="https://www.theguardian.com/tv-and-radio/2025/aug/02/goodbye-and-just-like-that-right-time-to-end-cursed-spin-off">2020s revival And Just Like That</a>. But this month sees the return of one of his most loved
SEAS Researchers Expose Hidden “Alignment Discretion” Shaping AI Behavior - Harvard School of Engineering and Applied Sciences
<a href="https://news.google.com/rss/articles/CBMiowFBVV95cUxONEZoeldndUE2ZUZsSkloQmZVMk1jZUhncUY0V3c5NDQ0TlNLVTluWkppejlpOVdpemxqNEVvaDAwbG43VWpxOFpuakFtRDNLUlVTbEwzR25kYnJieVdkSEs0MVRESHpPSHN6dEhXUk9qVGRlVjFhT2ZqVjlJV056MG94MDN3dWVCSUtoRWVDODF5bVVET2gxQW5DSk1oT1pUNlVB?oc=5" target="_blank">SEAS Researchers Expose Hidden “Alignment Discretion” Shaping AI Behavior</a> <font color="#6f6f6f">Harvard School of Engineering and Applied Sciences</font>
Knowledge Map
Connected Articles — Knowledge Graph
This article is connected to other articles through shared AI topics and tags.
More in Research Papers
Oracle Layoffs Recast Costs To Back US$50b AI Infrastructure Bet - simplywall.st
<a href="https://news.google.com/rss/articles/CBMivwFBVV95cUxQNWpZb2ZQVDBIOGVZTTBtLThzaGwxS3NkMnJBSS1wek5pQlJXRWdTOEh5aTdPTE9Cd3JHdjZDeWRtVzdMUUdESHJOQXZDdGNVdGZtTTBhanpfb3UxQnRobVlzNGdVUXJLZWptV2V6NXlNSWllX3FxOU5XYTF0RkM2TnJIaFJkcVBFOGc2alBSLTZEeU85QU1oTjBrMVZSTl84dm9GeFl5OGtUMjc3LVd1dS1fcHZ1RG9HcV82T2JFWdIBxAFBVV95cUxOSE5XVXh0QkM4Yi1WbXNhWkJ2Z2dLRlBGNjAwaTcyNFJWMWRPdXo5WjRQQkRGTG9IamxxbmdhMHpsaEJ6RDQwZl9ENGl5WDc5a2lrTXZ1bVpFbGdsdndHYjFINnZPSnNKX1dZamszUXByR1BlRXF6d1pKOHpBU3M5UFhUSldlUWtIMlRNQzdvTk9haEJKeDI1ZEg0WWQ1SXYzLUZCWElQc3pzR19ucGExdVpnc2hBQXlQNVpOZFVBVzRkLXFE?oc=5" target="_blank">Oracle Layoffs Recast Costs To Back US$50b AI Infrastructure Bet</a> <font color="#6f6f6f">simplywall.st</font>
Riyadh conference to discuss role of AI in media industry - Arab News PK
<a href="https://news.google.com/rss/articles/CBMiVEFVX3lxTE1jdFVMUFA3R2RXM19JR1M1NnpjX210dUZuNkI3VWdQc0tzVVBZaXR3ZlNqUVFyZlB5aTMxOGI3OXFpdGpQX2RsOXF3UU5kaXlma2VpTQ?oc=5" target="_blank">Riyadh conference to discuss role of AI in media industry</a> <font color="#6f6f6f">Arab News PK</font>
Losito named IBM Italia general manager - Telecompaper
<a href="https://news.google.com/rss/articles/CBMiigFBVV95cUxNRTQ0RzVrcHJsVXo0THF3UllROGwyam1FNl9RWlV2dzJFRGtGMktoTGlYVUR5dU1WX1JSTkExQlNSVEFSWktVQVJSazFUUTJyV2tadUlraVlGM3M3WHNZNFNodm5DeVBvTXFkaDNkNXJ4SzF0RnphNGxOYlFGaFRtR241R2M0NFhUakE?oc=5" target="_blank">Losito named IBM Italia general manager</a> <font color="#6f6f6f">Telecompaper</font>


Discussion
Sign in to join the discussion
No comments yet — be the first to share your thoughts!