Live
Black Hat USADark ReadingBlack Hat AsiaAI BusinessPOTS explained: The disorder that forced OpenAI exec Fidji Simo to take medical leaveBusiness InsiderWhat is POTS, the disorder that forced OpenAI exec Fidji Simo to take medical leave - Business InsiderGoogle News: OpenAIWhy Your Data Governance is Already ObsoleteAI YouTube Channel 35No Fooling, Spaceballs 2 Will Hit Theaters April 2027GizmodoF1 Built the Perfect Model. Then the Cars Went Racing.Medium AISpaceX IPO Access Reportedly Tied to xAI Grok Adoption by Major Banks - TipRanksGNews AI GrokMost People Use AI Every Day, But Don’t Understand These Simple ThingsMedium AII Let AI Make My Decisions for 7 Days. It Worked and That’s What Worried Me.Medium AII Built a Tiny Computer Inside a TransformerMedium AISteam could soon show estimated FPS based on crowd-sourced player dataTechSpotDesktop Canary v2.1.48-canary.36LobeChat ReleasesThe One Thing Most Python Tutorials Won’t Teach YouMedium AIBlack Hat USADark ReadingBlack Hat AsiaAI BusinessPOTS explained: The disorder that forced OpenAI exec Fidji Simo to take medical leaveBusiness InsiderWhat is POTS, the disorder that forced OpenAI exec Fidji Simo to take medical leave - Business InsiderGoogle News: OpenAIWhy Your Data Governance is Already ObsoleteAI YouTube Channel 35No Fooling, Spaceballs 2 Will Hit Theaters April 2027GizmodoF1 Built the Perfect Model. Then the Cars Went Racing.Medium AISpaceX IPO Access Reportedly Tied to xAI Grok Adoption by Major Banks - TipRanksGNews AI GrokMost People Use AI Every Day, But Don’t Understand These Simple ThingsMedium AII Let AI Make My Decisions for 7 Days. It Worked and That’s What Worried Me.Medium AII Built a Tiny Computer Inside a TransformerMedium AISteam could soon show estimated FPS based on crowd-sourced player dataTechSpotDesktop Canary v2.1.48-canary.36LobeChat ReleasesThe One Thing Most Python Tutorials Won’t Teach YouMedium AI
AI NEWS HUBbyEIGENVECTOREigenvector

Posterior Optimization with Clipped Objective for Bridging Efficiency and Stability in Generative Policy Learning

arXiv cs.ROby [Submitted on 2 Apr 2026]April 3, 20261 min read2 views
Source Quiz

arXiv:2604.01860v1 Announce Type: new Abstract: Expressive generative models have advanced robotic manipulation by capturing complex, multi-modal action distributions over temporally extended trajectories. However, fine-tuning these policies via RL remains challenging due to instability and sample inefficiency. We introduce Posterior Optimization with Clipped Objective (POCO), a principled RL framework that formulates policy improvement as a posterior inference problem tailored for temporal action chunks. Through an Expectation-Maximization procedure, POCO distills a reward-weighted implicit posterior into the policy without likelihood estimation. Furthermore, POCO adopts an offline-to-online paradigm that anchors online exploration to pre-trained priors, and its model-agnostic design scal

View PDF HTML (experimental)

Abstract:Expressive generative models have advanced robotic manipulation by capturing complex, multi-modal action distributions over temporally extended trajectories. However, fine-tuning these policies via RL remains challenging due to instability and sample inefficiency. We introduce Posterior Optimization with Clipped Objective (POCO), a principled RL framework that formulates policy improvement as a posterior inference problem tailored for temporal action chunks. Through an Expectation-Maximization procedure, POCO distills a reward-weighted implicit posterior into the policy without likelihood estimation. Furthermore, POCO adopts an offline-to-online paradigm that anchors online exploration to pre-trained priors, and its model-agnostic design scales to fine-tune large VLA models without architectural modifications. Evaluations across 7 simulation benchmarks and 4 contact-rich real-world tasks demonstrate that POCO prevents catastrophic policy collapse, outperforms SOTA baselines, and achieves a 96.7% success rate on real-world tasks. Videos are available at our project website this https URL.

Subjects:

Robotics (cs.RO)

Cite as: arXiv:2604.01860 [cs.RO]

(or arXiv:2604.01860v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2604.01860

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuhui Chen [view email] [v1] Thu, 2 Apr 2026 10:15:47 UTC (11,831 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by Eigenvector · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

Knowledge Map

Knowledge Map
TopicsEntitiesSource
Posterior O…modelbenchmarkannounceavailablevaluationpolicyarXiv cs.RO

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Building knowledge graph…

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!