Research Papers research paper arxiv ai artificial-intelligence

LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation

arXivMarch 31, 202610 min read0 views

arXiv:2603.27693v1 Announce Type: cross Abstract: Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely rely on implicit or indirect alignment signals and remain suboptimal for simultaneously supporting multimodal understanding and generation, particularly in settings that require fine-grained language-visual reasoning and controllable generation. In this work, we propose LVRPO, a language-visual reinforcement-based preference optimization framework that explicitly align — Shentong Mo, Sukmin Yun

View PDF HTML (experimental)

Abstract:Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely rely on implicit or indirect alignment signals and remain suboptimal for simultaneously supporting multimodal understanding and generation, particularly in settings that require fine-grained language-visual reasoning and controllable generation. In this work, we propose LVRPO, a language-visual reinforcement-based preference optimization framework that explicitly aligns language and visual representations using Group Relative Policy Optimization (GRPO). Instead of introducing additional alignment losses at the representation level, LVRPO directly optimizes multimodal model behaviors through preference-driven reinforcement signals, encouraging consistent and semantically grounded interactions between language and vision across both understanding and generation tasks. This formulation enables effective alignment without requiring auxiliary encoders or handcrafted cross-modal objectives, and naturally extends to diverse multimodal capabilities. Empirically, LVRPO consistently outperforms strong unified-pretraining baselines on a broad suite of benchmarks spanning multimodal understanding, generation, and reasoning.

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Multimedia (cs.MM)

Cite as: arXiv:2603.27693 [cs.CV]

(or arXiv:2603.27693v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.27693

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shentong Mo [view email] [v1] Sun, 29 Mar 2026 13:38:21 UTC (7,566 KB)

Original source

arXiv

https://arxiv.org/abs/2603.27693

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Research PapersFresh

Researchers Map Mycorrhizal Fungi Carbon Hotspots - Let's Data Science

Researchers Map Mycorrhizal Fungi Carbon Hotspots Let's Data Science

Google News: Machine Learning

1mabout 9 hours ago

ModelsFresh

[D] AI research on small language models

i'm doing research on some trending fields in AI, currently working on small language models and would love to meet people who are working in similar domains and are looking to write/publish papers! submitted by /u/StoicWithSyrup [link] [comments]

Reddit r/MachineLearning

1mabout 2 hours ago

CountriesFresh

Promising Signals on AI Governance from China

View the official memo here. China has consistently signaled a willingness to engage on global AI governance since at least 2017. This memo compiles key statements from the Chinese government and prominent figures demonstrating their desire to coordinate on the problem of AI. Chinese Vice Premier Ding Xuexiang, at the 2025 World Economic Forum, said: [ ] The post Promising Signals on AI Governance from China appeared first on Machine Intelligence Research Institute .

intelligence.org

1mabout 4 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 160 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersFresh

Researchers Map Mycorrhizal Fungi Carbon Hotspots - Let's Data Science

Researchers Map Mycorrhizal Fungi Carbon Hotspots Let's Data Science

Google News: Machine Learning

1mabout 9 hours ago

Research Papers

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI - WSJ

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI WSJ

GNews AI manufacturing

1mabout 1 month ago

Research Papers

AI Journey 2025 Conference: exploring the future of artificial intelligence - Азия-Плюс

AI Journey 2025 Conference: exploring the future of artificial intelligence Азия-Плюс

Google News - AI Tajikistan

1m5 months ago

Research Papers

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Vision Language Models struggle with fine-grained visual perception tasks due to their language-centric training approach, performing poorly on unnamed visual entities despite having relevant information in their representations. (1 upvotes on HuggingFace)

HuggingFace Papers

3m5 days ago