Live
Black Hat USAAI BusinessBlack Hat AsiaAI BusinessAfter a 23% Plunge in the First Quarter, Can Microsoft’s AI Story Continue? - NAI500GNews AI MicrosoftAI Video Generation Startup Runway Unveils $10 Mn VC Fund To Back Early-stage AI Startups: Report - bwdisrupt.comGNews AI startupsOracle layoffs: 12,000 jobs cut in India amid AI push, more layoffs likely - Storyboard18GNews AI IndiaIs Arista Networks (ANET) Becoming NVIDIA’s Go-To AI Network Spine or Just One Key Partner? - simplywall.stGNews AI NVIDIAZhipu's Stock Soars After Chinese AI Startup's Annual Revenue More Than Doubles - Yicai GlobalGNews AI ChinaAustralia signs AI MoU with Anthropic, flags data centre investment - W.MediaGNews AI AustraliaHong Kong hasn’t issued a single HKD stablecoin license after March targetCoinDesk AIBitcoin is closer to its 'buy zone' than it's been in three yearsCoinDesk AIRAG Web Browser: Give Your AI Real-Time Web Access Without HallucinationsDEV CommunityWhat Nobody Tells You About Building a Protocol for AI AgentsDEV CommunityHuawei highlights AI, HarmonyOS and auto momentum in 2025 annual report - TechNodeGNews AI HuaweiThe Evidence Is in the Phone. Most of It Never Makes It Into the Case.DEV CommunityBlack Hat USAAI BusinessBlack Hat AsiaAI BusinessAfter a 23% Plunge in the First Quarter, Can Microsoft’s AI Story Continue? - NAI500GNews AI MicrosoftAI Video Generation Startup Runway Unveils $10 Mn VC Fund To Back Early-stage AI Startups: Report - bwdisrupt.comGNews AI startupsOracle layoffs: 12,000 jobs cut in India amid AI push, more layoffs likely - Storyboard18GNews AI IndiaIs Arista Networks (ANET) Becoming NVIDIA’s Go-To AI Network Spine or Just One Key Partner? - simplywall.stGNews AI NVIDIAZhipu's Stock Soars After Chinese AI Startup's Annual Revenue More Than Doubles - Yicai GlobalGNews AI ChinaAustralia signs AI MoU with Anthropic, flags data centre investment - W.MediaGNews AI AustraliaHong Kong hasn’t issued a single HKD stablecoin license after March targetCoinDesk AIBitcoin is closer to its 'buy zone' than it's been in three yearsCoinDesk AIRAG Web Browser: Give Your AI Real-Time Web Access Without HallucinationsDEV CommunityWhat Nobody Tells You About Building a Protocol for AI AgentsDEV CommunityHuawei highlights AI, HarmonyOS and auto momentum in 2025 annual report - TechNodeGNews AI HuaweiThe Evidence Is in the Phone. Most of It Never Makes It Into the Case.DEV Community

MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models

arXivMarch 31, 20262 min read0 views
Source Quiz

arXiv:2603.24984v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) has emerged as an effective approach to reduce the computational overhead of Transformer architectures by sparsely activating a subset of parameters for each token while preserving high model capacity. This paradigm has recently been extended to Vision-Language Models (VLMs), enabling scalable multi-modal understanding with reduced computational cost. However, the widely adopted deterministic top-K routing mechanism may overlook more optimal expert combinations and lead to expert overfitting. To address this limitatio — Dohwan Ko, Jinyoung Park, Seoung Choi, Sanghyeok Lee, Seohyun Lee, Hyunwoo J. Kim

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) has emerged as an effective approach to reduce the computational overhead of Transformer architectures by sparsely activating a subset of parameters for each token while preserving high model capacity. This paradigm has recently been extended to Vision-Language Models (VLMs), enabling scalable multi-modal understanding with reduced computational cost. However, the widely adopted deterministic top-K routing mechanism may overlook more optimal expert combinations and lead to expert overfitting. To address this limitation and improve the diversity of expert selection, we propose MoE-GRPO, a reinforcement learning (RL)-based framework for optimizing expert routing in MoE-based VLMs. Specifically, we formulate expert selection as a sequential decision-making problem and optimize it using Group Relative Policy Optimization (GRPO), allowing the model to learn adaptive expert routing policies through exploration and reward-based feedback. Furthermore, we introduce a modality-aware router guidance that enhances training stability and efficiency by discouraging the router from exploring experts that are infrequently activated for a given modality. Extensive experiments on multi-modal image and video benchmarks show that MoE-GRPO consistently outperforms standard top-K routing and its variants by promoting more diverse expert selection, thereby mitigating expert overfitting and enabling a task-level expert specialization.

Comments: Accepted at CVPR 2026

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2603.24984 [cs.CV]

(or arXiv:2603.24984v2 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.24984

arXiv-issued DOI via DataCite

Submission history

From: Dohwan Ko [view email] [v1] Thu, 26 Mar 2026 03:23:45 UTC (3,700 KB) [v2] Sun, 29 Mar 2026 06:59:34 UTC (3,695 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by AI News Hub · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

Knowledge Map

Knowledge Map
TopicsEntitiesSource
MoE-GRPO: O…researchpaperarxivcomputer-vi…image-recog…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 96 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers