EVA: Efficient Reinforcement Learning for End-to-End Video Agent
arXiv:2603.22918v2 Announce Type: replace-cross Abstract: Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames. Existing approaches typically treat MLLMs as passive recognizers, processing entire videos or uniformly sampled frames without adaptive reasoning. Recent agent-based methods introduce external tools, yet still depend on manually designed workflows and perception-first strategies, resulting in inefficiency on long videos. We present EVA, an Eff — Yaolun Zhang, Ruohui Wang, Jiahao Wang, Yepeng Tang, Xuanyu Zheng, Haonan Duan, Hao Lu, Hanming Deng, Lewei Lu
View PDF HTML (experimental)
Abstract:Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames. Existing approaches typically treat MLLMs as passive recognizers, processing entire videos or uniformly sampled frames without adaptive reasoning. Recent agent-based methods introduce external tools, yet still depend on manually designed workflows and perception-first strategies, resulting in inefficiency on long videos. We present EVA, an Efficient Reinforcement Learning framework for End-to-End Video Agent, which enables planning-before-perception through iterative summary-plan-action-reflection reasoning. EVA autonomously decides what to watch, when to watch, and how to watch, achieving query-driven and efficient video understanding. To train such agents, we design a simple yet effective three-stage learning pipeline - comprising supervised fine-tuning (SFT), Kahneman-Tversky Optimization (KTO), and Group Relative Policy Optimization (GRPO) - that bridges supervised imitation and reinforcement learning. We further construct high-quality datasets for each stage, supporting stable and reproducible training. We evaluate EVA on six video understanding benchmarks, demonstrating its comprehensive capabilities. Compared with existing baselines, EVA achieves a substantial improvement of 6-12% over general MLLM baselines and a further 1-3% gain over prior adaptive agent methods.
Comments: CVPR2026
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2603.22918 [cs.CV]
(or arXiv:2603.22918v2 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2603.22918
arXiv-issued DOI via DataCite
Submission history
From: Yaolun Zhang [view email] [v1] Tue, 24 Mar 2026 08:06:29 UTC (17,676 KB) [v2] Thu, 26 Mar 2026 20:03:37 UTC (16,428 KB)
Sign in to highlight and annotate this article

Conversation starters
Daily AI Digest
Get the top 5 AI stories delivered to your inbox every morning.
More about
researchpaperarxiv
Europe urged to ‘learn to fight for itself’ in case US-China truce collapses
European governments breathed a sigh of relief in October when the US and China sealed a fragile trade truce that paused more sweeping Chinese rare earth restrictions and papered over a Sino-Dutch row over chipmaker Nexperia. Now, however, the European Union is being urged to come up with a battle plan should the ceasefire fail or expire. A spike in superpower tensions could expose the EU to Chinese export controls, potentially pulverising its military support for Ukraine, its own efforts to...
[D] TurboQuant author replies on OpenReview
<!-- SC_OFF --><div class="md"><p>I wanted to follow up to <a href="https://www.reddit.com/r/MachineLearning/comments/1s7m7rn/comment/odaect4/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button">yesterday's thread</a> and see if anyone wanted to weigh in on it. This work is far outside of my niche, but it strikes me as an attempt to reframe the issue instead of addressing concerns head on. </p> <p>OpenReview link for reference: <a href="https://openreview.net/forum?id=tO3ASKZlok">https://openreview.net/forum?id=tO3ASKZlok</a></p> <blockquote> <p>In response to recent commentary regarding our paper, "TurboQuant," we provide the following technical clarifications to correct the record.</p> <p>TurboQuant did not derive its core method from

Five CMU Faculty Members Named 2026 Sloan Research Fellows
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/sloan-collage-2000%20copy.jpg.webp?itok=Ih-KH3Na" width="900" height="508" alt="2026 Sloan Awardees"> </p> Five Carnegie Mellon University faculty members are among the 126 recipients of 2026 Sloan Research Fellowships, which honor early career scholars whose achievements put them among the best scientific minds working today.
Knowledge Map
Connected Articles — Knowledge Graph
This article is connected to other articles through shared AI topics and tags.
More in Research Papers
[D] TurboQuant author replies on OpenReview
<!-- SC_OFF --><div class="md"><p>I wanted to follow up to <a href="https://www.reddit.com/r/MachineLearning/comments/1s7m7rn/comment/odaect4/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button">yesterday's thread</a> and see if anyone wanted to weigh in on it. This work is far outside of my niche, but it strikes me as an attempt to reframe the issue instead of addressing concerns head on. </p> <p>OpenReview link for reference: <a href="https://openreview.net/forum?id=tO3ASKZlok">https://openreview.net/forum?id=tO3ASKZlok</a></p> <blockquote> <p>In response to recent commentary regarding our paper, "TurboQuant," we provide the following technical clarifications to correct the record.</p> <p>TurboQuant did not derive its core method from

Carnegie Mellon Researchers Rethink Chronic Pain
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/260123D_WTM_Yttri_Lab_RD002.jpg.webp?itok=DozEPS-g" width="900" height="508" alt="Eric Yttri"> </p> Across neuroscience, biomedical engineering and artificial intelligence, researchers from Carnegie Mellon University are exploring how pain is measured, understood and treated to support safer, more effective care.

Five CMU Faculty Members Named 2026 Sloan Research Fellows
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-02/sloan-collage-2000%20copy.jpg.webp?itok=Ih-KH3Na" width="900" height="508" alt="2026 Sloan Awardees"> </p> Five Carnegie Mellon University faculty members are among the 126 recipients of 2026 Sloan Research Fellowships, which honor early career scholars whose achievements put them among the best scientific minds working today.
Discussion
Sign in to join the discussion
No comments yet — be the first to share your thoughts!