Live
Black Hat USADark ReadingBlack Hat AsiaAI BusinessMassachusetts Sen. Ed Markey is putting AV firms on blast for using human staffersFast Company TechST’s smart IMU bolsters Qualcomm’s monster AI chip for wearablesFierce ElectronicsRound three: More Rising Stars 2026Fierce ElectronicsQ/A: How engineers must design AVs to drive safelyFierce ElectronicsBosch’s pressure sensor is part of Qualcomm’s new wearables chipFierce ElectronicsQ/A: Lumotive CTO talks software-defined optical sensingFierce ElectronicsOpenAI contract with U.S. Cyber Command went unnoticed amid degradation of transparency and veracity of U.S. procurement database - All-Source Intelligence | Jack PoulsonGoogle News: OpenAIEDITORIAL: Benefits of generative AI do not outweigh drawbacks - The Daily TargumGoogle News: Generative AIHere's the severance package Oracle offered laid-off US employeesBusiness InsiderTeenager died after asking ChatGPT for ‘most successful’ way to take his life, inquest toldThe Guardian AITeenager died after asking ChatGPT for ‘most successful’ way to take his life, inquest told - The GuardianGoogle News: ChatGPTBig Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.Dev.to AIBlack Hat USADark ReadingBlack Hat AsiaAI BusinessMassachusetts Sen. Ed Markey is putting AV firms on blast for using human staffersFast Company TechST’s smart IMU bolsters Qualcomm’s monster AI chip for wearablesFierce ElectronicsRound three: More Rising Stars 2026Fierce ElectronicsQ/A: How engineers must design AVs to drive safelyFierce ElectronicsBosch’s pressure sensor is part of Qualcomm’s new wearables chipFierce ElectronicsQ/A: Lumotive CTO talks software-defined optical sensingFierce ElectronicsOpenAI contract with U.S. Cyber Command went unnoticed amid degradation of transparency and veracity of U.S. procurement database - All-Source Intelligence | Jack PoulsonGoogle News: OpenAIEDITORIAL: Benefits of generative AI do not outweigh drawbacks - The Daily TargumGoogle News: Generative AIHere's the severance package Oracle offered laid-off US employeesBusiness InsiderTeenager died after asking ChatGPT for ‘most successful’ way to take his life, inquest toldThe Guardian AITeenager died after asking ChatGPT for ‘most successful’ way to take his life, inquest told - The GuardianGoogle News: ChatGPTBig Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.Dev.to AI

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

arXivMarch 31, 20262 min read0 views
Source Quiz

arXiv:2603.27259v1 Announce Type: new Abstract: Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse summarization, offering limited insight into temporal understanding over long contexts. In this work, we define a scene as a coherent segment of a video in which both visual and semantic contexts remain consistent, aligning with human perception. This leads us to a key question: can current VLMs reason effectively over — Seng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao, Jinping Wang, Yu Zhang, Chao Li

View PDF

Abstract:Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse summarization, offering limited insight into temporal understanding over long contexts. In this work, we define a scene as a coherent segment of a video in which both visual and semantic contexts remain consistent, aligning with human perception. This leads us to a key question: can current VLMs reason effectively over long, scene-level contexts? To answer this, we introduce a new benchmark, SceneBench, designed to provide scene-level challenges. Our evaluation reveals a sharp drop in accuracy when VLMs attempt to answer scene-level questions, indicating significant forgetting of long-range context. To further validate these findings, we propose Scene Retrieval-Augmented Generation (Scene-RAG), which constructs a dynamic scene memory by retrieving and integrating relevant context across scenes. This Scene-RAG improves VLM performance by +2.50%, confirming that current models still struggle with long-context retention. We hope SceneBench will encourage future research toward VLMs with more robust, human-like video comprehension.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2603.27259 [cs.CV]

(or arXiv:2603.27259v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.27259

arXiv-issued DOI via DataCite (pending registration)

Journal reference: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

Submission history

From: Hao Chen Calvin [view email] [v1] Sat, 28 Mar 2026 12:44:19 UTC (15,241 KB)

Original source

arXiv

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by AI News Hub · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
Seeing the …researchpaperarxivcomputer-vi…image-recog…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 182 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers