Live
Black Hat USADark ReadingBlack Hat AsiaAI BusinessWhy OpenAI Buying TBPN Matters More Than It LooksDev.to AII Built a Governance Layer That Works Across Claude Code, Codex, and Gemini CLIDev.to AIOpenAI’s AGI boss is taking a leave of absenceThe VergeGoogle's Gemma 4 AI can run on smartphones, no Internet requiredTechSpotb8656llama.cpp ReleasesThe future of RealSense 3D vision with Chris Matthieu - The Robot ReportGoogle News - AI roboticsThe future of RealSense 3D vision with Chris MatthieuThe Robot ReportLinkerbot’s Linker Hand L30 Can Tighten Screws in Seconds - TechEBlog -Google News - AI roboticsv0.20.1-rc0Ollama ReleasesPasta-like robot muscles powered by air can lift 100x their weight - Interesting EngineeringGoogle News - AI roboticsAssessing Marvell Technology (MRVL) After Nvidia’s US$2b AI Partnership And Connectivity Push - simplywall.stGNews AI NVIDIAb8651llama.cpp ReleasesBlack Hat USADark ReadingBlack Hat AsiaAI BusinessWhy OpenAI Buying TBPN Matters More Than It LooksDev.to AII Built a Governance Layer That Works Across Claude Code, Codex, and Gemini CLIDev.to AIOpenAI’s AGI boss is taking a leave of absenceThe VergeGoogle's Gemma 4 AI can run on smartphones, no Internet requiredTechSpotb8656llama.cpp ReleasesThe future of RealSense 3D vision with Chris Matthieu - The Robot ReportGoogle News - AI roboticsThe future of RealSense 3D vision with Chris MatthieuThe Robot ReportLinkerbot’s Linker Hand L30 Can Tighten Screws in Seconds - TechEBlog -Google News - AI roboticsv0.20.1-rc0Ollama ReleasesPasta-like robot muscles powered by air can lift 100x their weight - Interesting EngineeringGoogle News - AI roboticsAssessing Marvell Technology (MRVL) After Nvidia’s US$2b AI Partnership And Connectivity Push - simplywall.stGNews AI NVIDIAb8651llama.cpp Releases
AI NEWS HUBbyEIGENVECTOREigenvector

Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly

arXivby [Submitted on 30 Apr 2024 (v1), last revised 27 Mar 2026 (this version, v3)]March 30, 20262 min read1 views
Source Quiz

arXiv:2405.00181v3 Announce Type: replace-cross Abstract: Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: "what anomaly occurred?", "why did it happen?", and "how severe is this abnormal event?". In pursuit of these answers, we present a comprehensive benchmark for Causat — Hang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan, Jiayang Zhang, Junrui Xu, Hangyu Liu, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Jing Feng, Linli Chen, Can Zhang, Xuhuan Li, Hao Zhang, Jianhang Chen, Qimei Cui, Xiaofeng Tao

Authors:Hang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan, Jiayang Zhang, Junrui Xu, Hangyu Liu, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Jing Feng, Linli Chen, Can Zhang, Xuhuan Li, Hao Zhang, Jianhang Chen, Qimei Cui, Xiaofeng Tao

View PDF HTML (experimental)

Abstract:Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: "what anomaly occurred?", "why did it happen?", and "how severe is this abnormal event?". In pursuit of these answers, we present a comprehensive benchmark for Causation Understanding of Video Anomaly (CUVA). Specifically, each instance of the proposed benchmark involves three sets of human annotations to indicate the "what", "why" and "how" of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. In addition, we also introduce MMEval, a novel evaluation metric designed to better align with human preferences for CUVA, facilitating the measurement of existing LLMs in comprehending the underlying cause and corresponding effect of video anomalies. Finally, we propose a novel prompt-based method that can serve as a baseline approach for the challenging CUVA. We conduct extensive experiments to show the superiority of our evaluation metric and the prompt-based approach. Our code and dataset are available at this https URL.

Comments: Accepted in CVPR2024, Codebase: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as: arXiv:2405.00181 [cs.CV]

(or arXiv:2405.00181v3 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2405.00181

arXiv-issued DOI via DataCite

Submission history

From: Binzhu Xie [view email] [v1] Tue, 30 Apr 2024 20:11:49 UTC (20,803 KB) [v2] Mon, 6 May 2024 14:57:50 UTC (20,803 KB) [v3] Fri, 27 Mar 2026 09:28:23 UTC (12,595 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by Eigenvector · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
Uncovering …researchpaperarxivaiartificial-…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 108 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers