Research Papers research paper arxiv computer-vision image-recognition

Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformers

arXivMarch 31, 20262 min read0 views

arXiv:2601.14959v2 Announce Type: replace Abstract: Existing video frame interpolation (VFI) methods often adopt a frame-centric approach, processing videos as independent short segments (e.g., triplets), which leads to temporal inconsistencies and motion artifacts. To overcome this, we propose a holistic, video-centric paradigm named Local Diffusion Forcing for Video Frame Interpolation (LDF-VFI). Our framework is built upon an auto-regressive diffusion transformer that models the entire video sequence to ensure long-range temporal coherence. To mitigate error accumulation inherent in auto-re — Xinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng, Yaoming Wang, Xin Chen, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong

View PDF HTML (experimental)

Abstract:Existing video frame interpolation (VFI) methods often adopt a frame-centric approach, processing videos as independent short segments (e.g., triplets), which leads to temporal inconsistencies and motion artifacts. To overcome this, we propose a holistic, video-centric paradigm named Local Diffusion Forcing for Video Frame Interpolation (LDF-VFI). Our framework is built upon an auto-regressive diffusion transformer that models the entire video sequence to ensure long-range temporal coherence. To mitigate error accumulation inherent in auto-regressive generation, we introduce a novel skip-concatenate sampling strategy that effectively maintains temporal stability. Furthermore, LDF-VFI incorporates sparse, local attention and tiled VAE encoding, a combination that not only enables efficient processing of long sequences but also allows generalization to arbitrary spatial resolutions (e.g., 4K) at inference without retraining. An enhanced conditional VAE decoder, which leverages multi-scale features from the input video, further improves reconstruction fidelity. Empirically, LDF-VFI achieves state-of-the-art performance on challenging VFI benchmarks, demonstrating superior per-frame quality and temporal consistency, especially in scenes with large motion. The source code is available at this https URL.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2601.14959 [cs.CV]

(or arXiv:2601.14959v2 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2601.14959

arXiv-issued DOI via DataCite

Submission history

From: Xinyu Peng [view email] [v1] Wed, 21 Jan 2026 12:58:52 UTC (1,633 KB) [v2] Mon, 30 Mar 2026 08:09:27 UTC (24,420 KB)

Original source

arXiv

https://arxiv.org/abs/2601.14959

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

ModelsLive

What Karpathy's Autoresearch Unlocked for Me

I'm not a data scientist. I've trained a few models before — simple classification problems, with AI writing the Python and me running the iterations. It worked. I got confident. Then a friend asked for help with something harder. <h2> Three Weeks at 0.58 </h2> The problem involved predicting an outcome from a mix of CRM data and call recordings. Not trivial, but not exotic either. Quick primer on AUC — the metric I'll use throughout. Imagine your model looks at two random people: one where the answer is yes, one where it's no. AUC measures how often the model correctly ranks the yes above the no. Score of 0.5 means random guessing. Score of 1.0 means always right. I tried everything I knew: XGBoost, feature engineering, extracting features from transcripts u

DEV Community

3m24 minutes ago

Research PapersFresh

Is AI's visual understanding mostly a 'mirage'? New research suggests so - inkl

<a href="https://news.google.com/rss/articles/CBMimwFBVV95cUxNUjItTURkaldjcXhpNjA0b2dSVUhWNGRwaEkzclBMM0FFMk8wNTNsMWFDc1hmVU9jYU5jbzBKSG42WGZpbTI5TUs0R0w5Q2QxX0NCdjgtT2JtUUNsWHkxUFE1MzJjN1RxMmdmQjc5YlBsVTdfVU02Z1FZSlFrLU1hdmxibjBGcUV3STlvc0JnaG9PcXRDZVhQamdmYw?oc=5" target="_blank">Is AI's visual understanding mostly a 'mirage'? New research suggests so</a> inkl

GNews AI multimodal

1mabout 5 hours ago

Self-Evolving AI

Google DeepMind Introduces Aletheia: The AI Agent Moving from Math Competitions to Fully Autonomous Professional Research Discoveries - marktechpost.com

<a href="https://news.google.com/rss/articles/CBMigwJBVV95cUxPOWhkeUl4RkMxb1IzY1NKcUlVeFpYQ3NWc0ZLWVo2OWRESkdsTkZYYlpaamJ6WjE0Z3RaUkFJVEhIQnBFSHdjV2tEd3R1dFVGS28tYXlFNlZwZnVZSnV2TFlNNDFrNDFIdGJ6VzlwbENMZ2x3dHdFNXFWdzlWLWt2OEZQcW1WbTExSUdOVnNjbktiQURwQXRLUFZVeXp5WjZhbTY4dXhpdlphNWl2THNRNGxqbVcxcDlDbW5US3VBLUNvR1owSHNIUE5xMktmcVVDTjl4dXpOMmVPTUdQWkx0aF9yRXowU0NxQ2lHc0VMRzlaNDEyU0lLY0lSdWpLUndFUll30gGIAkFVX3lxTE0tR0JPRll6R09pU1d3NzVSQ0YwSVRJS1Q5YVF6THpfRUhEZ2EyVGJBNi1XX1ZOT19zZkU1WDlqelRzMm5NWjU3VTR2WC1LZ0drMTUyaHZVWFNLa1MwbEJ1OUZYM2R5cWFza0hJbllFSHhPc0tWYTNMbU1PMmw4T0RtLVpZLXBfbERKRUR0LTF6bl94S1FJRDBweVpxRGpTU242a25lMTVHdG5pTXBxUVgtaHp3MG9yX1NEd3p6Z0lKaWxLTVJGNWQxVkxVRFdZbERzSFBJUjUwQkRPNENYWVVrdlM4dTljODRxeWhDMDdLN1czT0tadUM1YmtCcENBcWU4NngwcUI3ZA?oc=5" target="_blank">Google DeepMind Intro

Google News: DeepMind

1m19 days ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 139 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersFresh

Is AI's visual understanding mostly a 'mirage'? New research suggests so - inkl

GNews AI multimodal

1mabout 5 hours ago

Research PapersLive

Google backs UH Mānoa AI, robotics research - University of Hawaii System

<a href="https://news.google.com/rss/articles/CBMickFVX3lxTE04TGNzcGpVeFNibkdwMzJIOFdrMHYtSS1OUDdFZkR6RFRtUU4yYWN1MlYtRGZQazl5Y1k3SklTYURwYVBycG5NZm5zbzNfRjY0SF9WNVJfM2tTU1BOX0xfb2ZtaWFVVFp3cFg3WXJmNnQwQQ?oc=5" target="_blank">Google backs UH Mānoa AI, robotics research</a> University of Hawaii System

GNews AI Google

1mabout 1 hour ago

Research PapersLive

A Retrospective on the ICLR 2026 Review Process

The selection of papers for ICLR 2026 has fully concluded. We extend our congratulations to the authors whose work will appear at the conference. Creating ICLR’s technical program requires immense effort from the authors, reviewers, and area chairs, and we thank you for your contributions and service. For researchers whose work was rejected, we hope […]

blog.iclr.cc

1mabout 2 hours ago

Research Papers

Vector Researchers present papers at ACL 2024

Vector researchers will be well represented at the 62nd Annual Meeting of the Association for Computational Linguistics in Bangkok, Thailand this year. 14 papers co-authored by Vector-affiliated researchers are being […] The post Vector Researchers present papers at ACL 2024 appeared first on Vector Institute for Artificial Intelligence .

Vector Institute

1mover 1 year ago