Research Papers research paper arxiv computer-vision image-recognition

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

arXivMarch 26, 202610 min read0 views

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables interactive storytelling and efficient on-the-fly frame generation. By reformulating the task as next-shot generation conditioned on historical context, ShotStream allows users to dynamically instruct ongoing narratives via streaming prompts. We achieve this by first fine-tuning a text-to-video model into a bidirectional next-shot generator, which is then dis — Yawen Luo, Xiaoyu Shi, Junhao Zhuang

View PDF HTML (experimental)

Abstract:Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables interactive storytelling and efficient on-the-fly frame generation. By reformulating the task as next-shot generation conditioned on historical context, ShotStream allows users to dynamically instruct ongoing narratives via streaming prompts. We achieve this by first fine-tuning a text-to-video model into a bidirectional next-shot generator, which is then distilled into a causal student via Distribution Matching Distillation. To overcome the challenges of inter-shot consistency and error accumulation inherent in autoregressive generation, we introduce two key innovations. First, a dual-cache memory mechanism preserves visual coherence: a global context cache retains conditional frames for inter-shot consistency, while a local context cache holds generated frames within the current shot for intra-shot consistency. And a RoPE discontinuity indicator is employed to explicitly distinguish the two caches to eliminate ambiguity. Second, to mitigate error accumulation, we propose a two-stage distillation strategy. This begins with intra-shot self-forcing conditioned on ground-truth historical shots and progressively extends to inter-shot self-forcing using self-generated histories, effectively bridging the train-test gap. Extensive experiments demonstrate that ShotStream generates coherent multi-shot videos with sub-second latency, achieving 16 FPS on a single GPU. It matches or exceeds the quality of slower bidirectional models, paving the way for real-time interactive storytelling. Training and inference code, as well as the models, are available on our

Comments: Project Page: this https URL Code: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2603.25746 [cs.CV]

(or arXiv:2603.25746v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.25746

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yawen Luo [view email] [v1] Thu, 26 Mar 2026 17:59:59 UTC (3,169 KB)

Original source

arXiv

https://arxiv.org/abs/2603.25746v1

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Self-Evolving AILive

How AIRA2 breaks AI research bottlenecks

While we've seen remarkable progress in AI for coding and mathematics, creating agents that can navigate the messy, open-ended nature of real research (where things break for no obvious reason) has proven far more challenging.

AI Accelerator Institute

1m28 minutes ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxQeVRmNi1ObGwyVWR3SURoakFvaDJ4bzBiaW9uRkZoVEp3WU96U05CeGd0LUF0SjhfMnRqajduN21hX1lXMVFHNmx5N0Z4Z1FmX0tpRHNqWVN0Wm9wQWFXeXgwRm9GTm1BRW1wSUR5WlptdF9tSGpWcktrb1NXMFRtMGRJaTNuYkk3ZFVTUF9nQ2ZHYUM0TWFaNDBiMG9NRFVGaFdHLUdiTkMxSldyaXBhZUI1V2wzc3BGZnlQVEgzTU1vMEoxcGtuOG9Zd0VkZW9zOXZXRWVKTGVIWUVEOEt5UVdFOUlWLUZ5ZFpYU3NqbUVUSVF3dXlIUkx0dl85cUM5cGVENS1jRS0wNGRkbTEyTXZUSmw1QTltdzR5ZlFnMV9XU3pueHF2TlJTZnhrSERmRmI5LWRtZFZyUzZXVnZVUDNWNzA1a3ctMEZ1THR1clRyT1Ywd2daTDlVS1RJdkxXdkNEdWtMbi1HRXVVRllqWVBEMFpHWXV4MzE1QXpqZnhfSFVEalkzUjVxd2dDWXltUlBhdWl6UXo5cVdxYU1OZ3JrNUNUcTBycjNBVFhFVHVoc1M0ZUFlNA?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 20 hours ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNWjFiT2ZQN1ZYQ1QxUUpxc2UzZjduQktEaW9DVjBib1hING0xOFptTUVBTUJQMFVVOUJ5eFptREliVkVtMXo1MlhwQy11c01YTlUwNWUzRjJ4dDM1T1hpOUdrcEdBR2czaDZvZ2V5Y1ZuRzFWSnlZQTNCOXR0d2ZXY015YjUya09FeXFHV2Fqd1htdlVwSDBBOWZhcTZmSmpfTjY3TVdfTWllV1RDQUd3a0dCT0NVUmdNSnF1Q2trM2xLdWdhcGx0aC1KRHMtcGJkSGFmTjZaNDNYZVQ5NnFpTk9wY1NkRkItRWZBVWJPQVdLcDhhYUdQaE1DMFdWbkp4VDd6a3dkSHVpVmhLZmItaUJTcWhQTWMtWlhfamVYT1FBQnBDS1VpWDFZZ3hnaFN0Qy1Ha2tUS2V1ZFJDYS1HczZjWFRRNkI1SlNxVFFNYzVwS2JWaGNQT3JXanFUNXZrdUw1UnFmMVAzaHpyUTI5QlBMdVI5SlRnTjdqbVNKdExKWC1jdzdMQTVFQkFySmo3TjBNRVQ4dmREdHJkQVhqWE1hQm5JTXlSelV1Vkt4OWNDRk95RnJRRg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 20 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 188 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersFresh

AI Inspires New Research Topics In Materials Science - miragenews.com

<a href="https://news.google.com/rss/articles/CBMihwFBVV95cUxQRlVFdkRBaHRvYkJJdFRlMTZmajEzeFRPU0hGWWdfbi02V1FnTUdVQ2pmY2VZLUV2NlB4V3BFdEVlSVZkUlhRSTZaNWFKMmcyWXJYbnNqbUhMTmp0NnFtMEppOXlPZkJSNHJfck5VSEVYcmUtX1k2QkJlR1BvUEdTTkp3UmlYRkk?oc=5" target="_blank">AI Inspires New Research Topics In Materials Science</a> miragenews.com

Google News: Machine Learning

1mabout 5 hours ago

Research Papers

From brain scans to alloys: Teaching AI to make sense of complex research data - Penn State University

<a href="https://news.google.com/rss/articles/CBMiwAFBVV95cUxPZDFHdkptQ2VUM2hmWjhqQkxoRnBiTWoxMXRRR21MUG5TamdUMlFRWmhvYVNHaFVNREVKU3VmSnVOdDVZYnNLb2ppYXRVRTZmVFVMV1pLTlVhUm9ybTNZbGtvZTdIMnIyMHNpOEk5aU9TSmxxS2Y4V2MwazYwY3JlX1Axbk1nd3pfcWhFdUJaaDJWRXJaMFIyTTROcmFHeXI3ZzFudXJ2M1h6UHI1LW1Ca1dta2RkM3BiYndocGk3Yjg?oc=5" target="_blank">From brain scans to alloys: Teaching AI to make sense of complex research data</a> Penn State University

GNews AI materials

1m3 months ago

Research PapersFresh

Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work

arXiv:2505.24246v4 Announce Type: replace Abstract: As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection process

arXiv cs.HC

2mabout 10 hours ago

Research PapersFresh

Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability

arXiv:2505.01000v5 Announce Type: replace Abstract: Scheduling is a perennial-and often challenging-problem for many groups. Existing tools are mostly static, showing an identical set of choices to everyone, regardless of the current status of attendees' inputs and preferences. In this paper, we propose Togedule, an adaptive scheduling tool that uses large language models to dynamically adjust the pool of choices and their presentation format. With the initial prototype, we conducted a formative study (N=10) and identified the potential benefits and risks of such an adaptive scheduling tool. Then, after enhancing the system, we conducted two controlled experiments, one each for attendees and organizers (total N=66). For each experiment, we compared scheduling with verbal messages, shared c

arXiv cs.HC

2mabout 10 hours ago