Research Papers research paper arxiv computer-vision image-recognition

FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation

arXivMarch 31, 20262 min read0 views

arXiv:2603.27915v1 Announce Type: new Abstract: Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often rely on complex intermediate representations, which limits their flexibility and efficiency. In this work, we propose a novel pose-free framework for real-time sign language video generation. Our method eliminates the need for intermediate pose representations by directly mapping natural language text to sign language videos using a diffusion-based approach. We introduce — Liuzhou Zhang, Zeyu Zhang, Biao Wu, Luyao Tang, Zirui Song, Hongyang He, Renda Han, Guangzhen Yao, Huacan Wang, Ronghao Chen, Xiuying Chen, Guan Huang, Zheng Zhu

Authors:Liuzhou Zhang, Zeyu Zhang, Biao Wu, Luyao Tang, Zirui Song, Hongyang He, Renda Han, Guangzhen Yao, Huacan Wang, Ronghao Chen, Xiuying Chen, Guan Huang, Zheng Zhu

View PDF HTML (experimental)

Abstract:Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often rely on complex intermediate representations, which limits their flexibility and efficiency. In this work, we propose a novel pose-free framework for real-time sign language video generation. Our method eliminates the need for intermediate pose representations by directly mapping natural language text to sign language videos using a diffusion-based approach. We introduce two key innovations: (1) a pose-free generative model based on the a state-of-the-art diffusion backbone, which learns implicit text-to-gesture alignments without pose estimation, and (2) a Trainable Sliding Tile Attention (T-STA) mechanism that accelerates inference by exploiting spatio-temporal locality patterns. Unlike previous training-free sparsity approaches, T-STA integrates trainable sparsity into both training and inference, ensuring consistency and eliminating the train-test gap. This approach significantly reduces computational overhead while maintaining high generation quality, making real-time deployment feasible. Our method increases video generation speed by 3.07x without compromising video quality. Our contributions open new avenues for real-time, high-quality, pose-free sign language synthesis, with potential applications in inclusive communication tools for diverse communities. Code: this https URL.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2603.27915 [cs.CV]

(or arXiv:2603.27915v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2603.27915

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zeyu Zhang [view email] [v1] Mon, 30 Mar 2026 00:06:26 UTC (400 KB)

Original source

arXiv

https://arxiv.org/abs/2603.27915

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Self-Evolving AILive

How AIRA2 breaks AI research bottlenecks

While we've seen remarkable progress in AI for coding and mathematics, creating agents that can navigate the messy, open-ended nature of real research (where things break for no obvious reason) has proven far more challenging.

AI Accelerator Institute

1m26 minutes ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxQeVRmNi1ObGwyVWR3SURoakFvaDJ4bzBiaW9uRkZoVEp3WU96U05CeGd0LUF0SjhfMnRqajduN21hX1lXMVFHNmx5N0Z4Z1FmX0tpRHNqWVN0Wm9wQWFXeXgwRm9GTm1BRW1wSUR5WlptdF9tSGpWcktrb1NXMFRtMGRJaTNuYkk3ZFVTUF9nQ2ZHYUM0TWFaNDBiMG9NRFVGaFdHLUdiTkMxSldyaXBhZUI1V2wzc3BGZnlQVEgzTU1vMEoxcGtuOG9Zd0VkZW9zOXZXRWVKTGVIWUVEOEt5UVdFOUlWLUZ5ZFpYU3NqbUVUSVF3dXlIUkx0dl85cUM5cGVENS1jRS0wNGRkbTEyTXZUSmw1QTltdzR5ZlFnMV9XU3pueHF2TlJTZnhrSERmRmI5LWRtZFZyUzZXVnZVUDNWNzA1a3ctMEZ1THR1clRyT1Ywd2daTDlVS1RJdkxXdkNEdWtMbi1HRXVVRllqWVBEMFpHWXV4MzE1QXpqZnhfSFVEalkzUjVxd2dDWXltUlBhdWl6UXo5cVdxYU1OZ3JrNUNUcTBycjNBVFhFVHVoc1M0ZUFlNA?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 20 hours ago

ModelsRecent

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNWjFiT2ZQN1ZYQ1QxUUpxc2UzZjduQktEaW9DVjBib1hING0xOFptTUVBTUJQMFVVOUJ5eFptREliVkVtMXo1MlhwQy11c01YTlUwNWUzRjJ4dDM1T1hpOUdrcEdBR2czaDZvZ2V5Y1ZuRzFWSnlZQTNCOXR0d2ZXY015YjUya09FeXFHV2Fqd1htdlVwSDBBOWZhcTZmSmpfTjY3TVdfTWllV1RDQUd3a0dCT0NVUmdNSnF1Q2trM2xLdWdhcGx0aC1KRHMtcGJkSGFmTjZaNDNYZVQ5NnFpTk9wY1NkRkItRWZBVWJPQVdLcDhhYUdQaE1DMFdWbkp4VDd6a3dkSHVpVmhLZmItaUJTcWhQTWMtWlhfamVYT1FBQnBDS1VpWDFZZ3hnaFN0Qy1Ha2tUS2V1ZFJDYS1HczZjWFRRNkI1SlNxVFFNYzVwS2JWaGNQT3JXanFUNXZrdUw1UnFmMVAzaHpyUTI5QlBMdVI5SlRnTjdqbVNKdExKWC1jdzdMQTVFQkFySmo3TjBNRVQ4dmREdHJkQVhqWE1hQm5JTXlSelV1Vkt4OWNDRk95RnJRRg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 20 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 188 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersFresh

AI Inspires New Research Topics In Materials Science - miragenews.com

<a href="https://news.google.com/rss/articles/CBMihwFBVV95cUxQRlVFdkRBaHRvYkJJdFRlMTZmajEzeFRPU0hGWWdfbi02V1FnTUdVQ2pmY2VZLUV2NlB4V3BFdEVlSVZkUlhRSTZaNWFKMmcyWXJYbnNqbUhMTmp0NnFtMEppOXlPZkJSNHJfck5VSEVYcmUtX1k2QkJlR1BvUEdTTkp3UmlYRkk?oc=5" target="_blank">AI Inspires New Research Topics In Materials Science</a> miragenews.com

Google News: Machine Learning

1mabout 5 hours ago

Research Papers

From brain scans to alloys: Teaching AI to make sense of complex research data - Penn State University

<a href="https://news.google.com/rss/articles/CBMiwAFBVV95cUxPZDFHdkptQ2VUM2hmWjhqQkxoRnBiTWoxMXRRR21MUG5TamdUMlFRWmhvYVNHaFVNREVKU3VmSnVOdDVZYnNLb2ppYXRVRTZmVFVMV1pLTlVhUm9ybTNZbGtvZTdIMnIyMHNpOEk5aU9TSmxxS2Y4V2MwazYwY3JlX1Axbk1nd3pfcWhFdUJaaDJWRXJaMFIyTTROcmFHeXI3ZzFudXJ2M1h6UHI1LW1Ca1dta2RkM3BiYndocGk3Yjg?oc=5" target="_blank">From brain scans to alloys: Teaching AI to make sense of complex research data</a> Penn State University

GNews AI materials

1m3 months ago

Research PapersFresh

Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work

arXiv:2505.24246v4 Announce Type: replace Abstract: As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection process

arXiv cs.HC

2mabout 10 hours ago

Research PapersFresh

Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability

arXiv:2505.01000v5 Announce Type: replace Abstract: Scheduling is a perennial-and often challenging-problem for many groups. Existing tools are mostly static, showing an identical set of choices to everyone, regardless of the current status of attendees' inputs and preferences. In this paper, we propose Togedule, an adaptive scheduling tool that uses large language models to dynamically adjust the pool of choices and their presentation format. With the initial prototype, we conducted a formative study (N=10) and identified the potential benefits and risks of such an adaptive scheduling tool. Then, after enhancing the system, we conducted two controlled experiments, one each for attendees and organizers (total N=66). For each experiment, we compared scheduling with verbal messages, shared c

arXiv cs.HC

2mabout 10 hours ago