{"ID":2885067,"CreatedAt":"2026-06-01T04:54:23.091178241Z","UpdatedAt":"2026-06-01T04:54:23.091178241Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2508.05115","arxiv_id":"2508.05115","title":"RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer","abstract":"Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.","short_abstract":"Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computation...","url_abs":"https://arxiv.org/abs/2508.05115","url_pdf":"https://arxiv.org/pdf/2508.05115v2","authors":"[\"Fangyu Du\",\"Taiqing Li\",\"Qian Qiao\",\"Tan Yu\",\"Ziwei Zhang\",\"Dingcheng Zhen\",\"Xu Jia\",\"Yang Yang\",\"Shunshun Yin\",\"Siyuan Liu\"]","published":"2025-08-07T07:47:16Z","proceeding":"cs.GR","tasks":"[\"cs.GR\",\"cs.CV\",\"cs.SD\",\"eess.AS\"]","methods":"[\"Diffusion Model\",\"Transformer\"]","has_code":false}
