{"ID":2830332,"CreatedAt":"2026-06-01T04:54:23.091178241Z","UpdatedAt":"2026-06-01T04:54:23.091178241Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2512.10778","arxiv_id":"2512.10778","title":"Building Audio-Visual Digital Twins with Smartphones","abstract":"Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that constructs editable audio-visual digital twins using only commodity smartphones. AV-Twin combines mobile RIR capture and a visual-assisted acoustic field model to efficiently reconstruct room acoustics. It further recovers per-surface material properties through differentiable acoustic rendering, enabling users to modify materials, geometry, and layout while automatically updating both audio and visuals. Together, these capabilities establish a practical path toward fully modifiable audio-visual digital twins for real-world environments.","short_abstract":"Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that constructs editable audio-visual digital twins using only commodity smartphones. AV-Twin combines mobile RIR capture and a visual-assisted acou...","url_abs":"https://arxiv.org/abs/2512.10778","url_pdf":"https://arxiv.org/pdf/2512.10778v1","authors":"[\"Zitong Lan\",\"Yiwei Tang\",\"Yuhan Wang\",\"Haowen Lai\",\"Yiduo Hao\",\"Mingmin Zhao\"]","published":"2025-12-11T16:14:32Z","proceeding":"cs.SD","tasks":"[\"cs.SD\",\"cs.MM\",\"eess.AS\"]","methods":"[]","has_code":false}
