{"ID":23475006,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.19688","arxiv_id":"2609.19688","title":"LYRIC: Language-Driven Physics-Based Character Control for Contact-Rich Whole-Body Object Interaction","abstract":"We present LYRIC, a generative flow-matching controller for language-driven physics-based contact-rich interaction control, that enables simulated characters to perform contact-rich whole-body object interactions from a free-form language instruction and a sparse terminal object goal. To obtain reliable expert trajectories from imperfect motion-capture references, a single tracking policy is trained using geometry-conditioned interaction rewards and relaxed reference tracking near hand-object contact. To guide interaction progress without prescribing a full-body kinematic reference, we factorize the controller into a task-level planner that predicts short-horizon object and humanoid-root trajectories, and an action generator that resolves whole-body motion and contacts in closed loop. After behavior cloning, we freeze the planner and post-tune the action generator on policy using the planner's predictions as stable supervision for intermediate task progression. In a controlled OMOMO evaluation, our tracker achieves 64.3% success compared with 53.2% for an InterMimic reimplementation, while a unified policy achieves 76.5% on the full OMOMO dataset. On the held-out split, LYRIC achieves 90.3% task success, compared with 74.2% for the strongest matched kinematic-planner baseline, with better semantic alignment and motion quality. Without retraining, the controller also supports test-time object-waypoint guidance. Qualitative results further demonstrate robust, natural contact-rich interactions and zero-shot transfer to novel object shapes. The webpage is available at https://neu-vi.github.io/LYRIC/","short_abstract":"We present LYRIC, a generative flow-matching controller for language-driven physics-based contact-rich interaction control, that enables simulated characters to perform contact-rich whole-body object interactions from a free-form language instruction and a sparse terminal object goal. To obtain reliable expert trajecto...","url_abs":"https://arxiv.org/abs/2609.19688","url_pdf":"https://arxiv.org/pdf/2609.19688v1","authors":"[\"Zeyu Han\",\"Zichong Meng\",\"Julian Tanke\",\"Minami Matsumoto\",\"Sergey Bashkirov\",\"Yingruo Fan\",\"Selim Engin\",\"Dongseok Shim\",\"Takashi Shibuya\",\"Yuki Mitsufuji\",\"Huaizu Jiang\"]","published":"2026-09-17T04:34:15Z","proceeding":"cs.RO","tasks":"[\"cs.RO\",\"cs.GR\"]","methods":"[]","has_code":false}
