{"ID":23475078,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.19830","arxiv_id":"2609.19830","title":"Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization","abstract":"Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.","short_abstract":"Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggrega...","url_abs":"https://arxiv.org/abs/2609.19830","url_pdf":"https://arxiv.org/pdf/2609.19830v1","authors":"[\"Yingxuan Zhuang\",\"Binhe Yu\",\"Jingxiao Yang\",\"Ruopei Sun\",\"Ziting Li\",\"Cheng Tan\",\"Xuhong Zhang\",\"Jianwei Yin\",\"Jintao Chen\"]","published":"2026-09-17T07:35:02Z","proceeding":"cs.AI","tasks":"[\"cs.AI\",\"stat.ML\"]","methods":"[\"Reinforcement Learning\",\"Large Language Model\"]","has_code":false}
