{"ID":23507429,"CreatedAt":"2026-09-18T02:21:44.056544415Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.20756","arxiv_id":"2609.20756","title":"OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher","abstract":"As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\\times$ and 9.5$\\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: https://01dami23.github.io/opted/","short_abstract":"As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop de...","url_abs":"https://arxiv.org/abs/2609.20756","url_pdf":"https://arxiv.org/pdf/2609.20756v1","authors":"[\"Damiano Da Col\",\"Maximilian Igl\",\"Peter Karkus\",\"Kashyap Chitta\",\"Boris Ivanovic\",\"Marco Pavone\",\"Konrad Schindler\",\"Christos Sakaridis\"]","published":"2026-09-17T17:41:51Z","proceeding":"cs.RO","tasks":"[\"cs.RO\",\"cs.CV\",\"cs.LG\"]","methods":"[\"Reinforcement Learning\"]","has_code":false}
