{"ID":23475145,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.19936","arxiv_id":"2609.19936","title":"VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes","abstract":"Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scaling as a Markov Decision Process and trains an agent on 3353 real Prometheus traces. Evaluated on a live Google Kubernetes Engine cluster using LAMMPS, graph analytics, in-memory analytics, and MLPerf 3D-UNet, the RL agent reclaims 31.6% of the available memory headroom and incurs at most one OOM event while VPA reclaims -7.9% over the same runs, raising memory provisioning, and its recommendation would have been insufficient to avoid OOM in 30 runs. The results demonstrate that an observation-driven RL recommender could outperform retrospective heuristics for dynamic memory scaling.","short_abstract":"Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scali...","url_abs":"https://arxiv.org/abs/2609.19936","url_pdf":"https://arxiv.org/pdf/2609.19936v1","authors":"[\"Ade Pramono\",\"Jie Ren\",\"Ivy Peng\"]","published":"2026-09-17T09:10:15Z","proceeding":"cs.DC","tasks":"[\"cs.DC\"]","methods":"[\"Reinforcement Learning\"]","has_code":false}
