{"ID":23507374,"CreatedAt":"2026-09-18T02:21:44.056544415Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.20625","arxiv_id":"2609.20625","title":"Chronicle: Cut-Point Replay for Regression Testing of LLM Agents","abstract":"Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing agent tooling records runs only to trace or score them, not to test a code change against them. We present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, serves a chosen subset of boundaries from the record and executes the complementary subset live with new code, turning a recorded incident into a regression test that runs in continuous integration. On a benchmark of 6 recorded failures with simulated model boundaries, recording adds 23 μs per crossing (0.008% of an assumed 300 ms model call), full replay issues zero model calls and is bit-stable across 20 repetitions, and cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents. In a mutation study of the guarded tools, cut-point tests catch every mutant that lets the recorded unsafe action through, while a baseline that stubs every boundary, using the same assertion, catches none. Chronicle and the benchmark are publicly available at https://github.com/theagentplane/chronicle.","short_abstract":"Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing...","url_abs":"https://arxiv.org/abs/2609.20625","url_pdf":"https://arxiv.org/pdf/2609.20625v1","authors":"[\"Tisha Chawla\",\"Susheem Koul\"]","published":"2026-09-17T16:09:57Z","proceeding":"cs.CL","tasks":"[\"cs.CL\",\"cs.AI\"]","methods":"[\"Large Language Model\",\"Language Model\"]","has_code":false,"code_links":[{"ID":639818,"CreatedAt":"2026-09-18T02:21:44.056544415Z","UpdatedAt":"2026-09-18T02:21:44.056544415Z","DeletedAt":null,"paper_id":23507374,"paper_url":"https://arxiv.org/abs/2609.20625","paper_title":"Chronicle: Cut-Point Replay for Regression Testing of LLM Agents","repo_url":"https://github.com/theagentplane/chronicle","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"github_stars":0}]}
