{"ID":23475237,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.20110","arxiv_id":"2609.20110","title":"Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents","abstract":"Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of \u003c10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.","short_abstract":"Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability f...","url_abs":"https://arxiv.org/abs/2609.20110","url_pdf":"https://arxiv.org/pdf/2609.20110v1","authors":"[\"Yichao Jin\",\"Yushuo Wang\",\"Yuxuan Han\",\"Kwan Ching Yee Sonia\",\"Weiyang Song\",\"Chiu Jin-Chun Kent\",\"Wong Chong Hwee\",\"Wong Tiong Kiat\",\"Kenneth Zhu Ke\",\"Jingyuan Zhao\"]","published":"2026-09-17T12:09:21Z","proceeding":"cs.AI","tasks":"[\"cs.AI\"]","methods":"[\"Language Model\"]","has_code":false}
