{"ID":22918785,"CreatedAt":"2026-09-17T01:02:08.507062015Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.17842","arxiv_id":"2609.17842","title":"Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents","abstract":"Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluating these multimodal outputs is challenging: curated reference benchmarks are costly to author, cannot comprehensively capture the space of valid responses, and are unavailable in production. Building on the Lexara evaluation framework, we introduce Lexara-RF, a reference-free set of metrics that scores CVA outputs using only the prompt, data, and model response. We reformulate evaluation as verification: 13 metrics operationalize visualization design theory and Gricean cooperative principles as computable consistency, intent-alignment, and design validity checks. On a human-rated corpus of CVA test-cases, Lexara-RF achieves alignment comparable to reference-based formulations, outperforms surface-similarity NLG baselines, and localizes structurally grounded failures with high accuracy.","short_abstract":"Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluating these multimodal outputs is challenging: curated reference benchmarks are costly to author, cannot comprehensively capture the space of valid respon...","url_abs":"https://arxiv.org/abs/2609.17842","url_pdf":"https://arxiv.org/pdf/2609.17842v1","authors":"[\"Srishti Palani\",\"Vidya Setlur\"]","published":"2026-09-15T21:02:24Z","proceeding":"cs.HC","tasks":"[\"cs.HC\",\"cs.AI\"]","methods":"[\"Language Model\"]","has_code":false}
