{"ID":23507417,"CreatedAt":"2026-09-18T02:21:44.056544415Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.20732","arxiv_id":"2609.20732","title":"Q\u0026A on Any Spreadsheet Requires Interpreting Its Grid Structure","abstract":"Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework beats the state of the art, yet it faces a hard ceiling. Spreadsheets are fundamentally two-dimensional unstructured data with continuous relationships and infinite potential cell roles. Because classification models are restricted to finite, pre-defined classes, they cannot perfectly capture this structural nuance, even with human-level annotation. We show that addressing the spreadsheet-to-LLM bottleneck requires moving beyond discrete cell classification. Instead, the field must develop dimensionality-reduction techniques to directly flatten 2D unstructured spreadsheets into 1D unstructured text. Text chunks would be easier for downstream RAG to interpret and generate from.","short_abstract":"Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework...","url_abs":"https://arxiv.org/abs/2609.20732","url_pdf":"https://arxiv.org/pdf/2609.20732v1","authors":"[\"Zofia Smoleń\"]","published":"2026-09-17T17:22:43Z","proceeding":"cs.AI","tasks":"[\"cs.AI\",\"cs.SE\"]","methods":"[\"Large Language Model\"]","has_code":false}
