[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83105-en":3,"doc-seo-83105-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83105,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Data Analysis in the Wild Benchmarking Large Language Models Against Real World Data Complexities","Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often miss real-world complexity. They emphasize fact retrieval from small tables and largely ignore problems from large multi-tabular datasets, the integration of external knowledge, and exploratory insight discovery. DataGovBench is introduced as a benchmark built from governmental open data to evaluate practical performance. It contains Table QA for decomposable table questions and Table Insight for open-ended expert-level analysis, supported by comprehensive experiments.","arXiv :2607 .06482v 1 [ cs .CL] 7 Jul 2026  \nData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data  \nComplexities  \nSo Hasegawa 1 , Shailaja Keyur Sampat 1 , Lei Liu 1 , and Wei-Peng Chen 1 Fujitsu Research of America, Santa Clara CA 95054, [USA](USA shasegawa@fujitsu.com)[ shasegawa@fujitsu.com](USA shasegawa@fujitsu.com)  \nAbstract. Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that evaluates the ability of models to generate expert-level findings through exploratory data analysis. Comprehensive experiments with state-of-the-art LLMs, both with and without agentic frameworks, reveal significant performance gaps across both tasks. These results suggest that current LLM-based systems remain far from satisfying the demands of real-world data analytics. DataGovBench provides a challenging benchmark for advancing research on LLMs capable of both answering analytical queries and discovering insights from data. Code and sample data are available at [https://github.com/SoHasegawa/datagovbench](https://github.com/SoHasegawa/datagovbench).  \n1 Introduction  \nThe ability to reason over structured data is a cornerstone of modern data science and a long-standing challenge in artificial intelligence. With the advent of Large Language Models (LLMs), we have witnessed a paradigm shift in how humans interact with complex information [17,25,32,8] . These models have led to the development of sophisticated agents designed to democratize data analysis, promising a future where any user can pose natural language questions to a dataset and receive accurate answers [9,24] . The ultimate vision is an autonomous system that not only retrieves information but also uncovers the knowledge hidden within raw data.  \nHowever, a significant gap persists between this vision and reality, as the benchmarks lack real-world complexity. While datasets like WikiTableQuestions [20] and Spider [33] propelled research in semantic parsing and text-to-SQL, their controlled environments use small-scale tables. They largely neglect practical challenges such as massive table scales, the need to merge multiple tables,  \n2 So Hasegawa et al.  \nDatasets Question Answers  \nFig. 1: DataGovBench evaluates LLMs and agents on two table reasoning tasks using large, multi-table datasets supplemented with metadata and external knowledge. (a) The Table QA task requires models to answer simple or decomposable questions with textual or visual answers. (b) In contrast, the Table Insight task challenges models to perform open-ended exploratory analysis, proactively generating a list of insights and a summary without a specific user query.  \nand the essential role of metadata and external knowledge. Beyond these data limitations, existing benchmarks have focused on direct fact retrieval [31,9,34] . Tasks like QA and text-to-SQL are about retrieving information to a query, while missing the capability of proactive insight discovery that data analysts exhibit. Consequently, this discovery-oriented skill remains largely unevaluated, as few benchmarks have formalized insight generation as a primary task [22,23] .  \nTo bridge these notable gaps in both data realism and task scope, we introduce DataGovBench, a comprehensive benchmark sourced from public repositories [like Data.gov](like Data.gov) [28]. Our benchmark features two complementary tasks: Table Question Answering (Table QA) and Table Insight Generation (Table In","cbCail18Nltab6v1","https://ap.wps.com/l/cbCail18Nltab6v1","pdf",1321145,3,1,29,"English","en",105,"# Introduction\n# DataGovBench Tasks\n## Table QA\n## Table Insight\n# Evaluation Findings\n# Dataset and Method Details","[{\"question\":\"Why do existing LLM data-analysis benchmarks fall short for real-world scenarios?\",\"answer\":\"They mainly test fact retrieval on small tables, and they do not account for the complexity of large multi-tabular datasets, the need to merge tables, external knowledge integration, and exploratory insight discovery.\"},{\"question\":\"What is DataGovBench, and what data source does it use?\",\"answer\":\"DataGovBench is a benchmark derived from public governmental open data, designed to evaluate LLMs in practical data analysis settings.\"},{\"question\":\"How do the two DataGovBench tasks differ?\",\"answer\":\"Table QA evaluates answering decomposable questions with textual or visual outputs, while Table Insight assesses open-ended exploratory analysis by generating expert-level findings and summaries without a specific user query.\"}]",1784185283,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"data-analysis-in-the-wild-benchmarking-large-language-models-against-real-world-data-complexities","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/data-analysis-in-the-wild-benchmarking-large-language-models-against-real-world-data-complexities/83105/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing LLM data-analysis benchmarks fall short for real-world scenarios?","Question",{"text":75,"@type":76},"They mainly test fact retrieval on small tables, and they do not account for the complexity of large multi-tabular datasets, the need to merge tables, external knowledge integration, and exploratory insight discovery.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is DataGovBench, and what data source does it use?",{"text":80,"@type":76},"DataGovBench is a benchmark derived from public governmental open data, designed to evaluate LLMs in practical data analysis settings.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the two DataGovBench tasks differ?",{"text":84,"@type":76},"Table QA evaluates answering decomposable questions with textual or visual outputs, while Table Insight assesses open-ended exploratory analysis by generating expert-level findings and summaries without a specific user query.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]