[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-135057-en":3,"doc-seo-135057-105":31,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},135057,8796093062539,"8796093062539","",8,"Research & Report","The Illusion of Generalization in Tabular Language Models - arXiv:2602.04031v2 - systematic reevaluation and findings","Tabular Language Models (TLMs) are often claimed to generalize strongly for tabular prediction, yet evidence may be distorted by evaluation design. This study systematically reevaluates Tabula-8B using 165 datasets from the UniPredict benchmark and reports three key results: weak lift for binary/categorical tasks, reliance on quartile classification for aggregate gains, widespread train-test contamination and task-level leakage, and substantial recovery under instruction-tuning without tabular exposure, narrowed by format familiarity. The results indicate generalization artifacts and motivate stronger TLM evaluation.","The Illusion of Generalization in Tabular Language Models  \nAditya Gorla 1 Ratish Puduppully 2  \narXiv :2602 .04031v2 [ cs .LG] 29 May 2026  \nAbstract  \nTabular Language Models (TLMs) have been claimed to achieve strong generalization for tabular prediction. We conduct a systematic reevaluation of Tabula-8B as a representative TLM, utilizing 165 datasets from the UniPredict benchmark. Our investigation reveals three findings.  \nFirst, binary and categorical classification achieve near-zero median lift over majority-class baselines and strong aggregate performance is driven entirely by quartile classification tasks. Second, top-performing datasets exhibit pervasive contamination, including complete train-test overlap and task-level leakage that evades standard deduplication. Third, instruction-tuning without tabular exposure recovers 92.2% of standard classification performance and on quartile classification, format familiarity closes 71.3% of the gap with the residual attributable to contaminated datasets. These findings suggest claimed generalization likely reflects evaluation artifacts rather than learned tabular reasoning. We conclude with recommendations for strengthening TLM evaluation.1  \n1. Introduction  \nLarge Language Models (LLMs) have transformed natural language processing (Zhao et al., 2025 ; Minaee et al., 2025), and their success has inspired a new wave of methods for tabular data (Fang et al., 2024 ; Gardner et al., 2024 ; Hegselmann et al., 2023) . Beyond naive approaches such as text serialization (Dinh et al., 2022 ; Lee et al., 2025), a new class of language models specifically (pre-)trained or fine-tuned for tabular domain have emerged (Fang et al., 2024 ; Hegselmann et al., 2023 ; Gardner et al., 2024 ; Wang et al., 2023 ; Sun et al., 2024) . We refer to these as Tabular Language  \n1University of California, Los Angeles, USA 2IT University of Copenhagen, Denmark. Correspondence to: Aditya Gorla \u003C[adityagorla@ucla.edu](adityagorla@ucla.edu) >, Ratish Puduppully \u003C[rapu@itu.dk](rapu@itu.dk)>.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n1Code and artifacts are available at [https://github](https://github) . com/ratishsp/tlm-illusion.  \nModels (TLMs) . The putative hypothesis behind TLMs mirrors that of LLMs and Vision-Language Models (VLMs): with sufficient scale, these models may learn to generalize across the structure, invariances, and patterns inherent in tabular data, enabling zero-and few-shot prediction, imputation, and synthetic generation across binary, categorical, and continuous data types (Hegselmann et al., 2023 ; Wanget al., 2023 ; Gorla et al., 2025 ; Borisov et al., 2023) .  \nTabular data, however, possesses three properties that jointly distinguish it from text and images (Fang et al., 2024 ; Borisov et al., 2024 ; Grinsztajn et al., 2022) . First, tabular data is row-permutation invariant: the ordering of samples carries no semantic meaning. Second, and more critically, tabular data is column-permutation invariant. Unlike text (which has sequential structure) or images (which have spatial locality), features in a table have no a priori or universal ordering2. Third, tabular data is heterogeneous, spanning binary, categorical, and continuous types. This is in stark contrast to text (discrete tokens from a fixed vocabulary) or images (bounded pixel intensities) . One can reasonably argue these properties explain why gradient-boosted trees (GBTs) and Prior-Fitted Networks (PFNs) have excelled on tabular tasks (Grinsztajn et al., 2022 ; Shwartz-Ziv & Armon, 2022 ; Hollmann et al., 2023 ; 2025) . GBTs naturally handle irregular, discontinuous patterns without assuming feature relationships (Chen & Guestrin, 2016 ; Grinsztajn et al., 2022), while PFNs encode inductive biases through synthetic (heterogeneous) data priors and leverage transformer architectures without fixed positional encodings to r","cbCaici8dMYTpw28","https://ap.wps.com/l/cbCaici8dMYTpw28","pdf",560740,3,1,28,"English","en",105,"# Abstract\n# Introduction\n## Tabular data properties and inductive biases\n## Prior work and the central question\n## Empirical verification approach\n## Contributions and key findings","[{\"question\":\"What question does the paper investigate about tabular language models?\",\"answer\":\"It asks whether TLMs truly generalize to tabular data, and if so, what mechanism enables that generalization.\"},{\"question\":\"What three main findings does the study report from reevaluating Tabula-8B?\",\"answer\":\"The paper finds near-zero gains for binary/categorical tasks with aggregate performance driven by quartile tasks, pervasive contamination including train-test overlap and task-level leakage, and that instruction-tuning without tabular exposure recovers most standard performance while format familiarity closes much of the remaining gap.\"},{\"question\":\"How do the authors interpret the claimed generalization of TLMs?\",\"answer\":\"They conclude that the reported generalization likely reflects evaluation artifacts rather than learned tabular reasoning, and they recommend improving TLM evaluation practices.\"}]","The Illusion of Generalization in Tabular Language Models - arXiv:2602.04031v2 - systematic reevaluation and findings | PDF",1787303980,71,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":10,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":29},"the-illusion-of-generalization-in-tabular-language-models-arxiv260204031v2-systematic-reevaluation-and-findings",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/the-illusion-of-generalization-in-tabular-language-models-arxiv260204031v2-systematic-reevaluation-and-findings/135057/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-26","2026-08-21",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper investigate about tabular language models?","Question",{"text":75,"@type":76},"It asks whether TLMs truly generalize to tabular data, and if so, what mechanism enables that generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What three main findings does the study report from reevaluating Tabula-8B?",{"text":80,"@type":76},"The paper finds near-zero gains for binary/categorical tasks with aggregate performance driven by quartile tasks, pervasive contamination including train-test overlap and task-level leakage, and that instruction-tuning without tabular exposure recovers most standard performance while format familiarity closes much of the remaining gap.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the authors interpret the claimed generalization of TLMs?",{"text":84,"@type":76},"They conclude that the reported generalization likely reflects evaluation artifacts rather than learned tabular reasoning, and they recommend improving TLM evaluation practices.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]