[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121267-en":3,"doc-seo-121267-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121267,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Text Serialization and Their Relationship with the Conventional Paradigms of Tabular Machine Learning - Abstract","Recent research examines how Language Models (LMs) can support feature representation and prediction in tabular machine learning through text serialization and supervised fine-tuning (SFT), yet reliability and practical applicability remain unclear. This study compares emerging LM-based approaches with conventional tabular paradigms, analyzing both data-level choices for serialized tabular representation and the effect of serialization plus LMs on classification under challenges such as class imbalance, distribution shift, biases, and high dimensionality. Results show current pre-trained models should not replace traditional methods.","Text Serialization and Their Relationship with the Conventional Paradigms of  \nTabular Machine Learning  \nKyoka Ono 1 2 * Simon A. Lee 3 *  \narXiv :2406 . 13846v1 [ cs .CL] 19 Jun 2024  \nAbstract  \nRecent research has explored how Language Models (LMs) can be used for feature representation and prediction in tabular machine learning tasks.  \nThis involves employing text serialization and supervised fine-tuning (SFT) techniques. Despite the simplicity of these techniques, significant gaps remain in our understanding of the applicability and reliability of LMs in this context. Our study assesses how emerging LM technologies compare with traditional paradigms in tabular machine learning and evaluates the feasibility of adopting similar approaches with these advanced technologies. At the data level, we investigate various methods of data representation and curation of serialized tabular data, exploring their impact on prediction performance. At the classification level, we examine whether text serialization combined with LMs enhances performance on tabular datasets (e.g. class imbalance, distribution shift, biases, and high dimensionality), and assess whether this method represents a state-ofthe-art (SOTA) approach for addressing tabular machine learning challenges. Our findings reveal current pre-trained models should not replace conventional approaches.  \n1. Introduction  \nIn the field of natural language processing (NLP), a paradigm shift has occurred, driven by the emergence of Language Models (LM) technologies rooted in the transformer architecture (Vaswani et al., 2017) . These advancements have led to immense progress across various domains  \n*Equal contribution 1Department of Statistics and Data Science, University of California, Los Angeles 2Department of Natural Sciences, International Christian University, Mitaka, Tokyo, Japan 3Department of Computational Medicine University of California, Los Angeles Los Angeles, California, USA 90095 . Correspondence to: Simon Lee \u003C[simonlee711@g.ucla.edu](simonlee711@g.ucla.edu) >.  \nProceedings of the 41 st International Conference on Machine Learning, AI4Science Workshop, Vienna, Austria. Copyright 2024 by the author(s) .  \nof machine learning (ML) and artificial intelligence (AI) . Leveraging sophisticated techniques such as transfer learning (Weiss et al., 2016) and attention mechanisms (Bahdanau et al., 2014), LMs have demonstrated exceptional capabilities in tasks encompassing language understanding (Devlin et al., 2018), translation (Lewis et al., 2019), and text generation (Radford et al., 2018), thereby significantly influencing applications within the field of NLP. However, researchers from various fields have discovered that these LMs are not limited to conventional tasks. Consequently, there has been a surge of research into other areas and domains, such as question-answering (Radford et al., 2019 ; Suet al., 2019) and mathematical reasoning (Trinh et al., 2024 ; Wang et al., 2023 ; Imani et al., 2023), among others.  \nTherefore, in this paper, we focus on the ability of LMs to solve tabular machine learning tasks as introduced by (Hegselmann et al., 2023 ; Sahakyan et al., 2021 ; Dinh et al., 2022 ; Fang et al., 2024) . These studies utilize text serialization—converting tabular data into natural language representations—combined with supervised fine-tuning (SFT) to evaluate LMs’ capability on supervised machine learning tasks. Yet, current papers do not explore whether this process or these LMs could represent a state-of-the-art (SOTA) approach in machine learning. This oversight is especially significant in light of previous assertions that gradient boosting methods outperform deep learning strategies (Grinsztajnet al., 2022) .  \nThese previous works also did not determine whether various data curation measures are required for obtaining accurate results and how to adequately handle the common data preparation practices commonly used in tabular machine learning (e.g. m","cbCaifLCNwOWmVt8","https://ap.wps.com/l/cbCaifLCNwOWmVt8","pdf",1039901,1,30,"English","en",105,"# Abstract\n# Introduction\n# Related Works\n## Text Serialization","[{\"question\":\"What is the main focus of the study?\",\"answer\":\"The study evaluates how text serialization combined with language models and supervised fine-tuning relates to conventional paradigms in tabular machine learning.\"},{\"question\":\"How does the research evaluate impact at the data level?\",\"answer\":\"It investigates methods for data representation and data curation for serialized tabular data, measuring how these choices affect prediction performance.\"},{\"question\":\"Do the findings suggest using current pre-trained language models in place of conventional tabular methods?\",\"answer\":\"No. The results reveal that current pre-trained models should not replace conventional approaches.\"}]","Text Serialization and Their Relationship with the Conventional Paradigms of Tabular Machine Learning - Abstract | PDF",1785734789,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"text-serialization-and-their-relationship-with-the-conventional-paradigms-of-tabular-machine-learning-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/text-serialization-and-their-relationship-with-the-conventional-paradigms-of-tabular-machine-learning-abstract/121267/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main focus of the study?","Question",{"text":75,"@type":76},"The study evaluates how text serialization combined with language models and supervised fine-tuning relates to conventional paradigms in tabular machine learning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the research evaluate impact at the data level?",{"text":80,"@type":76},"It investigates methods for data representation and data curation for serialized tabular data, measuring how these choices affect prediction performance.",{"name":82,"@type":73,"acceptedAnswer":83},"Do the findings suggest using current pre-trained language models in place of conventional tabular methods?",{"text":84,"@type":76},"No. The results reveal that current pre-trained models should not replace conventional approaches.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]