[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121450-en":3,"doc-seo-121450-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121450,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Performance Benchmarking of Traditional Machine Learning and Transformer Models for Multi-Class Text - Classification","Text classification is a core NLP task powering spam detection, sentiment analysis, and content organization. This study compares traditional machine learning methods—Random Forest, XGBoost, Support Vector Machine, and Naive Bayes—against a custom transformer trained from scratch and transformer-based transfer learning using pretrained checkpoints (BERT, DistilBERT, RoBERTa, ELECTRA). Traditional baselines reach up to 90.47% accuracy, while scratch transformers achieve about 91%. Transfer learning delivers the best results, with RoBERTa at 94.54% and BERT-family models exceeding 93%, underscoring contextual embeddings and large-scale pretraining.","Performance Benchmarking of Traditional Machine Learning and Transformer Models for Multi-Class Text  \nClassification  \nOmar El Khatiba *, Nabeel Alkhatibb  \na, bMath. and Computer Science Dept., Loyola University New Orleans, New Orleans, 70118 USA  \naEmail: [oelkhat@loyno.edu](oelkhat@loyno.edu)  \nbEmail: [nalkhati@my.loyno.edu](nalkhati@my.loyno.edu)  \nAbstract  \nText classification is a fundamental task in natural language processing (NLP), widely applied in areas such as spam detection, sentiment analysis, and text categorization. This study presents a comparative analysis of three distinct machine learning paradigms—traditional machine learning algorithms (like Random Forest, XGBoost, support vector machine and Naive Bayes), a custom-built transformer architecture, and transfer learning or pretrained transformer models (BERT, DistilBERT, RoBERTa, ELECTRA)—on the multi-class news classification dataset. While traditional models provided competitive baselines with up to 90.47% accuracy, modern transformer architecture surpassed them, achieving 91% accuracy when trained from scratch. The highest performance was observed with transfer learning using pre-trained models, where RoBERTa achieved 94.54% accuracy, DistillBERT achieved 94.32% accuracy, BERT achieved 94.07% accuracy and ELECTRA achieved 93.66% . These findings highlight the significance of contextual embeddings and large-scale pretraining in advancing text classification performance.  \nKeywords: NLP Multi-Class Classification; Transformer; NLP Transfer Learning; Text Classification.  \n1. Introduction  \nText classification is a fundamental task in natural language processing (NLP) with broad applications ranging from sentiment analysis and spam detection to information retrieval and text categorization. Accurate news classification, in particular, plays a critical role in organizing digital content, enabling personalized news feeds, and filtering misinformation.  \nReceived: 5/15/2025  \nAccepted: 7/1/2025  \nPublished: 7/14/2025  \n* Corresponding author.  \nTraditional machine learning (ML) models such as Random Forest, XGBoost, Support Vector Machine and Naïve Bayes, when combined with handcrafted features like bag-of-words or TF-IDF, have proven computationally efficient and interpretable. However, these models often struggle to capture long-range dependencies and semantic nuances in textual data.  \nRecent advances in deep learning, particularly transformer-based architecture, have significantly improved the ability of models to understand contextual meaning through mechanisms such as self-attention. Furthermore, transfer learning via pre-trained language models like BERT, DistilBERT, RoBERTa, ELECTRA … etc have enabled state-of-the-art performance on a wide range of NLP tasks with minimal labeled data.  \nThis study conducts an empirical investigation to compare the effectiveness of traditional machine learning classifiers and transformer-based models—both trained from scratch and fine-tuned from pre-trained checkpoints—on the AG News dataset [1] . The objective is to evaluate the relative strengths and weaknesses of these approaches in terms of accuracy, training time, and resource efficiency. A unified experimental framework for comparing traditional machine learning, scratch-trained transformers, and transfer learned transformers on a common benchmark.  \nThe remainder of this paper is organized as follows: Section 2 presents a literature review of traditional and transformer-based models. Section 3 describes methodology and experimental setup. Section 4 provides evaluation results and visualizations. Section 5 offers a discussion of the findings, and Section 6 concludes the paper with future research directions.  \n2. Literature Review  \nText classification has evolved significantly over the past two decades, transitioning from traditional statistical models to neural networks [2] and, more recently, transformer-based architectures [3] . This section provides an ove","cbCaitjYRUgcfzJW","https://ap.wps.com/l/cbCaitjYRUgcfzJW","pdf",1401756,1,15,"English","en",105,"# Introduction\n## Traditional Machine Learning for Text Classification\n# Literature Review\n## Traditional Machine Learning for Text Classification","[{\"question\":\"Which models are compared for multi-class news text classification?\",\"answer\":\"The study compares traditional ML classifiers (Random Forest, XGBoost, Support Vector Machine, Naive Bayes) with a transformer trained from scratch and transfer learning models using pretrained checkpoints such as BERT, DistilBERT, RoBERTa, and ELECTRA.\"},{\"question\":\"How do traditional machine learning models perform compared with transformer models?\",\"answer\":\"Traditional models reach up to 90.47% accuracy, while a transformer trained from scratch surpasses them, achieving around 91% accuracy.\"},{\"question\":\"Which approach delivers the highest accuracy and what does it imply?\",\"answer\":\"Transfer learning with pretrained transformers delivers the top performance, with RoBERTa at 94.54% accuracy and other models above 93%. Results highlight the importance of contextual embeddings and large-scale pretraining for improved classification performance.\"}]","Performance Benchmarking of Traditional Machine Learning and Transformer Models for Multi-Class Text - Classification | PDF",1785735716,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"performance-benchmarking-of-traditional-machine-learning-and-transformer-models-for-multi-class-text-classification","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/performance-benchmarking-of-traditional-machine-learning-and-transformer-models-for-multi-class-text-classification/121450/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which models are compared for multi-class news text classification?","Question",{"text":75,"@type":76},"The study compares traditional ML classifiers (Random Forest, XGBoost, Support Vector Machine, Naive Bayes) with a transformer trained from scratch and transfer learning models using pretrained checkpoints such as BERT, DistilBERT, RoBERTa, and ELECTRA.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do traditional machine learning models perform compared with transformer models?",{"text":80,"@type":76},"Traditional models reach up to 90.47% accuracy, while a transformer trained from scratch surpasses them, achieving around 91% accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"Which approach delivers the highest accuracy and what does it imply?",{"text":84,"@type":76},"Transfer learning with pretrained transformers delivers the top performance, with RoBERTa at 94.54% accuracy and other models above 93%. Results highlight the importance of contextual embeddings and large-scale pretraining for improved classification performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]