[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82902-en":3,"doc-seo-82902-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82902,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Choosing a Parallel Heterogeneous Ensemble Method for Tabular Classification","Parallel ensemble methods are evaluated on 56 small-to-medium tabular classification tasks from OpenML CC18, leading to a set of “best practice” recommendations for using ensembles effectively. The recommendations are then validated on 28 additional tasks with TabArena’s precomputed data, where the guidance significantly outperforms Single Best and matches or exceeds multiple individual ensemble approaches. Key findings include independent inconsistencies for Blending versus Stacking, and strong probabilistic performance of Robust Soft Voting, especially for multiclass problems.","arXiv :2607 .05 103v 1 [ cs .LG] 6 Jul 2026  \nChoosing a parallel heterogeneous ensemble method for tabular classiﬁcation  \nVassili Maillet1 ,2  \n[0009−0001−7827−0989]  \nGustavo (Jesús) Angulo 1  \n[0000−0001−6106−2692]  \nPierre Jouvelot1  \n[0000−0002−6783−5796]  \n1 Mines Paris, PSL University, France, [vassili.maillet@minesparis.psl.eu](vassili.maillet@minesparis.psl.eu)  \n2 Alten Labs, France  \nAbstract. Parallel ensemble methods were compared on 56 small-tomedium tabular classiﬁcation tasks drawn from OpenML CC18 . A set of “best practice” recommendations on the use of ensemble methods was derived from these observations. It was later validated on 28 additional tasks using TabArena’s precomputed data, where the recommendation set signiﬁcantly outperformed Single Best and matched or exceeded individual ensemble methods.  \nTwo key observations were made. First, Blending and Stacking are inconsistent, but their inconsistencies are independent and happen on diﬀerent tasks. Second, while Hard Voting’s probabilistic classiﬁcation is rather weak, a consequence of using vote proportions as posterior estimates, Robust Soft Voting’s probabilistic classiﬁcation is particularly successful, especially in the multiclass case.  \nKeywords: Ensemble Learning · Tabular Classiﬁcation · Benchmark  \n1 Introduction  \nEnsemble methods improve the performance of individual models, called base learners, by combining them. The set of base learners of an ensemble model, its pool, is homogeneous if every base learner is of the same family, and heterogeneous otherwise. Homogeneous tree-based ensemble models have been considered some of the best performing models in tabular classiﬁcation for years [15] . If much has been written about their performance, the capabilities of heterogeneous ensemble models’ have been less studied (see Section 2), despite their prevalence in machine-learning competitions such as on Kaggle.  \nIn this work, we provide the following contributions:  \n– a thorough experimental study of the performances of 9 diﬀerent parallel ensemble methods using a diverse and heterogeneous pool;  \n2 V. Maillet et al.  \n– a new set of “best practice” recommendations on the use of such ensemble methods, for small to large datasets;  \n– a validation of these recommendations using TabArena’s precomputed data [12];  \n– and two key observations, namely (1) that Blending and Stacking are inconsistent, but their inconsistencies are independent and happen on diﬀerent tasks and (2), while Hard Voting’s probabilistic classiﬁcation is rather weak, Robust Soft Voting’s own probabilistic classiﬁcation is particularly successful, especially in the multiclass case.  \nThe structure of the paper is as follows. After surveying the related work (Section 2), we evaluate ensemble methods in a ﬁrst exploratory study on a simple pool (Section 3) . Then, we make recommendations (Section 4) that we validate in a second study (Section 5), closer to real-life conditions, using TabArena’s pre-computed results, before concluding (Section 6) .  \n2 Related Work  \nPrevious studies have often focused on homogeneous ensemble models [21] [7], but works comparing heterogeneous ensemble models do exist. Creators of ensemble methods and AutoML researchers interested in post-hoc ensembling have compared the results of diﬀerent ensemble methods in diﬀerent ways.  \nIn the ﬁrst case, studies have focused on a speciﬁc subset of ensemble methods, such as Stacking [11], or were done using a few datasets [5] . The work most similar to ours’ is [11], which compared many types of Stacking techniques with both Hard Voting and Single Best on 30 diﬀerent datasets. Our work diﬀers in terms of ensemble methods, pools, datasets, and objective, as we have access to more modern classiﬁers, test a larger curated set of datasets [2][12] and don’t focus only on Stacking.  \nIn the second case, the comparison of diﬀerent ensemble methods was often not the focus. The closest work to ours lies in the appendi","cbCaihAiTdOlKxrZ","https://ap.wps.com/l/cbCaihAiTdOlKxrZ","pdf",244935,1,17,"English","en",105,"# Abstract\n## Introduction\n## Related Work\n## Exploratory Experiments\n## Ensemble Methods and Base Learners","[{\"question\":\"What are the two main observations about blending, stacking, and voting?\",\"answer\":\"Blending and Stacking are inconsistent, but their inconsistencies are independent and occur on different tasks. Hard Voting’s probabilistic estimates are weak, while Robust Soft Voting performs particularly well, especially in multiclass settings.\"}]",1784183821,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"choosing-a-parallel-heterogeneous-ensemble-method-for-tabular-classification","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/choosing-a-parallel-heterogeneous-ensemble-method-for-tabular-classification/82902/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two main observations about blending, stacking, and voting?","Question",{"text":75,"@type":76},"Blending and Stacking are inconsistent, but their inconsistencies are independent and occur on different tasks. Hard Voting’s probabilistic estimates are weak, while Robust Soft Voting performs particularly well, especially in multiclass settings.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]