[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128617-en":3,"doc-seo-128617-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},128617,962084925636,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","An Improved Predictive Accuracy Bound for Averaging Classifiers","An improved bound is developed for the gap between training error and test (predictive) error of voting/averaging classifiers. The result strengthens theoretical justification for widely used averaging-based methods including Bayesian classification, Maximum Entropy discrimination, Winnow, and Bayes point machines. The new analysis yields an optimization perspective: predictive error benefits from averaging many hypotheses while maintaining large margins, rather than relying solely on empirical margin maximization.","An Improved Predictive Accuracy Bound for Averaging Classi􀀌ers  \nJohn Langford School of Computer Science Carnegie Mellon [jcl@cs.cmu.edu](jcl@cs.cmu.edu)[ ](jcl@cs.cmu.edu)[http://www.cs.cmu.edu/](http://www.cs.cmu.edu/)~jcl  \nMatthias Seeger Institute for Adaptive and Neural Computation University of Edinburgh [seeger@dai.ed.ac.uk](seeger@dai.ed.ac.uk)[ ](seeger@dai.ed.ac.uk)[http://www.dai.ed.ac.uk/homes/seeger/](http://www.dai.ed.ac.uk/homes/seeger/)  \nNimrod Megiddo IBM Almaden Research Center [megiddo@almaden.ibm.com](megiddo@almaden.ibm.com)[http://theory.stanford.edu/~megiddo/](http://theory.stanford.edu/~megiddo/)  \nAbstract  \nWe present an improved bound on the difference between training and test errors for voting classi􀀌ers. This improved averaging bound provides a theoretical justi􀀌cation for popular averaging techniques such as Bayesian classi􀀌cation, Maximum Entropy discrimination, Winnow and Bayes point machines and has implications for learning algorithm design.  \n1. Introduction  \nAveraging is a standard technique in applied machine learning for combining multiple classi􀀌ers to achieve greater accuracy. Examples include Bayesian classi􀀌 -cation [4], boosting [7], bagging [2], Winnow [13], Maximum Entropy discrimination [11], and Bayes point machines [9] . Despite the prevalence of this technique there is only weak theoretical justi􀀌cation so far for the practice. This paper provides a new stronger theoretical justi􀀌cation for the practice of averaging. In particular, we state and prove a bound on the gap between the training set error rate and the predictive error rate which improves as more hypotheses are averaged over.  \nUntil 1998, theoretical bounds such as the Occam's razor bound [3] suggested that averaging was wrong because it increased the description length of the resulting hypothesis.1 The Occam's razor bound only suggests that averaging may be bad since there is no corresponding lower bound. Schapire, Freund, Bartlett and Lee [16] showed a great improvement on the naive  \n1We note, however, that it is the minimum description length that should be used in the bound.  \nbound for an average-of-classi􀀌ers hypothesis. Loosely speaking, their margin bound states that if the average has a small empirical error rate (i.e. , it is accurate on most training examples) and has a large \\margin\"(de􀀌ned in 2.1), then its true error rate is also small. The proof itself works in a very intuitive manner by showing that the accuracy of a large margin classi􀀌eris close to the accuracy of a simple classi􀀌er, for which standard bounds are tight.  \nThe problem with this result is that the value of the bound depends only on the empirical margin which does not necessarily improve with an average over a larger number of hypotheses. Thus, the only design principle to be inferred from this bound is the simple criterion: choose the average to maximize the margin. However, empirical results [8] indicate that this procedure is not optimal, being prone to over􀀌tting and behaving somewhat non-robust in the presence of outliers.  \nIn this paper, we prove a new bound on the true error rate, which suggests a new optimization criterion, namely, optimize for a large margin and for a uniform average over as many hypotheses as possible.  \nThe layout of this paper is as follows:  \n1. Discussion of the relationship with prior relevant results.  \n2. Development of a simple improved theoretical bound.  \n3. A proof of the bound.  \n4. An example of the bene􀀌t of the new bound on a toy problem.  \n5. Discussion of implications of the new bound on prior work.  \n2. The setting and important earlier results  \n2.1 The setting  \nWe 􀀌rst explain the setting, which is the same as the one used in [16] .  \nAn input space X is given, where the members of X are also referred to as examples. The set X 􀀂 f􀀀1; 1gis the space of labeled examples. A base hypothesis his a mapping from the input space X into f􀀀1; 1g. A (possibly in􀀌nite) space H (the hypothesis s","cbCaipgMLBZTREsk","https://ap.wps.com/l/cbCaipgMLBZTREsk","pdf",201951,1,"English","en",105,"# Introduction\n## Averaging in applied machine learning\n## Prior margin bound limitations\n# The setting and important earlier results\n## The setting\n## Quantities used in the bound","[{\"question\":\"What problem does the paper address for averaging classifiers?\",\"answer\":\"It addresses the weak theoretical justification for averaging by deriving a stronger bound relating training error to predictive (test) error.\"},{\"question\":\"Which methods does the improved averaging bound support?\",\"answer\":\"The bound provides justification for averaging techniques such as Bayesian classification, Maximum Entropy discrimination, Winnow, and Bayes point machines.\"},{\"question\":\"How does the new result change the optimization criterion for averaging?\",\"answer\":\"It suggests optimizing for both a large margin and a uniform average over as many hypotheses as possible, rather than only maximizing empirical margin.\"}]","An Improved Predictive Accuracy Bound for Averaging Classifiers | PDF",1786002129,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"an-improved-predictive-accuracy-bound-for-averaging-classifiers","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/an-improved-predictive-accuracy-bound-for-averaging-classifiers/128617/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address for averaging classifiers?","Question",{"text":75,"@type":76},"It addresses the weak theoretical justification for averaging by deriving a stronger bound relating training error to predictive (test) error.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which methods does the improved averaging bound support?",{"text":80,"@type":76},"The bound provides justification for averaging techniques such as Bayesian classification, Maximum Entropy discrimination, Winnow, and Bayes point machines.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the new result change the optimization criterion for averaging?",{"text":84,"@type":76},"It suggests optimizing for both a large margin and a uniform average over as many hypotheses as possible, rather than only maximizing empirical margin.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":28,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":28,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]