[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128614-en":3,"doc-seo-128614-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128614,962084925502,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","稀疏概率分类器","Support vector machine outputs are often reused as confidence measures, yet no theory guarantees they truly correspond to classification uncertainty. When uncertainty matters, conditional probability estimators are more principled. This work targets ambiguity near the decision boundary, introducing a sparse probabilistic classifier via an adaptation of maximum likelihood implemented with logistic regression. The method produces calibrated probabilities only within a user-defined interval, improves computational efficiency, and yields better results than standard logistic regression with performance similar to SVMs.","Sparse Probabilistic Classi􀀌ers  \nRomain H􀀓erault [romain.herault@hds.utc.fr](romain.herault@hds.utc.fr)  \nHEUDIASYC, Universit􀀓e de Technologie de Compi􀀒egne, BP 20529, 60205 Compi􀀒egne cedex, France  \nYves Grandvalet [yves.grandvalet@idiap.ch](yves.grandvalet@idiap.ch)  \nIDIAP, Rue du Simplon 4, Case Postale 592, CH-1920 Martigny, Switzerland  \nAbstract  \nThe scores returned by support vector machines are often used as a con􀀌dence measures in the classi􀀌cation of new examples.  \nHowever, there is no theoretical argument sustaining this practice. Thus, when classi-􀀌cation uncertainty has to be assessed, it is safer to resort to classi􀀌ers estimating conditional probabilities of class labels. Here, we focus on the ambiguity in the vicinity of the boundary decision. We propose an adaptation of maximum likelihood estimation, instantiated on logistic regression. The model outputs proper conditional probabilities into a user-de􀀌ned interval and is less precise elsewhere. The model is also sparse, in the sense that few examples contribute to the solution. The computational e􀀎ciency is thus improved compared to logistic regression. Furthermore, preliminary experiments show improvements over standard logistic regression and performances similar to support vector machines.  \n1. Motivation  \nThere have been several attempts to turn the scores returned by support vector machines (SVMs) into probabilistic assignments (Platt, 2000; Grandvalet et al. , 2006) . However, there is no guaranty that these scores re􀀍ect a classi􀀌cation con􀀌dence; we even know that the conditional probabilities of class labels cannot be recovered unambiguously except at the decision boundary (Bartlett & Tewari, 2004) . Thus, when the classi􀀌cation uncertainty has to be assessed, estimating conditional probabilities is better motivated.  \nAppearing in Proceedings of the 24 th International Conference on Machine Learning, Corvallis, OR, 2007 . Copyright 2007 by the author(s)/owner(s) .  \nWe propose to build probabilistic classi􀀌ers that are accurate on the “gray zone”, where class labels switch. Well-calibrated probabilities in this area allow to assess classi􀀌cation uncertainty. The classi􀀌er also provides relevant decision rules for the set of corresponding asymmetric misclassi􀀌cation losses, or equivalently, for the corresponding cone of the ROC curve. Focusing on a small range of conditional probabilities instead of estimating them on their full span has two advantages. First, the training objective is closer to the ultimate goal of minimizing the misclassi􀀌cation risk, and second, inaccuracy outside of the focus range is a key element for kernelized models, since Bartlett and Tewari (2004) proved that sparsity does not occur when the conditional probabilities can be unambiguously estimated everywhere. Sparsity refers here to the limited number of non-zero elements in a kernel expansion. It implies that many training examples have no in􀀍uence in the training process, thus improving its computational e􀀎ciency.  \nThe assessment of classi􀀌cation uncertainty and the sparsity of the model are important issues for the class imbalance problem, which is our original motivation for this work. When a vast majority of examples belong to the negative “uninteresting” class, and only a few interesting examples are available, learning tends to be biased towards the recognition of the majority class. This problem can be addressed by rebalancing the training distribution, either by over-sampling the minority class, or generating arti􀀌cial examples of the minority class (Chawla et al., 2002), or by downsampling the majority class. However, undersampling may discard relevant pieces of information and oversampling is not computationally e􀀎cient. Another tactic consists in post-processing standard classi􀀌cation techniques, such as tuning a bias term after learning to correct for the original decision bias, but this scheme fails to discover the changes in the shape or orientation of","cbCaikTKHMVEVx7E","https://ap.wps.com/l/cbCaikTKHMVEVx7E","pdf",204252,2,1,"English","en",105,"# Motivation\n## From SVM scores to probabilistic assignments\n## Gray zone and calibrated decision rules\n## Motivation from class imbalance\n# Learning criterion\n## Bayes decision rule\n## Maximum likelihood","[{\"question\":\"为什么不能直接把SVM的分数当作置信度？\",\"answer\":\"因为SVM分数与分类置信度之间缺乏理论保证，除决策边界附近外，类别条件概率也不能被唯一恢复。\"},{\"question\":\"文中提出的模型如何处理决策边界附近的不确定性？\",\"answer\":\"聚焦于“灰区”，即类别切换附近的条件概率歧义，通过在该区间内输出经过校准的条件概率来评估不确定性。\"},{\"question\":\"该方法为何被称为“稀疏”？\",\"answer\":\"训练过程中只有少量样本对解产生贡献，使核展开中的非零元素数量受限，从而提升计算效率。\"}]","稀疏概率分类器 | PDF",1786002113,20,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sparse-probabilistic-classifiers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/sparse-probabilistic-classifiers/128614/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-25","2026-08-06",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么不能直接把SVM的分数当作置信度？","Question",{"text":75,"@type":76},"因为SVM分数与分类置信度之间缺乏理论保证，除决策边界附近外，类别条件概率也不能被唯一恢复。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文中提出的模型如何处理决策边界附近的不确定性？",{"text":80,"@type":76},"聚焦于“灰区”，即类别切换附近的条件概率歧义，通过在该区间内输出经过校准的条件概率来评估不确定性。",{"name":82,"@type":73,"acceptedAnswer":83},"该方法为何被称为“稀疏”？",{"text":84,"@type":76},"训练过程中只有少量样本对解产生贡献，使核展开中的非零元素数量受限，从而提升计算效率。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":29,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":29,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]