[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84776-en":3,"doc-seo-84776-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84776,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Geometry-Aware Bayesian Quantification via Compositional Data Analysis","Accurately estimating an unknown target label distribution is essential for adapting to label shift, a problem known as quantification or class prevalence estimation. Recent KDE-based quantifiers model densities of multiclass classifier posteriors, but these posterior vectors are compositional and reside on the probability simplex. Euclidean Gaussian kernels ignore simplex geometry and can assign mass outside valid boundaries. The work proposes a geometry-aware KDE using log-ratio representations and Aitchison geometry, plus shrinkage regularization near the boundary, yielding point and Bayesian inference. Experiments on 42 datasets across tabular, text, and image domains demonstrate competitive performance and improvements over standard KDE baselines.","GEOMETRY-AWARE BAYESIAN QUANTIFICATION VIACOMPOSITIONAL DATA ANALYSIS  \narXiv :2607 .04977v 1 [ cs .LG] 6 Jul 2026  \nAlejandro Moreo  \nIstituto di Scienza e Tecnologie dell’Informazione Consiglio Nazionale delle Ricerche Pisa, Italy [alejandro.moreo@isti.cnr.it](alejandro.moreo@isti.cnr.it)  \nPablo González, Juan José del Coz  \nArtificial Intelligence Center University of Oviedo Asturias, Spain  \n{gonzalezgpablo,[juanjo}@uniovi.es](juanjo}@uniovi.es)  \nJuly 7, 2026  \nABSTRACT  \nAccurately estimating the unknown target label distribution is the critical first step for adapting to label shift. This task, widely known as quantification or class prevalence estimation, has recently seen significant advances through continuous KDE-based methods which model the density of multiclass classifier posteriors. Posterior vectors might be regarded as compositional data, since they lie on the probability simplex. However, existing KDE-based quantifiers typically rely on Euclidean Gaussian kernels, which ignore simplex geometry and incorrectly assign probability mass outside its boundaries. We introduce a geometry-aware KDE model for multiclass quantification based on log-ratio representations and Aitchison geometry, together with a shrinkage regularization that improves robustness near the simplex boundary. Combined with a maximum-likelihood interpretation of KDE-based quantification, we derive both point-estimation and Bayesian inference procedures for class prevalences. Experiments on 42 datasets across tabular, text, and image domains show that the proposed method is competitive with state-of-the-art quantifiers, often improving over standard KDE-based baselines, while also yielding strong results among Bayesian quantification methods.  \n1 Introduction  \nMachine learning models deployed in the wild frequently encounter label shift [41] . Adapting to this shift requires accurately estimating the unknown target label distribution. This problem, called quantification [16] or class prevalence estimation [31], is a critical standalone task in domains where aggregate population trends matter more than individual predictions. Furthermore, it serves as the essential foundation for label shift adaptation, as accurate prevalence estimates are strictly required to compute importance weights for classifier retraining [3] .  \nMany prominent methods, such as Maximum Likelihood for Label Shift (MLLS) [38] and Black-Box Shift Estimation (BBSE) [31], operate directly on the posterior probabilities produced by a base classifier. In multiclass quantification, modeling the continuous distribution of these posteriors using density-based methods like KDEy [35] has emerged as a powerful alternative that naturally preserves the dependence structure across classes. Consequently, KDEy is increasingly adopted as a robust solution in a variety of application domains, such as fairness monitoring in IR [27], graph-related applications [8, 33], classifier accuracy prediction [44], and medical imaging for healthcare [20] .  \nHowever, modeling classifier posteriors via continuous density estimation introduces a fundamental geometric mismatch. Posterior probability vectors lie on the probability simplex and therefore constitute compositional data. Existing KDEbased quantifiers rely on Gaussian kernels in Euclidean space, which ignore this relative structure and inevitably assign probability mass outside the simplex boundaries. This mismatch becomes particularly problematic near the boundary of the simplex, where posterior probabilities heavily concentrate in many real-world settings.  \nIn this work, we revisit multiclass quantification from the perspective of compositional data analysis (CoDA) [1] to resolve this bottleneck. We introduce a geometry-aware kernel based on log-ratio transformations, enabling continuous density estimation in Aitchison geometry. While this representation is more consistent with the underlying simplex  \nBayesian Quantification via Compositional","cbCainn2BxVcQ7oI","https://ap.wps.com/l/cbCainn2BxVcQ7oI","pdf",2499364,1,22,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What is quantification in the context of label shift adaptation?\",\"answer\":\"Quantification refers to estimating the unknown target label distribution (class prevalence) under label shift. It serves as the foundation for computing importance weights needed for classifier retraining.\"},{\"question\":\"Why do Euclidean KDE-based multiclass quantifiers fail on the probability simplex?\",\"answer\":\"Posterior probability vectors lie on the probability simplex and form compositional data. Euclidean Gaussian kernels ignore this geometry and can assign probability mass outside the simplex, especially near its boundary.\"},{\"question\":\"How does the proposed method address simplex geometry and improve robustness?\",\"answer\":\"It introduces a geometry-aware KDE built on log-ratio representations and Aitchison geometry, together with shrinkage regularization to stabilize estimation near the simplex boundary. This enables both point estimation and Bayesian inference for class prevalences.\"}]",1784198170,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"geometry-aware-bayesian-quantification-via-compositional-data-analysis","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/geometry-aware-bayesian-quantification-via-compositional-data-analysis/84776/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is quantification in the context of label shift adaptation?","Question",{"text":75,"@type":76},"Quantification refers to estimating the unknown target label distribution (class prevalence) under label shift. It serves as the foundation for computing importance weights needed for classifier retraining.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do Euclidean KDE-based multiclass quantifiers fail on the probability simplex?",{"text":80,"@type":76},"Posterior probability vectors lie on the probability simplex and form compositional data. Euclidean Gaussian kernels ignore this geometry and can assign probability mass outside the simplex, especially near its boundary.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed method address simplex geometry and improve robustness?",{"text":84,"@type":76},"It introduces a geometry-aware KDE built on log-ratio representations and Aitchison geometry, together with shrinkage regularization to stabilize estimation near the simplex boundary. This enables both point estimation and Bayesian inference for class prevalences.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]