[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83701-en":3,"doc-seo-83701-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83701,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","CUBAS: Information Geometric Curvature-Based Adaptive Sampling for Supervised Classification","In supervised classification, the value of a training set depends as much on its informativeness as on its size, yet most sampling methods ignore the intrinsic geometry of the data distribution. CuBAS (Curvature-Based Adaptive Sampling) introduces an information-geometric framework that uses curvature computed from a q-state Potts Markov random field to guide adaptive selection. Labeled data are treated as a statistical manifold, where local curvature scores partition k-NN graphs into smooth regions and decision-boundary regions. The approach yields compact, maximally informative training subsets with linear-time efficiency across diverse benchmarks.","arXiv :2607 .03 145v 1 [ cs .LG] 3 Jul 2026  \nCUBAS: INFORMATION GEOMETRIC CURVATURE-BASED ADAPTIVE SAMPLING FOR SUPERVISED CLASSIFICATION  \nAlexandre Luis Magalhães Levada  \nFederal University of São Carlos  \n13565-905, São Carlos-SP, Brazil  \n[alexandre.levada@ufscar.br](alexandre.levada@ufscar.br)  \nJuly 7, 2026  \nABSTRACT  \nThe informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geometry of the data distribution. We address this gap by introducing CuBAS (Curvature-Based Adaptive Sampling), an information-geometric framework for adaptive data selection in supervised classification, grounded in the q-state Potts Markov random field (MRF) model. The central insight is that a labeled dataset can be viewed as a statistical manifold, on which local curvature, estimated via the ratio of second-to first-order observed Fisher information, faithfully encodes the geometric complexity of the underlying data distribution. Concretely, we construct ak-nearest-neighbor graph over the labeled data and derive a closed-form curvature score at each vertex from the Potts sufficient statistics, avoiding costly eigendecompositions or kernel density estimates.  \nThis curvature signal naturally partitions the graph into two complementary regimes: low-curvature regions, corresponding to smooth, homogeneous clusters that are efficiently represented by a small number of prototypical samples, and high-curvature regions, concentrated around decision boundaries and topologically complex structures that are disproportionately informative for classification. By selecting nodes from both regimes in a principled, geometry-aware manner, CuBAS constructs compact yet maximally informative training subsets. Extensive empirical evaluation across 30 benchmark datasets, spanning tabular, image, and biological domains, demonstrates consistent and statistically significant improvements in classification accuracy over random sampling and uncertaintybased baselines, across a wide range of labeling budgets and classifier architectures. Our method is computationally efficient (linear in the number of edges of the k-NN graph), theoretically grounded in the differential geometry of statistical manifolds, and directly interpretable in terms of the local shape operator of the data manifold. CuBAS thus offers a principled, scalable, and geometry-aware alternative to heuristic sampling for supervised learning.  \n1 Introduction  \nThe remarkable success of modern supervised learning has been driven not only by increasingly sophisticated learning algorithms, but also by the availability of large annotated datasets. Nevertheless, the effectiveness of a classifier depends far more on the quality and representativeness of its training samples than on their sheer quantity. Large datasets often contain substantial redundancy, noisy observations, class imbalance, and numerous samples that contribute little to defining the underlying decision boundaries. This observation naturally motivates a fundamental question: how can we identify the most informative samples in a labeled dataset for classification?  \nAddressing this question is becoming increasingly important as datasets continue to grow in size and complexity. Although considerable effort has been devoted to designing more expressive classifiers, comparatively less attention has been paid to principled mechanisms for selecting informative training samples in a way that reflects the intrinsic geometric organization of the data Wilson [1972] . Traditional sampling and instance reduction techniques, including random subsampling, prototype selection, and instance weighting, are typically heuristic and operate primarily in the input space, making limited use of the statistical relationships among neighboring samples Garcia et al. [2012] . As a  \nPREPRINT-JULY 7, 2026  \nconsequence, they often fail to distinguish truly informative boundary samples from redundan","cbCaib90mGP52dYW","https://ap.wps.com/l/cbCaib90mGP52dYW","pdf",1301359,4,1,30,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does CuBAS address in supervised classification sampling?\",\"answer\":\"CuBAS targets the gap that most sampling strategies rely on dataset size or heuristics and do not account for the intrinsic geometry of the data distribution when selecting training samples.\"},{\"question\":\"How does CuBAS compute informativeness for labeled data?\",\"answer\":\"CuBAS builds a k-nearest-neighbor graph over labeled samples and derives a closed-form curvature score at each vertex using q-state Potts Markov random field statistics, avoiding costly eigendecompositions or kernel density estimates.\"},{\"question\":\"What is the practical effect of curvature on which samples are selected?\",\"answer\":\"Low-curvature regions correspond to smooth homogeneous clusters and can be represented with few prototypes, while high-curvature regions concentrate near decision boundaries and complex structures and are selected to maximize discriminative value.\"}]",1784189828,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cubas-information-geometric-curvature-based-adaptive-sampling-for-supervised-classification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/cubas-information-geometric-curvature-based-adaptive-sampling-for-supervised-classification/83701/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CuBAS address in supervised classification sampling?","Question",{"text":75,"@type":76},"CuBAS targets the gap that most sampling strategies rely on dataset size or heuristics and do not account for the intrinsic geometry of the data distribution when selecting training samples.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CuBAS compute informativeness for labeled data?",{"text":80,"@type":76},"CuBAS builds a k-nearest-neighbor graph over labeled samples and derives a closed-form curvature score at each vertex using q-state Potts Markov random field statistics, avoiding costly eigendecompositions or kernel density estimates.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the practical effect of curvature on which samples are selected?",{"text":84,"@type":76},"Low-curvature regions correspond to smooth homogeneous clusters and can be represented with few prototypes, while high-curvature regions concentrate near decision boundaries and complex structures and are selected to maximize discriminative value.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":22,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]