[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86258-en":3,"doc-seo-86258-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86258,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers","Post-hoc calibration is used to correct probability estimates from trained classifiers, yet aggregate evaluation can conceal how calibration behavior changes under different operating regimes within the same dataset. This paper presents a pre-registered, condition-stratified robustness study comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1–C4). Four hypothesis groups assess discrimination deltas, Brier differences, calibration slopes, and AUROC differences, revealing metric- and condition-dependent robustness with no external transportability claims.","Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers  \nGurdeep Singh Virdee  \nFergana State Technical University, Uzbekistan  \narXiv :2607 . 1 1542v 1 [ cs .LG] 13 Jul 2026  \nAbstract—Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-registered, condition-stratified robustness analysis comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1–C4). Four hypothesis groups are evaluated: discrimination deltas with Holm-corrected multiplicity control (H1), Brierscore differences (H2), calibration slope outcomes (H3), and AUROC differences under best-condition setups (H4). TEMPminus-ISO discrimination deltas remain small across all conditions (−0 .0155 to 0.0139), with Holm-adjusted p-values of 0.9895 everywhere. TEMP Brier differences are consistently negative (C1:−0 .0002 through C4: −0 .0074), while ISO shows sign reversals. TEMP calibration slopes stay closer to unity in every condition (range 0.7597–0.9493) than ISO slopes (0 .1364–0.2726). AUROC differences shift from near zero in C1 (−0 .0004) to positive in C4 (0 .0264). These results establish that in-dataset robustness is condition-dependent and metric-specific. No claim of external transportability is made.  \nIndex Terms—post-hoc calibration, temperature scaling, isotonic regression, model reliability, distribution shift, calibration robustness, probabilistic classification  \nI. INTRODUCTION  \nA classifier’s ranking quality can stay intact even when its predicted probabilities become unreliable—a gap that matters whenever those probabilities drive thresholds, triage decisions, or risk communication downstream [1]–[3] . Two standard correctives exist: temperature scaling adjusts the logit scale with a single parameter [4], while isotonic regression fits a stepwise non-parametric map to a held-out calibration set [5] . Both are easy to implement. Neither comes with a guarantee that the correction will remain stable when operating conditions change inside the same study corpus.  \nThat gap between aggregate headline numbers and conditionlevel behavior is what motivated this work. Consider a concrete scenario: a model that reports 0.85 AUROC across a pooled test set may still emit poorly calibrated probabilities in one subgroup while performing well in another. If an operator relies on the pooled metric alone, interventions can be mistimed and confidence overstated for exactly the cases where stakes are highest [6] . Pooled summaries hide sign changes. We saw  \nthis firsthand—our initial analysis aggregated results across conditions C1–C4 and failed to reveal that ISO Brier differences reverse direction in C3 relative to the other strata.  \nThis paper reports a bounded, in-dataset robustness analysis of TEMP and ISO under four pre-defined condition shifts.  \nThree qualities separate the study from prior calibration benchmarks. First, every quantitative claim is traced to a locked registry of verified experimental outputs; no post-hoc recomputation was performed during manuscript preparation. Second, multiplicity control via Holm adjustment is applied to H1 discrimination deltas, keeping family-wise error in check across conditions. Third, scope boundaries are stated at the outset: we evaluate in-dataset condition robustness only, and no external transportability statement is made [7] .  \nThe core question is direct. Do TEMP and ISO behave consistently across conditions C1–C4, or does their relative advantage change with the operating regime? The answer, asthe data show, depends on which metric you examine.  \nFour hypothesis groups structure the evaluation:  \n• H1: Condition-wise TEMP-minus-ISO discrimination deltas under Holm-adjusted multiplicity control.  \n• ","cbCaibanv1r3yBi0","https://ap.wps.com/l/cbCaibanv1r3yBi0","pdf",1181081,5,1,6,"English","en",105,"# Introduction\n## Problem Motivation and Gap\n## Study Contributions and Scope\n# Problem Formulation\n## Calibration Mapping and Ideal Condition\n## Temperature Scaling and Isotonic Regression\n## Condition-Stratified Evaluation Domain\n# Hypotheses and Evaluation Metrics\n## Discrimination Deltas\n## Brier Score Differences\n## Calibration Slope Outcomes\n## AUROC Differences","[{\"question\":\"What is the main research question of the study?\",\"answer\":\"Whether temperature scaling (TEMP) and isotonic regression (ISO) behave consistently across four in-dataset operating conditions (C1–C4), or whether their relative advantage changes depending on the condition and the metric examined.\"},{\"question\":\"Which calibration methods are compared and how do they work at a high level?\",\"answer\":\"Temperature scaling applies a monotonic logit transformation with a single learned parameter fit on a held-out calibration partition, while isotonic regression learns a stepwise non-decreasing mapping from calibration pairs.\"},{\"question\":\"What findings show that robustness is metric- and condition-dependent?\",\"answer\":\"Discrimination deltas between TEMP and ISO stay small across all conditions with consistently non-significant Holm-adjusted p-values; Brier differences are consistently negative for TEMP but show sign reversals for ISO; calibration slopes stay closer to unity for TEMP than ISO; AUROC differences shift from near zero in C1 to positive in C4.\"}]",1784209861,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"condition-stratified-robustness-analysis-of-post-hoc-calibration-methods-for-probabilistic-classifiers","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/condition-stratified-robustness-analysis-of-post-hoc-calibration-methods-for-probabilistic-classifiers/86258/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main research question of the study?","Question",{"text":76,"@type":77},"Whether temperature scaling (TEMP) and isotonic regression (ISO) behave consistently across four in-dataset operating conditions (C1–C4), or whether their relative advantage changes depending on the condition and the metric examined.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which calibration methods are compared and how do they work at a high level?",{"text":81,"@type":77},"Temperature scaling applies a monotonic logit transformation with a single learned parameter fit on a held-out calibration partition, while isotonic regression learns a stepwise non-decreasing mapping from calibration pairs.",{"name":83,"@type":74,"acceptedAnswer":84},"What findings show that robustness is metric- and condition-dependent?",{"text":85,"@type":77},"Discrimination deltas between TEMP and ISO stay small across all conditions with consistently non-significant Holm-adjusted p-values; Brier differences are consistently negative for TEMP but show sign reversals for ISO; calibration slopes stay closer to unity for TEMP than ISO; AUROC differences shift from near zero in C1 to positive in C4.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]