[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86513-en":3,"doc-seo-86513-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86513,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","When Does Distribution Shift Break Graph Neural Networks Calibration?","Graph neural networks are increasingly used in real settings where the test graph differs from training, yet the effect of distribution shift on calibration—alignment between predictive confidence and actual accuracy—remains unclear. This work provides a first closed-form theory of GNN calibration under shift, showing it is controlled by a single scalar dependent on structural changes between source and target graphs and feature quality. It determines when models become over-confident, under-confident, or remain calibrated, yields the optimal temperature scaling, and extends results to symmetric graph convolution, multiclass settings, and covariate shift with an expected calibration error bound.","arXiv :2607 . 10804v 1 [ cs .LG] 12 Jul 2026  \nWhen does distribution shift break graph neural networks calibration?  \nAbderaouf Bahia  \na Computer Science and Applied Mathematics Laboratory (LIMA), Faculty of Science and Technology, Chadli Bendjedid University, El  \nTarf 36000, Algeria  \nAbstract  \nGraph neural networks (GNNs) are increasingly deployed in real-world applications where distribution shift is unavoidable. However, how such shifts affect model calibration, defined as the agreement between predictive confidence and actual accuracy, remains poorly understood, and existing graph calibration methods typically rely on labeled validation data from the deployment distribution. In this work, I present the first closed-form theoretical characterization of GNN calibration under distribution shift. I show that calibration is governed by a single scalar quantity that explicitly depends on structural changes between the source and target graphs, as well as feature quality. This characterization precisely identifies when a model becomes over-confident, under-confident, or remains calibrated, and directly yields the optimal temperature scaling strategy. I further extend the analysis to graph convolutional networks with symmetric normalization, multi-class classification, and covariate shift, and derive a theoretical upper bound on the expected calibration error. My analysis also reveals that, under homogeneous distribution shift, a single global temperature is theoretically optimal, providing a principled explanation for why more complex nodewise recalibration methods offer no additional benefit. Building on these theoretical insights, I propose STAC, a source-free, label-free calibration method. Experiments on synthetic benchmarks demonstrate substantial calibration improvements, while evaluations on five real-world graph datasets show that reliable calibration without target labels remains challenging despite the strong predictive power of the theory. These findings identify label-free accuracy estimation under distribution shift as the central unresolved challenge for practical calibration of GNNs and establish a theoretical foundation for future research on trustworthy graph learning. Project details are available at [https://github.com/AraoufBh/GNN-Calibration](https://github.com/AraoufBh/GNN-Calibration).  \nKeywords: graph neural networks, calibration, distribution shift, homophily, temperature scaling, uncertainty quantification  \n1. Introduction  \nGraph neural networks (GNNs) [1] are now routinely deployed in settings where the graph seen at test time differs from the one used for training: social and financial networks evolve, sensor and infrastructure graphs are rewired after failures or upgrades, and citation, molecular or recommendation graphs used at inference are drawn from different sub-populations than those used to train the model. In many of these settings raw predictive accuracy is not the primary bottleneck; what determines whether a GNN can be trusted inside a fraud-review, clinical-triage, or infrastructure-monitoring pipeline is whether its confidence scores are calibrated: whether a prediction reported with 90% confidence is, in fact, correct 90% of the time. Miscalibration is not cosmetic: it silently corrupts every downstream mechanism that consumes probability outputs, from selective prediction and human-in-the-loop triage to risk-weighted decision rules and model ensembling.  \nTwo observations make GNN calibration under distribution shift both practically important and theoretically awkward. (i) GNNs are miscalibrated in structure-dependent ways; unlike image classifiers (typically over-confident), they are often under-confident in-distribution, and the size of the gap correlates with graph properties such as homophily and node degree [2, 3] . (ii) Standard calibrators need labels from the test distribution: temperature scaling [4] and its graph-specific variants [2, 3 , 5–7] all fit a temperature (o","cbCaighhSrrMIoaV","https://ap.wps.com/l/cbCaighhSrrMIoaV","pdf",1262708,3,1,19,"English","en",105,"# Introduction\n# Contributions","[{\"question\":\"What is calibration in graph neural networks, and why does distribution shift threaten it?\",\"answer\":\"Calibration measures whether reported confidence matches true accuracy. Under distribution shift, the confidence scores no longer align with accuracy, corrupting downstream decisions that rely on probabilistic outputs.\"},{\"question\":\"What determines whether a GNN becomes over-confident, under-confident, or stays calibrated under shift?\",\"answer\":\"A single scalar quantity governs calibration behavior, explicitly depending on structural changes between source and target graphs and feature quality. This directly identifies the direction (over/under) and magnitude of miscalibration.\"},{\"question\":\"How does the paper handle temperature scaling when target labels are unavailable?\",\"answer\":\"It derives a theoretical basis for the optimal temperature scaling under shift and proposes STAC, a source-free, label-free calibration method. Experiments show calibration improvements on synthetic benchmarks, while real datasets still make label-free accurate estimation challenging.\"}]",1784212304,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-does-distribution-shift-break-graph-neural-networks-calibration","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-does-distribution-shift-break-graph-neural-networks-calibration/86513/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is calibration in graph neural networks, and why does distribution shift threaten it?","Question",{"text":75,"@type":76},"Calibration measures whether reported confidence matches true accuracy. Under distribution shift, the confidence scores no longer align with accuracy, corrupting downstream decisions that rely on probabilistic outputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What determines whether a GNN becomes over-confident, under-confident, or stays calibrated under shift?",{"text":80,"@type":76},"A single scalar quantity governs calibration behavior, explicitly depending on structural changes between source and target graphs and feature quality. This directly identifies the direction (over/under) and magnitude of miscalibration.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper handle temperature scaling when target labels are unavailable?",{"text":84,"@type":76},"It derives a theoretical basis for the optimal temperature scaling under shift and proposes STAC, a source-free, label-free calibration method. Experiments show calibration improvements on synthetic benchmarks, while real datasets still make label-free accurate estimation challenging.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]