[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84360-en":3,"doc-seo-84360-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84360,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Eigenvalue Calibration for Semantic Embeddings of Large Language Models","Uncertainty quantification underpins reliable deployment of large language models (LLMs), and recent work shows that eigenvalues of semantic embeddings can serve as a key mechanism. Existing calibration methods for classification probabilities cannot be transferred directly to eigenvalues. This study proposes a framework calibrating eigenvalues of semantic embedding–based density matrix predictors via temperature scaling. The work derives entropy–risk equivalence, a central eigenvalue calibration inequality, and proves temperature-scaled eigenvalues minimize proper scoring risk. Experiments confirm systematic overconfidence in real-world LLMs.","Eigenvalue Calibration for Semantic Embeddingsof Large Language Models  \nSebastian G. Gruber †1 Nassim Walha †2,3,4 Francis Bach6 Florian Buettner2,3,4,5  \n1ESAT-PSI, KU Leuven, Belgium  \n2 German Cancer Research Center (DKFZ), Heidelberg, Germany  \n3 German Cancer Consortium (DKTK), Germany  \n4 Goethe University Frankfurt, Germany  \n5Frankfurt Cancer Institute, Germany  \n6PSL Research University / Inria , France.  \nAbstract  \narXiv :2607 .08377v 1 [ cs .LG] 9 Jul 2026  \nUncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods.  \nHowever, conventional calibration results developed for classification probabilities cannot be directly transferred to eigenvalues. We address this gap by proposing a novel framework for calibrating the eigenvalues of semantic embeddings. We interpret LLMs combined with semantic embeddings of their generated answers as density matrix predictors, and we propose a novel approach to calibrate density matrix predictors by applying temperature scaling to their eigenvalues. We establish entropy–risk equivalence under calibration, derive a central calibration inequality specific to eigenvalues, and prove that temperature-scaled eigenvalues optimize calibration when minimizing proper score risks. Experiments on a variety of real-world settings show that current LLMs are systematically overconfident, and validate our theoretical findings.  \nTogether, these results advance the foundations and practice of uncertainty quantification for semantic embeddings.  \n1 INTRODUCTION  \nUncertainty quantification has become a cornerstone for assessing the reliability of modern machine learning models, particularly large language models (LLMs) [Shorinwa et al., 2025] . In high-stakes applications, calibrated uncertainty estimates are essential for downstream decisionmaking, model comparison, and human–AI collaboration  \n†Equal contribution. Alphabetical order.  \n(b) Eigenvalues as Probabilities  \n(c) Before temperature scaling (d) After temperature scaling  \nFigure 1: Normalised semantic embeddings reside in a hypersphere (Figure 1a) . We interpret the eigenvalues of the respective density matrix as probabilities of latent outcomes (Figure 1b) . Large language models are overconfident in their predicted maximum eigenvalue (Figure 1c), which can be adjusted via temperature scaling (Figure 1d) resulting ina lower expected calibration error (ECE), and, thus, more reliable uncertainties.  \n[Silva Filho et al., 2023, Maier-Hein et al., 2024] . Recent advances have shown that semantic embeddings ofLLM outputs provide a powerful basis for uncertainty quantification, enabling fine-grained measures of predictive confidence [Gruber and Buettner, 2024, Nikitin et al., 2024, Walha et al., 2026] . In particular, the eigenvalues constructed from embeddings capture uncertainty through entropy quantities, and have already been adopted in state-of-the-art LLM uncertainty quantification methods [Nikitin et al., 2024, Walha et al., 2026] .  \nDespite this progress, a fundamental gap remains: conventional calibration results developed for classification probabilities cannot be directly transferred to eigenvalues. In classification, calibration aligns predicted probabilities with empirical frequencies, providing interpretable confidence estimates [Murphy, 1973] . In contrast, the eigenvalues of density matrices (constructed from embeddings) encode latent outcome probabilities in a different mathematical space, raising the question of how calibration should be defined, analyzed, and optimized in this setting. Addressing this gap is essential to ensure that embedding-based uncertainty quantification methods are both theoretically principled and practically reliable.  \nIn this paper, we provide the first comprehensive study of eigenvalue calibration for density matrix predictors in general and semantic embed","cbCaimx1WW2NDLhX","https://ap.wps.com/l/cbCaimx1WW2NDLhX","pdf",2534419,5,1,21,"English","en",105,"# Abstract\n# 1 INTRODUCTION\n## Eigenvalues as Probabilities\n# 2 BACKGROUND","[{\"question\":\"What do the experiments show about large language models’ uncertainty estimates?\",\"answer\":\"Across a variety of real-world settings, current LLMs are systematically overconfident. Temperature calibration reduces overconfidence by lowering expected calibration error, aligning predicted eigenvalue uncertainty more reliably with risk.\"}]",1784195091,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"eigenvalue-calibration-for-semantic-embeddings-of-large-language-models","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/eigenvalue-calibration-for-semantic-embeddings-of-large-language-models/84360/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What do the experiments show about large language models’ uncertainty estimates?","Question",{"text":76,"@type":77},"Across a variety of real-world settings, current LLMs are systematically overconfident. Temperature calibration reduces overconfidence by lowering expected calibration error, aligning predicted eigenvalue uncertainty more reliably with risk.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":20,"slug":130},19,"General","general"]