[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81886-en":3,"doc-seo-81886-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81886,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language","Large language models are increasingly used to explain probabilistic AI outputs, yet it remains uncertain whether they can communicate likelihood and uncertainty in natural language in a consistent and calibrated way. This study evaluates nine LLMs using a two-stage pipeline where an upstream model’s Beta-sampled predictions are verbalized by prompted models across six contexts and ten temperatures. Results show general consistency but systematic miscalibration, with weaker performance for uncertainty. Precomputed summary inputs reduce framing sensitivity but not the verbalization error, limiting zero-shot risk communication reliability.","arXiv :2607 .03882v2 [ cs .CL] 10 Jul 2026  \nConsistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language  \nDiego Cerda-Mardini1,2,3 , Sarath Chandar2,3,4,5 , and Sreenath Madathil1,3  \n1Faculty of Dental Medicine and Oral Health Sciences, McGill University, Montréal, QC, Canada, 2 Chandar Research Lab, Polytechnique Montréal, Montréal, H3T 1J4, Canada, 3 Mila – Québec Artificial Intelligence Institute, Montréal, H2S 3H1, Canada, 4 Département de Génie Informatique et Génie Logiciel (GIGL), Polytechnique Montréal, Montréal, H3T 1J4, Canada, 5 Canada CIFAR AI Chair  \nLLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select descriptions that accurately reflect the magnitude of the underlying numerical quantities. We evaluate whether nine LLMs meet these requirements within a two-stage prediction pipeline, in which an upstream model has produced probabilistic outputs characterized by their likelihood and uncertainty, and LLMs are tasked with selecting an appropriate verbal descriptor for each. We simulate predictions from an upstream model by taking samples from a Beta distribution parameterized by its mode and prior sample size. We then prompt LLMs to explain these predictions under six domain contexts and with ten temperature settings, and repeating each experiment ten times. We find that LLMs are generally consistent but miscalibrated, with substantially weaker performance on uncertainty than on likelihood tasks. Providing models with precomputed summary statistics (mode and prior sample size) reduced sensitivity to contextual framing but did not resolve the underlying miscalibration, suggesting that the bottleneck resides in the verbalization step itself. These findings indicate that current LLMs do not yet constitute reliable zero-shot standalone risk communication tools for probabilistic predictions.  \n Code Repository: github/CTruAI/risk_communication_llm  \n1. Introduction  \nLarge Language Models (LLMs) are increasingly used in high-stakes domains, such as healthcare, finance and law, where the cost of incorrect predictions can be severe [Maity and Jyoti, 2025, Dehghani et al., 2025, Nie et al., 2024] . In such settings, it is not enough for a model to be accurate; users also need to knowhow confident the model is in its predictions [Band et al., 2024] . Calibrated uncertainty scores address this need by indicating when to trust a model’s prediction [Minderer et al., 2021, Madsen et al., 2024] . Yet even when such scores are available, correctly interpreting them poses a separate challenge for end users. Concepts like “statistical uncertainty” are frequently misinterpreted, even by skilled professionals like doctors [Gigerenzer et al., 2007] . Predictive distributions can be characterized by two properties: the likelihood of an event occurring and the uncertainty in its prediction [Tyralis and Papacharalampous, 2024] . How laypeople intuitively understand probability often di-  \nverges from its formal statistical definition [Hashim, 2024], adding to the misunderstanding of risk. As reliance on AI for decision-making grows, there is need to convey predictive concepts like likelihood and uncertainty to audiences with varying statistical literacy. LLMs are a promising vehicle for this, as they maybe able to translate numerical estimates into Natural Language Explanations (NLEs) that are more interpretable to non-expert users [Kayser et al., 2022, 2021, Stern et al., 2024] . This raises two questions: whether LLMs have a coherent internal mapping of concepts like likelihood and uncertainty, and whether they can express these concepts in language that is consistent and calibrated to the underlying probabilistic information. To address this, we evaluated","cbCaiuA5RKr0MP9T","https://ap.wps.com/l/cbCaiuA5RKr0MP9T","pdf",5670353,4,1,22,"English","en",105,"# Introduction\n# Definitions\n# Related works\n## The challenge of risk communication","[{\"question\":\"What problem does the paper address about LLMs in risk communication?\",\"answer\":\"It tests whether LLMs can reliably translate probabilistic quantities—both likelihood and uncertainty—into natural-language descriptors that are consistent and calibrated to the underlying numbers.\"},{\"question\":\"How is the evaluation set up for measuring consistency and calibration?\",\"answer\":\"The study uses a two-stage pipeline: an upstream model’s probabilistic predictions are simulated with Beta distributions, then nine LLMs are prompted to choose verbal descriptors across multiple domain contexts and temperature settings, repeated for reliability.\"},{\"question\":\"What are the key findings about LLM performance?\",\"answer\":\"LLMs are generally consistent but miscalibrated, especially for uncertainty compared with likelihood. Supplying precomputed summary statistics improves framing robustness but does not fix the core miscalibration introduced during verbalization.\"}]","Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language | PDF",1784176879,55,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"consistent-but-miscalibrated-evaluating-llm-limitations-for-risk-communication-in-natural-language","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/consistent-but-miscalibrated-evaluating-llm-limitations-for-risk-communication-in-natural-language/81886/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address about LLMs in risk communication?","Question",{"text":76,"@type":77},"It tests whether LLMs can reliably translate probabilistic quantities—both likelihood and uncertainty—into natural-language descriptors that are consistent and calibrated to the underlying numbers.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the evaluation set up for measuring consistency and calibration?",{"text":81,"@type":77},"The study uses a two-stage pipeline: an upstream model’s probabilistic predictions are simulated with Beta distributions, then nine LLMs are prompted to choose verbal descriptors across multiple domain contexts and temperature settings, repeated for reliability.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the key findings about LLM performance?",{"text":85,"@type":77},"LLMs are generally consistent but miscalibrated, especially for uncertainty compared with likelihood. Supplying precomputed summary statistics improves framing robustness but does not fix the core miscalibration introduced during verbalization.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]