[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81963-en":3,"doc-seo-81963-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81963,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Estimating Uncertainty from Reasoning A Large-Scale Study of Multi and Crosslingual MCQA Performance in LLMs","Uncertainty estimation (UE) helps LLM-powered systems decide when to abstain, yet prior work largely centers on English. This study evaluates UE methods across 22 languages, covering high-, mid-, and low-resource settings, using two human-curated Q&A datasets. Nine open- and closed-box UE approaches are compared under long-form reasoning while avoiding LLM-as-judge and embedding-based scoring. Findings show that reasoning in English improves multilingual UE, generation language matters more than question language, method choice depends on model scale, and threshold calibration guides selective prediction.","Estimating Uncertainty from Reasoning: A Large-Scale Study of Multiand Crosslingual MCQA Performance in LLMs  \nAndrea Bacciu1 Andrea Alfarano2 * Saab Mansour1 Amin Mantrach1 Marcello Federico1  \n1Amazon 2INSAIT, Sofia  \n{andbac, saabm, mantrach, [marcfede}@amazon.com](marcfede}@amazon.com)  \narXiv :2607 .06327v2 [ cs .CL] 10 Jul 2026  \nAbstract  \nUncertainty estimation (UE) enables LLMpowered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across  \n22 languages, spanning high-, mid-, and lowresource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-asa-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, openbox probability-based methods outperform alternatives; at larger scales, closed-box selfverbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.  \n1 Introduction  \nLarge Language Models (LLMs) have changed how people access and interact with information, supporting tasks from everyday planning to complex question-answering (Bommasani, 2021) . Recent research has primarily focused on improving  \n*Work done during internship at Amazon.  \ntask performance, for example through chain-ofthought prompting (Wei et al., 2022) and instruction tuning (Ouyang et al., 2022), enabling models to solve increasingly complex problems. However, high accuracy alone is insufficient: downstream systems must recognize when an LLM’s answer is not grounded in its knowledge, enabling them to abstain, defer to humans, or fall back to safer behavior. This has motivated a parallel line of work on uncertainty estimation (UE), which seeks to determine when models know the answer, enabling LLM-powered systems to recognize and communicate their lack of knowledge (Kuhn et al., 2023) . Existing research on LLM uncertainty has predominantly focused on English, leaving limited evidence on whether UE methods maintain their efficacy in other languages, especially in low-resource settings (Kuhn et al. (2023); Kossen et al. (2024); Santilliet al. (2025); Cecere et al. (2025), inter alia) . The only dedicated multilingual UE study we are aware of (Xue et al., 2025) evaluates just three methodson five languages, relies on machine-translated data without human post-editing, and uses a multiplechoice, short-answer format rather than open-ended generation. Their dataset yields a median answer length of just one word, meaning this setup primarily evaluates how language affects question comprehension while offering limited evidence about uncertainty during longer, more linguistically rich generation. Establishing multilingual trustworthiness through UE is also methodologically challenging because standard metrics such as AUROC rely on ground truth labels. When ground truth is approximated via LLM-as-judge, BERTScore, or n-gram overlap, these proxies introduce noise that can misrank uncertainty methods (Santilli et al., 2025) . In multilingual settings, this problem is compounded: neural judges may behave inconsistently across languages, introducing la","cbCainWOmQkUukES","https://ap.wps.com/l/cbCainWOmQkUukES","pdf",3148169,5,1,16,"English","en",105,"# Abstract\n# Introduction\n# Related Work and Background\n## Uncertainty Estimation (UE) in LLMs","[{\"question\":\"What problem does uncertainty estimation (UE) address in LLM systems?\",\"answer\":\"UE enables systems to recognize when an LLM’s answer is not grounded in its knowledge so they can abstain, defer, or behave more safely.\"},{\"question\":\"How does the study evaluate UE methods across languages?\",\"answer\":\"It compares nine open- and closed-box UE methods across 22 languages using two human-curated Q\\u0026A datasets, eliciting long-form reasoning while avoiding LLM-as-judge and embedding-based scoring.\"},{\"question\":\"What key conclusion is drawn about reasoning language versus question language?\",\"answer\":\"Prompting models to reason in English substantially improves UE, and generation language has a larger effect on UE performance than the language of the questions, especially across resource levels.\"}]","Estimating Uncertainty from Reasoning A Large-Scale Study of Multi and Crosslingual MCQA Performance in LLMs | PDF",1784177307,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"estimating-uncertainty-from-reasoning-a-large-scale-study-of-multi-and-crosslingual-mcqa-performance-in-llms","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/estimating-uncertainty-from-reasoning-a-large-scale-study-of-multi-and-crosslingual-mcqa-performance-in-llms/81963/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-30","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does uncertainty estimation (UE) address in LLM systems?","Question",{"text":77,"@type":78},"UE enables systems to recognize when an LLM’s answer is not grounded in its knowledge so they can abstain, defer, or behave more safely.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the study evaluate UE methods across languages?",{"text":82,"@type":78},"It compares nine open- and closed-box UE methods across 22 languages using two human-curated Q&A datasets, eliciting long-form reasoning while avoiding LLM-as-judge and embedding-based scoring.",{"name":84,"@type":75,"acceptedAnswer":85},"What key conclusion is drawn about reasoning language versus question language?",{"text":86,"@type":78},"Prompting models to reason in English substantially improves UE, and generation language has a larger effect on UE performance than the language of the questions, especially across resource levels.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]