[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85542-en":3,"doc-seo-85542-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85542,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models","Vision-language models (VLMs) are used increasingly in high-stakes scenarios where reliable uncertainty quantification (UQ) is as critical as predictive accuracy. Although chain-of-thought (CoT) prompting and reasoning-trained models are widely adopted, their impact on UQ reliability is insufficiently understood. Results show that reasoning can degrade the ranking quality of uncertainty estimates even while improving task accuracy. The main mechanism is implicit answer conditioning: reasoning traces drive token probabilities toward consistency with the model’s own trace rather than uncertainty about correctness, yielding overconfident answers. Agreement-based consistency remains robust and often improves UQ.","arXiv :2603 . 16728v2 [ cs .LG] 11 Jul 2026  \nThe Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models  \nRobert Welch 1 ,2 , Emir Konuk 1 ,2 , and Kevin Smith 1 ,2  \n1 KTH Royal Institute of Technology, Stockholm, Sweden  \n2 Science for Life Laboratory, Stockholm, Sweden  \n[rwe2@kth.se](rwe2@kth.se)  \nAbstract. Vision-language models (VLMs) are increasingly deployed in high-stakes settings where reliable uncertainty quantification (UQ) is as important as predictive accuracy. Extended reasoning via chainof-thought (CoT) prompting or reasoning-trained models has become ubiquitous in modern VLM pipelines, yet its effect on UQ reliability remains poorly understood. Our results show that reasoning tends to degrade the quality of many uncertainty estimates, even when it improves task accuracy. We identify implicit answer conditioning as the primary mechanism: as reasoning traces converge on a conclusion before the final answer is generated, token probabilities increasingly reflect consistency with the model’s own reasoning trace rather than uncertainty about correctness. In effect, the model becomes overconfident in its answer. In contrast, agreement-based consistency remains robust and often improves under reasoning, making it a practical choice for uncertainty estimation in reasoning-enabled VLMs.  \nKeywords: Vision-Language Models · Chain-of-Thought · Uncertainty Quantification  \n1 Introduction  \nVision-language models (VLMs) are increasingly deployed in high-stakes settings where reliable predictions are as important as accurate ones, from interpreting medical images [15] to autonomous navigation [29] . Modern VLM pipelines increasingly rely on chain-of-thought (CoT) prompting and reasoningoriented models that generate intermediate, step-by-step inference before producing a final answer. Reasoning substantially improves performance on challenging multimodal benchmarks requiring multi-step visual and mathematical reasoning [17, 32] .  \nHowever, accuracy alone is insufficient for safe deployment. Selective generation addresses this issue by ensuring that the models can abstain when uncertain [2] . In this setting, uncertainty quality is determined by how well confidence scores rank correct predictions above incorrect ones, and is typically evaluated using ranking-based metrics such as the Prediction Rejection Ratio (PRR) [19, 20] and AUGRC [25] .  \n2 R. Welch et al.  \nUser Prompt  \n3 \u003C/answer>  \nWith CoT Reasoning  \n\u003Cthought>  \n,  \n. Therefore  \n\u003C/thought>\u003Canswer>  \n,  \nis aligned  \n0-inch  \ntip extends  \n3-inch  \n4-inch  \nthe nearest  \ninch gives  \n3 inches  \n3 \u003C/answer>  \nPredicted Answer  \n3  \nGround Truth  \n2  \n0  \n1  \nFig. 1: Example of implicit answer conditioning: intermediate reasoning progressively commits the model to its eventual answer before that answer is explicitly generated. Colour indicates the model’s confidence in its final predicted answer aˆ on the same input, without reasoning (P (aˆ | x,θ)) and at each reasoning token position t with chainof-thought reasoning (P (aˆ | x, r ≤t,θ)), where x is the input and r ≤t is the reasoning trace up to token t. Although the prediction is incorrect in both cases (3 inches; ground truth: 2), confidence rises sharply as the reasoning trace converges on the wrong answer, exceeding the no-CoT confidence.  \nnail  \nThe  \n ’s  \nhead  \nwith  \nthe  \nmark  \nand  \n its  \npast  \nthe  \nmark  \nbut  \nnot  \nto  \nthe  \nmark  \nrounding  \nto  \nAnswer likelihood  \nGiven the widespread adoption of reasoning, a natural question arises. How does reasoning affect uncertainty estimation in multi-modal generation? One might expect that deliberate reasoning improves both accuracy and UQ reliability. However, across the VLM families and benchmarks we evaluated, we find that this intuition often fails for many uncertainty estimates. While reasoning typically improves task accuracy, it tends to degrade the ranking quality of uncertainty estimates that rely on answer-toke","cbCaiclLeFU7i2om","https://ap.wps.com/l/cbCaiclLeFU7i2om","pdf",1203458,6,1,32,"English","en",105,"# Introduction\n## Uncertainty in high-stakes VLM deployment\n## Selective generation and uncertainty metrics\n# Mechanism: implicit answer conditioning","[{\"question\":\"What problem does the document address for vision-language models in high-stakes use?\",\"answer\":\"It addresses uncertainty quantification reliability—how well uncertainty estimates reflect correctness when VLMs are deployed in safety-critical settings.\"},{\"question\":\"How does chain-of-thought reasoning affect uncertainty estimation according to the results?\",\"answer\":\"Reasoning tends to degrade many uncertainty estimates’ ranking quality even when it improves task accuracy.\"},{\"question\":\"What mechanism causes the model to become overconfident when reasoning is used?\",\"answer\":\"Implicit answer conditioning: as the reasoning trace converges on a conclusion, token probabilities increasingly reflect consistency with the model’s own reasoning trace instead of uncertainty about correctness.\"}]",1784204323,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"the-cost-of-reasoning-chain-of-thought-induces-overconfidence-in-vision-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-cost-of-reasoning-chain-of-thought-induces-overconfidence-in-vision-language-models/85542/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the document address for vision-language models in high-stakes use?","Question",{"text":76,"@type":77},"It addresses uncertainty quantification reliability—how well uncertainty estimates reflect correctness when VLMs are deployed in safety-critical settings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does chain-of-thought reasoning affect uncertainty estimation according to the results?",{"text":81,"@type":77},"Reasoning tends to degrade many uncertainty estimates’ ranking quality even when it improves task accuracy.",{"name":83,"@type":74,"acceptedAnswer":84},"What mechanism causes the model to become overconfident when reasoning is used?",{"text":85,"@type":77},"Implicit answer conditioning: as the reasoning trace converges on a conclusion, token probabilities increasingly reflect consistency with the model’s own reasoning trace instead of uncertainty about correctness.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]