[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82561-en":3,"doc-seo-82561-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82561,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models","Large Reasoning Models (LRMs) succeed on complex tasks by producing long chain-of-thought (CoT) trajectories, but they often overthink simple queries, causing excessive token cost and lower inference efficiency. Existing compression methods usually apply uniform reduction or coarse difficulty estimation, which can degrade performance on hard problems. Confidence-Adaptive Thinking (CAT) uses the model’s intrinsic self-certainty signals to guide preference optimization, autonomously shortening reasoning on easy inputs and preserving exploration on uncertain ones. Experiments show improved accuracy across multiple benchmarks.","CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large  \nReasoning Models  \nQizhi Jiang 1 Shuo Wang 1 Pei Ke 1 ,2 ,∗ Yuhang Song 1 Ke Qin 1 ,2  \n1Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China, Chengdu, China  \n2Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province {jiangqizhi, [202422900227}@std.uestc.edu.cn](202422900227}@std.uestc.edu.cn) , [kepei@uestc.edu.cn](kepei@uestc.edu.cn)  \n[songyuhang@std.uestc.edu.cn](songyuhang@std.uestc.edu.cn) , [qinke@uestc.edu.cn](qinke@uestc.edu.cn)  \narXiv :2607 .00862v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nLarge Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet they frequently exhibit overthinking on simple queries, resulting in significant token overhead and reduced inference efficiency. However, existing compression methods predominantly apply uniform length reduction or rely on coarse-grained difficulty estimation, often leading to performance degradation on difficult problems. To address this limitation, we propose Confidence-Adaptive Thinking (CAT), a framework that incorporates the model’s intrinsic self-certainty signals as confidence into the preference optimization process, which autonomously modulates reasoning lengths based on problem difficulty. Experimental results show that CAT consistently outperforms state-of-the-art baselines on reasoning accuracy across multiple benchmarkson different base models. Our work enables LRMs to effectively compress confident responses while deliberating on uncertain ones, offering a potentially robust solution for balancing accuracy and latency in practical industrial scenarios.  \n1 Introduction  \nRecently, large reasoning models (LRMs) have rapidly emerged and made substantial progress on complex natural language processing (NLP) tasks, as exemplified by OpenAI-o1 (OpenAI, 2024) and DeepSeek-R1 (DeepSeek-AI, 2025) . These models are equipped with the ability to generate long reasoning chains, demonstrating strong potential on challenging reasoning problems such as mathematical competitions (Xu et al., 2025) . However, while LRMs heavily rely on long chain-of-thought (CoT) traces to perform well on difficult tasks, they tend  \n∗ Corresponding author.  \nto produce redundant reasoning and self-reflection for simple inputs, incurring pronounced overthinking and token overhead (Chen et al., 2024 ; Feng et al., 2025 ; Liu et al., 2025 ; Sui et al., 2025) . This behavior leads to verbose thought chains that increase computation cost and reduce overall inference efficiency. Accordingly, how to enable LRMs to dynamically adjust token consumption based on the input difficulty has attracted increasing attention, determining the practical industrial usability of LRMs in terms of the balance between accuracy and latency (Shen et al., 2025a) .  \nMost of the existing approaches focus on reasoning compression and length control predominantly, which treat shortening reasoning chains as the primary objective (Qu et al., 2025) and apply a uniform reduction of reasoning tokens to all the queries (Xia et al., 2025 ; Chen et al., 2024 ; Ma et al., 2025 ; Munkhbat et al., 2025) . While such methods can substantially decrease generation length, they often incur non-trivial performance degradation on difficult problems, since complex tasks still require sufficient reasoning depths and lengths to sustain accurate answers (Muennighoff et al., 2025 ; Zeng et al., 2024) . Another line of work resorts to difficulty-adaptive reasoning to mitigates the imbalance between overthinking for easier instances and underthinking for harder ones. This category of methods tends to dynamically adjust the budget of output tokens based on the model performance (Shen et al., 2025a) .  \nHowever, existing works on adaptive reasoning still face a severe challenge of coarse-grained difficulty estimation. Current m","cbCaigsIXBnoub1R","https://ap.wps.com/l/cbCaigsIXBnoub1R","pdf",902566,2,1,13,"English","en",105,"# Introduction\n## Problem: Overthinking and token overhead\n## Prior work: Uniform compression and coarse difficulty adaptation\n## Proposed method: CAT driven by intrinsic confidence\n## Method components: self-certainty estimation and confidence-weighted optimization","[{\"question\":\"Why do large reasoning models waste tokens on simple questions?\",\"answer\":\"LRMs often generate redundant reasoning and self-reflection even for easy inputs, which creates long CoT chains, increased computation cost, and reduced inference efficiency.\"},{\"question\":\"What limitation do existing adaptive reasoning methods have?\",\"answer\":\"They often rely on coarse-grained difficulty estimation using output accuracy, which depends on external labels and evaluates only the answer quality rather than the overall reasoning quality.\"},{\"question\":\"How does CAT use confidence to control reasoning length?\",\"answer\":\"CAT treats the model’s intrinsic self-certainty as confidence to estimate the quality of reasoning trajectories, then constructs confidence-informed preference data. It further applies a confidence-weighted preference optimization objective to encourage shorter reasoning when confidence is high while retaining necessary exploration when confidence is low.\"}]",1784181540,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cat-confidence-adaptive-thinking-for-efficient-reasoning-of-large-reasoning-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cat-confidence-adaptive-thinking-for-efficient-reasoning-of-large-reasoning-models/82561/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do large reasoning models waste tokens on simple questions?","Question",{"text":75,"@type":76},"LRMs often generate redundant reasoning and self-reflection even for easy inputs, which creates long CoT chains, increased computation cost, and reduced inference efficiency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation do existing adaptive reasoning methods have?",{"text":80,"@type":76},"They often rely on coarse-grained difficulty estimation using output accuracy, which depends on external labels and evaluates only the answer quality rather than the overall reasoning quality.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CAT use confidence to control reasoning length?",{"text":84,"@type":76},"CAT treats the model’s intrinsic self-certainty as confidence to estimate the quality of reasoning trajectories, then constructs confidence-informed preference data. It further applies a confidence-weighted preference optimization objective to encourage shorter reasoning when confidence is high while retaining necessary exploration when confidence is low.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]