[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160432-en":3,"doc-seo-160432-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160432,962084928904,"Jake","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Nice Fold or Hero Call - Learning Budget-Efficient Thinking for Adaptive Reasoning","Large reasoning models (LRMs) improve problem solving through extended reasoning, yet frequently misallocate test-time compute. Existing efficiency approaches reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, but largely ignore solvability. This work frames adaptive reasoning as a computational investment under uncertainty, where budget follows expected reasoning return. Budget-Efficient Thinking (BET) uses a two-stage scheme with behavioral cold-start and GRPO plus an investment-cost-aware reward. Across seven benchmarks and three base models, BET reduces reasoning tokens by about 55% on average while improving overall performance, transferring zero-shot gains from math to scientific QA and logical reasoning efficiently.","arXiv :2605 . 11625v1 [ cs .AI] 12 May 2026  \nNice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning  \nZhaomeng Zhou♠ Lan Zhang♠∗ Junyang Wang♠ Mu Yuan♡ Junda Lin♠♠University of Science and Technology of China ♡The Chinese University of Hong Kong  \n[zhouzhm@mail.ustc.edu.cn](zhouzhm@mail.ustc.edu.cn) , [zhanglan@ustc.edu.cn](zhanglan@ustc.edu.cn) , [muyuan@cuhk.edu.hk](muyuan@cuhk.edu.hk)[ ](muyuan@cuhk.edu.hk)Project Page: [https://github.com/houqiii/BET](https://github.com/houqiii/BET)  \nAbstract  \nLarge reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet largely overlook solvability. As a result, they may spend large budgets on queries beyond the model’s capability while compressing hard-but-solvable queries that require deeper reasoning. In this work, we formulate adaptive reasoning as a computational investment under uncertainty, where budget should follow the expected return of reasoning rather than perceived difficulty alone. To instantiate this principle, we propose Budget-Efficient Thinking (BET), a two-stage framework that combines behavioral cold-start with GRPO under an investment-cost-aware reward. By aligning solve-or-fold decisions with rollout-derived solvability, BET learns three behaviors: (1) short solve, answering easy queries concisely; (2) nice fold, abstaining early when continued reasoning has near-zero expected return;  \nand (3) hero call, preserving sufficient compute for hard-but-solvable queries.  \nAcross seven benchmarks and three base models, BET reduces reasoning tokens by ∼55% on average while achieving overall performance improvements, and transfers zero-shot from mathematical reasoning to scientific QA and logical reasoning with comparable efficiency gains.  \n1 Introduction  \nLRMs such as OpenAI-o1 [12] and DeepSeek-R1 [15] have advanced complex problem solving through extended reasoning with multi-step verification [42, 48] . As these models become central to real-world agentic systems [41, 40, 52], their inference-time cost is becoming unsustainable [35, 13] . Crucially, not all of this cost is wasteful: some reflects the genuine price of solving hard problems, while the rest stems from overthinking on easy queries and futile exploration beyond the model’s competence [8, 24, 45] . The central challenge is not simply to shorten reasoning, but to allocate test-time compute according to its expected return without sacrificing problem-solving capability. Recent work addresses this burden in two main ways: uniform reasoning compression reduces generation cost through concise-chain distillation [8], length shaping [32, 37], or target-chain supervision [9, 3]; while difficulty-conditioned control acts more finely by injecting early-exit signals during decoding [47] or learning query-aware depth policies [53, 19, 7] . These methods reduce token usage by compressing reasoning globally or scaling reasoning depth according to perceived difficulty, turning adaptive reasoning into a problem of how much compute to spend on each query.  \n∗ Corresponding author.  \nPreprint.  \n(a) Comparison of model reasoning behaviors on different tasks. (b) Performance of adaptive methods. Figure 1: Reasoning behavior and accuracy-efficiency trade-offs on Omni-Math. (a) Easy, Worthy, and Unsolvable correspond to 16/16, 1–15/16, and 0/16 correct vanilla rollouts. Vanilla LRMs overthink on easy queries and over-allocate on unsolvable ones, while prior adaptive methods often curtail worthy reasoning prematurely. (b) BET lies on the Pareto frontier, preserving worthy reasoning while reducing waste on unsolvable queries.  \nIn realistic workloads, however, difficulty is not solvability. Problems with similar nominal difficulty can have opposite return profiles under the current policy: one may yield to deeper reason","cbCaii6B3Zwe67Iu","https://ap.wps.com/l/cbCaii6B3Zwe67Iu","pdf",1359745,2,1,24,"English","en",105,"# Introduction\n## Adaptive reasoning as compute investment\n## Limitations of difficulty-conditioned methods\n## Budget-Efficient Thinking (BET) framework\n## Experimental evaluation and results","[{\"question\":\"Why do large reasoning models misallocate test-time compute?\",\"answer\":\"They often spend excessive budget on unsolvable queries and over-compress hard-but-solvable queries, because existing methods focus on perceived difficulty rather than solvability.\"},{\"question\":\"How does BET define adaptive reasoning?\",\"answer\":\"BET formulates it as a computational investment under uncertainty, allocating budget according to expected return of continued reasoning instead of difficulty alone.\"},{\"question\":\"What behaviors does BET learn?\",\"answer\":\"BET learns three behaviors: short solve (concise answers for easy queries), nice fold (early abstention when expected return is near zero), and hero call (sufficient compute for hard-but-solvable queries).\"}]","Nice Fold or Hero Call - Learning Budget-Efficient Thinking for Adaptive Reasoning | PDF",1788063354,60,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"nice-fold-or-hero-call-learning-budget-efficient-thinking-for-adaptive-reasoning","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/nice-fold-or-hero-call-learning-budget-efficient-thinking-for-adaptive-reasoning/160432/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-30",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do large reasoning models misallocate test-time compute?","Question",{"text":76,"@type":77},"They often spend excessive budget on unsolvable queries and over-compress hard-but-solvable queries, because existing methods focus on perceived difficulty rather than solvability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does BET define adaptive reasoning?",{"text":81,"@type":77},"BET formulates it as a computational investment under uncertainty, allocating budget according to expected return of continued reasoning instead of difficulty alone.",{"name":83,"@type":74,"acceptedAnswer":84},"What behaviors does BET learn?",{"text":85,"@type":77},"BET learns three behaviors: short solve (concise answers for easy queries), nice fold (early abstention when expected return is near zero), and hero call (sufficient compute for hard-but-solvable queries).","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":30,"slug":109},5,"Comic","comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]