[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83493-en":3,"doc-seo-83493-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83493,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking","Reasoning language models often overthink by producing long, unproductive behavior chains such as hedging, abandoning approaches, and self-contradiction, which consume tokens without improving final answers. Evidence shows that even with fixed response length, incorrect reasoning traces contain more unproductive self-reflection than correct ones. DASH identifies when self-reflection helps or harms using intermediate answer commitments inside traces, then assigns segment-level credit based on drift toward or away from correctness. On competition math benchmarks, DASH improves accuracy in high-overthinking settings and enables more productive self-correction.","Know When to Stop: Segment-Level Credit Assignment for Reducing  \nOverthinking  \nChia-Hsuan Lee, Sihui Dai, Mingyang Zhou, Isha Slavin Shi-Xiong Zhang, Sambit Sahu, William Campbell  \nCapital One  \narXiv :2607 .00482v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nReasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without improving answers. We show that these behaviors are not merely a consequence of length;  \neven when controlling for response length, incorrect traces exhibit higher rates of unproductive self-reflection than correct ones. Addressing this requires identifying where selfreflection helps vs hurts, but obtaining these step-level annotations is costly. We observe that intermediate answer commitments within reasoning traces can provide a cheap proxy: by comparing each final answer candidate in the trace to the ground truth, we can determine whether subsequent reflection is productive without any additional supervision. Building on this insight, we propose DASH (Drift Aware advantage SHaping), which assigns segmentlevel credit based on whether each reasoning segment leads toward or away from correctness. On competition-level math benchmarks, DASH achieves the highest accuracy where overthinking is prevalent (AIME25: 50.8% vs. 45.4% GRPO) while reducing overthinking behaviorsand achieving more productive self-correction than baselines.  \n1 Introduction  \nReasoning-focused language models, such as DeepSeek-R1 (DeepSeek-AI, 2025), achieve strong performance through extended chains of thought. However, longer reasoning does not always help: models frequently exhibit overthinking behaviors such as hedging, re-verifying, or switching approaches (Wang et al., 2025b ; Peng et al., 2025) which can lead the model to an incorrect final answer.  \nThis motivates a natural question: can we train models to retain productive self-reflection while surpressing unproductive self-reflection? A signif-  \nicant challenge is cost: identifying whether selfreflection is helpful at each reasoning step would require step-wise labels via a process reward model (Lightman et al., 2024 ; Wang et al., 2024), LLMas-a-judge, or manual annotation (Lightman et al., 2024) . In this work, we propose a cheaper alternative.  \nOur key observation is that reasoning models can commit to intermediate answers within their thinking traces–for example, writing \"the answer is X\" or boxing a result before continuing to reason. These commitments provide verifiable demonstrations of productive and unproductive self-reflection: by comparing each to the ground truth, we know whether subsequent reflection improved or degraded the answer, without any external supervision. When a model reaches a correct intermediate answer and then reflects its way to an incorrect one, we have direct evidence that this self-reflection was harmful.  \nBased on this intuition, we propose DASH (Drift-Aware advantage SHaping), which converts traces where the answer drifts from correct to incorrect intermediate examples from wasted negatives into informative training examples. Rather than assigning a single scalar advantage to the entire rollout, DASH divides each trace into segments bounded by consecutive answer checkpoints and assigns advantages based on whether each segment moves towards or away from the correct answer. A single drift trace simultaneously teaches the model to reinforce the reasoning that found the correct answer and to suppress the overthinking that abandoned it–extracting dual training signal from what GRPO would treat as a flat negative example.  \nWe complement DASH with six lightweight linguistic overthinking signals—repetition, hedging, abandonment, contradiction, recomputation, and length outlier—that characterize reasoning quality without requiring intermediate answer extraction. These signals serve as evaluation metrics to verify  \nFigure 1: Overview of DASH. W","cbCaid7anOJI3z7J","https://ap.wps.com/l/cbCaid7anOJI3z7J","pdf",1645111,3,1,15,"English","en",105,"# Introduction\n## Analyzing Overthinking in Reasoning Traces","[{\"question\":\"What problem does the paper address in reasoning language models?\",\"answer\":\"It addresses overthinking behaviors where models generate extended, token-consuming chains like hedging, approach abandonment, and self-contradiction that often lead to incorrect final answers without improving correctness.\"},{\"question\":\"How does DASH determine whether self-reflection is productive?\",\"answer\":\"DASH uses intermediate answer commitments within reasoning traces and compares each final answer candidate in the trace to ground truth, inferring whether subsequent reflection moved the reasoning toward or away from correctness without extra step-level labels.\"},{\"question\":\"What training signal does DASH use instead of a single rollout-level advantage?\",\"answer\":\"DASH divides each reasoning trace into segments bounded by consecutive answer checkpoints and assigns advantages per segment based on drift, so informative segments reinforce correct reasoning and harmful segments suppress overthinking.\"}]",1784188414,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"know-when-to-stop-segment-level-credit-assignment-for-reducing-overthinking","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/know-when-to-stop-segment-level-credit-assignment-for-reducing-overthinking/83493/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in reasoning language models?","Question",{"text":75,"@type":76},"It addresses overthinking behaviors where models generate extended, token-consuming chains like hedging, approach abandonment, and self-contradiction that often lead to incorrect final answers without improving correctness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DASH determine whether self-reflection is productive?",{"text":80,"@type":76},"DASH uses intermediate answer commitments within reasoning traces and compares each final answer candidate in the trace to ground truth, inferring whether subsequent reflection moved the reasoning toward or away from correctness without extra step-level labels.",{"name":82,"@type":73,"acceptedAnswer":83},"What training signal does DASH use instead of a single rollout-level advantage?",{"text":84,"@type":76},"DASH divides each reasoning trace into segments bounded by consecutive answer checkpoints and assigns advantages per segment based on drift, so informative segments reinforce correct reasoning and harmful segments suppress overthinking.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]