[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81661-en":3,"doc-seo-81661-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81661,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Latent Reward Steering: Adaptive Inference-Time Framework for Cognitive Behaviors in Reasoning LLMs","Reasoning ability in large language models depends on both knowledge and the timely deployment of cognitive behaviors during generation. Existing approaches control behaviors through explicit, predefined instructions or fixed representation steering, which lacks adaptivity when errors and needed corrections differ across reasoning states, tasks, and model backbones. Latent Reward Steering (LRS) optimizes sparse autoencoder latent states using a learned latent reward model from reasoning traces to provide state-specific corrections gated by reward and confidence, improving performance across benchmarks.","Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs  \nJiakang Li*1 , Guanyu Zhu*2 , Can Jin*1 , Chenxi Huang3 , Dexu Yu4 , Ronghao Chen5 Yang Zhou 1 , Hongwu Peng6 , Xuanqi Lan7 , Dimitris N. Metaxas†1 , Youhua Li†8  \n1Rutgers University 2 South China Agricultural University 3 Columbia University  \n4Fenz.AI 5 QuantaAlpha 6Adobe  \n7 Santa Clara University 8 City University of Hong Kong  \n[Contact:](Contact: {jiakang.li@rutgers.edu})[ {jiakang.li@rutgers.edu}](Contact: {jiakang.li@rutgers.edu})  \n*Equal contribution. †Equal corresponding authors.  \narXiv :2606 .00726v2 [ cs .AI] 10 Jul 2026  \nAbstract  \nStrong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures and required corrections vary across reasoning states, tasks, and models. To this end, we propose Latent Reward Steering (LRS), an adaptive inferencetime framework that promotes cognitive behaviors by optimizing the sparse-autoencoder (SAE) latent states that implicitly carry them.  \nRather than relying on predefined cognitive behaviors or steering directions derived from them, LRS trains a latent reward model on reasoning traces by final answer correctness to estimate the quality of intermediate latent states. During inference, reward gradients provide state-specific correction directions for fragile latent states, while a reward and confidence gate restricts intervention to states the reward signal flags as fragile. Experiments on multiple reasoning LLM backbones and benchmarks show that LRS consistently improves performance over various baselines, and posthoc analyses further indicate that LRS implicitly promotes good cognitive behaviors that fix the original reasoning errors. Code is available at: [https://github.com/jiakanglee/](https://github.com/jiakanglee/)[ ](https://github.com/jiakanglee/)Latent-Reward-Steering.  \n1 Introduction  \nPerforming step-by-step reasoning to solve complex problems has become a central research focus in large language models (Wei et al., 2022 ; Kojima et al., 2022) . Yet even strong reasoning models remain brittle: a single early mistake, such as a flawed assumption or a skipped verification step, can gradually derail an otherwise promising reasoning chain (Gan et al., 2025 ; Huang et al., 2023 ; Tyen et al., 2024) . Recent work highlights cognitive behaviors such as verification, backtracking,  \nFigure 1: Motivation. A fragile reasoning state can derail reasoning, while explicit behavior-level control may suffer from fixed labels and directions. LRS instead optimizes fragile latent states with a learned reward signal and implicitly promotes useful cognitive behaviors.  \nand subgoal setting as important ingredients of successful reasoning (Gandhi et al., 2025) . This suggests that some reasoning failures are not purely failures of model knowledge, but failures to induce cognitive behaviors at the right moments within an ongoing reasoning chain.  \nThe importance of cognitive behaviors in reasoning LLMs has motivated a line of work on controlling such behaviors, most of which involve explicit behavior-level control. Prompt-based methods elicit desired cognitive behaviors through textual instructions, from few-shot in-context learning (Brown et al., 2020) and chain-of-thought (COT) prompting (Wei et al., 2022 ; Kojima et al., 2022) to more specific behaviors such as sub-goal decomposition (Zhou et al., 2022 ; Wang et al., 2023), strategic planning (Zheng et al., 2024), and verification (Weng et al., 2023 ; Miao et al., 2023 ; Dhuliawala et al., 2024) . Representation-level steering methods instead intervene directly on latent states (Turner et al., 2023 ; Zou et al., 2023), by associating cognitive behaviors with steering directions (Chen et al., 2025), head-specific interventions (Zha","cbCaibQIqPqp2qXI","https://ap.wps.com/l/cbCaibQIqPqp2qXI","pdf",2619286,4,1,20,"English","en",105,"# Abstract\n# Introduction\n## Motivation\n## Related work on cognitive behavior control\n## Limitation of explicit steering paradigms\n## Research question and approach preview","[{\"question\":\"What problem does Latent Reward Steering (LRS) target in reasoning LLMs?\",\"answer\":\"LRS targets brittleness in step-by-step reasoning where a single early mistake can derail the chain. It focuses on promoting cognitive behaviors at the right moments rather than relying only on model knowledge.\"},{\"question\":\"How does LRS promote cognitive behaviors during inference?\",\"answer\":\"LRS trains a latent reward model on reasoning traces labeled by final answer correctness, then uses reward gradients to generate state-specific correction directions for fragile latent states. A reward-and-confidence gate limits intervention to states flagged as fragile.\"},{\"question\":\"Why are existing behavior-control methods considered insufficiently adaptive?\",\"answer\":\"Most methods rely on predefined cognitive behaviors and fixed steering objects or directions, which may not apply uniformly across models or align with the local reasoning state where an error occurs. This mismatch prevents consistent improvements when failure modes vary.\"}]",1784175265,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"latent-reward-steering-adaptive-inference-time-framework-for-cognitive-behaviors-in-reasoning-llms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/latent-reward-steering-adaptive-inference-time-framework-for-cognitive-behaviors-in-reasoning-llms/81661/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Latent Reward Steering (LRS) target in reasoning LLMs?","Question",{"text":75,"@type":76},"LRS targets brittleness in step-by-step reasoning where a single early mistake can derail the chain. It focuses on promoting cognitive behaviors at the right moments rather than relying only on model knowledge.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LRS promote cognitive behaviors during inference?",{"text":80,"@type":76},"LRS trains a latent reward model on reasoning traces labeled by final answer correctness, then uses reward gradients to generate state-specific correction directions for fragile latent states. A reward-and-confidence gate limits intervention to states flagged as fragile.",{"name":82,"@type":73,"acceptedAnswer":83},"Why are existing behavior-control methods considered insufficiently adaptive?",{"text":84,"@type":76},"Most methods rely on predefined cognitive behaviors and fixed steering objects or directions, which may not apply uniformly across models or align with the local reasoning state where an error occurs. This mismatch prevents consistent improvements when failure modes vary.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]