[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82037-en":3,"doc-seo-82037-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82037,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","HALO: Hybrid Adaptive Latent Reasoning for Language Models","HALO (Hybrid Adaptive Latent Refinement) improves frozen pretrained language models by adding a small amount of adaptive inference-time computation instead of uniformly refining all tokens. A coarse refinement stage is followed by a selective second-stage latent refinement on a token subset chosen via token scoring and monotonic token halting. Skipped tokens bypass the expensive path and are merged back afterward. On benchmarks combining MMLU-Pro and GPQA-Diamond, HALO attains the best overall public average while using fewer executed refinement steps than fixed baselines.","arXiv :2607 .08775v1 [ cs .CL] 3 May 2026  \nHALO: Hybrid Adaptive Latent Reasoning for  \nLanguage Models  \nMicah Zhang  \nLockheed Martin  \n[micah.t.zhang@lmco.com](micah.t.zhang@lmco.com)  \nAbstract  \nWe study how to improve a frozen pretrained language model with a small amount of adaptive extra computation. A simple approach is to add additional refinement steps on top of the backbone hidden states, but fixed extra refinement can be wasteful: a one-step refinement head may be too weak, while forcing a second full-sequence refinement step everywhere can increase compute without improving transfer. We introduce HALO, a hybrid adaptive latent-refinement method that combines a coarse refinement stage with selective second-stage latent refinement on a subset of tokens chosen by token scoring and monotonic token halting. On the main public benchmark comparison built from MMLU-Pro and GPQA-Diamond, HALO achieves the best overall average among the paper-facing methods, outperforming the frozen backbone, fixed-1, and fixed-2 . Internal analysis further shows that HALO reaches nearly the same token-accuracy level as fixed-2 while using fewer average applied refine steps than fixed-1 and far fewer than fixed-2 . These results suggest that the key advantage is not simply more refinement, but a better allocation of refinement: HALO achieves the strongest paper-facing result while also using less measured controller compute than either fixed baseline.  \n1 Introduction  \nFrozen pretrained language models are strong general-purpose systems, but some reasoning problems benefit from additional inference-time computation. A natural way to add such computation is torefine backbone hidden states after the main forward pass. This can improve reasoning behavior without full-model finetuning, but the extra computation must be allocated carefully.  \nWe introduce HALO (Hybrid Adaptive Latent reasOning), a lightweight adaptive latent-refinement method for frozen language models. HALO uses a coarse refinement stage together with token scoring and monotonic token halting to choose a subset of tokens for a second-stage latent refinement block. Skipped tokens bypass this expensive path and are merged back afterward. The resulting design concentrates the more expensive second-stage computation on only part of the sequence rather than spending a uniform refinement budget everywhere. Importantly, the coarse refinement stage isan architectural path, whereas the paper’s compute metric counts the refinement updates actually executed by the controller at runtime. In the winning configuration, these are not the same thing: the initial adaptive keep decision is made token by token from cheap gate features computed from the current logits and, when available, the change from the previous step’s logits, so the controller can skip refinement before the later budgeted token-halting stage is reached.  \nWe study a focused question: when extending a frozen language model with extra latent computation, is selective second-step refinement a better use of compute than either stopping after one refinement step or forcing a second refinement step everywhere? We compare HALO against three paper-facing baselines: the frozen backbone, a one-step full-sequence refinement baseline (fixed-1), and a two-step full-sequence refinement baseline (fixed-2) .  \nPreprint.  \nOur results support a clear answer. On the public comparison built from MMLU-Pro and GPQADiamond, HALO achieves the best overall average among the paper-facing methods. Internally, it reaches nearly the same token-accuracy level as fixed-2 while using fewer average applied refine steps than fixed-1 and far fewer than fixed-2 . In other words, HALO achieves the strongest paper-facing performance while also using less measured controller compute than either fixed baseline. The gain is not uniform across benchmarks—HALO’s strongest advantage comes from GPQA-Diamond while MMLU-Pro remains strongest for the frozen backbone—s","cbCaihovmqjTJLYN","https://ap.wps.com/l/cbCaihovmqjTJLYN","pdf",272813,4,1,15,"English","en",105,"# Introduction\n# Method\n## Problem setup\n## HALO","[{\"question\":\"What problem does HALO address in frozen pretrained language models?\",\"answer\":\"HALO targets reasoning tasks where adding inference-time computation can help, but where uniformly applying refinement to every token wastes compute and may not improve transfer.\"},{\"question\":\"How does HALO decide which tokens receive the second-stage refinement?\",\"answer\":\"HALO uses token-level scoring via gate features from current logits (and when available, changes from previous-step logits) to compute keep probabilities, then applies refinement only when thresholds and stability conditions are met.\"},{\"question\":\"How does HALO perform compared with fixed-1 and fixed-2 refinement baselines?\",\"answer\":\"On the public benchmark built from MMLU-Pro and GPQA-Diamond, HALO achieves the best overall average, reaching nearly the same token accuracy level as fixed-2 while executing fewer refinement steps than fixed-1 and far fewer than fixed-2.\"}]","HALO: Hybrid Adaptive Latent Reasoning for Language Models | PDF",1784177730,38,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"halo-hybrid-adaptive-latent-reasoning-for-language-models","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/halo-hybrid-adaptive-latent-reasoning-for-language-models/82037/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does HALO address in frozen pretrained language models?","Question",{"text":76,"@type":77},"HALO targets reasoning tasks where adding inference-time computation can help, but where uniformly applying refinement to every token wastes compute and may not improve transfer.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does HALO decide which tokens receive the second-stage refinement?",{"text":81,"@type":77},"HALO uses token-level scoring via gate features from current logits (and when available, changes from previous-step logits) to compute keep probabilities, then applies refinement only when thresholds and stability conditions are met.",{"name":83,"@type":74,"acceptedAnswer":84},"How does HALO perform compared with fixed-1 and fixed-2 refinement baselines?",{"text":85,"@type":77},"On the public benchmark built from MMLU-Pro and GPQA-Diamond, HALO achieves the best overall average, reaching nearly the same token accuracy level as fixed-2 while executing fewer refinement steps than fixed-1 and far fewer than fixed-2.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]