[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85121-en":3,"doc-seo-85121-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85121,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Depth-Entropy Guided Sampling for Training-Free LLM Reasoning","Reinforcement learning (RL) improves large language model reasoning but depends on expensive posttraining, curated data, and reliable reward signals. Depth-Entropy Guided Sampling (DEGS) introduces a training-free test-time approach that reweights candidate sequences using layer-wise entropy collapse. Stronger reasoners show a distinctive “late collapse,” quantified as a per-sequence collapse depth and integrated with a power-sampling MCMC objective. Evaluated on multiple openweight models and benchmarks, DEGS achieves near-chance-to-state-of-the-art training-free accuracy with small wall-clock overhead and robust out-of-domain gains.","arXiv :2607 .09693v1 [ cs .LG] 19 Jun 2026  \nDEPTH-ENTROPY GUIDED SAMPLING FOR TRAININGFREE LLM REASONING  \nZibin Meng 1 Peng Xie 1 Kani Chen 1∗  \n{zmengal, [pxieaf}@connect.ust.hk](pxieaf}@connect.ust.hk) [makchen@ust.hk](makchen@ust.hk)[ ](makchen@ust.hk)1The Hong Kong University of Science and Technology  \nABSTRACT  \nReinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals. Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-layer likelihoods and ignore the transformer’s internal forward-pass dynamics. We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal. We observe that stronger reasoners—including RL-posttrained variants—exhibit a distinctive “late collapse”: logit-lens–decoded entropy stays elevated until deeper layers before converging. We define a persequence collapse depth D (x) and a joint objective π(x) ∝ p(x)α exp􀀀β D(x)􀀁 that combines sequence likelihood with this depth-entropy structure, instantiated inside an MCMC power-sampling framework (DEGS-MCMC) . Across three openweight models and four reasoning benchmarks, this near-chance per-candidate signal compounds over the sampling trajectory into state-of-the-art training-free accuracy, with gains largest out of domain and on the harder splits—exactly where likelihood alone falls short—at single-digit-percent wall-clock overhead. DEGS narrowly trails an in-house GRPO reference on the math splits GRPO was trained for, yet surpasses it out of domain on GPQA for all three models, without any training, reward model, or labeled data.  \n1 INTRODUCTION  \nReinforcement learning (RL) with verifiable rewards—e.g., Group Relative Policy Optimization (GRPO) (Shao et al., 2024)—is the dominant paradigm for improving the reasoning capabilities of large language models (LLMs), posttraining frontier models to sizeable gains on mathematical, scientific, and coding benchmarks (Guo et al., 2025) . Yet these gains are costly: RL posttraining requires curated data, extensive tuning, and—most critically—a reliable reward signal, which is unavailable in many domains of practical interest (Prabhudesai et al., 2025) . A growing body of evidence further suggests they may not reflect fundamentally new capabilities: RL-posttrained models concentrate probability mass on traces that already have high likelihood under the base model (He et al., 2025; Yue et al., 2025), and base models can even outperform them in the multi-shot regime due to degraded diversity (Song et al., 2025) . This distribution sharpening hypothesis (Shao et al., 2025) casts RL mainly as a filter, redistributing pass@k capability into single-shot performance rather than teaching genuinely novel reasoning.  \nThis has inspired training-free methods that sharpen the base distribution at inference time, part of abroader move toward scaling test-time computation (Snell et al., 2024; Welleck et al., 2024; Brown et al., 2024) . Karan and Du (2025) formalize this through the power distribution p(x)α (α > 1), targeted by Metropolis–Hastings (MH) sampling, to match GRPO without training, data, or reward models; Scalable Power Sampling (Ji et al., 2026) and Power-SMC (Azizi et al., 2026) improve efficiency via token-level approximation and a batch-parallel SMC scheme. Together they show base models are far more capable at single-shot reasoning than standard sampling reveals.  \n∗ Corresponding author.  \nDespite their success, all existing training-free methods rely exclusively on the output-layer likelihood p (x), ignoring the structured intermediate representations of the forward pass (Nostalgebraist, 2020; Belrose et al., 2023; Wendler et al., 2024) . This is a missed opportunity: Wendler e","cbCaijAtJzcCXXvi","https://ap.wps.com/l/cbCaijAtJzcCXXvi","pdf",2705102,1,21,"English","en",105,"# Abstract\n# Introduction\n# Contributions","[{\"question\":\"What problem does DEGS address in training-free LLM reasoning?\",\"answer\":\"Existing training-free methods mostly rely on output-layer likelihoods and ignore structured transformer dynamics during the forward pass. DEGS targets this gap by using layer-wise entropy behavior as an intrinsic quality signal at test time.\"},{\"question\":\"What is “late entropy collapse” and how is it used by DEGS?\",\"answer\":\"Stronger reasoners exhibit a distinctive late collapse, where decoded entropy stays elevated until deeper layers before converging. DEGS defines a per-sequence collapse depth D(x) and uses it to reweight samples alongside sequence likelihood in a joint objective.\"},{\"question\":\"How is DEGS implemented and what does it require?\",\"answer\":\"DEGS integrates into a Metropolis–Hastings (MH) power-sampling framework (DEGS-MCMC) without any parameter updates. It does not require reward models or labeled data, relying on the base model’s internal entropy trajectories during inference.\"}]",1784201232,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"depth-entropy-guided-sampling-for-training-free-llm-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/depth-entropy-guided-sampling-for-training-free-llm-reasoning/85121/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DEGS address in training-free LLM reasoning?","Question",{"text":75,"@type":76},"Existing training-free methods mostly rely on output-layer likelihoods and ignore structured transformer dynamics during the forward pass. DEGS targets this gap by using layer-wise entropy behavior as an intrinsic quality signal at test time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is “late entropy collapse” and how is it used by DEGS?",{"text":80,"@type":76},"Stronger reasoners exhibit a distinctive late collapse, where decoded entropy stays elevated until deeper layers before converging. DEGS defines a per-sequence collapse depth D(x) and uses it to reweight samples alongside sequence likelihood in a joint objective.",{"name":82,"@type":73,"acceptedAnswer":83},"How is DEGS implemented and what does it require?",{"text":84,"@type":76},"DEGS integrates into a Metropolis–Hastings (MH) power-sampling framework (DEGS-MCMC) without any parameter updates. It does not require reward models or labeled data, relying on the base model’s internal entropy trajectories during inference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]