[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85846-en":3,"doc-seo-85846-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85846,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","When Does Depth Survive Composition? Compute-Quality Regimes in Latent World Models","Adaptive-compute latent world models that use early exits or mixture-of-depth predictors rely on two premises: deeper per-step prediction improves rollout quality, and depth can be routed adaptively. The study tests whether this per-step precision survives autoregressive composition using a single pre-registered shallow penalty metric across matched single-step (K=1) and multi-step (K=4) training. Results show three regimes across nine DeepMind Control tasks: intrinsic benefit, inversion where shallow beats full depth, and a flat case, with causal ablation and planning transfer analyses.","arXiv :2607 . 10203v 1 [ cs .LG] 11 Jul 2026  \nWhen Does Depth Survive Composition? Compute–Quality Regimes in Latent World Models  \nAchyuthan Sivasankar  \nNew York University  \n[as21154@nyu.edu](as21154@nyu.edu)  \nAbstract  \nAdaptive-compute world models—early-exit or mixture-of-depths predictors that spend variable depth per step—rest on two assumptions: that more predictor depth buys better predictions, and that depth can be routed adaptively. In autoregressive rollouts (the planning regime, where the predictor is called many times per encoded observation), the first assumption requires depth’s per-step precision to survive composition. We test this directly. Using a single pre-registered instrument—the shallow penalty ρ = err(shallowest-exit rollout)/err(full-depth rollout)—we classify nine DeepMind Control tasks under matched single-step (K=1) and multi-step (K=4 latent-overshooting) training, three seeds each. We find three regimes: on 6/9 tasks depth genuinely helps rollouts (an intrinsic tradeoff, ρ up to 4.7×), on 2/9 tasks the shallow exits beat the full stack (inversion, ρ down to 0.85×), and one is flat. We then show that the robust inversion (cheetah) is not a property of the dynamics but is created by training: a causal ablation that super  \nvises the early exits only at the first rollout step erases it (ρ : 0 .87 → 1.18 over n=8 seeds, ∆=+0 .31, distributions non-overlapping), while a task with an intrinsic tradeoff is unaffected—a double dissociation we call the routability catch-22, because the per-step deep supervision that makes early exits usable for routing is exactly what trains them to out-roll the full stack. The mechanism is task-specific:  \na second, marginal inversion task is unaffected by the ablation. A simple analysis suggests the regime is partly predictable a priori—observation and action dimensionality and one-step model error each correlate with ρ at |Spearman| ≈ 0.75 (n=9) . We also close the loop to control: inside a CEM planner, ρ predicts whether planning benefits from depth on the tasks where the regime and reward model are robust—most sharply, on the inversion task shallow planning yields higher return than deep. Finally, three cautions for how latent world models are evaluated: a task’s regime depends on the metric space (humanoid is flat in latent space but inverted in observation space), the rollout horizon (the inversion is a short-horizon effect), and the encoder (state-space intrinsic tradeoffs flatten under pixel encoding, one at matched model quality) . All thresholds and decision gates were fixed before the compute campaign; the study includes an airtight pre-registered negative for the hypothesis that motivated it.  \n1 Introduction  \nA latent world model encodes an observation to a latent z and predicts future latents zt+1 = f (zt , at); planning and model-based RL then roll f forward autoregressively, calling it tens of times per encoded observation [Ha and Schmidhuber, 2018, Hafner et al., 2019, Hansen et al., 2024] . Because most of a world model’s compute is spent inside these rollouts, adaptive compute—routing predictor depth per step so that “easy” transitions use fewer blocks—is an attractive way to make planning cheaper. This idea inherits the early-exit / mixture-of-depths (MoD) machinery from su  \nPreprint. Under review.  \npervised and language models [Teerapittayanon et al., 2016, Graves, 2016, Raposo et al., 2024], and it has an implicit premise: that a deeper predictor produces a better next-latent, so that trading depth trades quality, and a router can pick the depth each step needs.  \nWe ask a prior question that this premise takes for granted: does per-step depth precision survive autoregressive composition? If a d-block prediction is more accurate than a 2-block prediction fora single step, is the d-block rollout more accurate than the 2-block rollout after ten composed steps? Only if the answer is yes does routing depth in a planner have anything to route.  \nW","cbCaicvjDvbbuZyl","https://ap.wps.com/l/cbCaicvjDvbbuZyl","pdf",474939,2,1,15,"English","en",105,"# Introduction\n## Contributions\n# Evaluation setup and shallow penalty metric\n## Compute–quality regimes across tasks\n## Causal ablation and routability catch-22\n## Evaluation cautions and metric/horizon/encoder effects\n## Planning transfer with CEM","[{\"question\":\"What question does the paper investigate about adaptive compute in latent world models?\",\"answer\":\"It asks whether additional predictor depth per step continues to improve accuracy after many composed steps in autoregressive rollouts, which determines whether routing depth in planning can work meaningfully.\"},{\"question\":\"How is the paper’s main metric defined and interpreted?\",\"answer\":\"It uses the shallow penalty ρ = err(shallowest-exit rollout) / err(full-depth rollout). Values above 1 indicate depth improves rollout quality, around 1 indicate no advantage, and below 1 indicate shallow outperforms full depth.\"},{\"question\":\"What causes the observed inversion regime where shallow exits beat full depth?\",\"answer\":\"The paper shows the inversion is induced by training: a causal ablation that supervises shallow exits only at the first rollout step removes the inversion on cheetah, while an intrinsic-tradeoff task remains unchanged (a double dissociation), leading to the routability catch-22 explanation.\"}]",1784206674,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-does-depth-survive-composition-compute-quality-regimes-in-latent-world-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-does-depth-survive-composition-compute-quality-regimes-in-latent-world-models/85846/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper investigate about adaptive compute in latent world models?","Question",{"text":75,"@type":76},"It asks whether additional predictor depth per step continues to improve accuracy after many composed steps in autoregressive rollouts, which determines whether routing depth in planning can work meaningfully.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the paper’s main metric defined and interpreted?",{"text":80,"@type":76},"It uses the shallow penalty ρ = err(shallowest-exit rollout) / err(full-depth rollout). Values above 1 indicate depth improves rollout quality, around 1 indicate no advantage, and below 1 indicate shallow outperforms full depth.",{"name":82,"@type":73,"acceptedAnswer":83},"What causes the observed inversion regime where shallow exits beat full depth?",{"text":84,"@type":76},"The paper shows the inversion is induced by training: a causal ablation that supervises shallow exits only at the first rollout step removes the inversion on cheetah, while an intrinsic-tradeoff task remains unchanged (a double dissociation), leading to the routability catch-22 explanation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]