[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83421-en":3,"doc-seo-83421-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83421,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Latent Memory Palace Reasoning for Control as Autoregressive Variational Inference","Human decision-making mixes immediate actions with longer deliberation, and large language models can adapt “reasoning” by generating intermediate tokens. Extending similar iterative computation to continuous control is difficult because language-space reasoning lacks spatial granularity for precise motion. Latent Memory Palace (LMP) organizes control-relevant information in an autoregressive latent space with iterative retrieval. LMP formulates reasoning as variational inference and trains with a latent-space RL objective, producing strong simulation and real-world performance, interpretable adaptive test-time compute allocation, and a variable-length tokenizer variant, LMP-tok, that boosts downstream autoregressive policies.","Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference  \nChuning Zhu  \nUniversity of Washington  \nEva Xu Jose Barreiros  \nUniversity of Washington Toyota Research Institute  \nKrishnan Srinivasan Paarth Shah Abhishek Gupta  \nToyota Research Institute Toyota Research Institute University of Washington  \narXiv :2607 .08724v 1 [ cs .LG] 9 Jul 2026  \n[https://weirdlabuw.github.io/lmp/](https://weirdlabuw.github.io/lmp/)  \nAbstract: Human decision-making is highly flexible – some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive “reasoning.” However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive. Our method, Latent Memory Palace (LMP), formulates reasoning as variational inference with an autoregressive latent distribution. We derive a latent-space reinforcement learning technique to tractably optimize its variational lower bound. The resulting policy, LMP-π, achieves strong empirical performance in simulation and real-world domains while exhibiting interpretable, adaptive allocation of test-time compute. We further show that the same framework yields a variable-length action tokenizer, LMP-tok, which significantly improves the performance of downstream autoregressive policies. Together, these results present a new perspective on latent reasoning for control through the lens of variational inference.  \n1 Introduction  \nHuman decision-making ranges from the reflexive (e.g. walking) to the deliberate (e.g. playing chess), with variability in the time and effort devoted to each decision. This process is both iterative and adaptive, proceeding through a chain of logical steps that scales with the complexity of the decision. Modern large language models (LLMs) exhibit a similar pattern: a broad class of “reasoning” LLMs achieves substantially improved performance by generating intermediate tokens before the final answer [1–3] .  \nWe ask: can iterative, adaptive computation benefit sequential decision making problems such as robotics? For robotic policies, we posit that the benefits are twofold: efficient decision making under task variability, and improved generalization. However, directly transferring methods from language models is unlikely to suffice, as language tokens may not capture the nuances required for spatial understanding and precise motions [4] . We therefore propose that robotic policies reason iterativelyin a latent space learned end-to-end for action prediction. This preserves the benefits of iterative, adaptive computation while allowing the intermediate representations to capture control-relevant information at the appropriate granularity.  \nTo enable latent-space reasoning in robotic policies, we build on variational inference, a principled framework for learning latent-variable models by approximating posterior distributions over latent variables [5] . Applied to control, this perspective interprets an action not as a direct prediction from an observation, but as the outcome of an intermediate latent computation. However, standard variational inference typically uses fixed-dimensional latent variables, unable to adapt their computation to each input. To address this, we propose Latent Memory Palace (LMP), a formulation of reasoning as  \nobs act  \n(a) (b)  \nFigure 1: (a) Latent Memory Palace (LMP) formulates iterative and adaptive reasoning as variational inference with a variable-length autoregressive latent distribution. (b) Applying LMP to decisionmaking results in control policies that adaptively allocate their test-time computation.  \nvariational inference with ","cbCaisP2GYwtMflF","https://ap.wps.com/l/cbCaisP2GYwtMflF","pdf",13416112,4,1,26,"English","en",105,"# Introduction\n## Iterative reasoning for robotics\n## Variational inference for latent control\n## Latent Memory Palace (LMP)\n## Training and test-time inference\n## Experiments and results","[{\"question\":\"What problem does Latent Memory Palace (LMP) address in transferring language-model reasoning to robotics?\",\"answer\":\"Directly transferring language-model reasoning is challenging because language tokens do not provide the spatial granularity needed for precise motion. LMP aims to enable iterative, adaptive computation for sequential control by reasoning in a latent space learned end-to-end for action prediction.\"},{\"question\":\"How does LMP formulate reasoning for control?\",\"answer\":\"LMP formulates reasoning as variational inference using an autoregressive latent distribution. It encodes observation–action information into a variable-length sequence of latent tokens, with an EOS mechanism enabling retrieval to stop when enough information has been gathered.\"},{\"question\":\"What improvements does the framework provide beyond the main control policy (LMP-π)?\",\"answer\":\"Dropping observation conditioning yields LMP-tok, a variable-length sequential action tokenizer. LMP-tok compresses continuous actions into discrete latent tokens, significantly improving downstream autoregressive policies.\"}]",1784187482,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"latent-memory-palace-reasoning-for-control-as-autoregressive-variational-inference","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/latent-memory-palace-reasoning-for-control-as-autoregressive-variational-inference/83421/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Latent Memory Palace (LMP) address in transferring language-model reasoning to robotics?","Question",{"text":75,"@type":76},"Directly transferring language-model reasoning is challenging because language tokens do not provide the spatial granularity needed for precise motion. LMP aims to enable iterative, adaptive computation for sequential control by reasoning in a latent space learned end-to-end for action prediction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LMP formulate reasoning for control?",{"text":80,"@type":76},"LMP formulates reasoning as variational inference using an autoregressive latent distribution. It encodes observation–action information into a variable-length sequence of latent tokens, with an EOS mechanism enabling retrieval to stop when enough information has been gathered.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does the framework provide beyond the main control policy (LMP-π)?",{"text":84,"@type":76},"Dropping observation conditioning yields LMP-tok, a variable-length sequential action tokenizer. LMP-tok compresses continuous actions into discrete latent tokens, significantly improving downstream autoregressive policies.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]