[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81772-en":3,"doc-seo-81772-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81772,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning","Multimodal Large Language Models often face a language-space bottleneck: complex visual reasoning must pass through discrete tokens, losing perceptual nuance and causing hallucinations in fine-grained, multi-step tasks. Continuous latent reasoning can address this by learning implicit reasoning pathways, but introduces train-inference mismatch where a posterior conditioned on the ground-truth answer may exploit shortcuts. Asymmetric Mutual Variational Learning (AMVL) resolves the mismatch with a bidirectional dual-KL calibration, formalizes answer leakage, and improves performance on the BLINK benchmark.","arXiv :2607 .0046 1v 1 [ cs .CV] 1 Jul 2026  \nMultimodal Continuous Reasoning via Asymmetric Mutual Variational Learning  \nShijie Li 1 ,2 ∗ Yilin Gao2 ∗ Siyuan Yang2 Tieyuan Chen 1 Chaofan Gan 1 Zhihao He 1 Zicheng Zhao 1 Yuyu Guo2† Weiyao Lin 1† Hang Yu2†  \n1 Shanghai Jiao Tong University  \n2 Ant Group  \n{shijieli, [wylin}@sjtu.edu.cn](wylin}@sjtu.edu.cn)  \n{fhlyhv, [yuyuguo1994}@gmail.com](yuyuguo1994}@gmail.com)  \nAbstract  \nMultimodal Large Language Models (MLLMs) are often constrained by a languagespace bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuance. A promising alternative is continuous latent reasoning, where the goal is to discover implicit reasoning pathways that bridge the multimodal query and the final answer. However, this introduces a severe train-inference mismatch: a training-time posterior, conditioned on the ground-truth answer, can exploit answer-dependent shortcuts. Standard variational training then forces the inference-time prior to mimic a posterior that has access to information unavailable at test time, leading to poor performance. To address this, we propose Asymmetric Mutual Variational Learning (AMVL), a framework that resolves this mismatch via a bidirectional calibration objective. A forward KL divergence trains the target-agnostic prior to match the posterior, while a novel reverse KL divergence simultaneously regularizes the posterior, preventing it from collapsing into inference-incompatible regions and mitigating this “answer leakage”. We provide theoretical analysis formalizing this leakage as prior contamination and prove that our dual-KL objective reduces it. We instantiate AMVL in a latentintegrated MLLM and show that it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.  \n1 Introduction  \nHuman reasoning is inherently multimodal. When we solve a visual puzzle or interpret a complex diagram, we think directly in perceptual, spatial, and abstract representations that language alone cannot fully capture [1, 2] . This poses a fundamental challenge for Multimodal Large Language Models (MLLMs): if the intermediate reasoning process is constrained to the discrete token space of natural language, the model is forced to verbalize visual concepts that are intrinsically continuous and high-dimensional. The result is a systematic language-space bottleneck—reasoning quality is limited not by the model’s representational capacity, but by the expressive constraints of the discrete language space through which all intermediate thought must pass. This bottleneck is particularly acute in vision-language tasks demanding fine-grained spatial abstraction or multi-step planning, where text-based Chain-of-Thought (CoT) can cause models to drift from the visual input, introduce hallucinations, and lose precise perceptual grounding [3, 4] .  \n∗ Equal contribution.  \n†Corresponding author.  \nPreprint.  \nThese limitations have motivated a growing body of work on latent visual reasoning, where models perform intermediate steps directly in a continuous embedding space [5–7] . Recent methods, including LVR [5], Monet [8], and Mull-Tokens [9], have shown promise by replacing discrete reasoning tokens with continuous latent states. However, these pioneering methods share a critical and underexplored limitation: they all rely on explicit, hand-crafted supervision signals—such as reconstruction objectives or alignment losses—to shape what these latent states should encode. The latent reasoning process is thus constrained to encode whatever the designer has pre-specified as important, rather than being free to discover the representations most useful for bridging the input question and the final answer.  \nWe argue that a more principled alternative lies in form","cbCaisoUNPdzXmOh","https://ap.wps.com/l/cbCaisoUNPdzXmOh","pdf",2547620,3,1,34,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why do discrete token-based reasoning methods limit multimodal large language models?\",\"answer\":\"They force intermediate visual reasoning into discrete natural-language tokens, creating a language-space bottleneck that can drift from the visual input and degrade perceptual grounding.\"},{\"question\":\"What train-inference mismatch does continuous latent reasoning introduce?\",\"answer\":\"A training-time posterior conditioned on the ground-truth answer can use answer-dependent shortcuts that are unavailable during inference, leading to miscalibrated priors.\"},{\"question\":\"How does AMVL address the train-inference mismatch?\",\"answer\":\"AMVL introduces a dual-KL calibration with a forward KL term aligning a target-agnostic prior to the posterior and a reverse KL term regularizing the posterior to prevent leakage and collapse into incompatible regions.\"}]",1784176057,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multimodal-continuous-reasoning-via-asymmetric-mutual-variational-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/multimodal-continuous-reasoning-via-asymmetric-mutual-variational-learning/81772/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do discrete token-based reasoning methods limit multimodal large language models?","Question",{"text":75,"@type":76},"They force intermediate visual reasoning into discrete natural-language tokens, creating a language-space bottleneck that can drift from the visual input and degrade perceptual grounding.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What train-inference mismatch does continuous latent reasoning introduce?",{"text":80,"@type":76},"A training-time posterior conditioned on the ground-truth answer can use answer-dependent shortcuts that are unavailable during inference, leading to miscalibrated priors.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AMVL address the train-inference mismatch?",{"text":84,"@type":76},"AMVL introduces a dual-KL calibration with a forward KL term aligning a target-agnostic prior to the posterior and a reverse KL term regularizing the posterior to prevent leakage and collapse into incompatible regions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]