[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86461-en":3,"doc-seo-86461-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86461,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning","Unified multimodal models (UMMs) that interleave textual and visual reasoning traces can greatly support visual mathematical reasoning, yet intermediate visual steps are not always beneficial. Self-generated visuals may introduce erroneous evidence that misleads later reasoning, and frequent triggering of visual steps increases computation and memory overhead, lowering inference efficiency. AdaViG leverages early internal signals—Generation Intent and Visual Fidelity—to adaptively gate each visual step, aborting weak or ungrounded generations. Experiments show up to 5.7% accuracy gains and 25.0%–91.0% fewer visual-generation FLOPs, with 15.4%–45.6% lower latency.","arXiv :2607 . 10004v1 [ cs .CV] 10 Jul 2026  \nModel Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning  \nWenxi Gao1 , Guanxi Lu1 , Didi Zhu1 , Hao Mark Chen1 , Quan Deng1,2 , Zhican Wang1 , Jiankang Deng1 , Hongxiang Fan1  \n1Imperial College London, 2Tsinghua University  \nAbstract  \nUnified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, we identify a key insight in this paradigm: generating intermediate visual reasoning steps is not always beneficial and can even be harmful, as self-generated visual steps may introduce erroneous visual evidence that misleads subsequent reasoning.  \nMoreover, frequently triggering visual steps during reasoning incurs substantial computational and memory overhead, degrading inference efficiency. To address these accuracy and efficiency challenges, we observe that the model’s internal signals can indicate whether a visual step will benefit reasoning before the entire visual generation is completed. Specifically, this work identifies two internal signals: i) Generation Intent, which reflects whether the model has a concrete textual plan for what to draw, and ii) Visual Fidelity, which measures whether the visual generation remains grounded in the original input image. Leveraging these internal signals, we propose AdaViG, a training-free adaptive visual gating method for unified multimodal reasoning. AdaViG dynamically evaluates each triggered visual step at an early visual generation stage and aborts it when both signals are weak, thereby preventing misleading visual evidence from entering the reasoning trace while avoiding unnecessary computation. Comprehensive experiments demonstrate that AdaViG improves accuracy by up to 5.7% while reducing visual generation FLOPs by 25 .0%–91 .0% and wall-clock latency by 15.4%–45 .6% .  \n1 Introduction  \nRecent advances in multimodal large language models (MLLMs) [1, 42] have shown promising performance in mathematical reasoning, particularly through multimodal chain-of-thought (MCoT) [36, 27, 37] . However, existing MCoT methods with text-only rationales [36, 27] are limited for vision math problems that require explicit visual construction as intermediate reasoning steps, such as drawing auxiliary lines or transforming diagrams. To overcome this limitation, MCoT has extended from text-only rationales to interleaved multimodal rationales [14, 33], where textual and visual steps are interleaved so the model can reason with images.  \nAmong various interleaved MCoT methods, tool-driven approaches [14, 33, 24, 29] operate over a fixed toolkit, such as zoom-in or cropping, limiting their flexibility. Code-based approaches [7, 11, 31, 38] can generate more diverse visual aids but depend on reliable code synthesis and rendering. In contrast, intrinsic approaches [17, 10, 8, 16, 25] enable the model itself to generate visual steps, showing great potential for interleaved reasoning. This paradigm is primarily realized by unified multimodal models (UMMs) [6, 32, 3, 35], with fine-tuned variants such as Bagel-Zebra-CoT [16] and MathCanvas [25] that natively generate visual steps as part of the reasoning.  \nPreprint.  \nQuestion:  \nA right-angled triangle with side lengths a=8, b=15 and c=17 is given. How big is the radius r of the inscribed semicircle shown?  \nFigure 1: A representative harmful generation case from MathVision under the fixed MathCanvas paradigm exhibiting two failure modes (left), and how AdaViG adaptively aborts it (right) .  \nHowever, the decision of whether to emit a textual or visual step in UMMs with CoT is primarily determined by the next action token. This paradigm lacks an explicit mechanism for assessing whether a visual step would actually improve reasoning quality. We identify two systemic issues with this design: (1) Harmful Generation Err","cbCaihfz6CbtBZiN","https://ap.wps.com/l/cbCaihfz6CbtBZiN","pdf",6366815,4,1,17,"English","en",105,"# Introduction\n## Unified multimodal interleaved reasoning\n## Limitations: harmful generation and inefficiency\n## AdaViG: adaptive visual gating signals and early abort","[{\"question\":\"Why can intermediate visual reasoning steps be harmful in unified multimodal reasoning?\",\"answer\":\"Self-generated visual steps may be redundant, incorrect, or misleading, injecting erroneous visual evidence that misleads subsequent text reasoning.\"},{\"question\":\"What efficiency problem does AdaViG address?\",\"answer\":\"Frequent triggering of visual steps substantially increases computational and memory costs, accounting for large portions of inference latency and KV-cache usage.\"},{\"question\":\"How does AdaViG decide whether to continue or abort a visual step?\",\"answer\":\"AdaViG uses internal signals—Generation Intent to check whether there is a concrete textual plan, and Visual Fidelity to measure whether the generation stays grounded in the input image—then aborts when both signals are weak early in generation.\"}]",1784211866,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"model-guides-you-how-to-draw-adaptive-visual-gating-for-unified-multimodal-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/model-guides-you-how-to-draw-adaptive-visual-gating-for-unified-multimodal-reasoning/86461/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can intermediate visual reasoning steps be harmful in unified multimodal reasoning?","Question",{"text":75,"@type":76},"Self-generated visual steps may be redundant, incorrect, or misleading, injecting erroneous visual evidence that misleads subsequent text reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What efficiency problem does AdaViG address?",{"text":80,"@type":76},"Frequent triggering of visual steps substantially increases computational and memory costs, accounting for large portions of inference latency and KV-cache usage.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AdaViG decide whether to continue or abort a visual step?",{"text":84,"@type":76},"AdaViG uses internal signals—Generation Intent to check whether there is a concrete textual plan, and Visual Fidelity to measure whether the generation stays grounded in the input image—then aborts when both signals are weak early in generation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]