[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81995-en":3,"doc-seo-81995-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81995,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning","Current vision-language models (VLMs) often fail on complex visual tasks that demand consistent, fine-grained reasoning. Existing self-reflection approaches can improve reasoning but typically require large annotated datasets and do not provide explicit reflective behavior during test time. Inspired by neuroscience, this work studies whether mainstream VLMs support backward prediction and introduces an annotation-free framework, BUS, to verify and strengthen reflective reasoning via backward prediction signals. As a model-agnostic plug-in, BUS integrates with SFT and RL, improving multiple benchmarks without labeled self-reflection data.","BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning  \nJiacheng Yang 1 , Tongying Xiao 1 , Yunkai Dang 1 , Cong Wang 1, 2 , Yuekun Yang 1 , Qi Fan 1 , Tianyu Ding3 , Wenbin Li 1, 4∗, Feng Miao2 , Yang Gao 1  \n1 State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China  \n2Institute of Brain-inspired Intelligence, Nanjing University, Nanjing, China  \n3Microsoft, Redmond, WA, USA  \n4 Shenzhen Research Institute of Nanjing University, Shenzhen, China  \n[liwenbin@nju.edu.cn](liwenbin@nju.edu.cn)  \narXiv :2607 .07361v2 [ cs .CV] 10 Jul 2026  \nAbstract  \nCurrent Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and finegrained reasoning. Recent methods aim to train models to facilitate self-reflective reasoning, i.e., reviewing and improving the generated reasoning. However, they require large volumes of annotated data and lack explicit reflective behavior during test time. By contrast, humans perform explicit and efficient self-reflection through mechanisms such as backward prediction, i.e., predicting which current states are likely to precede a given future state. Inspired by neuroscience, this work proposes a novel solution to address these challenges. We first observe and investigate the phenomenon that mainstream VLMs can perform backward prediction, similar to the human brain. A label-free training framework named Brain-inspired Unsupervised Self-reflection (BUS) is proposed to leverage and exploit backward prediction capability to enhance reflective reasoning in complex visual tasks. BUS enables self-verification of reflective reasoning based on backward prediction, providing explicit learning signals under unsupervised conditions. In this way, BUS eliminates reliance on annotated data while improving reasoning performance. Designed as a model-agnostic plug-in, our framework is compatible with popular fine-tuning methods, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) . Initialized from Qwen3-VL-8B, it improves HR-Bench- 8K (+8.0%), HR-Bench-4K (+7.7%), V* Bench (+6.3%), and MME-RealWorld-Lite (+5.8%), proving backward prediction is key to advancing reflective reasoning.  \n1 Introduction  \nRecent breakthroughs in perception and understanding capabilities of Vision-Language Models (VLMs) have shown promise in performing various vision–language tasks, such as visual search (Team et al. 2026; OpenAI 2025; Bai et al. 2025a) and visual question answering (Fan et al. 2026; Google DeepMind 2025; Wang et al. 2025e) . However, improving multimodal reasoning in complex real-world scenarios still presents significant challenges (Wei et al. 2026a) . A primary reason is the presence of inconsistent, unreliable, and incorrect reasoning paths, which negatively affect final performance (Wan et al. 2025) . Recent studies propose leveraging self-reflection strategies to boost reasoning performance, enabling VLMs to review and improve their  \n∗Corresponding author.  \n| If it is a cat, then it should have some feline characteristics. |  |  |\n| --- | --- | --- |\n|  |  |  |\n\nFigure 1: Backward prediction process in question-answering scenarios. To answer the given question, the brain predicts which events are likely to precede an image that possibly contains a cat.  \nown reasoning process (Yang et al. 2026; Shi et al. 2026a) . Self-reflection capability promotes logical coherence, deeper understanding, and rigorous reasoning, which are essential for handling complex problems.  \nAlthough self-reflection is a highly desirable capability of reasoning VLMs, existing models have been found to perform it inefficiently (Yuan et al. 2025) . Cognitive biases in self-reflection strategies do not significantly improve reflective reasoning and can even degrade overall reasoning performance (Zhang et al. 2025a) . Some recent studies have attempted to enhance the self-reflection capability of reasoning VLMs through fin","cbCaiqjsuXY4LVjW","https://ap.wps.com/l/cbCaiqjsuXY4LVjW","pdf",4214689,6,1,9,"English","en",105,"# Abstract\n# Introduction\n## Challenges in multimodal reasoning\n## Limits of existing self-reflection methods\n## Neuroscience motivation: backward prediction\n## Research goals and proposed direction","[{\"question\":\"What is backward prediction and why is it important here?\",\"answer\":\"Backward prediction predicts which prior events or states could precede a future observed state. The work uses it as a mechanism for self-verification and improving reflective reasoning in multimodal tasks.\"}]",1784177473,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"brain-inspired-unsupervised-self-reflection-via-backward-prediction-for-multimodal-reasoning","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/brain-inspired-unsupervised-self-reflection-via-backward-prediction-for-multimodal-reasoning/81995/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What is backward prediction and why is it important here?","Question",{"text":76,"@type":77},"Backward prediction predicts which prior events or states could precede a future observed state. The work uses it as a mechanism for self-verification and improving reflective reasoning in multimodal tasks.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,107,112,115,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":99,"slug":129},19,"General","general"]