[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86035-en":3,"doc-seo-86035-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86035,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Mixture of Cognitive Experts in Large Vision-Language Models","Large Vision Language Models require strong reasoning over both visual and textual input, and recent research links improved performance to cognitive elements such as diverse representations and metacognition. Many perceptual functions are already available from specialized computer-vision experts, yet integrating their outputs into a trustworthy, interpretable, coherent representation remains difficult. A Bloom-inspired evidence-driven multimodal reasoning framework decomposes expert outputs into atomic evidence, produces a staged reasoning trace, and uses a quantitative trace module to clarify evidence usage and reasoning progression, improving perception and reasoning while reducing hallucination.","arXiv :2607 . 10796v1 [ cs .CV] 12 Jul 2026  \nMixture of Cognitive Experts in Large Vision-Language Models  \nRobert Wijaya and Ngai-Man Cheung  \nSingapore University of Technology and Design (SUTD), Singapore robert [wijaya@mymail.sutd.edu.sg](wijaya@mymail.sutd.edu.sg) , ngai-man [cheung@sutd.edu.sg](cheung@sutd.edu.sg)  \n[Abstract.](Abstract. Large Vision Language Models)[ Large Vision Language Models](Abstract. Large Vision Language Models) ([LVLMs](LVLMs)) [require strong rea](require strong rea)soning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations and metacognition, correlate with better performance. Many of the needed perceptual functions are already provided by specialized domain-specific computer vision models, which act as the perceptual subsystem for detecting objects, localizing them, inferring states, recovering spatial layout, and reading text. The key challenge is to integrate these multi-encoder experts into a trustworthy, interpretable, and coherent representation that improves verifiability and reduces hallucinations. This is difficult because vision-language questions span different cognitive levels, yet most LVLM pipelines apply the same perception-reasoning routing regardless of the query’s demands. We propose an evidence-driven multimodal reasoning framework that utilizes a Bloom-inspired taxonomy as a hierarchical reasoning protocol. The two-stage cognitive verbalization first produces a Literal Evidence Summary by decomposing expert outputs into short, atomic evidence statements. It then performs Bloom Verbalization to turn these evidence items into a staged reasoning trace, and a lightweight Reasoning Trace Module quantitatively analyzes the trace to make evidence usage and reasoning progression explicit. Through this integration, we observed several improvements in perception and reasoning abilities.  \nMoreover, the trace module provides quantitative evidence that different queries induce different cognitive entry levels and evidence-use trajectories that enable fine-grained analysis.  \nKeywords: Vision-Language Reasoning · Mixture of Experts · Cognitive Modeling  \n1 Introduction  \nFrom the earliest days of artificial intelligence, researchers and the public alike have been captivated by the idea of building machines that can think and act like humans [37,47,3,49] . Today, with the rapid advances of large language models (LLMs) [46,11,45,6], this long-standing vision feels closer than ever. The line separating human and machine abilities is rapidly fading [12] . These systems are now used routinely in everyday life, and people are increasingly depending on them for decision-making [18], medicine [28], and creative tasks [36] .  \n2 R. Wijaya and N.-M. Cheung  \nLarge vision-language models (LVLMs), in particular, demand strong reasoning across both visual and textual inputs [34,2] . Several studies in cognitive science and machine learning [42,20,15,7,22] suggests that fundamental scene perception emerges from multiple cognitive functions, including detecting objects, localizing them, inferring their states, modeling relationships, extracting spatial scene layout, and capturing non-object cues such as written text. Recent work further suggests that models incorporating cognitive elements tend to correlate with greater task success, where diverse internal representations and metacognitive abilities are particularly critical [22,7] . Luckily, many of these perceptual functions are already well supported by specialized computer vision (CV) models, since most contemporary CV models are domain-specific and designed to excel at one particular task [38,56,23,25,32,48,53] .  \nThe specialized CV models can be viewed as the perceptual part of a visionlanguage system, orchestrating different cognition functions, from detecting object presence, localizing positions, inferring states, extracting spatial layout, and even interpreting non-object cues","cbCaiiJtgifSrXlP","https://ap.wps.com/l/cbCaiiJtgifSrXlP","pdf",4365702,5,1,25,"English","en",105,"# Introduction\n## Cognitive role of specialized vision models\n## Bloom-inspired reasoning protocol","[{\"question\":\"What problem does the paper address for large vision-language models?\",\"answer\":\"It addresses how to integrate multi-encoder specialized vision experts into a representation that is trustworthy, interpretable, coherent, and leads to more verifiable reasoning with fewer hallucinations.\"},{\"question\":\"How does the proposed framework structure multimodal reasoning?\",\"answer\":\"It uses a Bloom-inspired, two-stage cognitive verbalization: first generating a Literal Evidence Summary from decomposed expert outputs, then performing Bloom Verbalization to produce a staged reasoning trace.\"},{\"question\":\"What role does the Reasoning Trace Module play?\",\"answer\":\"It quantitatively analyzes the reasoning trace to make evidence usage and reasoning progression explicit, providing evidence that different queries trigger different cognitive entry levels and trajectories.\"}]",1784207990,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"mixture-of-cognitive-experts-in-large-vision-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mixture-of-cognitive-experts-in-large-vision-language-models/86035/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address for large vision-language models?","Question",{"text":76,"@type":77},"It addresses how to integrate multi-encoder specialized vision experts into a representation that is trustworthy, interpretable, coherent, and leads to more verifiable reasoning with fewer hallucinations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed framework structure multimodal reasoning?",{"text":81,"@type":77},"It uses a Bloom-inspired, two-stage cognitive verbalization: first generating a Literal Evidence Summary from decomposed expert outputs, then performing Bloom Verbalization to produce a staged reasoning trace.",{"name":83,"@type":74,"acceptedAnswer":84},"What role does the Reasoning Trace Module play?",{"text":85,"@type":77},"It quantitatively analyzes the reasoning trace to make evidence usage and reasoning progression explicit, providing evidence that different queries trigger different cognitive entry levels and trajectories.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]