[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85098-en":3,"doc-seo-85098-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85098,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Cognitive-structured Multimodal Agent for Multimodal Understanding Generation and Editing","Unified multimodal models can jointly perform vision-language understanding and image generation/editing, but monolithic designs scale poorly for long-horizon multimodal dialogue due to visual token explosion and unreliable cross-turn visual referencing. This work introduces a Cognitive-structured Multimodal Agent that externalizes visuals into an Episodic Visual Memory and selectively reactivates relevant episodes. It includes perceptual abstraction, cognitive retrieval, and an executive controller, plus a scenario engine with retrieval annotations for reinforcement learning.","Cognitive-structured Multimodal Agent for Multimodal Understanding,  \nGeneration, and Editing  \nFeng Wang 1 *, Canmiao Fu2 , Zhipeng Huang2 , Chen Li2 , Jing LYU2 , Ge Li 1  \n1Peking University 2WeChat Vision, Tencent Inc.  \narXiv :2607 .08497v 1 [ cs .CV] 9 Jul 2026  \nAbstract  \nRecent unified multimodal models demonstrate that a single architecture can jointly perform vision/language understanding and image generation/editing. However, these monolithic designs rely on repeatedly feeding all historical visual and textual inputs into a shared context window, limiting scalability in long-horizon multimodal dialogue due to visual token explosion and unreliable cross-turn visual referencing. In this work, we propose a Cognitive-structured Multimodal Agent that externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant visual episodes during reasoning. The agent consists of a Perceptual Abstraction Engine for structured visual abstraction, a Cognitive Retrieval Engine for cross-turn memory retrieval, and a Multimodal Executive Controller for autonomous task inference and action planning. To address the lack of turn-level retrieval supervision in existing multimodal dialogue datasets, we further develop a Unified Scenario Engine that programmatically generates structured multiturn conversations with fine-grained retrieval annotations, enabling reinforcement learning to optimize perceptual abstraction and retrieval policies. We additionally construct a long-horizon visual-dialogue benchmark and stratify it by difficulty to evaluate episodic visual recall. Extensive experiments show that our 8B agent achieves 91 .4% retrieval accuracy over 20-turn sessions, surpassing 32B baselines by +8.2%, while nearly halving per-turn inference time (23.1s → 12.7s). We further present the Cognitive-structured Multimodal Agent Harness (CMA-Harness), a tool-augmented deployment of the same cognitive structure that integrates persistent multimodal memory, web access, image generation/editing/composition tools, and OpenAI-compatible serving. These results suggest that structured memory and modular decision-making provide a more scalable and efficient paradigm for long-horizon multimodal agents than monolithic parameter scaling. Our code, dataset, and project page ([caseclose.github.io/cma-harness](caseclose.github.io/cma-harness)) will all be released.  \n*Work done during an internship at WeChat Vision, Tencent Inc.  \n1. Introduction  \nUnified multimodal models have recently shown that a single architecture can jointly perform vision-language understanding and image generation and editing. Recent approaches [6, 7, 12, 22, 34, 39] consolidate perception and generation within a shared parameter space and achieve strong performance across diverse tasks. These systems typically treat multimodal interaction as autoregressive prediction over a unified token sequence.  \nWhile effective for short-context interaction, this unified paradigm exhibits structural limitations in long-horizon multimodal dialogue. In practical image-text conversations, users frequently reference images introduced many turns earlier, request iterative modifications, or switch between understanding and generation without explicit task indicators. Figure 1 shows one such session, spanning 20 turns and four distinct topics. Under such settings, repeatedly injecting all historical visual tokens into the context window leads to two major issues. First, visual tokens are substantially more expensive than text tokens; as dialogue length increases, token usage grows rapidly and crowds out the context budget available for reasoning. Second, reliance on implicit attention over extended contexts weakens cross-turn visual referencing, resulting in retrieval errors and semantic drift. These challenges indicate that parameter scaling alone is insufficient for sustained multimodal interaction. Indeed, a strong unified model (BAGEL [7]) retrieves the correct v","cbCairSM2yqcxv75","https://ap.wps.com/l/cbCairSM2yqcxv75","pdf",6762600,3,1,16,"English","en",105,"# Abstract\n# Introduction\n## Long-horizon limitations of unified multimodal models\n## Proposed memory-structured agent approach\n## Benchmarking and results","[{\"question\":\"Why do unified multimodal models struggle in long-horizon multimodal dialogue?\",\"answer\":\"They must repeatedly inject all historical visual tokens into a shared context, which rapidly consumes the context budget and weakens cross-turn visual referencing, leading to retrieval errors and semantic drift.\"},{\"question\":\"What is the core idea behind the Cognitive-structured Multimodal Agent?\",\"answer\":\"It externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant visual episodes during reasoning, instead of treating images as persistent context tokens.\"},{\"question\":\"How does the paper improve training supervision for turn-level visual retrieval?\",\"answer\":\"It develops a Unified Scenario Engine that programmatically generates structured multiturn conversations with fine-grained retrieval annotations for reinforcement learning to optimize abstraction and retrieval policies.\"}]",1784201097,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cognitive-structured-multimodal-agent-for-multimodal-understanding-generation-and-editing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cognitive-structured-multimodal-agent-for-multimodal-understanding-generation-and-editing/85098/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do unified multimodal models struggle in long-horizon multimodal dialogue?","Question",{"text":75,"@type":76},"They must repeatedly inject all historical visual tokens into a shared context, which rapidly consumes the context budget and weakens cross-turn visual referencing, leading to retrieval errors and semantic drift.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea behind the Cognitive-structured Multimodal Agent?",{"text":80,"@type":76},"It externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant visual episodes during reasoning, instead of treating images as persistent context tokens.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper improve training supervision for turn-level visual retrieval?",{"text":84,"@type":76},"It develops a Unified Scenario Engine that programmatically generates structured multiturn conversations with fine-grained retrieval annotations for reinforcement learning to optimize abstraction and retrieval policies.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]