[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83469-en":3,"doc-seo-83469-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83469,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Rosetta Composable Native Multimodal Pretraining","Rosetta introduces a composable native multimodal pretraining framework that expands generative modalities without destroying previously learned language and visual understanding. The approach preserves core foundation knowledge inside global shared experts while assigning modality-specific capabilities to plug-and-play experts. Momentum-Anchored Orthogonal Projection (MAOP) uses optimizer momentum as an implicit semantic anchor to neutralize conflicting gradient components from new modalities. Compared with MoE and MoT baselines under equal active parameters, Rosetta mitigates catastrophic forgetting and improves image generation while enabling cross-modal synergy.","arXiv :2607 .00293v 1 [ cs .CV] 1 Jul 2026  \nRosetta: Composable Native Multimodal Pretraining  \nXiangyue Liu1 Zijian Zhang2 Miles Yang2 Zhao Zhong2 Liefeng Bo2 Ping Tan 1 ∗  \n1HKUST 2Tencent Hunyuan  \nAbstract  \nAchieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete understanding tasks causes severe gradient conflicts. Existing architectures, including standard Mixtureof-Experts (MoE), are highly susceptible to representation overwriting. Even structurally partitioned paradigms like Mixture-of-Transformers (MoT) remain vulnerable to catastrophic forgetting, severely impeding multimodal scalability.  \nIn this work, we introduce Rosetta, a composable native multimodal pretraining framework designed for seamless and non-destructive modality expansion. Rosetta adopts a modular paradigm where core foundational knowledge is preserved within global shared experts, while modality-specific capabilities are distributed across plug-and-play experts. To guarantee non-destructive composition, we propose Momentum-Anchored Orthogonal Projection (MAOP) . MAOP leverages the optimizer’s momentum state as an implicit semantic anchor, selectively neutralizing conflicting gradient components from new modalities while preserving synergistic updates. To strictly isolate the architectural impact, we evaluate Rosetta against standard MoE and MoT baselines under strict active parameter parity. All models are trained from scratch within the Transfusion framework, using discrete next-token prediction for language and continuous visual diffusion. Extensive evaluations demonstrate that, while standard MoE and MoT architectures suffer catastrophic forgetting of previously acquired knowledge, Rosetta robustly preserves established language and visual understanding. Furthermore, it delivers superior image generation and unlocks cross-modal synergy, paving the way for truly composable and unified multimodal foundation models. To facilitate further multimodal research, we release our code and checkpoints to the community.  \nProject page at [https://rosetta-lmm.github.io/](https://rosetta-lmm.github.io/) .  \n1 Introduction  \nThe evolution toward general-purpose AI necessitates foundation models capable of natively integrating diverse modalities within a singular architecture [1, 58], spanning from discrete autoregressive language comprehension to continuous diffusion (or flow matching) visual synthesis [72, 66] . However, integrating these disparate training objectives intrinsically triggers severe gradient conflicts. Specifically, the high-variance gradients from generative tasks tend to overwrite the established representations of language modeling, creating a critical optimization bottleneck.  \nScaling unified models effectively points toward Sparse Mixture-of-Experts (MoE) [22, 9] . Yet, standard MoE architectures typically deploy modality-agnostic routing mechanisms. When exposed to heterogeneous multimodal signals, this unconstrained routing leads to a catastrophic routing collapse: aggressive generative gradients monopolize and irreversibly overwrite the pre-established experts, severely degrading the model’s foundational language (as shown in Fig. 1 Left) and visual  \n∗Corresponding author.  \nPreprint.  \nFigure 1: Escaping the Forgetting-Synergy Dilemma. (Left) Performance dynamics on MMLU benchmark across composable pretraining stages. While standard MoE and structurally isolated MoT suffer from catastrophic routing collapse and degradation upon the integration of continuous generative objectives (+T2I), our Rosetta architecture acts as a robust semantic anchor, maintaining a highly stable foundation. (Right) Qualitative results of Rosetta, demonstrating that the preservation of foundational knowledge seamlessly unlocks superior visual generation capabilities.  \nunderstanding capabilitie","cbCairaun5Jm3o9U","https://ap.wps.com/l/cbCairaun5Jm3o9U","pdf",3670637,3,1,19,"English","en",105,"# Abstract\n# Introduction\n## Forgetting-Synergy Dilemma\n## Rosetta Framework Overview\n## Momentum-Anchored Orthogonal Projection","[{\"question\":\"What problem does Rosetta address in multimodal foundation model training?\",\"answer\":\"Rosetta targets severe gradient conflicts that arise when adding continuous generative objectives alongside discrete understanding tasks, which can overwrite previously learned representations and cause catastrophic forgetting.\"},{\"question\":\"How does Rosetta enable modality expansion without destructive interference?\",\"answer\":\"Rosetta uses a modular expert design: core knowledge is preserved in global shared experts, while modality-specific skills are distributed across plug-and-play experts, enabling non-destructive composition.\"},{\"question\":\"What is MAOP and how does it help with gradient conflicts?\",\"answer\":\"MAOP leverages the optimizer’s momentum state as an implicit semantic anchor to selectively neutralize conflicting gradient components from newly added modalities while preserving synergistic updates.\"}]",1784188177,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rosetta-composable-native-multimodal-pretraining","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rosetta-composable-native-multimodal-pretraining/83469/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Rosetta address in multimodal foundation model training?","Question",{"text":75,"@type":76},"Rosetta targets severe gradient conflicts that arise when adding continuous generative objectives alongside discrete understanding tasks, which can overwrite previously learned representations and cause catastrophic forgetting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Rosetta enable modality expansion without destructive interference?",{"text":80,"@type":76},"Rosetta uses a modular expert design: core knowledge is preserved in global shared experts, while modality-specific skills are distributed across plug-and-play experts, enabling non-destructive composition.",{"name":82,"@type":73,"acceptedAnswer":83},"What is MAOP and how does it help with gradient conflicts?",{"text":84,"@type":76},"MAOP leverages the optimizer’s momentum state as an implicit semantic anchor to selectively neutralize conflicting gradient components from newly added modalities while preserving synergistic updates.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]