[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82153-en":3,"doc-seo-82153-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82153,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","MOSAIC Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models","Vision-Language Models (VLMs) have progressed using homogeneous Transformer designs that process visual and textual data. Research on Large Multimodal Models shows that heterogeneous structures mixing efficient operators (e.g., linear attention) can improve both quality and inference latency, yet existing variants depend on static, handcrafted mixing patterns that poorly match target hardware. MOSAIC addresses this by hardware-aware multi-objective search, transforming homogeneous models into optimized heterogeneous architectures under latency constraints, using multi-objective MIP and two-stage parameter recovery with distillation.","MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models  \nYuncheng Yang* Feiyang Ye* Shixian Luo† Yinna Zhu Lianlei Shan Wangcai Zhao Kuo Zhang Yan Chen Yong Wu† Yan Xie  \nLiAuto Inc.  \narXiv :2607 .09029v1 [ cs .CV] 10 Jul 2026  \nAbstract  \nVision-Language Models (VLMs) have achieved significant success by employing homogeneous Transformer architectures topro cess multimedia information, specifically visual and textual data. Recent studies on Large Multimodal Models (LMMs) indicate that heterogeneous structures interleaving efficient mechanisms, such as linear attention, have demonstrated improvements in both performance and inference latency compared to homogeneous designs. However, these efforts rely on handcrafted designs with static mixing patterns, which are inherently sub-optimal and difficult to adapt to specific hardware deployment targets. To bridge this gap, we propose Multi-Objective Search for Adaptive Interlayer Composition (MOSAIC), a hardware-aware search method that automatically transforms homogeneous models into optimized heterogeneous architectures. MOSAIC integrates diverse efficiency mechanisms, including linear, sparse, and low-rank operators, into a unified search space. By formulating the selection process as a multi-objective Mixed Integer Programming (MIP) problem, our method identifies optimal configurations that maximize downstream performance under strict hardware latency constraints. To mitigate the performance degradation arising from structural transitions, we introduce a two-stage parameter recovery process. We first perform global off-policy distillation to stabilize the model’s internal representations, followed by a dualteacher on-policy distillation strategy that leverages a 235B oracle teacher for knowledge expansion while utilizing the original 4B teacher to maintain distributional stability. We validate the effectiveness of MOSAIC through MOSAIC-4B, a heterogeneous model derived from Qwen3-VL-4B-Instruct. Experimental results demonstrate that MOSAIC-4B matches the performance of the Qwen3-VL-4B-Instruct baseline across multiple benchmarks while requiring less than 2% of the training cost of the original model. Furthermore, MOSAIC-4B substantially improves inference efficiency, achieving a 1.76 × prefilling speedup and 2.54 × decoding acceleration. Our MOSAIC-4B model is publicly available at [https://huggingface.co/LiAuto-DSR/MOSAIC-4B](https://huggingface.co/LiAuto-DSR/MOSAIC-4B).  \n*Equal contribution.†Corresponding author.  \nAverage Performance (%)  \n\n|  |  | 2.54× Speedup |  |  |\n| --- | --- | --- | --- | --- |\n|  |  |  |  |  |\n| \u003Cbr>Qwen3-VL-4B | \u003Cbr>-Instruct Int | ernVL3.5-4B | MOSA (O | \u003Cbr>IC-4Burs) |\n| Molmo2-4 | Qwen B | 3-VL-2B-Instru | ct\u003Cbr>\u003Cbr>InternVL3.5 | -2B |\n|  |  Smol | VLM2-2.2B |  |  |\n|  |  |  |  |  |\n\n80  \n70  \n60  \n50  \n0.5× 1 .0× 1 .5× 2.0× 2.5×  \nDecoding Acceleration  \nFigure 1 . Average performance on image understanding benchmarks vs. decoding acceleration, measured by time per output token (TPOT) . MOSAIC-4B matches the performance of the teacher model (Qwen3-VL-4B-Instruct) while achieving a 2.54 × TPOT speedup.  \n1. Introduction  \nVision-Language Models (VLMs) [20, 38, 49, 60, 72] playa pivotal role in modern real-world multimodal systems, ranging from autonomous driving [33, 69, 76] to embodied AI [7, 26, 63, 73] . Standard VLM designs predominantly rely on homogeneous architectures, which feature a fixed sequential interleaving of dense self-attention and multi-layer perceptron (MLP) layers. However, the scalability of these architectures is significantly challenged by the quadratic time complexity of standard attention mechanisms [56] . This complexity results in prohibitive computational overhead, particularly in long-context scenarios [19, 34] or under specific hardware deployment constraints.  \nRecent studies on LMMs indicate that heterogeneous structures, which interleave efficient mechanisms such as linear att","cbCaies9QT5uBxC2","https://ap.wps.com/l/cbCaies9QT5uBxC2","pdf",732906,1,17,"English","en",105,"# Abstract\n# Introduction\n## Limitations of existing heterogeneous architectures\n# (Proposed) MOSAIC method\n## Multi-objective search formulation\n## Two-stage parameter recovery","[{\"question\":\"What problem does MOSAIC address in vision-language models?\",\"answer\":\"MOSAIC targets the gap between improved heterogeneous VLM designs and the need for hardware-adaptive, non-handcrafted architectures that meet strict latency constraints.\"},{\"question\":\"How does MOSAIC search for an optimized heterogeneous architecture?\",\"answer\":\"MOSAIC formulates inter-layer composition selection as a multi-objective Mixed Integer Programming (MIP) problem, searching within a unified space of efficiency mechanisms such as linear, sparse, and low-rank operators.\"},{\"question\":\"What technique does MOSAIC use to reduce performance degradation after structural transitions?\",\"answer\":\"It introduces a two-stage parameter recovery pipeline: global off-policy distillation to stabilize internal representations, followed by dual-teacher on-policy distillation using a 235B oracle teacher for knowledge expansion while retaining stability with the original 4B teacher.\"}]",1784178479,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mosaic-adaptive-inter-layer-composition-for-efficient-heterogeneous-vision-language-models","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mosaic-adaptive-inter-layer-composition-for-efficient-heterogeneous-vision-language-models/82153/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MOSAIC address in vision-language models?","Question",{"text":75,"@type":76},"MOSAIC targets the gap between improved heterogeneous VLM designs and the need for hardware-adaptive, non-handcrafted architectures that meet strict latency constraints.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MOSAIC search for an optimized heterogeneous architecture?",{"text":80,"@type":76},"MOSAIC formulates inter-layer composition selection as a multi-objective Mixed Integer Programming (MIP) problem, searching within a unified space of efficiency mechanisms such as linear, sparse, and low-rank operators.",{"name":82,"@type":73,"acceptedAnswer":83},"What technique does MOSAIC use to reduce performance degradation after structural transitions?",{"text":84,"@type":76},"It introduces a two-stage parameter recovery pipeline: global off-policy distillation to stabilize internal representations, followed by dual-teacher on-policy distillation using a 235B oracle teacher for knowledge expansion while retaining stability with the original 4B teacher.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]