[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83342-en":3,"doc-seo-83342-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83342,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Dive Into the Implicit Biases of Low-rank Vision-language Alignment","Vision-language alignment connects pretrained vision encoders to large language models and is commonly treated like full-parameter pretraining. This work challenges that assumption by applying low-rank adaptation to the LLM during alignment and analyzing the implicit biases it introduces. Low-rank alignment reduces compute costs and improves performance across most benchmarks and model scales. It shifts behavior from hallucinatory to conservative, preserves per-token linear separability, and avoids premature entity-level semantic fusion, dubbed LS-curse, supported by geometric and theoretical analysis.","arXiv :2607 .08 194v 1 [ cs .CV] 9 Jul 2026  \nDive into the implicit biases of low-rank vision-language alignment  \nMingjia Shi Shuo Wang 1‡ Xiaobo Wang2† Sifan Zhou 1 Kai Wang3 Tianyu Fu 1 Chenxu Zhao 1† Anyang Su 1 Ping Jiang 1 Minghui Wu 1  \n1 Mininglamp  \n2 Shenzhen University of Advanced Technology  \n3 National University of Singapore  \n{wangshuo.e,[zhaochenxu}@mininglamp.com](zhaochenxu}@mininglamp.com) , [3101ihs@gmail.com](3101ihs@gmail.com) ,  \n[wangxiaobo@suat-sz.edu.cn](wangxiaobo@suat-sz.edu.cn)  \nAbstract. Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM during this stage instead. We find that low-rank alignment not only reduces computational costs but also outperforms full-parameter alignment on most benchmarks. To understand this phenomenon, we systematically characterize the implicit biases introduced by low-rank adaptation during alignment. Empirically, we find that low-rank alignment shifts model behavior from hallucinatory to conservative and preserves per-token linear separability of visual features that full-parameter alignment disrupts, a phenomenon we term LS-curse. Geometrically, lowrank aligned models exhibit more homogeneous and structurally stable visual representations, maintaining modality-specific knowledge rather than prematurely fusing entity-level semantics. Theoretically, we establish two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for the observed structure-preserving behavior. Extensive experiments cover ablation over  \n100 alignment configurations, three families of low-rank operators, and various rank, encoder, and other settings.  \n1 Introduction  \nThe construction of vision-language models (VLMs) hinges on a critical stage: vision-language alignment, where pretrained vision encoders are bridged with large language models (LLMs) through adapter modules and joint training [5, 8, 25] . As illustrated in Figure 1, this stage precedes instruction tuning and establishes the cross-modal feature correspondence upon which all downstream  \ncapabilities depend. Standard practice updates the full LLM parameters during alignment, treating it as a form of pretraining. However, this view constitutes a 3 † Corresponding author. ‡ Project lead.  \n2 M. Shi et al.  \nFig. 1: Vision-language alignment framework (default): a minimum achievable VLM with a cross-modal adapter (MLP) . How does low-rank alignment shape visual tokens?.  \nmisconception: unlike LLM pretraining, which builds linguistic representations from scratch, vision-language alignment operates atop pretrained knowledge and reasoning priors: it is, in fact, supervised fine-tuning, precisely the regime for which low-rank adaptation methods were designed [10, 12, 14, 55] . Yet, applying low-rank methods during this stage remains largely unexplored.  \nWhen low-rank adaptation is applied to the LLM during alignment (i. e. , low-rank alignment ), we observe that not only does the training cost drop substantially, but downstream performance empirically improves across all involved model scales (from 1.4B to 14B) . This improvement holds across diverse benchmarks spanning perception, knowledge, reasoning, and hallucination, and generalizes across three families of low-rank operators (LoRA, LoHa, LoKr) . This finding raises a natural question: how do the low-rank methods reshape the visual features to yield superior representations, and technically, what implicit biases does the low-rank alignment introduce?  \nEmpirical observations. We investigate this question first at the behavior and feature levels. At the behavior level, hypothesis testing on visual perception benchmarks reveals th","cbCaifRsqLJSdd5W","https://ap.wps.com/l/cbCaifRsqLJSdd5W","pdf",2117270,3,1,36,"English","en",105,"# Abstract\n# Introduction\n## Empirical observations\n## Geometric characterization\n## Theoretical analysis","[{\"question\":\"What is LS-curse, and how is it related to visual features?\",\"answer\":\"LS-curse refers to the disruption of per-token linear separability of visual tokens caused by full-parameter alignment, while low-rank alignment preserves the token-wise knowledge structure from the pretrained vision tower.\"}]",1784186885,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"dive-into-the-implicit-biases-of-low-rank-vision-language-alignment","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/dive-into-the-implicit-biases-of-low-rank-vision-language-alignment/83342/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is LS-curse, and how is it related to visual features?","Question",{"text":75,"@type":76},"LS-curse refers to the disruption of per-token linear separability of visual tokens caused by full-parameter alignment, while low-rank alignment preserves the token-wise knowledge structure from the pretrained vision tower.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]