[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84986-en":3,"doc-seo-84986-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84986,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces","Adversarial vulnerability in deep neural networks (DNNs) is examined through the spectral structure of intermediate linear transformations that carry information across modern models, instead of focusing only on decision boundaries, feature robustness, Jacobians, or end-to-end behavior. The work targets transformer-based vision–language models (VLMs), leveraging interpretable spectral decompositions of their linear layers. A white-box spectral-subspace-guided attack (SSGRA) aligns intermediate representations with the subspace spanned by bottom right singular vectors. Experiments demonstrate stronger attack effectiveness than baselines and provide a spectral interpretation for improving robustness.","arXiv :2607 .07375v 1 [ cs .LG] 8 Jul 2026  \nOn Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces  \nChethan Krishnamurthy Ramanaik 1 , Tobias Callies 1 , Michael Hecht2 , and  \nEirini Ntoutsi 1  \n1 University of the Bundeswehr Munich, Germany {chethan.krishnamurthy,tobias.callies,[eirini.ntoutsi}@unibw.de](eirini.ntoutsi}@unibw.de)  \n2 University of Wrocław, Poland  \n[michael.hecht@math.uni.wroc.pl](michael.hecht@math.uni.wroc.pl)  \nAbstract. Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input–output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear transformations that propagate information through modern DNNs, an unexplored mechanism of adversarial vulnerability. Specifically, we investigate transformer-based vision–language models, whose linear layers admit interpretable spectral decompositions and whose widespread adoption makes understanding their robustness increasingly important.  \nWe propose a white-box spectral-subspace-guided attack (SSGRA) that aligns intermediate representations with the subspace spanned by the bottom right singular vectors. Our experiments show improved attack effectiveness over existing baselines. In addition, SSGRA offers a spectral interpretation of adversarial vulnerability in VLMs, providing insights for improving their robustness.  \nKeywords: Adversarial attacks · Spectral analysis · VLMs  \n1 Introduction  \nDeep neural networks, including modern vision–language models (VLMs), are known to be vulnerable to adversarial perturbations that substantially alter model predictions while remaining visually imperceptible [25, 57 , 62] . Despite extensive research, the mechanisms governing how such perturbations propagate through deep networks remain incompletely understood. Existing explanations primarily analyze adversarial vulnerability from input-space [16, 22 , 24 , 38 , 55 , 56 , 64] or end-to-end perspectives, including decision-boundary geometry [49], robust and non-robust features [33], Jacobian analysis [31, 36 , 39 , 51], inverse problem instability and Lipschitz properties [2, 3 , 26] . Despite these advances, existing theories predominantly explain adversarial vulnerability from the input space or through end-to-end network properties, leaving the spectral behavior of intermediate linear transformations largely unexplored.  \n2 C.K. Ramanaik et al.  \nFig. 1: Overview of the proposed spectral framework.  \nTransformer-based VLMs provide a natural setting for such an analysis. Their architectures comprise numerous learnable intermediate linear transformations including the projection matrices in self-attention, feed-forward networks, and multimodal fusion modules [4, 14 , 21 , 58], making spectral decomposition a principled tool for studying representations. Because singular values govern how different representation directions (singular vectors) are amplified or suppressed by each linear transformation, they naturally provide a lens for studying how adversarial signals propagate through transformer layers. Moreover, their widespread adoption makes them an important testbed for adversarial robustness.  \nThis motivates us to investigate adversarial vulnerability through the singularvector basis of intermediate linear transformations. Inspired by the instability of ill-posed inverse problems, where near-null singular directions govern information loss, we study how adversarial intermediate representations align with top and bottom singular-vector subspaces during attack optimization. Guided by this perspective, we formulate a spectral-guidance principle and instantiate it through a white-box attack that serves to empirically validate the proposed mechanism. Our results suggest that, beyond constraining large singular values, explicitly controlling near-null singular directions m","cbCailI6BNlpUO6C","https://ap.wps.com/l/cbCailI6BNlpUO6C","pdf",20815927,1,31,"English","en",105,"# Introduction\n## Motivation and Contributions\n# Related Work\n## Theoretical Perspectives\n## Manifolds & Decision Boundaries\n## Non-robust Features\n## Linearity Approximation","[{\"question\":\"What spectral property of transformer-based VLMs does the paper focus on?\",\"answer\":\"It focuses on the spectral structure of intermediate linear transformations, specifically how representations align with singular-vector subspaces during adversarial optimization.\"},{\"question\":\"What is SSGRA and how does it guide the attack?\",\"answer\":\"SSGRA is a white-box spectral-subspace-guided attack that aligns intermediate representations with the subspace spanned by the bottom right singular vectors.\"},{\"question\":\"How do the paper’s experiments position SSGRA relative to existing attack baselines?\",\"answer\":\"Experiments show improved attack effectiveness over existing baselines, and the approach provides a spectral interpretation of adversarial vulnerability in VLMs.\"}]",1784200057,78,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"on-adversarial-vulnerability-of-vision-language-models-through-the-lens-of-intermediate-spectral-subspaces","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/on-adversarial-vulnerability-of-vision-language-models-through-the-lens-of-intermediate-spectral-subspaces/84986/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What spectral property of transformer-based VLMs does the paper focus on?","Question",{"text":75,"@type":76},"It focuses on the spectral structure of intermediate linear transformations, specifically how representations align with singular-vector subspaces during adversarial optimization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is SSGRA and how does it guide the attack?",{"text":80,"@type":76},"SSGRA is a white-box spectral-subspace-guided attack that aligns intermediate representations with the subspace spanned by the bottom right singular vectors.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the paper’s experiments position SSGRA relative to existing attack baselines?",{"text":84,"@type":76},"Experiments show improved attack effectiveness over existing baselines, and the approach provides a spectral interpretation of adversarial vulnerability in VLMs.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]