[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85567-en":3,"doc-seo-85567-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85567,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Semantic Richness or Geometric Reasoning? The Fragility of VLM’s Visual Invariance","This work investigates the fundamental fragility of state-of-the-art vision-language models under basic geometric transformations. While modern VLMs perform well on semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they fail to maintain robust spatial invariance and equivariance needed to judge object identity under rotation, scaling, and identity matching. A systematic evaluation across symbolic sketches, natural photographs, and abstract art shows sharp performance drops as semantic content becomes sparse, revealing a gap between semantic understanding and spatial reasoning.","arXiv :2604 .01848v3 [ cs .CV] 10 Jul 2026  \nSemantic Richness or Geometric Reasoning? The Fragility of VLM’s Visual Invariance  \nJason Qiu1∗ Zachary Meurer1∗ Xavier Thomas1∗† Deepti Ghadiyaram1 1 Boston University  \n{jasonq, zmeurer, xthomas, [dghadiya](dghadiya}@bu.edu)[}](dghadiya}@bu.edu)[@bu.edu](dghadiya}@bu.edu)[ ](dghadiya}@bu.edu)∗ Equal contribution. †Corresponding author.  \nIf I rotate the ﬁrst image, can I get the second image?  \nDo the two images depict the same character/object but of a different size, scale, or resolution?  \nAre these two images the same characters/ objects?  \n Rotation  Scale  Identity  \nGemini-2.5-Pro  \nGemini-2.5-Pro  \n0 ~~ ~~  \nGemini-2.5-Pro  \nGemini-2.5-Pro  \nGemini-2.5-Pro  \nFigure 1: Failure of visual transformation reasoning across visual domains. Given a pair of images, models are asked to determine whether they depict the same object under transformations of rotation, scale, or identity. While performance remains near-perfect on natural images (Art, Photo), accuracy drops sharply on abstract and symbolic images (Symbolic and Semantic Sketches), particularly for rotation. Results shown are for Gemini-2.5-Pro (Comanici et al., 2025), with similar trends across evaluated MLLMs.  \nAbstract  \nThis work investigates the fundamental fragility of state-of-the-art VisionLanguage Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they exhibit systematic failures at a more fundamental level: lack of robust spatial invariance and equivariance required to reliably determine object identity under simple rotations, scaling, and identity transformations. We demonstrate this limitation through a systematic evaluation across diverse visual domains, including symbolic sketches, natural photographs, and abstract art. Performance drops sharply as semantic content becomes sparse, and this behavior is observed across architectures, model capacities, and prompting strategies. Overall, our results reveal a systematic gap between semantic understanding and spatial reasoning in current VLMs, highlighting the need for stronger geometric grounding in future multimodal systems. Code is available at: [xthomasbu.github.io/visual](xthomasbu.github.io/visual) invariance  \n1 Introduction  \nImagine being presented with a sentence in an unfamiliar script, such as Glagolitic. To identify repeating characters, we tend to rely solely on rigorous geometric analysis – matching curves, angles, and topology of characters. Now, consider performing the same task on your native script. The task becomes almost trivial due to semantic familiarity with the characters bypassing the need for shape reasoning. This human ability to fluidly switch between geometric reasoning and semantic recognition as needed raises a critical question: do present day vision-language models (VLMs) possess similar robustness?  \nTo study this, we evaluate models across a spectrum of semantic granularity, ranging from sparse symbolic sketches and handwritten scripts to texture-rich photographs (Figure 1) . Within these domains, we test three fundamental transformations: rotation, scaling, and identity matching. A VLM truly possessing geometric reasoning should identify if two images depict the same content regardless of the transformation applied. Crucially, this capability should remain consistent across both familiar (e.g., Latin) and unfamiliar (e.g., Glagolitic) scripts, semantic sketches, and real photos, as the underlying reasoning remains identical. If, however, a model’s apparent robustness is merely a byproduct of data familiarity or context, the performance should collapse on semantically sparser content. Our in-depth analysis confirms this suspicion: the performance of even top-tier closed-and open-sourced VLMs collapses on semantically sparse content under basic geometric transformations. We demonstrate that the apparen","cbCaictHRxVXGjiR","https://ap.wps.com/l/cbCaictHRxVXGjiR","pdf",9089319,3,1,34,"English","en",105,"# Abstract\n# Introduction\n## Motivation and Key Question\n## Evaluation Setup and Transformations\n## Summary of Observed Failure Modes","[{\"question\":\"Why does performance drop in the experiments?\",\"answer\":\"Performance drops sharply when semantic content becomes sparse, such as in symbolic sketches and abstract/symbolic domains, indicating reliance on semantic/object labels rather than underlying geometry.\"}]",1784204647,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"semantic-richness-or-geometric-reasoning-the-fragility-of-vlms-visual-invariance","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/semantic-richness-or-geometric-reasoning-the-fragility-of-vlms-visual-invariance/85567/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Why does performance drop in the experiments?","Question",{"text":75,"@type":76},"Performance drops sharply when semantic content becomes sparse, such as in symbolic sketches and abstract/symbolic domains, indicating reliance on semantic/object labels rather than underlying geometry.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]