[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86380-en":3,"doc-seo-86380-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86380,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs","Unimodal vision models are known to rely on spurious correlations, but the persistence of similar biases in Multimodal Large Language Models (MLLMs) remains insufficiently understood. This paper studies spurious bias in MLLMs and presents SpurLens, an automated pipeline that uses GPT-4 together with open-set object detectors to identify spurious visual cues without human supervision. Results show two dominant failure modes: accuracy drops when spurious cues are removed, and object hallucinations increase by over 10× when spurious cues are present. The work evaluates multiple MLLMs and datasets with robustness checks, explores mitigation via prompt ensembling and reasoning-based prompting, and includes ablations.","arXiv :2503 .08884v3 [ cs .CV] 13 Jul 2026  \nSpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs  \nParsa Hosseini∗ Department of Computer Science University of Maryland  \nSumit Nawathe∗ Department of Computer Science University of Maryland  \nMazda Moayeri  \nDepartment of Computer Science University of Maryland  \nSriram Balasubramanian  \nDepartment of Computer Science University of Maryland  \nSoheil Feizi  \nDepartment of Computer Science University of Maryland  \n[phoseini@umd. edu](phoseini@umd. edu)  \n[snawathe@umd. edu](snawathe@umd. edu)  \n[mmoayeri@umd. edu](mmoayeri@umd. edu)  \n[sriramb@umd. edu](sriramb@umd. edu)  \n[sfeizi@umd. edu](sfeizi@umd. edu)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= jxw6t5VVNL](https: // openreview. net/ forum? id= jxw6t5VVNL)  \nAbstract  \nUnimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we investigate spurious bias in MLLMs and introduce SpurLens, a pipeline that leverages GPT-4 and open-set object detectors to automatically identify spurious visual cues without human supervision. Our findings reveal that spurious correlations cause two major failure modes in MLLMs: (1) over-reliance on spurious cues for object recognition, where removing these cues reduces accuracy, and (2) object hallucination, where spurious cues amplify the hallucination by over 10x. We investigate various MLLMs and datasets, and validate our findings with multiple robustness checks. Beyond diagnosing these failures, we explore potential mitigation strategies, such as prompt ensembling and reasoning-based prompting, and conduct ablation studies to examine the root causes of spurious bias in MLLMs. By exposing the persistence of spurious correlations, our study calls for more rigorous evaluation methods and mitigation strategies to enhance the reliability of MLLMs.  \nCode: [https://github.com/sudoparsa/SpurLens](https://github.com/sudoparsa/SpurLens)  \n1 Introduction  \nMultimodal large language models (MLLMs) (Wang et al., 2024; Liu et al. , 2024a; Meta, 2024; OpenAI, 2024a) have seen rapid advances in recent years. These models leverage the powerful capabilities of large language models (LLMs) (OpenAI, 2024b; Touvron et al., 2023) to process diverse modalities, such as images  \n∗ Equal Contribution  \nFigure 1: Some failures of GPT-4o-mini identified by SpurLens. (Left) The model fails to recognize objects in the absence of spurious cues. (Right) Spurious cues trigger hallucinations.  \nand text. They have demonstrated significant proficiency in tasks such as image perception, visual question answering, and instruction following.  \nDespite these advancements, MLLMs still exhibit critical visual shortcomings (Tong et al., 2024a;b) . One such failure is object hallucination (Li et al., 2023; Hu et al., 2023; Lovenia et al., 2023; Leng et al. , 2024), where MLLMs generate semantically coherent but factually incorrect content, falsely detecting objects that are not present in the input images, as exemplified in Figure 1 . We hypothesize that many of these failures stem from a well-known robustness issue in deep learning models: spurious bias – the tendency to rely on non-essential input attributes rather than truly recognizing the target object (Ye et al., 2024) . While spurious correlation reliance has been well-documented in single-modality image classifiers, the extent to which it persists in MLLMs remains unclear.  \nUnderstanding spurious correlations is crucial because they can lead to systematic failures. Consider the case of a fire hydrant (Figure 2) . When a fire hydrant appears in a street scene, Llama-3.2 (Meta, 2024) correctly recognizes it 96% of the time. However, when placed in unusual contexts, such as a warehouse, accuracy drops to 83%, suggesting an over-reliance on contextual cues rather than the object itself. Conv","cbCaisACRjviVHuQ","https://ap.wps.com/l/cbCaisACRjviVHuQ","pdf",38146101,4,1,46,"English","en",105,"# Abstract\n# Introduction\n## Spurious bias in MLLMs\n## SpurLens pipeline\n## Experimental findings and failure modes\n## Advantages over prior work","[{\"question\":\"What is SpurLens designed to do in multimodal LLMs?\",\"answer\":\"SpurLens automatically identifies spurious visual cues in MLLMs without human supervision, using GPT-4 to propose candidate cues and open-set object detectors to assess their presence in images.\"},{\"question\":\"What two major failure modes does the paper attribute to spurious correlations in MLLMs?\",\"answer\":\"The paper reports (1) over-reliance on spurious cues for object recognition—removing them reduces accuracy—and (2) object hallucination, where spurious cues amplify hallucinations by over 10×.\"},{\"question\":\"How are the paper’s findings validated?\",\"answer\":\"The study investigates various MLLMs and datasets and validates results using multiple robustness checks, along with analyses of SpurLens components and human-study alignment.\"}]",1784211362,116,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"spurlens-automatic-detection-of-spurious-cues-in-multimodal-llms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/spurlens-automatic-detection-of-spurious-cues-in-multimodal-llms/86380/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is SpurLens designed to do in multimodal LLMs?","Question",{"text":75,"@type":76},"SpurLens automatically identifies spurious visual cues in MLLMs without human supervision, using GPT-4 to propose candidate cues and open-set object detectors to assess their presence in images.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What two major failure modes does the paper attribute to spurious correlations in MLLMs?",{"text":80,"@type":76},"The paper reports (1) over-reliance on spurious cues for object recognition—removing them reduces accuracy—and (2) object hallucination, where spurious cues amplify hallucinations by over 10×.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the paper’s findings validated?",{"text":84,"@type":76},"The study investigates various MLLMs and datasets and validates results using multiple robustness checks, along with analyses of SpurLens components and human-study alignment.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]