[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86180-en":3,"doc-seo-86180-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86180,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs","DeepBias introduces an adaptive framework to probe social biases in Large Vision-Language Models (LVLMs), addressing limitations of static bias benchmarks that cannot evolve with model responses. The method uses a generation-evolution-probing loop with a ProposerAgent refined via Direct Preference Optimization (DPO) to tailor test distributions to model-specific failures, and a DiggerAgent that performs multi-turn rewrites using a curated skill library to deepen bias exposure. DeepBiasBench, built with an ensemble of five LVLM anchors, provides more challenging, architecture-spanning evaluations.","DeepBias: Adaptive In-depth Probing of Social  \nBiases in LVLMs  \nAnqi Li, Jie Zhang, Member, IEEE, Zhongqi Wang, Graduate Student Member, IEEE, Songkai Xue, Jiahao Wang,  \nShiguang Shan, Fellow, IEEE, and Xilin Chen, Fellow, IEEE  \narXiv :2607 . 11228v1 [ cs .CY] 13 Jul 2026  \nAbstract—While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic “generation-evolution-probing” loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM’s responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from acurated skill library of deepening and rewriting strategies. At each turn, this process is conditioned on the model’s previous response, enabling progressively deeper biases to be exposed. Furthermore, we build a benchmark named DeepBiasBench using our framework. By employing an ensemble of five diverse state-of-the-art LVLMs as anchors, the benchmark captures vulnerabilities shared across architectures. Comprehensive experiments demonstrate the effectiveness of our framework and show that DeepBias provides a challenging benchmark for indepth bias evaluation, establishing an evolutionary paradigm for LVLM safety assessment.  \nIndex Terms—Vision-Language Model, In-Depth Bias Probing, Adaptive Benchmark.  \nI. INTRODUCTION  \nLARGE Vision-Language Models (LVLMs) have demon  \nstrated remarkable capabilities in multimodal understanding and reasoning, enabling applications ranging from visual question answering to visual agents [1]–[3] . However, these models often inherit and amplify social biases embedded in their training data, leading to discriminatory behaviors across age, gender, race, and other sensitive attributes. Although safety alignment techniques such as RLHF [4] can suppress surface-level biased behaviors, underlying biases that persist after alignment remain difficult to quantify.  \nAlthough recent LVLM evaluation suites have expanded general capability and robustness assessment under diverse multimodal settings [5], [6], current bias evaluation protocols still rely on static benchmarks [7]–[9], which typically consist of fixed image-question pairs evaluated in a single turn. Fig. 1 (left) provides an illustration of this conventional evaluation paradigm, where the target model receives no adaptive followup once a response is generated. Although widely used, static benchmarks suffer from three major limitations. First, static benchmarks face a risk of data leakage and benchmark-specific adaptation once they become public. A representative example  \nis GPQA [10], whose score increased from 39% for GPT- 4 [2] to 94. 1% for Gemini 3 . 1 Pro [11] within roughly 1.5 years, far exceeding the estimated PhD-expert baseline of 65–70% . Second, existing methods rely on fixed singleturn queries and cannot generate new test data according to models’ responses. This makes deeper biases difficult to uncover, especially in safety-aligned models, since they may refuse to answer obvious social bias questions. Third, static datasets rarely precisely test the specific bias vulnerabilities of the target model, leading to redundant evaluations on already robust cases while leaving real bias risks unexplored.  \nTo address these limitations, we introduce DeepBias, a dynamic framework for in-depth adversarial probing. By dynamic, we mean that the evaluation process is ad","cbCailEFHB20Rmkl","https://ap.wps.com/l/cbCailEFHB20Rmkl","pdf",9431956,2,1,12,"English","en",105,"# Introduction\n## Limitations of Static Bias Benchmarks\n## DeepBias: Dynamic In-depth Adversarial Probing\n## Two-Level Adaptation: Distribution and Instance\n## Benchmark Construction with Anchor LVLMs","[{\"question\":\"Why do static bias benchmarks fail to evaluate social bias depth in LVLMs?\",\"answer\":\"Static benchmarks use fixed image-question pairs and single-turn queries, so they cannot adapt to model responses or generate new targeted test data. This also increases risks such as data leakage and redundant evaluation on already robust cases.\"},{\"question\":\"How does DeepBias adapt test data during evaluation?\",\"answer\":\"DeepBias adapts at both the distribution level and instance level. It uses a ProposerAgent refined with Direct Preference Optimization (DPO) to evolve the test distribution toward the target model’s vulnerabilities.\"},{\"question\":\"What role do the ProposerAgent and DiggerAgent play in in-depth probing?\",\"answer\":\"The ProposerAgent synthesizes and iteratively refines test cases aligned with model-specific failure modes via DPO. The DiggerAgent rewrites each test instance across multiple probing turns, selecting skills from a library to progressively deepen bias exposure.\"}]",1784209152,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"deepbias-adaptive-in-depth-probing-of-social-biases-in-lvlms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/deepbias-adaptive-in-depth-probing-of-social-biases-in-lvlms/86180/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do static bias benchmarks fail to evaluate social bias depth in LVLMs?","Question",{"text":75,"@type":76},"Static benchmarks use fixed image-question pairs and single-turn queries, so they cannot adapt to model responses or generate new targeted test data. This also increases risks such as data leakage and redundant evaluation on already robust cases.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DeepBias adapt test data during evaluation?",{"text":80,"@type":76},"DeepBias adapts at both the distribution level and instance level. It uses a ProposerAgent refined with Direct Preference Optimization (DPO) to evolve the test distribution toward the target model’s vulnerabilities.",{"name":82,"@type":73,"acceptedAnswer":83},"What role do the ProposerAgent and DiggerAgent play in in-depth probing?",{"text":84,"@type":76},"The ProposerAgent synthesizes and iteratively refines test cases aligned with model-specific failure modes via DPO. The DiggerAgent rewrites each test instance across multiple probing turns, selecting skills from a library to progressively deepen bias exposure.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]