[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83477-en":3,"doc-seo-83477-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83477,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision Language Models","Vision–Language Models (VLMs) are increasingly used in safety-critical applications, yet remain vulnerable to backdoor attacks. Existing approaches mainly manipulate final outputs, often causing reasoning traces that are inconsistent or detectable. ReShift introduces an aha-moment-driven reasoning-level backdoor framework that redirects internal chain-of-thought (CoT) trajectories while preserving surface coherence. It uses a Poisoned Reasoning-Aware Data Construction pipeline and a Supervised–Reinforcement Joint Optimization strategy, guided by Entropy Rebound with theoretical guarantees. Extensive experiments show high attack success with maintained clean-task performance.","arXiv :2607 .0036 1v 1 [ cs .CR] 1 Jul 2026  \nReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models  \nZhihao Dou 1†, Qinjian Zhao2† , Zhiqiang Gao2, and Sumon Biswas 1  \n1 Case Western Reserve University, Cleveland, OH, USA  \n2 Kean University, Union, NJ, USA  \nAbstract. Vision–Language Models (VLMs) are increasingly deployed in safety-critical applications, yet remain vulnerable to backdoor attacks.  \nExisting methods primarily manipulate final outputs, often producing reasoning traces that are inconsistent or easily detectable. In this paper, we propose ReShift, the novel aha-moment-driven reasoning-level backdoor framework that explicitly redirects the internal chain-of-thought (CoT) trajectory while preserving surface-level coherence. ReShift introduces a Poisoned Reasoning-Aware Data Construction (PRDC) pipeline and a Supervised–Reinforcement Joint Optimization (SRJO) strategy to induce stable trigger-conditioned reasoning shifts. We further formalize Entropy Rebound as a principled signal for characterizing reasoning redirection and provide theoretical guaranties linking entropy gaps to trajectory-level divergence. Extensive experiments demonstrate that ReShift achieves high attack success rates while maintaining cleantask performance and realistic reasoning traces, substantially improving stealthiness against existing defenses. Code can be found at [https:](https:)//[github.com/AlbertZhaoCA/ReShift](github.com/AlbertZhaoCA/ReShift).  \nKeywords: Vision–Language Models · LLM Reasoning · Backdoor attack  \n1 Introduction  \nVision–Language Models (VLMs), including Qwen2.5-VL [3], Gemini [32], and GPT-4v [23], have achieved remarkable progress in multimodal tasks such as image captioning [45, 51], visual question answering (VQA) [1, 29], and visual grounding [42], demonstrating strong fine-grained perception and cross-modal reasoning capabilities. These advances, together with reasoning-oriented prompting techniques such as Chain-of-Thought (CoT) [40], have gradually established a unified paradigm for visual–linguistic understanding and interaction [4, 6 , 15] . Benefiting from these advances, VLMs have been increasingly deployed in highrisk scenarios, including autonomous driving [56], medical image analysis [31], and embodied AI [8, 10 , 38] . By exposing token-level reasoning instead of operating as black-box systems, modern VLMs enhance interpretability and performance, yet simultaneously expand the attack surface, as the reasoning process itself becomes manipulable.  \n†These authors contributed equally.  \n2 Z. Dou et al.  \nExisting backdoor attacks [12, 49] against VLMs primarily operate at the output level. Attacks based on token injection or response rewriting directly override the final answer once the trigger is activated. Although these approaches can effectively control the target output, they neglect the internal reasoning dynamics. As a result, the generated CoT often remains logically aligned with the correct answer, while the final prediction is forcibly altered, leading to internal inconsistencies. Such anomalies may be detected either through human inspection  \nVLM  \nFig. 1: Illustration of (a) BadToken and (b) Our ReShift. Unlike BadToken, which directly overrides the final answer, our ReShift induces a reasoninglevel turn during generation, redirecting the trajectory before convergence and leading to a different final decision, thereby making the backdoor attack more stealthy and harder to detect. More examples can be seen in Supplement D.  \nor by uncertainty-based monitoring mechanisms. For example, Fig. 1(a) shows a reasoning process poisoned by one representative attack, e.g., BadToken [49] that triggers predefined target outputs without altering the model’s internal reasoning trajectory. Although the poisoned model outputs the target answer, the reasoning process is entirely inconsistent with the preceding steps, because it merely forcefully inserts the target answer by c","cbCaimVgCRp4W1gz","https://ap.wps.com/l/cbCaimVgCRp4W1gz","pdf",4055636,4,1,34,"English","en",105,"# Introduction\n## Motivation and threat model\n## Existing output-level backdoor limitations\n## Proposed reasoning-level approach (ReShift)\n## Key objectives and requirements\n# Method overview\n## Poisoned Reasoning-Aware Data Construction (PRDC)\n## Supervised–Reinforcement Joint Optimization (SRJO)\n## Entropy Rebound signal and theory\n# Experiments and results\n## Attack success and stealth evaluation\n## Clean-task performance preservation\n# Code availability","[{\"question\":\"What training components and signals does ReShift use to achieve stable reasoning redirection?\",\"answer\":\"ReShift employs Poisoned Reasoning-Aware Data Construction (PRDC) and a Supervised–Reinforcement Joint Optimization (SRJO) strategy to condition reasoning on the trigger. It also formalizes Entropy Rebound as a principled signal, with theoretical guarantees linking entropy gaps to trajectory divergence.\"}]",1784188280,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"reshift-aha-moment-driven-reasoning-level-backdoor-attacks-on-vision-language-models","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/reshift-aha-moment-driven-reasoning-level-backdoor-attacks-on-vision-language-models/83477/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What training components and signals does ReShift use to achieve stable reasoning redirection?","Question",{"text":75,"@type":76},"ReShift employs Poisoned Reasoning-Aware Data Construction (PRDC) and a Supervised–Reinforcement Joint Optimization (SRJO) strategy to condition reasoning on the trigger. It also formalizes Entropy Rebound as a principled signal, with theoretical guarantees linking entropy gaps to trajectory divergence.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]