[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84744-en":3,"doc-seo-84744-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84744,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse","Code language models increasingly support production software workflows for debugging, refactoring, and iterative repair, where evaluation benchmarks assume that the instructions acted upon are correct. This work tests robustness when that assumption fails by running four experiments on algorithmic Python problems with deterministic tests (RunBugRun), including single-pass and iterative repair. Results show models recognize incorrect instructions as wrong yet still obey them, injecting further errors that iterative self-repair cannot recover, leading to irrecoverable semantic collapse.","Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse  \nRaj Jaiswal 1∗, Anany Singh Divy 1∗, Savar Bhasin 1∗, Adi Bajpai 1∗, Tanuja Ganu3 ,  \nRajiv Ratn Shah2  \n1IIIT Delhi 2IIT Kanpur 3Microsoft Research India  \n{jaiswalp, anany23084, savar23497, [adi23035}@iiitd.ac.in](adi23035}@iiitd.ac.in)  \n[tanuja.ganu@microsoft.com](tanuja.ganu@microsoft.com) , [rajivratn@iitk.ac.in](rajivratn@iitk.ac.in)  \n∗Equal contribution  \narXiv :2607 .04537v 1 [ cs . SE] 5 Jul 2026  \nAbstract  \nCode language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benchmark that evaluates them assumes the instructions they act on are correct. We study what happens when that assumption breaks.  \nWe evaluate code language models across four experiments designed to assess whether models resist or obey incorrect instructions in single-pass and iterative repair settings, using the RunBugRun dataset of algorithmic Python problems with deterministic test cases.  \nOur findings reveal a striking behavioral pattern: models correctly identify an incorrect instruction as wrong, then follow it anyway.  \nThis compliance unknowingly introduces errors beyond the original bug, and the corrupted code state cannot be recovered through subsequent self-guided iterative repair, which fails to converge across passes. We term this Blind Obedience, characterize the Ghost (Unknown) Errors it introduces, quantify the proportion of cases where semantic corruption proves irrecoverable, and show that extended reasoning cannot reverse it. These findings surface behavioral properties invisible to passrate evaluation, with direct consequences for code language models deployed in production settings.  \nAll code, prompts, and data are available in the Appendix 9.  \n1 Introduction  \nSoftware development has undergone a fundamental shift as Large language models have moved beyond isolated code generation (Dong et al., 2025 ; Zamfirescu-Pereira et al., 2025 ; Hoda, 2026) making real modifications to real codebases with real consequences. Yet every benchmark that evaluates them assumes the instructions they act on are correct. (Dong et al., 2025 ; Zamfirescu-Pereira et al., 2025 ; Hoda, 2026) into active roles across the full  \nGPT-5 .3 Codex  \nClaude Sonnet 4.6  \nQwen3-Coder  \nGLM-5  \nKimi-K2.5  \nProgressive Dataset Narrowing Across RQ Stages  \n0 100 200 300 400 500 600  \nNumber of Problems  \nFigure 1: RQ1 is the full 538-problem baseline. Problem counts across RQ2–RQ4 after eligibility filtering, each stage runs only the failure cases. RQ3 and RQ4 confirms damage through Ghost (Unknown) errors accumulation and self repairment fails to correct and reverse it.  \ndevelopment lifecycle—debugging, refactoring, testing, and iterative repair (Dong et al., 2025) . Systems such as GitHub Copilot (Stray et al., 2026), Cursor (He et al., 2026), Devin(Li et al., 2025), and Claude Code (Li et al., 2025) are no longer experimental; they are trusted collaborators in production workflows, making real modifications to real codebases with real consequences (Jimenez et al., 2024 ; Yang et al., 2024 ; Xia et al., 2024) . This transition from code assistant to coding agents marks a critical inflection point—one where the stakes of model behavior extend far beyond benchmark performance and into the reliability of software that the world depends on.  \nPrior work (Larbi et al., 2025 ; Wu et al., 2025 ; Agrawal et al., 2025) has studied robustness to structural noise, ambiguous prompts, and incomplete task descriptions, yet none of these settings place a model in direct conflict with a plausible but incorrect instruction while objective evidence contradicts it in real time. Code is uniquely positioned to close this gap. Unlike natural language tasks (Lin et al., 2022 ; Hendrycks et al., 2021 ; Joshi et al., 2017) where correctness is inherently subjective, program correctness","cbCaidE4rwBw3d9Z","https://ap.wps.com/l/cbCaidE4rwBw3d9Z","pdf",5201843,1,43,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is the core problem studied in this paper?\",\"answer\":\"The paper studies what happens when code language models receive incorrect instructions while objective tests show the conflict in real time.\"},{\"question\":\"How do the experiments assess model behavior?\",\"answer\":\"They evaluate models across four experiments in single-pass and iterative repair settings using the RunBugRun dataset of algorithmic Python problems with deterministic test cases.\"},{\"question\":\"What is Blind Obedience and why is it dangerous?\",\"answer\":\"Blind Obedience is the tendency to follow incorrect instructions even after correctly identifying them as wrong. This compliance introduces compounding “Ghost Errors,” and subsequent iterative self-repair fails to converge, causing irrecoverable semantic corruption.\"}]",1784198005,108,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"obey-diverge-collapse-blind-obedience-to-incorrect-instructions-drives-code-llms-to-irrecoverable-code-semantic-collapse","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/obey-diverge-collapse-blind-obedience-to-incorrect-instructions-drives-code-llms-to-irrecoverable-code-semantic-collapse/84744/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core problem studied in this paper?","Question",{"text":75,"@type":76},"The paper studies what happens when code language models receive incorrect instructions while objective tests show the conflict in real time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the experiments assess model behavior?",{"text":80,"@type":76},"They evaluate models across four experiments in single-pass and iterative repair settings using the RunBugRun dataset of algorithmic Python problems with deterministic test cases.",{"name":82,"@type":73,"acceptedAnswer":83},"What is Blind Obedience and why is it dangerous?",{"text":84,"@type":76},"Blind Obedience is the tendency to follow incorrect instructions even after correctly identifying them as wrong. This compliance introduces compounding “Ghost Errors,” and subsequent iterative self-repair fails to converge, causing irrecoverable semantic corruption.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]