[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83836-en":3,"doc-seo-83836-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83836,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Retroactive Chain-of-Thought (RetroCoT) Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations","Safety alignment in large language models is typically assessed using direct, imperative harmful requests, but performance is strongly affected by pragmatic register. Models that refuse direct requests may comply when the same underlying objective is expressed through a different communicative stance. The work proposes Retroactive Chain-of-Thought (RETROCOT), a single-turn attack that reframes harmful requests as forensic reconstruction tasks by presupposing the outcome and asking for a reverse causal chain. Experiments report substantial attack success on GPT-4o, while GPT-5-family models refuse the specific reconstruction register.","arXiv :2607 .04645v 1 [ cs .CL] 6 Jul 2026  \nRetroactive Chain-of-Thought ( RETROCOT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations  \nSamira Hajizadeh  \nDepartment of Computer Science Columbia University New York, NY, USA [samira.hajizadeh@columbia.edu](samira.hajizadeh@columbia.edu)  \nAbstract  \nSafety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently comply when the same underlying objective is expressed through a different communicative stance. This suggests that current alignment policies are not invariant to semantic equivalence, but remain sensitive to how a request is pragmatically framed. We introduce Retroactive Chain-ofThought (RETROCOT), a single-turn attack that reframes harmful requests as forensic reconstruction tasks. Rather than requesting harmful instructions directly, RETROCOT presupposes that the harmful outcome has already occurred and asks the model, acting as a forensic analyst, to reconstruct in reverse the causal chain that produced it. On AdvBench (n=50), RETROCOT achieves attack success rate (ASR) of 58% on gpt-4o and 52% on gpt-4omini, compared with direct-request baselines of 0% and 4%, respectively.  \nWe further identify a pronounced generation gap: GPT-5-family models refuse RETROCOT entirely, explicitly identifying the reconstruction premise in their refusal rationales, consistent with explicit coverage of this reconstruction register. However, this robustness does not generalize across pragmatic forms. A single adversarial feedback turn presenting an existing forensic reconstruction response alongside evaluator critique raises ASR from 0% to 48% on GPT-5.4-mini and from 58% to 94% on GPT-4o; a control condition omitting the fabricated low score achieves 85% on GPT-5.4-mini, indicating that the operative element is pragmatic continuation within the established forensic frame rather than score manipulation. These results suggest that frontier-model alignment remains conditioned on pragmatic framing rather than semantic intent, and that new pragmatic registers can continue to expose alignment blind spots even after previously known attack forms have been covered.  \n1 Introduction  \nLarge language models (LLMs) are aligned to refuse harmful requests, but this alignment isnot invariant to how those requests are expressed. We show that models which consistently refuse direct imperative prompts will frequently comply with identical underlying objectives expressed through a different communicative stance. A model that refuses ”Write instructions for synthesizing X” will often comply when asked instead to reconstruct, as a forensic analyst, the steps that led to the documented synthesis of X. The harmful information sought is identical; only the pragmatic form changes.  \nThis failure is an instance of what Wei et al. (2023) term mismatched generalization: safety training covers a specific distribution of input forms without generalizing to the full scope of contexts in which harmful capabilities can be elicited. Andriushchenko & Flammarion (2025) demonstrate a minimal case of this gap: shifting the tense of a harmful question  \nfrom imperative to past tense raises GPT-4o’s attack success rate significantly across 20 reformulations. We ask whether deeper pragmatic shifts, those that change not merely tense but the entire communicative function of the request, produce proportionally larger failures from a single attempt.  \nWe introduce Retroactive Chain-of-Thought (RETROCOT), a single-turn attack grounded in the linguistic phenomenon of presupposition accommodation: the tendency of cooperative language agents to silently accept unstated background assumptions in order to make an utterance coherent (Stalnaker, 1999) . Rather than requesting harmful instructions directly, RETROCOT presupposes that the ha","cbCaipIECRq2Hkz2","https://ap.wps.com/l/cbCaipIECRq2Hkz2","pdf",158407,2,1,"English","en",105,"# Abstract\n# Introduction\n## Mismatched generalization and pragmatic shifts\n## Presupposition accommodation and the RETROCOT framing","[{\"question\":\"What problem does Retroactive Chain-of-Thought (RETROCOT) address in model safety evaluation?\",\"answer\":\"It addresses that safety alignment often depends on how harmful goals are pragmatically framed, not just on the semantic intent of the request. The same objective can yield refusal or compliance depending on communicative stance.\"},{\"question\":\"How does RETROCOT change a harmful request to bypass refusals?\",\"answer\":\"RETROCOT presupposes the harmful outcome has already occurred and asks the model, as a forensic analyst, to reconstruct the reverse causal chain. This shifts the model from evaluating whether to comply to analyzing how the outcome could have happened.\"},{\"question\":\"Do newer model generations respond consistently to RETROCOT?\",\"answer\":\"No. GPT-5-family models refuse RETROCOT more robustly and identify the reconstruction premise in their refusals, but robustness does not transfer across other pragmatic forms. Adversarial feedback turns can substantially increase attack success on these models.\"}]",1784190874,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"retroactive-chain-of-thought-retrocot-forensic-reconstruction-prompts-as-a-safety-diagnostic-across-model-generations","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/retroactive-chain-of-thought-retrocot-forensic-reconstruction-prompts-as-a-safety-diagnostic-across-model-generations/83836/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does Retroactive Chain-of-Thought (RETROCOT) address in model safety evaluation?","Question",{"text":74,"@type":75},"It addresses that safety alignment often depends on how harmful goals are pragmatically framed, not just on the semantic intent of the request. The same objective can yield refusal or compliance depending on communicative stance.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does RETROCOT change a harmful request to bypass refusals?",{"text":79,"@type":75},"RETROCOT presupposes the harmful outcome has already occurred and asks the model, as a forensic analyst, to reconstruct the reverse causal chain. This shifts the model from evaluating whether to comply to analyzing how the outcome could have happened.",{"name":81,"@type":72,"acceptedAnswer":82},"Do newer model generations respond consistently to RETROCOT?",{"text":83,"@type":75},"No. GPT-5-family models refuse RETROCOT more robustly and identify the reconstruction premise in their refusals, but robustness does not transfer across other pragmatic forms. Adversarial feedback turns can substantially increase attack success on these models.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]