[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83305-en":3,"doc-seo-83305-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83305,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs","Large language models (LLMs) are powerful yet vulnerable to adversarial prompts and jailbreak attacks, with existing analyses limited to input-output behavior or coarse attribution. A mechanistic framework is introduced using paired internal computation graphs that encode prompt-specific inference as causal interactions among latent features. Aligning graphs for clean versus attacked prompts shows systematic changes: suppression of safety-relevant components, emergence of attack-specific features, and rerouting of computation paths. Experiments across open-source LLMs and multiple jailbreak benchmarks link structural deviations to unsafe outputs, and causal interventions on identified motifs improve robustness.","arXiv :2607 .07903v 1 [ cs .CR] 8 Jul 2026  \nMechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs  \nAnupam Wagle 1 , Ifrat Ikhtear Uddin 1 , Chaowei Zhang2 , Longwei Wang 1†,  \n[anupam.wagle@coyotes.usd.edu](anupam.wagle@coyotes.usd.edu) , [ifratikhtear.uddin@coyotes.usd.edu](ifratikhtear.uddin@coyotes.usd.edu) , [cwzhang@yzu.edu.cn](cwzhang@yzu.edu.cn) ,[longwei.wang@usd.edu](longwei.wang@usd.edu)  \n1Department of Computer Science, University of South Dakota, USA  \n2 School of Information and Artificial Intelligence, Yangzhou University, Yangzhou, China  \nAbstract  \nLarge language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model’s internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulnerabilities using paired internal computation graphs, which represent prompt-specific inference as structured causal interactions among latent features.  \nBy constructing and aligning computation graphs for clean and attacked prompts, we reveal that adversarial attacks induce systematic transformations of internal reasoning, including suppression of safety-relevant components, emergence of attack-specific features, and rerouting of computation paths. Building on this representation, we propose a unified framework that (i) decomposes computation into invariant, suppressed, and emergent structures,(ii) identifies recurring vulnerability motifs associated with failure modes, and (iii) performs causal interventions on nodes, paths, and subgraphs to directly evaluate their contributions to attack success.  \nThis enables a transition from descriptive attribution to causal diagnosis of model failures. Experiments across multiple open-source LLMs and diverse adversarial and jailbreak benchmarks demonstrate that structural deviations in internal computation graphs strongly correlate with unsafe behaviors. Furthermore, targeted interventions on identified vulnerability motifs improve model robustness, establishing internal computation graphs as a principled foundation for understanding, diagnosing, and mitigating LLM vulnerabilities.  \n1 Introduction  \nLarge language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, including reasoning, coding, and scientific analysis [1, 2] . Despite these advances, they remain highly vulnerable to adversarial perturbations and jailbreak attacks that can induce unsafe, unfaithful, or incorrect outputs [3–5] . Small, often imperceptible modifications to input prompts can lead to disproportionate changes in model behavior, bypassing alignment safeguards or triggering hallucinated responses. These vulnerabilities pose significant challenges for deploying LLMs in safety-critical and decision-support applications.  \nExisting approaches to mitigating these failures primarily operate at the input-output level. Techniques  \nsuch as adversarial training [4], alignment tuning [6], prompt filtering, and retrieval augmentation aim † Corresponding Authors.  \n40th Conference on Neural Information Processing Systems (NeurIPS 2026) .  \nto constrain undesirable behaviors without explicitly modeling how such behaviors arise internally [7–23, 19, 24–29] . While these methods have achieved partial success, they remain fundamentally reactive and offer limited insight into the underlying mechanisms of model failure. In particular, they do not explain how adversarial or malicious inputs alter the model’s internal computation to produce erroneous or unsafe outputs.  \nA growing body of work in mechanistic interpretability suggests that neural network inference can be understood as structured computation over latent","cbCaiqCQQMQ8JvEs","https://ap.wps.com/l/cbCaiqCQQMQ8JvEs","pdf",7981943,1,33,"English","en",105,"# Introduction\n## Vulnerability challenges in LLMs\n## Limitations of input-output and attribution-only methods\n## Mechanistic interpretability and circuit-level analysis\n# Proposed framework\n## Internal attribution graphs for clean vs attacked prompts\n## Aligning paired computation graphs\n## Decomposition and vulnerability motifs\n## Causal intervention and mechanistic validation\n# Experimental evaluation\n## Correlation with unsafe behaviors\n## Robustness gains from targeted interventions","[{\"question\":\"What problem does the paper address in LLM jailbreak research?\",\"answer\":\"It addresses the lack of understanding of how adversarial prompts change a model’s internal reasoning to produce unsafe or unfaithful outputs.\"},{\"question\":\"How does the proposed internal computation graph framework model prompt inference?\",\"answer\":\"Prompt-specific inference is represented as a causal attribution graph where nodes are internal features and edges encode directed influence relationships.\"},{\"question\":\"What kinds of internal changes do the authors observe between clean and attacked prompts?\",\"answer\":\"Adversarial attacks suppress safety-relevant components, induce attack-specific features, and reroute computation paths through structured deviations in internal graphs.\"}]",1784186633,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mechanistic-interpretability-of-llm-jailbreaks-via-internal-attribution-graphs","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mechanistic-interpretability-of-llm-jailbreaks-via-internal-attribution-graphs/83305/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in LLM jailbreak research?","Question",{"text":75,"@type":76},"It addresses the lack of understanding of how adversarial prompts change a model’s internal reasoning to produce unsafe or unfaithful outputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed internal computation graph framework model prompt inference?",{"text":80,"@type":76},"Prompt-specific inference is represented as a causal attribution graph where nodes are internal features and edges encode directed influence relationships.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of internal changes do the authors observe between clean and attacked prompts?",{"text":84,"@type":76},"Adversarial attacks suppress safety-relevant components, induce attack-specific features, and reroute computation paths through structured deviations in internal graphs.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]