[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86485-en":3,"doc-seo-86485-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86485,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Gradient-Skipping Relevance Propagation for Efficient Explainability of Vision Transformers","Vision Transformers (ViTs) are challenging to interpret because relevance propagation and attention-flow methods often ignore architectural realities like heterogeneous attention-head importance and residual (skip) connections. Existing techniques frequently assume uniform head contributions and treat skip paths as identities, which causes misassigned relevance. GradSkip introduces adaptive head weighting and skip-aware propagation, redistributing relevance between attention and residual pathways. Experiments on ImageNet1K and BloodMNIST yield state-of-the-art faithfulness with over 14× fewer GFLOPs, and segmentation evaluations improve localization against ground-truth regions.","arXiv :2607 . 10365v1 [ cs .CV] 11 Jul 2026  \nGradient-Skipping Relevance Propagation for Efficient Explainability  \nof Vision Transformers  \nChristopher Buratti 1 , Michele Marchetti 1 ,Federica Parlapiano 1 , Davide Traini 1 ,2 , Domenico  \nUrsino 1 , and Luca Virgili 1∗  \n1 DII, Polytechnic University of Marche  \n2 CHIMOMO, University of Modena and Reggio Emilia  \n∗ Contact Author  \n[c.buratti@pm.univpm.it](c.buratti@pm.univpm.it) ; [michele.marchetti@univpm.it](michele.marchetti@univpm.it) ; [f.parlapiano@pm.univpm.it](f.parlapiano@pm.univpm.it) ;  \n[davide.traini@unimore.it](davide.traini@unimore.it) ; [d.ursino@univpm.it](d.ursino@univpm.it) ; [luca.virgili@univpm.it](luca.virgili@univpm.it)  \nAbstract  \nVision Transformers (ViTs) are difficult to interpret because current methods of relevance propagation and attention flow do not fully consider some key architectural features, such as the uneven importance of attention heads and residual connections. Prior approaches typically assume uniform importance across attention heads; furthermore, they model skip connections as identity paths, leading to inaccurate relevance attribution. To address these issues, we introduce GradSkip, a novel relevance propagation method for ViTs based on adaptive head weighting and skip-aware propagation. GradSkip models the different importance of the attention heads and dynamically distributes relevance between the attention and residual paths. Experiments on ImageNet1K and BloodMNIST demonstrate a state-of-the-art faithfulness of GradSkip while requiring over 14 times fewer GFLOPs than the best-performing existing approaches. Additional evaluations using transformerbased segmentation confirm improved localization and alignment with ground-truth regions.  \nKeywords: Vision Transformers; Explainable AI; Gradient based Explanation; Relevance Propagation Method; Computer Vision  \n1 Introduction  \nVision Transformers (ViTs) have emerged as a powerful alternative to Convolutional Neural Networks (CNNs) for many computer vision tasks, achieving state-of-the-art results on several benchmarks [41] . Unlike CNNs, which rely on local receptive fields, ViTs process images as sequences of patch tokensand use self-attention to capture long-range dependencies [38, 16] . While this design improves accuracy and scalability, it introduces challenges regarding interpretability. It is therefore crucial to identify which image regions influence a model’s predictions in order to make it possible to trust, debug, and understand learned representations. This need is particularly important in high-stakes fields, such as  \nmedical imaging, where understanding the rationale behind a prediction is essential for clinical trust and accountability, and where an opaque decision can have serious consequences.  \nViT explainability methods can be categorized into three approaches: gradient-based, attentionbased, and perturbation-based. Gradient-based approaches estimate feature importance using backpropagated gradients [32, 34, 4] . Attention-based approaches exploit self-attention weights to trace information flow across layers [19, 17] . Perturbation-based approaches evaluate feature relevance by masking or modifying input regions and measuring the impact on model predictions [14, 43, 29] . While these approaches provide useful insights, they have limitations. These include unstable gradient attributions (for gradient-based approaches), limited causal reliability (for attention-based approaches), and high computational costs (for perturbation-based approaches) . Furthermore, most approaches treat attention heads uniformly and overlook residual connections, which hinders accurate relevance propagation in transformer architectures.  \nFundamentally, these limitations stem from a shared issue in the propagation of relevance through the two mechanisms that define a transformer block, i.e. , multi-head self-attention and residual connections [21] . Attention heads are known to pl","cbCaipzQUP0gkIXo","https://ap.wps.com/l/cbCaipzQUP0gkIXo","pdf",27468678,4,1,29,"English","en",105,"# Introduction\n## Explainability approaches for ViTs\n## Limitations from attention-head uniformity and residual paths\n## GradSkip contributions","[{\"question\":\"Why are current ViT explainability methods insufficient for relevance propagation?\",\"answer\":\"They often assume uniform importance across attention heads and model skip connections as identity paths, which blurs how much relevance comes from learned transformations versus residual preservation.\"},{\"question\":\"What are the two main ideas behind GradSkip?\",\"answer\":\"GradSkip uses adaptive attention-head weighting combining gradient-based flow metrics with Gini-based sparsity, and it performs skip-connection-aware relevance propagation that explicitly models attention and residual paths during backward computation.\"},{\"question\":\"How does GradSkip perform in experiments and what efficiency gain does it claim?\",\"answer\":\"On ImageNet1K and BloodMNIST it reports state-of-the-art faithfulness while using over 14× fewer GFLOPs than the best existing methods, and segmentation evaluations show improved localization and alignment with ground-truth regions.\"}]",1784212077,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"gradient-skipping-relevance-propagation-for-efficient-explainability-of-vision-transformers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/gradient-skipping-relevance-propagation-for-efficient-explainability-of-vision-transformers/86485/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are current ViT explainability methods insufficient for relevance propagation?","Question",{"text":75,"@type":76},"They often assume uniform importance across attention heads and model skip connections as identity paths, which blurs how much relevance comes from learned transformations versus residual preservation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two main ideas behind GradSkip?",{"text":80,"@type":76},"GradSkip uses adaptive attention-head weighting combining gradient-based flow metrics with Gini-based sparsity, and it performs skip-connection-aware relevance propagation that explicitly models attention and residual paths during backward computation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does GradSkip perform in experiments and what efficiency gain does it claim?",{"text":84,"@type":76},"On ImageNet1K and BloodMNIST it reports state-of-the-art faithfulness while using over 14× fewer GFLOPs than the best existing methods, and segmentation evaluations show improved localization and alignment with ground-truth regions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]