[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85899-en":3,"doc-seo-85899-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85899,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Vertical Fusion Condensing Internal Representations for Robust ViT Classification","Vision Transformers (ViTs) expose rich intermediate representations, yet downstream use typically treats them as black-box feature extractors relying almost exclusively on the last layer. The work introduces recoverability: the ability of intermediate representations to correct failures from the last-layer classifier. Across 16 datasets, independent probes at every ViT depth correctly classify 18%–76% of samples misclassified by the last layer. Gains stem from redundancy–correctness correspondence rather than predictive diversity. VFusion aggregates internal hierarchy via a learnable low-dimensional latent mapping, outperforming baselines in both in-distribution and out-of-distribution settings.","arXiv :2607 . 10391v1 [ cs .CV] 11 Jul 2026  \nVertical Fusion: Condensing Internal Representations for Robust ViT Classification  \nFrancesco Di Salvoa , Shyam Nandan Raia , Hamed Damirchib , Ignacio Meza De la Jarab , Sebastian Doerricha , Marco Lentsa , Christian Lediga  \na University of Bamberg, Bamberg, Bavaria, Germany bAIML, The University of Adelaide, Adelaide, South Australia, Australia  \nAbstract  \nDespite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks. We challenge this convention by introducing the notion of recoverability: the capacity of intermediate representations to correct lastlayer failures. By evaluating independent classification probes at every model depth across 16 datasets, we observe that intermediate probes correctly classify 18% to 76% of samples that the last-layer probe misclassifies. We show that these gains are not primarily driven by predictive diversity, but by a redundancy-correctness correspondence, where the internal hierarchy acts as a series of stable, redundant probes of a shared discriminative signal. While established horizontal ensemble strategies (i.e., across multiple models) can improve performance, they incur high computational cost and ignore this vertical signal within a single model. To bridge this gap, we propose VFusion, a principled vertical aggregation strategy employing a learnable mapping into a low-dimensional latent space that synthesizes features across the internal ViT hierarchy. VFusion substantially outperforms established aggregation baselines in both in-distribution and out-of-distribution settings, notably closing 45% of the accuracy gap between the best individual layer and a theoretical oracle performance. Our gains consistently generalize across model sizes and pre-training regimes, confirming that VFusion offers a robust and efficient alternative to horizontal ensemble methods.  \nEmail address: [francesco.di-salvo@uni-bamberg.de](francesco.di-salvo@uni-bamberg.de) (Francesco Di Salvo)  \nThe code is available [at github.com](at github.com/francescodisalvo05/vit-vertical-fusion)[/](at github.com/francescodisalvo05/vit-vertical-fusion)[francescodisalvo05](at github.com/francescodisalvo05/vit-vertical-fusion)[/](at github.com/francescodisalvo05/vit-vertical-fusion)[vit-vertical-fusion](at github.com/francescodisalvo05/vit-vertical-fusion). Keywords: Vertical Ensembling, Internal Representations, Recoverability, ViT  \n1. Introduction  \nVision Transformers (ViTs) have achieved remarkable performance across a wide range of computer vision tasks [1, 2, 3], enabled by large-scale pre-training. Despite these advances, their reliability in real-world deployment remains a critical concern [4] . In safety-critical domains such as healthcare [5] and autonomous driving [6], robustness to distribution shifts and inherent sample difficulty is fundamental. Current paradigms in uncertainty quantification [7], calibration [8, 9], and robustness evaluation [10, 11, 9] share a common implicit assumption: the last ViT layer provides the most informative representation for downstream applications. Consequently, predictive signals are almost exclusively derived from last-layer features. While effective, this reliance on the last-layer representation assumes the model has optimally retained and condensed all task-relevant information at the last layer, a design choice that remains largely undiscussed. Unlike the clear hierarchical abstraction found in convolutional neural networks, ViTs maintain a more uniform representational profile across depth [12] . Nevertheless, downstream applications still treat them as naive feature extractors, ignoring the potential of their rich internal representations.  \nRecent works in out-of-distribution (OOD) detection across vision [13], language [14], and vision-language models [15] suggest that this assum","cbCaie5CQSYrdytL","https://ap.wps.com/l/cbCaie5CQSYrdytL","pdf",549795,3,1,40,"English","en",105,"# Introduction\n## Recoverability and Internal Corrective Capacity\n## Limitations of Last-Layer Reliance\n## VFusion: Vertical Aggregation Strategy","[{\"question\":\"What is VFusion and how does it improve classification?\",\"answer\":\"VFusion is a vertical aggregation method that learns a mapping into a low-dimensional latent space to synthesize features across the internal ViT hierarchy. It outperforms aggregation baselines in both in-distribution and out-of-distribution settings and closes part of the accuracy gap toward an oracle performance.\"}]",1784207037,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"vertical-fusion-condensing-internal-representations-for-robust-vit-classification","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/vertical-fusion-condensing-internal-representations-for-robust-vit-classification/85899/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is VFusion and how does it improve classification?","Question",{"text":75,"@type":76},"VFusion is a vertical aggregation method that learns a mapping into a low-dimensional latent space to synthesize features across the internal ViT hierarchy. It outperforms aggregation baselines in both in-distribution and out-of-distribution settings and closes part of the accuracy gap toward an oracle performance.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":22,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]