[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85618-en":3,"doc-seo-85618-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85618,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","The Shared Substrate of Modern Encoders","Different vision neural networks trained for classification, self-supervised contrastive learning, masked reconstruction, or image–text matching are expected to form distinct internal representations. This work finds they converge: the leading sixteen principal variation directions in fourteen modern vision encoders collapse onto the same sixteen-dimensional geometric object, termed the cross-architecture substrate. The invariant survives Gröger-calibrated CKA across heterogeneous domains and across training, enabling frozen low-shot features, domain detection, distillation gains, and calibrated forensics.","arXiv :2606 .07882v2 [ cs .CV] 12 Jul 2026  \nThe Shared Substrate of Modern Encoders: A Calibration-Surviving Geometric Invariant Across  \nVision  \nand Language, and the One Training-Time Tool That  \nExploits It  \nYousef Radwan  \nKAUST  \n[yousef.radwan@kaust.edu.sa](yousef.radwan@kaust.edu.sa)  \nAbstract  \nDifferent vision neural networks are trained to do very different things—classify ImageNet labels, contrast augmented crops, fill in masked pixels, or match images to captions—and we would expect their internal representations to look correspondingly different. We report that they do not. After training, the top sixteen principal directions of variation inside fourteen modern vision encoders (12 discriminative + 2 MAE controls) converge to the same sixteen-dimensional geometric object, in the same way that independently trained machine-translation systems converge on a shared notion of word meaning. We call this object the cross-architecture substrate and study it with three tools: principal-component analysis to find directions of variation; centred kernel alignment (CKA), the standard measure of how similarly two networks represent a fixed set of images; and the Gröger 2026 calibration, which subtracts the baseline CKA value expected from random data, a known confound in earlier work. The substrate transports across four heterogeneous visual domains (natural photographs, medical CT, satellite RGB, microscopy) at median Procrustes-CKA 0.679, and across eight domains (adding hand-drawn sketches, depth maps, thermal infrared, telescope images of galaxies) at 0.604, with every cross-domain pair ≥ 0.40. The substrate survives Gröger et al.’s calibration both globally (7 .4× separation between classification-style encoders and masked-reconstruction encoders at n=13,394) and under the harder local nearestneighbour-recall variant (4 .82–5.30×, p \u003C 10 −44) . It is not pixel statistics (0 .263), not a random sixteen-dimensional slice (median 0.19 over 50 orthonormal seeds), not driven by any single encoder (±0 .027 when any one of five is removed), and emerges in the first 10% of training while accuracy keeps climbing. We deliver four uses plus bounded scope: a frozen feature space for low-shot learning (16 dimensions beat 768-dimensional DINOv2 features by +3 .78pp at N=50 labels per class); a four-way domain detector (99 .6%); a knowledge-distillation auxiliary loss that beats cross-entropy by +5 . 14pp at epoch 100 and +1 . 19pp at epoch 200 on CIFAR-100/RN-18 and by +5 .63pp on TinyImageNet/RN-50 (closing 64.3% of the trained-teacher gap), with no teacher forward pass at training time and a label-efficiency peak of +8 .35pp at 10% labels; and a Gröger-calibrated crossarchitecture forensic fingerprint—one primitive doing three jobs: provenance (model kind/architecture/clade at ROC-AUC 0.92, same-family vs. unrelated), transform-type classification (finetune/quantize architecture-transferable at LOPO 0.889), and deduplication (0 .986)—that beats raw CKA at low probe-n (+0 .043 at n=100), is robust to quantize/prune/fine-tune-proxy (recovery@1 = 1 .0), holds at REEF’s probe band (0 .916/0 .914 at n=200/300) and in a third (audio) modality  \n40th Conference on Neural Information Processing Systems (NeurIPS 2026) .  \n(0 .917), and clears the Gröger width-matched null (0 .92/0 .916/0 .986 vs. 95thpct ∼0.71); full-tree phylogeny is honestly scoped out (three reconstructions all \u003C 0.5), so it resolves architecture clades, not exact ancestry. The shared substrate is moreover intrinsically low-rank: its effective rank (participation ratio) is ∼3.5 in both vision (3 .52, CI [3 .28 , 3.57]) and audio (3 .56), far below the K=16 design choice, retro-justifying K=16 as capture-not-capacity (the companion LLM valence direction is the rank-1 limit of the same object) . We also tested a labelfree transferability score (subs-rank) as a replacement for LogME and report it as a null: substrate alignment does not predict downstream transfer accuracy","cbCaim5pJ14OqhOu","https://ap.wps.com/l/cbCaim5pJ14OqhOu","pdf",674122,2,1,28,"English","en",105,"# Abstract\n# Introduction\n## Shared geometric invariant across vision encoders\n## Evaluation across heterogeneous domains\n## Downstream tools and bounded scope\n## Extension to language encoders","[{\"question\":\"What is the cross-architecture substrate proposed in the paper?\",\"answer\":\"It is a shared sixteen-dimensional geometric object formed by the top sixteen principal directions of variation across different vision encoders’ penultimate-layer features.\"},{\"question\":\"How is the invariant measured and validated?\",\"answer\":\"The paper uses PCA for variation directions, centred kernel alignment (CKA) for representational similarity, and Gröger 2026 calibration that subtracts baseline CKA expected from random data.\"},{\"question\":\"What practical tools and limits does the paper claim for the substrate?\",\"answer\":\"It provides uses including a frozen low-shot feature space, a domain detector, an auxiliary loss for knowledge distillation, and a Gröger-calibrated forensic fingerprint. The scope is bounded: it does not extend across modalities (vision + audio fails) and does not predict downstream transfer quality.\"}]",1784204964,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-shared-substrate-of-modern-encoders","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/the-shared-substrate-of-modern-encoders/85618/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the cross-architecture substrate proposed in the paper?","Question",{"text":75,"@type":76},"It is a shared sixteen-dimensional geometric object formed by the top sixteen principal directions of variation across different vision encoders’ penultimate-layer features.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the invariant measured and validated?",{"text":80,"@type":76},"The paper uses PCA for variation directions, centred kernel alignment (CKA) for representational similarity, and Gröger 2026 calibration that subtracts baseline CKA expected from random data.",{"name":82,"@type":73,"acceptedAnswer":83},"What practical tools and limits does the paper claim for the substrate?",{"text":84,"@type":76},"It provides uses including a frozen low-shot feature space, a domain detector, an auxiliary loss for knowledge distillation, and a Gröger-calibrated forensic fingerprint. The scope is bounded: it does not extend across modalities (vision + audio fails) and does not predict downstream transfer quality.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]