[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83386-en":3,"doc-seo-83386-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83386,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Cross-seed explainability using Procrustes conditioned Joint End-to-end Top-K Sparse Autoencoders","A Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) extracts cross-seed universal features from independently trained BERT models, tackling mechanistic interpretability issues caused by non-convex dictionary learning and geometric misalignment. The method computes an orthogonal Procrustes rotation between seeds’ activation spaces prior to joint SAE training, combining Top-K sparsity, end-to-end downstream optimization, and an auxiliary dead-feature revival loss. Across five seed pairs on SST-2, Stanford Politeness, and TweetEval Emotion, the full pipeline yields more universal features (Pearson r ≥ 0.70) than post-hoc alignment baselines. Qualitative analysis links high-universality features to interpretable sociolinguistic patterns.","arXiv :2607 .08499v 1 [ cs .CL] 9 Jul 2026  \nCross-seed explainability using Procrustes conditioned Joint End-to-end Top-K Sparse Autoencoders  \nBendeg´uz V´aradi 1,2 and Zolt´an Kmetty 1,2  \n1 Centre for Social Sciences, CSS-RECENS Research Group, Budapest 1097, Hungary.  \n2 Department of Sociology, Faculty of Social Sciences, E¨otv¨os Lor´and University, Budapest, Hungary.  \nAbstract  \nWe present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained BERT models. Cross-seed feature universality is a fundamental challenge in mechanistic interpretability: because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical features may differ by random initialization. We address this by computing an orthogonal Procrustes rotation between seeds’ activation spaces before joint SAE training, combining Top-K sparsity, end-to-end downstream optimization, and an auxiliary dead-feature revival loss based on previous SAE literature.  \nEvaluating on five independent seed pairs (ten BERT models) across three benchmark datasets (SST-2, Stanford Politeness, TweetEval Emotion), our full pipeline produces more universal features (Pearson r ≥ 0.70 across seeds) than post-hoc alignment baselines on all three datasets. A minimal qualitative analysis confirms that high-universality features encode interpretable sociolinguistic patterns.  \nKeywords: Sparse Autoencoders, XAI, BERT, mechanistic interpretability  \n1 Introduction  \nIn recent years, a large body of research has been published on the interpretability of language models [1] . Much of the recent literature aims to solve two fundamental issues: addressing polysemanticity [2–6] and improving the structural fidelity and functional  \n1  \nrobustness of the extracted features [4, 7 , 8] . In this paper, we propose an End-toEnd pipeline that combines recent novel methodologies to enhance the reliability of extracting interpretable concepts from BERT models. To demonstrate the feature universalities, we use three benchmark corpora to measure the feature correlations.  \nSparse Autoencoders (SAEs) have emerged at the forefront of mechanistic interpretability for disentangling polysemantic representations in large language models [1] . Recent advancements have rapidly improved the structural fidelity of these extracted features. TopK SAEs [4] eliminate the shrinkage bias inherent in L1 regularization, End-to-End SAEs [7] prevent feature splitting by optimizing for downstream consistency, and Orthogonal SAEs [9](under review at the time of citation) mitigate feature absorption by applying competition-aware orthogonality constraints to disentangle coactivating concepts. Despite these single-model improvements, achieving cross-seed”universal interpretability” [10] remains a fundamental challenge. Because dictionary learning is non-convex, independently trained networks, even those sharing the same architecture and data, suffer from feature splitting, where identical semantic concepts map to entirely different latent dimensions. Recent solutions, such as the Feature Aligned SAE [8], attempt to mitigate this by training multiple SAEs in parallel with a Mutual Feature Regularization (MFR) penalty to encourage decoder similarity. However, relying solely on training penalties may fail to address the underlying geometric misalignment of the models’ native activation spaces. In this paper, we present a novel architecture that directly solves this spatial misalignment through a Procrustesconditioned Joint End-to-End Top-K Sparse Autoencoder. Rather than penalizing separate dictionaries to force alignment, we calculate an Orthogonal Procrustes rotation matrix to superimpose the activation space of different model seeds before extracting concepts with a single, joint SAE. The Top-K activation enforces sparsity structurally, avoiding the L1 shrinkage bias, the P","cbCais91zzqStRMn","https://ap.wps.com/l/cbCais91zzqStRMn","pdf",616138,3,1,17,"English","en",105,"# Introduction\n# Contributions\n# Methodology","[{\"question\":\"What problem does cross-seed universal interpretability address in BERT models?\",\"answer\":\"Independently trained networks learn misaligned feature spaces due to non-convex dictionary learning, so the same semantic concepts can map to different latent dimensions. This makes apparently identical features non-universal across random initializations.\"},{\"question\":\"How does the proposed method use Procrustes conditioning during training?\",\"answer\":\"An orthogonal Procrustes rotation matrix is computed to align the activation spaces of different seeds before training a single joint SAE. This avoids forcing alignment only through training penalties and targets geometric misalignment directly.\"},{\"question\":\"What empirical results demonstrate the effectiveness of the full pipeline?\",\"answer\":\"Evaluations on five independent seed pairs (ten BERT models) across SST-2, Stanford Politeness, and TweetEval Emotion show more universal features than post-hoc alignment baselines, with Pearson correlation r ≥ 0.70 across seeds. High-universality features also correspond to interpretable sociolinguistic patterns in qualitative analysis.\"}]",1784187139,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cross-seed-explainability-using-procrustes-conditioned-joint-end-to-end-top-k-sparse-autoencoders","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cross-seed-explainability-using-procrustes-conditioned-joint-end-to-end-top-k-sparse-autoencoders/83386/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does cross-seed universal interpretability address in BERT models?","Question",{"text":75,"@type":76},"Independently trained networks learn misaligned feature spaces due to non-convex dictionary learning, so the same semantic concepts can map to different latent dimensions. This makes apparently identical features non-universal across random initializations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method use Procrustes conditioning during training?",{"text":80,"@type":76},"An orthogonal Procrustes rotation matrix is computed to align the activation spaces of different seeds before training a single joint SAE. This avoids forcing alignment only through training penalties and targets geometric misalignment directly.",{"name":82,"@type":73,"acceptedAnswer":83},"What empirical results demonstrate the effectiveness of the full pipeline?",{"text":84,"@type":76},"Evaluations on five independent seed pairs (ten BERT models) across SST-2, Stanford Politeness, and TweetEval Emotion show more universal features than post-hoc alignment baselines, with Pearson correlation r ≥ 0.70 across seeds. High-universality features also correspond to interpretable sociolinguistic patterns in qualitative analysis.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]