[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84464-en":3,"doc-seo-84464-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84464,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data","Improving Sparse Autoencoders (SAEs) requires benchmarks that can validate architectural innovations with minimal noise. Current LLM-based SAE benchmarks are too noisy, while existing synthetic-data studies are too small, unstandardized, and unrealistic. SynthSAEBench introduces a scalable synthetic benchmark and toolkit with realistic feature properties—correlation, hierarchy, and superposition—plus ground-truth feature directions and firings. It reproduces known SAE phenomena, diagnoses failure modes, and highlights a new overfitting issue in Matching Pursuit SAEs that leverages superposition noise to boost reconstruction without learning true features.","SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data  \nDavid Chanin 1 2 3 Adri Garriga-Alonso 2  \narXiv :2602 . 14687v2 [ cs .LG] 13 Jul 2026  \nAbstract  \nImproving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. Current LLM-based SAE benchmarks are too noisy to differentiate architectural improvements, while commonly used synthetic-data experiments are too small-scale, unstandardized, and unrealistic to be meaningful. We introduce SynthSAEBench, a benchmark and toolkit for evaluating SAEs against large-scale synthetic data with realistic feature characteristics including correlation, hierarchy, and superposition, while providing ground-truth features and firings. SynthSAEBench acts as a controlled lower-bound test: SAE architectures that fail when the Linear Representation Hypothesis holds by construction have little hope on real LLMs. The benchmark reproduces known LLM SAE phenomena including the disconnect between reconstruction and latent quality, poor SAE probing, and a precision–recall trade-off mediated by L0, demonstrating that SynthSAEBench findings reproduce results on LLM SAEs. We further identify a novel failure mode: Matching Pursuit SAEs exploit superposition noise to improve reconstruction without learning ground-truth features, suggesting more expressive encoding procedures can easily overfit. SynthSAEBench complements LLM benchmarks with ground-truth features and controlled ablations for diagnosing SAE failure modes, while providing a clear target for SAE architecture work.  \n1. Introduction  \nLarge language models (LLMs) achieve remarkable performance but remain opaque, motivating interpretability  \n1University College London 2MATS 3Decode Research. Correspondence to: David Chanin \u003C[david.chanin.22@ucl.ac.uk](david.chanin.22@ucl.ac.uk)>.  \nMechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026 . Copyright 2026 by the author(s) .  \nSynthSAEBench feature characteristics  \nHierarchy  \nAnimal  \nDog Bird  \nPoodle Husky Eagle Hawk  \nSuperposition  \n\n|  |  |\n| --- | --- |\n|  |  |\n\nFigure 1. SynthSAEBench provides a large-scale synthetic data model with realistic feature characteristics including correlation, hierarchy, superposition and zipfian firing distributions, scalable to hundreds of thousands of features and realistic hidden dimension sizes.  \nresearch into how these models represent knowledge. The Linear Representation Hypothesis (LRH) (Park et al., 2024) posits that concepts (hereafter “features”) are represented as nearly-orthogonal linear directions. Models can represent many more features than dimensions by allowing non-orthogonal directions, a phenomenon known as superposition (Elhage et al., 2022) . Superposition is efficient but makes interpreting activations difficult, motivating the use of Sparse Autoencoders (SAEs) (Bricken et al., 2023 ; Cunningham et al., 2024) to recover underlying feature directions via sparse dictionary learning.  \nA key challenge in improving SAEs is that we lack groundtruth knowledge of the “true features” in an LLM. LLM  \nbenchmarks such as SAEBench (Karvonen et al., 2025) evaluate SAE performance on tasks like sparse probing (Gurnee et al., 2023 ; Kantamneni et al., 2025), concept disentanglement (Karvonen et al., 2024), and autointerpretability (Paulo et al., 2025) . However, SAEBench metrics exhibit substantial noise between runs (see Appendix J), making it difficult to evaluate small architectural improvements. Moreover, without ground-truth access, we cannot diagnose why SAEs score poorly, a critical obstacle given recent work showing that SAEs underperform supervised methods like logistic-regression probes (Kantamneni et al., 2025) .  \nOn the other extreme, SAE research already relies on bespoke synthetic models: toy experiments with fewer than 10 independent features (Song et al., 2025 ; Gribonval & Schnass, 2010 ; Elh","cbCaioCuEsMq2rw1","https://ap.wps.com/l/cbCaioCuEsMq2rw1","pdf",945797,1,19,"English","en",105,"# Introduction\n## Feature characteristics in SynthSAEBench\n## Linear Representation Hypothesis and superposition\n## Limitations of existing SAE benchmarks\n## SynthSAEBench design and controlled evaluation\n## Reproducing known SAE phenomena and identifying failures","[{\"question\":\"What novel failure mode is identified using SynthSAEBench?\",\"answer\":\"Matching Pursuit SAEs can exploit superposition noise to improve reconstruction without learning the ground-truth features, suggesting that more expressive encoding procedures may overfit to synthetic artifacts.\"}]",1784195805,48,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"synthsaebench-evaluating-sparse-autoencoders-on-scalable-realistic-synthetic-data","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/synthsaebench-evaluating-sparse-autoencoders-on-scalable-realistic-synthetic-data/84464/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What novel failure mode is identified using SynthSAEBench?","Question",{"text":75,"@type":76},"Matching Pursuit SAEs can exploit superposition noise to improve reconstruction without learning the ground-truth features, suggesting that more expressive encoding procedures may overfit to synthetic artifacts.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},"General","general"]