[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120977-en":3,"doc-seo-120977-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120977,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Data Compression and Inference in Cosmology with Self-Supervised Machine Learning","Massive volumes of data from current and upcoming cosmological surveys require compression methods that summarize observations while preserving essential physical information. The work presents a self-supervised machine learning approach that builds representative summaries of large datasets using simulation-based augmentations. Experiments on hydrodynamical cosmological simulations show highly informative summaries that support precise parameter inference. The method also yields summary representations robust to systematic effects, including baryonic physics influences.","arXiv :2308 .09751v2 [ astro-ph .CO] 14 Dec 2023  \nData Compression and Inference in Cosmology with Self-Supervised Machine Learning  \nAizhan Akhmetzhanova, 1★ Siddharth Mishra-Sharma2,3, 1†, and Cora Dvorkin1‡  \n1 Department of Physics, Harvard University, Cambridge, MA 02138, USA  \n2 The NSF AI Institute for Artificial Intelligence and Fundamental Interactions, Cambridge, MA 02139, USA  \n3 Center for Theoretical Physics, Massachusetts Institute of Technology, Cambridge, MA 02139, USA  \n18 December 2023  \nABSTRACT  \nThe influx of massive amounts of data from current and upcoming cosmological surveys necessitates compression schemes that can efficiently summarize the data with minimal loss of information. We introduce a method that leverages the paradigm of self-supervised machine learning in a novel manner to construct representative summaries of massive datasets using simulationbased augmentations. Deploying the method on hydrodynamical cosmological simulations, we show that it can deliver highly informative summaries, which can be used for a variety of downstream tasks, including precise and accurate parameter inference. We demonstrate how this paradigm can be used to construct summary representations that are insensitive to prescribed systematic effects, such as the influence of baryonic physics. Our results indicate that self-supervised machine learning techniques offer a promising new approach for compression of cosmological data as well its analysis.  \nKey words: methods: data analysis – cosmology: miscellaneous.  \n1 INTRODUCTION the sufficiency of manually-derived statistics (i.e., ability to compress all physically-relevant information) is the exception rather than the  \nOver the last few decades, cosmology has undergone a phenomenal  \nnorm.  \ntransformation from a ‘data-starved’ field of research to a precision  \nscience. Current and upcoming cosmological surveys such as those In addition, given the estimated sizes of the datasets from future conducted by the Dark Energy Spectroscopic Instrument (DESI) surveys, even traditional summary statistics might be too large for (Aghamousa et al. 2016), Euclid (Laureĳs et al. 2011), the Vera scalable data analysis. For instance, Heavens et al. (2017) estimates C. Rubin Observatory (LSST Dark Energy Science Collaboration that for surveys such as Euclid or LSST, with an increased number 2012), and the Square Kilometer Array (SKA) (Weltman et al. 2020), of tomographic redshift bins, the total number of data points of among others, will provide massive amounts of data, and making full summary statistics for weak-lens∼ ing data (such as shear correlation  \nuse of these complex datasets to probe cosmology is a challenging functions) could be as high as 104, which might be prohibitively task. The necessity to create and manipulate simulations correspond- expensive when the covariance matrices for the data need to being to these observations further exacerbates this challenge. evaluated numerically from complex simulations. In order to take  \nAnalyzing the raw datasets directly is a computationally expen- advantage of the recent advances in the field of simulation-based sive procedure, so they are typically first described in terms of a inference (SBI) (e.g. Cranmer et al. (2020); Alsing et al. (2019)), theset of informative lower-dimensional data vectors or summary statis- size of the summary statistic presents an important consideration duetics, which are then used for parameter inference and other down- to the curse of dimensionality associated with the comparison of the stream tasks. These summary statistics are often motivated by in- simulated data to observations in a high-dimensional space (Alsing ductive biases drawn from the physics of the problem at hand. Some & Wandelt 2019) .  \nwidely used classes of summary statistics include power spectra and A number of methods have been proposed to construct optimal  \nhigher-order correlation functions (Chen et al. 2021b; Gualdi et al.","cbCaiv5tXIZHz7iw","https://ap.wps.com/l/cbCaiv5tXIZHz7iw","pdf",3136892,1,23,"English","en",105,"# Abstract\n# Introduction\n## Motivation: large-scale cosmological surveys and dimensionality challenges\n## Existing summary statistics and their limitations\n## Fisher-information-preserving and optimal compression methods\n## Score-function and nuisance-hardened compression approaches","[{\"question\":\"Why is data compression necessary in modern cosmology surveys?\",\"answer\":\"Upcoming surveys will generate extremely large datasets, making direct analysis and full high-dimensional summary statistics computationally expensive. Compression is needed to retain physically relevant information with minimal loss.\"},{\"question\":\"What core idea does the method use to compress cosmological data?\",\"answer\":\"The method leverages self-supervised machine learning with simulation-based augmentations to construct representative summaries from massive datasets.\"},{\"question\":\"How does the approach handle systematic effects such as baryonic physics?\",\"answer\":\"It constructs summary representations designed to be insensitive to prescribed systematic effects, including the influence of baryonic physics, improving robustness for downstream inference tasks.\"}]","Data Compression and Inference in Cosmology with Self-Supervised Machine Learning | PDF",1785733148,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"data-compression-and-inference-in-cosmology-with-self-supervised-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/data-compression-and-inference-in-cosmology-with-self-supervised-machine-learning/120977/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is data compression necessary in modern cosmology surveys?","Question",{"text":75,"@type":76},"Upcoming surveys will generate extremely large datasets, making direct analysis and full high-dimensional summary statistics computationally expensive. Compression is needed to retain physically relevant information with minimal loss.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What core idea does the method use to compress cosmological data?",{"text":80,"@type":76},"The method leverages self-supervised machine learning with simulation-based augmentations to construct representative summaries from massive datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the approach handle systematic effects such as baryonic physics?",{"text":84,"@type":76},"It constructs summary representations designed to be insensitive to prescribed systematic effects, including the influence of baryonic physics, improving robustness for downstream inference tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]