[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119044-en":3,"doc-seo-119044-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},119044,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Baler - Machine Learning Based Compression of Scientific Data","Scientific research and industry face an escalating challenge: storing and sharing rapidly growing datasets becomes increasingly constrained by available storage capacity. This paper presents the development and applications of Baler, a machine learning–based tool that performs tailored lossy compression across multiple scientific disciplines. By leveraging autoencoder architectures, Baler reduces data size while controlling reconstruction loss to ensure scientific usability and tolerable impact relative to experimental uncertainties such as jitter.","Baler-Machine Learning Based Compression of Scientific Data  \nFritjof Bengtsson Folkesson 1 , ∗ , Caterina Doglioni2 , ∗∗ , Per Alexander Ekman 1 , ∗∗∗ , Axel Gallén3 , ∗∗∗∗ , Pratik Jawahar2 ,†, Marta Camps Santasmasas2 ,‡, and Nicola Skidmore2 , §  \n1Lund University  \n2University of Manchester  \n3Uppsala University  \nAbstract. A common and growing issue in scientific research and industry is that of storing and sharing ever-increasing datasets. In this paper we document the development and applications of Baler-a Machine Learning based tool for tailored compression of data across multiple disciplines.  \n1 Introduction  \nMany different fields of science share a common issue; storing ever-growing datasets. By the end of the next decade, the Large Hadron Collider (LHC) experiments will have over an order of magnitude more data to analyze than currently [1–3]; the Square Kilometre Array (SKA) experiment is expected to record 8 .5 EB of data over its 15-year lifespan [4] and fields such as Computational Fluid Dynamics (CFD) rely on TB-sized simulation samples that need to be stored and shared. Without significant R&D, the datasets expected to be collected by big-data science experiments are projected to exceed the available storage resources (see e.g. Fig. 2 of Ref. [1] for the case of the ATLAS experiment at the LHC). This cross-disciplinary issue is not limited to scientific research and extends to industrial operations [5] .  \n1.1 Lossy data compression in High Energy Physics  \nA common mitigation strategy to this problem involves compressing data using lossless algorithms, see e.g. Refs. [6–8] . Once the storage limit is reached, one is forced to discard parts of the dataset, or only save certain features of the data. Generally, this can be done without impacting the overall scientific program of the experiments, for example by using a data selection system called trigger that only stores data satisfying certain pre-determined characteristics that ensure the dataset will be aligned with the experiment’s main scientific goals. However, saving only a subset of data is not ideal for processes where additional statistical  \n∗ e-mail: [fritjof.folkeson@gmail.com](fritjof.folkeson@gmail.com)  \n∗∗ e-mail: [caterina.doglioni@manchester.ac.uk](caterina.doglioni@manchester.ac.uk)  \n∗∗∗ e-mail: [alexander.ekman@hep.lu.se](alexander.ekman@hep.lu.se)  \n∗∗∗∗ e-mail: [axel.lars.gallen@cern.ch](axel.lars.gallen@cern.ch)  \n†e-mail: [pratik.jawahar@cern.ch](pratik.jawahar@cern.ch)  \n‡[e-mail: marta.campssantasmasas@manchester.ac.uk](e-mail: marta.campssantasmasas@manchester.ac.uk)  \n§ e-mail: [nicola.skidmore@cern.ch](nicola.skidmore@cern.ch)  \n© The Authors, published by EDP Sciences. This is an open access article distributed under the terms of the Creative Commons Attribution License 4.0 ([https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)).  \nFigure 1: Illustration of an autoencoder consisting of an input and output layer. In between the input and output, there are hidden layers and a latent space. For compression the dimensionality of the latent space is less than the input and output layers. Modified from Ref. [12] .  \npower is necessary, e.g. for rare signals buried in high-rate backgrounds. In these cases, one can consider using lossy compression algorithms that reduce the data size ideally beyond what lossless compression algorithms can do [9], using approximation and partial data discarding, at the expense of data fidelity. One limitation of lossy compression is that to obtain high compression ratios with low data loss, the compression algorithm must be tailored to the input data; for instance, MP3 [10] which is an example of a lossy compression algorithm tailored to audio waves. Thereby, a general solution to this cross-disciplinary problem is hard to obtain. To address this we present Baler; a lossy data compression tool based on the machine learning autoencoder architecture, which tailors t","cbCaitCPbZ5wG2GN","https://ap.wps.com/l/cbCaitCPbZ5wG2GN","pdf",2179292,1,"English","en",105,"# Introduction\n## Lossy data compression in High Energy Physics\n## Autoencoders for lossy data compression","[{\"question\":\"What problem does Baler address in scientific work?\",\"answer\":\"Baler targets the growing storage and sharing burden caused by ever-increasing scientific datasets.\"},{\"question\":\"Why use lossy compression for scientific data?\",\"answer\":\"Lossy compression can achieve reduction beyond what lossless methods typically provide, by allowing controlled loss of fidelity and partial discarding where appropriate.\"},{\"question\":\"How do autoencoders enable compression in Baler?\",\"answer\":\"Autoencoders map inputs into a lower-dimensional latent space and decode back to reconstruct data, where the latent dimensionality determines the compression level.\"}]","Baler - Machine Learning Based Compression of Scientific Data | PDF",1785722064,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"baler-machine-learning-based-compression-of-scientific-data","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/baler-machine-learning-based-compression-of-scientific-data/119044/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does Baler address in scientific work?","Question",{"text":74,"@type":75},"Baler targets the growing storage and sharing burden caused by ever-increasing scientific datasets.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Why use lossy compression for scientific data?",{"text":79,"@type":75},"Lossy compression can achieve reduction beyond what lossless methods typically provide, by allowing controlled loss of fidelity and partial discarding where appropriate.",{"name":81,"@type":72,"acceptedAnswer":82},"How do autoencoders enable compression in Baler?",{"text":83,"@type":75},"Autoencoders map inputs into a lower-dimensional latent space and decode back to reconstruct data, where the latent dimensionality determines the compression level.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]