[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81855-en":3,"doc-seo-81855-105":31,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81855,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Automated Data Readiness for Scientific AI","Leadership computing facilities steward large-scale scientific datasets that require substantial transformation before serving as AI training data. Existing solutions do not fully unify automated transformation, readiness assessment, provenance tracking, and agent-native deployment. REDI, an open-source framework, introduces a unified five-stage pipeline—ingest, preprocess, transform, structure, and output—with per-stage instrumentation for reproducibility and agent-callable skills. Companion SetGo automates FAIR compliance and catalog publication. Evaluations across climate, proteomics, materials science, and nuclear fusion validate outputs against domain references and show near-ideal scaling.","Automated Data Readiness for Scientific AI  \nSean R. Wilkinson , Valentine G. Anantharaj , Jong Youl Choi , Ketan Maheshwari , Marshall McDonnell , Massimiliano Lupo Pasini , Polina Shpilker , Renan Souza , Patrick Widener , Sarp Oral , and Wesley Brewer   \nOak Ridge National Laboratory, Oak Ridge, TN, USA  \narXiv :2607 .0277 1v 1 [ cs .AI] 2 Jul 2026  \nAbstract—Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenanceinstrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.  \nIndex Terms—Data, Automation, Readiness, AI  \nI. INTRODUCTION  \nArtificial Intelligence (AI) is increasingly recognized asa transformative capability for scientific discovery, in part due to its ability to support data-driven modeling, surrogate simulation, uncertainty quantification, and the development of foundation models across a growing range of scientific domains. Augmenting traditional simulation and experimental workflows with learning-based approaches has the potential to accelerate discovery across scales and modalities.  \nRealizing this potential, however, depends critically on the availability of computable scientific data. Specifically, it requires datasets that are not only archived or preserved but also structured, validated, and semantically enriched such that they can be directly consumed by large-scale AI workflows. As scientific instruments and experiments produce ever-growing volumes of complex, heterogeneous data, the transition from raw data to AI-ready assets has emerged as a new and challenging component of the scientific data lifecycle.  \nTo this end, we present the Readiness Engine for Data Integration (REDI), 1 an open-source framework that automates AI data readiness through a unified five-stage pipeline, namely ingest, preprocess, transform, structure, and output. To our knowledge, REDI is the first framework to integrate  \nCorresponding author: Sean R. Wilkinson ([wilkinsonsr@ornl.gov](wilkinsonsr@ornl.gov)).  \n1 REDI source code: [https://doi.org/10.11578/dc.20260702.1](https://doi.org/10.11578/dc.20260702.1)  \nautomated transformation, readiness assessment, provenance tracking, and validation within a lifecycle-aware architecture for scientific computing.  \nA. The Data Readiness Challenge  \nThe term data readiness is inherently context-dependent because different computational workflows impose distinct requirements on data structure, semantics, and validation. It is therefore necessary to ask, “data readiness for what?” In this work, we primarily focus on readiness with respect to preparing data for training scientific AI models, particularly foundation models [1] . Within this scope, we define AI-ready data as machine-readable datasets that (a) have undergone domain-appropriate cleaning, validation, and feature engineering, and (b) include sufficient metadata to support reproducible model training and evaluat","cbCaicOxTp0Rod3t","https://ap.wps.com/l/cbCaicOxTp0Rod3t","pdf",516976,4,1,12,"English","en",105,"# Introduction\n## The Data Readiness Challenge\n## Cost and Complexity of Data Preparation","[{\"question\":\"What did evaluations show about scalability and performance costs?\",\"answer\":\"The climate-case results show near-ideal parallel scaling to 100 nodes on Frontier, while provenance-instrumented profiling identifies file I/O as the dominant pipeline cost and format selection as a key optimization lever.\"}]","Automated Data Readiness for Scientific AI | PDF",1784176660,30,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":29},"automated-data-readiness-for-scientific-ai","",{"@graph":37,"@context":78},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/automated-data-readiness-for-scientific-ai/81855/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What did evaluations show about scalability and performance costs?","Question",{"text":76,"@type":77},"The climate-case results show near-ideal parallel scaling to 100 nodes on Frontier, while provenance-instrumented profiling identifies file I/O as the dominant pipeline cost and format selection as a key optimization lever.","Answer","https://schema.org",{"og:url":53,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":47,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":47,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":30,"slug":114},"research-report",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":99,"slug":130},19,"General","general"]