[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118293-en":3,"doc-seo-118293-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118293,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Improving Radiography Machine Learning Workflows via Metadata Management for Training Data Selection","Machine learning models require repeated hyper-parameter tuning, feature engineering, and debugging to achieve strong results, and this complexity is increasingly difficult to manage as pipelines grow. In physical sciences, ongoing research generates large volumes of metadata that, when tracked and organized, can reduce redundant work, strengthen reproducibility, and improve feature engineering and training dataset selection. A case study presents a tool for machine-learning metadata management in dynamic radiography, evaluates it against an initial workflow, and outlines extensions for physical-science pipelines.","Improving Radiography Machine Learning Workflows via Metadata Management for Training Data Selection  \narXiv :2408 . 12655v1 [ cs .LG] 22 Aug 2024  \nMirabel Reid [mreid48@gatech.edu](mreid48@gatech.edu)[ ](mreid48@gatech.edu)Georgia Tech, LANL∗  \nChristine Sweeney [cahrens@lanl.gov](cahrens@lanl.gov)[ ](cahrens@lanl.gov)LANL  \nOleg Korobkin [korobkin@lanl.gov](korobkin@lanl.gov)[ ](korobkin@lanl.gov)LANL  \nSeptember 27, 2024  \nAbstract  \nMost machine learning models require many iterations of hyper-parameter tuning, feature engineering, and debugging to produce effective results. As machine learning models become more complicated, this pipeline becomes more difficult to manage effectively. In the physical sciences, there is an ever-increasing pool of metadata that is generated by the scientific research cycle. Tracking this metadata can reduce redundant work, improve reproducibility, and aid in the feature and training dataset engineering process. In this case study, we present a tool for machine learning metadata management in dynamic radiography. We evaluate the efficacy of this tool against the initial research workflow and discuss extensions to general machine learning pipelines in the physical sciences.1  \n1 Introduction  \nMachine learning (ML) is an increasingly integral part of scientific research. In the natural sciences, ML can recognize patterns in observational data and inform the development of scientific theories [Ros+20] . In applications such as materials science, an effective machine learning model can dramatically reduce the need for costly experiments and long development cycles [Wei+19; Car+19] . It is undeniable that ML can be an effective method for interpreting large volumes of observed or simulated data.  \nHowever, the introduction of a complex method like ML comes with challenges. Outside of simple models, most ML models require many iterations of hyper-parameter tuning, feature engineering, and debugging to produce effective results. Because of this, the workflow is often packaged into a pipeline: an automated or semi-automated program which prepares the data, trains, and evaluates the model in one swoop. As machine learning models become more complicated, the pipeline becomes more difficult to manage effectively. The creation of tools which automatically track metadata is an active area of development. Several industry leaders in ML such as Google [Goo] and Netflix [Net] have built and released their own open source metadata stores.  \nMost existing tools to manage the ML pipeline are aimed at commercial applications and webbased learning. There is a dearth of metadata management tools for the physical sciences. Figure 1 shows a generic pipeline which scientists may employ in the research cycle. While the scientific research cycle has similarities to the commercial development cycle, there are significant differences  \n∗Work completed while at Los Alamos National Laboratory  \n1 LA-UR-22-30449  \nthat often make commercial tools impractical to apply. For example, these tools tend to emphasize the capability to package and deploy models to production, a step which is often unnecessary in research in the physical sciences [Zah+18] . Additionally, commercial ML applications generally aim to improve model performance in order to maximize profits or improve the service [Zah+18], while scientific applications may also require that the model conforms to natural laws. In many cases, physical scientists use ML models as tools to gain scientific understanding or develop theories about causal relationships [Ros+20] . This leads to a different development cycle to which a pipeline management tool focused on monitoring performance may not adapt well. In addition, existing  \nFigure 1: A flowchart describing a generic pipeline for learning on simulation data. As the objectives change over time, each step may be repeated and fine-tuned. Ovals indicate the steps of machine learning training. Each step is marked with a cy","cbCaidzH3Vl721cD","https://ap.wps.com/l/cbCaidzH3Vl721cD","pdf",17125945,1,14,"English","en",105,"# Introduction\n# Contribution\n# Benefits of Metadata Tracking","[{\"question\":\"Why is metadata management important for training data selection in machine learning?\",\"answer\":\"Most ML workflows need many iterations of tuning and engineering, and physical-science projects generate extensive metadata during research and simulation. Tracking this metadata helps reduce redundant work, improves reproducibility, and supports better feature and training dataset engineering.\"},{\"question\":\"What tool is presented for dynamic radiography workflows?\",\"answer\":\"The document presents a tool to store and visualize metadata produced during a scientific project studying machine learning in dynamic radiography. It supports interactive exploration of the feature space, selection of training data via visual queries, and centralized tracking of training datasets with the parameters used to generate them.\"},{\"question\":\"Which capabilities were identified as essential beyond the previous workflow?\",\"answer\":\"Key capabilities include interactive investigation of the feature space to find degeneracies and multi-modal regions, training data selection based on visual queries, and centralized tracking of training datasets together with the query parameters used to generate them.\"}]","Improving Radiography Machine Learning Workflows via Metadata Management for Training Data Selection | PDF",1785682851,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-radiography-machine-learning-workflows-via-metadata-management-for-training-data-selection","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-radiography-machine-learning-workflows-via-metadata-management-for-training-data-selection/118293/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is metadata management important for training data selection in machine learning?","Question",{"text":75,"@type":76},"Most ML workflows need many iterations of tuning and engineering, and physical-science projects generate extensive metadata during research and simulation. Tracking this metadata helps reduce redundant work, improves reproducibility, and supports better feature and training dataset engineering.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What tool is presented for dynamic radiography workflows?",{"text":80,"@type":76},"The document presents a tool to store and visualize metadata produced during a scientific project studying machine learning in dynamic radiography. It supports interactive exploration of the feature space, selection of training data via visual queries, and centralized tracking of training datasets with the parameters used to generate them.",{"name":82,"@type":73,"acceptedAnswer":83},"Which capabilities were identified as essential beyond the previous workflow?",{"text":84,"@type":76},"Key capabilities include interactive investigation of the feature space to find degeneracies and multi-modal regions, training data selection based on visual queries, and centralized tracking of training datasets together with the query parameters used to generate them.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]