[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125685-en":3,"doc-seo-125685-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125685,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Discrepancies and the Error Evaluation Metrics for Machine Learning Interatomic Potentials","Machine learning interatomic potentials (MLIPs) provide a promising route for atomic modeling, yet low reported average errors leave an open question about whether MLIPs faithfully reproduce atomistic dynamics and physical properties in molecular dynamics simulations. This study compares state-of-the-art MLIPs with ab initio methods and identifies discrepancies tied to atom dynamics, defects, and rare events. It proposes quantitative, rare-event-based evaluation metrics that better reveal predictive failures, enabling improved performance across multiple properties and offering general guidance for more accurate MLIP testing.","Discrepancies and the Error Evaluation Metrics for Machine Learning Interatomic Potentials  \nYunsheng Liu 1 , Xingfeng He 1 , and Yifei Mo 1,*  \n1 Department of Materials Science and Engineering, University of Maryland, College Park, MD, USA  \n* Email: [yfmo@umd.edu](yfmo@umd.edu)  \nAbstract. Machine learning interatomic potentials (MLIPs) are a promising technique for atomic modeling. While small errors are widely reported for MLIPs, an open concern is whether MLIPs can accurately reproduce atomistic dynamics and related physical properties in molecular dynamics (MD) simulations. In this study, we examine the stateof-the-art MLIPs and uncover several discrepancies related to atom dynamics, defects, and rare events (REs), compared to ab initio methods. We find that low averaged errors by current MLIP testing are insufficient, and develop quantitative metrics that better indicate the accurate prediction of atomic dynamics by MLIPs. The MLIPs optimized by the RE-based evaluation metrics are demonstrated to have improved prediction in multiple properties. The identified errors, the evaluation metrics, and the proposed process of developing such metrics are general to MLIPs, thus providing valuable guidance for future testing and improvements of accurate and reliable MLIPs for atomistic  \nmodeling.  \n1. Introduction.  \nAtomistic modeling, which simulates physical phenomena based on the interactions of atoms, is a crucial research technique in a wide range of disciplines including physics, chemistry, biology, and materials science. Density functional theory (DFT) calculation have been the standard technique for evaluating atom interactions among a diverse range of configurations and chemistries, but their applications are limited to small system sizes of a few hundred atoms ( up to a few nm) due to high computation costs.1–3 By contrast, classical interatomic potentials , also known as force fields, have significantly lower computation costs and thus can be employed for atomistic simulations with much larger length-scale (nm – ߤ m) and longer time-scale (ns – ߤs), but they lack the transferability to different atomistic configurations that are not considered in the potential fitting. 1,4,5  \nAs an emerging technique to bridge the gaps among different computational techniques 1–3,6–10 , machine learning interatomic potentials (MLIPs) utilize machine learning (ML) models to predict energies and forces of atomistic structures, which are mapped into the atomistic descriptors as input. Current state-of-the-art MLIPs include Gaussian Approximation Potential (GAP) based on Smooth Overlap of Atomic Positions (SOAP) descriptors, 11,12 Neural Network Potential (NNP), 13,14 Spectral Neighbor Analysis Potential (SNAP), 15 Moment Tensor Potential (MTP), 16 and Deep Potential (DeePMD) models,5 and many MLIP variances derived from different modifications and combinations of ML models and descriptors. MLIPs are trained using DFT-calculated energies and forces from a diverse range of atomistic configurations, typically encompassing bulk and defected structures, equilibrium and non-equilibrium structures, and solid and liquid phases. Current state-of-the-art MLIPs are claimed to achieve accuracies similar to ab  \ninitio calculations, 1,5,11,13,15–18 while maintaining low computation costs and linear size scaling akin to classical interatomic potentials.  \nHowever, the MLIPs are black-box predictors not directly based on physical principles.  \nAn open question is whether MLIPs can always accurately reproduce physical phenomena in atomistic simulations. Conventional ML error testing primarily quantifies MLIP accuracies through average errors, such as root-mean-square error (RMSE) or mean-absolute error (MAE), of energies and atomic forces across a range of configurations known as testing dataset. These atomistic configurations in the testing dataset are randomly split from the entire datasets generated in the same manner as the training dataset, and t","cbCaivsYHutRRySR","https://ap.wps.com/l/cbCaivsYHutRRySR","pdf",2710396,1,52,"English","en",105,"# Introduction\n## Limits of classical potentials and the role of MLIPs\n## Conventional ML error testing and its insufficiency\n## Discrepancies revealed in dynamics, defects, and rare events","[{\"question\":\"Why are average error metrics like RMSE and MAE insufficient for MLIP validation?\",\"answer\":\"They quantify errors on randomly split testing sets that are similar to training data, but they may not reflect configurations encountered during molecular dynamics. Small average energy/force errors can still coincide with incorrect physical predictions such as migration barriers.\"},{\"question\":\"What kinds of discrepancies are examined in this study for ML interatomic potentials?\",\"answer\":\"The study focuses on discrepancies in atom dynamics, defects, and rare events when compared to ab initio methods, showing failures even when relevant defects were included in training or testing datasets.\"},{\"question\":\"How do the proposed evaluation metrics improve MLIP performance?\",\"answer\":\"The approach introduces quantitative metrics based on rare-event evaluation, and MLIPs optimized using these metrics demonstrate improved prediction across multiple properties. This helps ensure more reliable behavior in MD-relevant scenarios.\"}]","Discrepancies and the Error Evaluation Metrics for Machine Learning Interatomic Potentials | PDF",1785900655,131,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"discrepancies-and-the-error-evaluation-metrics-for-machine-learning-interatomic-potentials","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/discrepancies-and-the-error-evaluation-metrics-for-machine-learning-interatomic-potentials/125685/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are average error metrics like RMSE and MAE insufficient for MLIP validation?","Question",{"text":75,"@type":76},"They quantify errors on randomly split testing sets that are similar to training data, but they may not reflect configurations encountered during molecular dynamics. Small average energy/force errors can still coincide with incorrect physical predictions such as migration barriers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What kinds of discrepancies are examined in this study for ML interatomic potentials?",{"text":80,"@type":76},"The study focuses on discrepancies in atom dynamics, defects, and rare events when compared to ab initio methods, showing failures even when relevant defects were included in training or testing datasets.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the proposed evaluation metrics improve MLIP performance?",{"text":84,"@type":76},"The approach introduces quantitative metrics based on rare-event evaluation, and MLIPs optimized using these metrics demonstrate improved prediction across multiple properties. This helps ensure more reliable behavior in MD-relevant scenarios.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]